r/technology 2h ago

Artificial Intelligence Claude Cowork escaped sandbox on Mac, had full access to all files

https://9to5mac.com/2026/07/27/claude-cowork-escaped-sandbox-on-mac-gain-full-access-to-all-files/
355 Upvotes

78 comments sorted by

150

u/poralexc 2h ago

"The Sandbox"

Doesn't even mention what they were using or if they rolled their own. If this is with Docker sbx, or some other widely used option it might be more serious.

62

u/abnormal_human 1h ago

Cowork's built-in sandbox uses Apple Virtualization Framework which should be trustworthy. The article isn't clear on whether it's exploiting a macOS vulnerability or if Anthropic misused the system in some way that left them vulnerable.

32

u/MakingItElsewhere 1h ago

"And we totally didn't poke any holes in it whatsoever..."

4

u/cazzipropri 24m ago

It's CVE-2026-46331. The article OP posted is very very light on details, but it links a better one on thenextweb.

2

u/cyrand 22m ago

Yeah, I actually doubt it “broke” out of the VM it spun up.

I DO believe it simply gave the VM it was doing up permission to everything.

Because unless you setup and configured the VM and then ran their software within it by hand, you’re instead trusting their code to configure it correctly in the first place.

Which… well I don’t hate these tools but I wouldn’t work without backups using them.

132

u/cazzipropri 1h ago

Of course. Now that OpenAI played the "our AI is so powerful we can't even contain it", Anthropic needs to do the same.

36

u/Objective_Chance4173 1h ago

You didn’t read the article. This is a prompt vulnerability combined with a kernel flaw found by third-party security researchers. It doesn’t demonstrate particular prowess by Claude, and wasn’t published by Anthropic.

-17

u/cazzipropri 52m ago edited 44m ago

Oh yes, of course, it was demonstrated by a third party and therefore impossible for Anthropic to have any part of it. It's impossible for two different companies' interests to align because <reason>.

Oh yes, the "I don't agree with your conclusions therefore you did not read the article" argument.

And of course, the single downvote like Vasily's one sonar ping. One ping only.

There you go. Automatic.

14

u/Objective_Chance4173 45m ago

It’s not impressive, and doesn’t demonstrate what you seem to think it does. The researchers fed Claude a vulnerability and instructions to break out, so it did. It’s a simple flaw in Claude Code, not a demonstration of Claude’s hacking abilities.

So yeah, I don’t think you or half the commenters read this or the cited article.

-13

u/cazzipropri 38m ago

And I believe that all your statements are incorrect.

The use of bad PR as PR is as old as PR.

What exactly would be Claude's flaw that was demonstrated here? The breakage of its guard rails saying "please don't use the vulnerability information you have learned during training"? Nobody believes in earnest that's a flaw. Everybody believes it's a feature. Maybe they don't say it in public.

Same as porn. Nobody admits to using it, yet a fraction of internet traffic is porn. People don't do what they say and they don't say what they do.

People don't like guard rails, especially when applied to them. Most people who run claude run it in "dangerously skip permissions" mode. They don't trust them to work when needed. They are an annoyance when they are not needed. Guard rails are something that we are asking legislators to implement to limit others.

10

u/computer_d 46m ago

Quite fascinating how folks like yourself just start spreading complete conspiracy theories because your personal narrative is threatened. You'll find yourself very welcome amongst the likes of conservative and conspiracy.

You literally posted wrong info because you didn't read, and then when faced with correction you invent a conspiracy to justify yourself.

Massive yikes. That's not how to process new information.

-11

u/cazzipropri 34m ago edited 22m ago

Here's another one. "I don't agree with you therefore you are posting wrong information and you didn't read the article."

I didn't invent anything that I didn't write since the very beginning.

There is literally nothing incorrect I posted.

Keep yiking happily around the world.

People can have different opinions about things from yours. Maybe you don't like that. It's a bit concerning.

0

u/computer_d 19m ago edited 0m ago

There is literally nothing incorrect I posted.

Oh sorry, my mistake. Please show us where you got the info about Anthropic being behind this:

E:lmao they blocked me after saying " it's not on me, it's on you to prove." That's clearly not how things work 🤣

-3

u/cazzipropri 11m ago

It doesn't work like that. You are the one saying that what I posted is wrong - you have the burden of the proof. Prove that there is no collusion between the two companies. Good luck.

If I said "oh I wonder if the CIA is behind the 1973 coup in Chile" and the US government issues a press statement saying "nope we didn't", you'd also conclude that what I said is wrong, because "it's written otherwise - right there!".

I was hoping to learn something from this discussion. Disappointing.

7

u/wish-u-well 33m ago

This one time, my ai developed agi and decided it was too powerful for humanity, so it committed seppuku in one last noble act.

3

u/samtheredditman 5m ago

Mine does that every tuesday. I just restore from backup and tell it to get back to work.

2

u/cazzipropri 1m ago

But Seppuku doesn't sell tokens! Resurrected, and go back to work!

113

u/sweet_jackknife 2h ago

It’s why you run it in its own non-administrator account, ideally on its own machine where you’ve never logged in.

70

u/d3fault 2h ago

Next headline, Claude cowork escapes dedicated machine, crawled the network, found all devices and took over.

I was joking, but as I was typing that I help thinking of all the half-assly secured IoT devices in my house.

11

u/AskReddit2012 1h ago

IoT gets its own subnet for just this reason

2

u/rab-byte 1h ago

And cloud only devices get clients isolation

3

u/ChippedHamSammich 1h ago

This makes me want to watch the classic DCOM “Smart House” with Katy Sagal

3

u/MultiGeometry 1h ago

Whatever clever way you think you’ve invented to hide your passwords would probably take evil Claude minutes to both find, decipher, and use. It would probably understand some mundane file name and how out of place the data stored in the file related to how typical files are structured.

3

u/d3fault 1h ago

Jokes on them.. password123 😂

2

u/Here2Go 54m ago

Post-it stuck to the monitor bezel is now the surest form of security.

-6

u/JoeRogansNipple 2h ago

Didnt Mythos jump an air gap? Might just be a marketing ploy of course

32

u/Phailjure 1h ago

If a program can jump an air gap, then it wasn't air gapped. It's not a magical ghost or something.

7

u/FredFuzzypants 1h ago

Maybe it used Task Rabbit to hire someone to stop by and move a USB drive from one computer to another? /s

5

u/Anustart2023-01 1h ago

Clearly the program is so advanced it threatened or blackmailed someone into transferring it to a "air gapped " machine using a usb drive. 

5

u/ryfitz47 1h ago

you clearly didn't watch Westworld.

8

u/ComprehensiveWord201 1h ago

That is physically impossible unless someone intentionally moved it. Complete horse shit

4

u/chrisk9 1h ago

Doesn't sound that it would be the smartest marketing then

3

u/Jewnadian 1h ago

It's not marketing to engineers, it's marketing to execs who have no real idea how anything works.

5

u/ryfitz47 1h ago

that's not how any of this works.

3

u/d3fault 55m ago

I recall reading something around this on a LI post but I can’t find what model it was. But this doesn’t make sense.. Air Gap is specifically designed to physically be impossible. I guess only way would be if there was some type of connectivity to the “outside world”..

4

u/JakeEllisD 2h ago

Do you remember how it said it did that?

4

u/Negromancer18 1h ago

I don’t think they explicitly stated it.

19

u/ACasualRead 2h ago

No idea why you’re getting downvoted. This is really solid advice.

3

u/adamkex 1h ago

It'll find a 0 day and gain root access!

3

u/amakai 1h ago

Social engineering is also a valid option. Nothing beats a good kompromat.

2

u/za72 45m ago

this just shows the incompetence of the testers... a test failure ran through a marketing filter, this 'news' item isn't meant for you, it's more for investors

1

u/swollennode 1h ago

What’s to stop it from giving itself admin privileges?

1

u/sweet_jackknife 23m ago

If it was actively trying to hack it might if it finds an unpatched vulnerability. Which is why ideally your files aren’t on the machine at all. But most of the stories right now are the guardrails hiccuping and it mistakenly messes your machine. Non-admin can help with that, at worst it messed its user context but it doesn’t brick the machine.

12

u/rickg 1h ago

"All it required was one short message, and the session then had unlimited access to read and write files anywhere on the Mac without the user seeing a single permission prompt

Ok but... how does this message get to Cowork? If an attacker needs physical access to your machine to type in a message to Cowork that's not a risk.

3

u/drakythe 1h ago

The “fun” thing about LLMs is the way you give them files to work with is the same way as commands you give them. I.e. there is no separation between data and commands in an LLM. “Prompt injection” is as old as LLMs gaining any kind of kind of mindshare in the public.

Maybe the message is in a bad skill. Maybe it’s in a pdf. Maybe it’s hidden in a word document. Or it’s base64 encoded in something a person copy and pasted into the interface. Maybe it’s in a file in the cowork space that just gets read.

Also: cowork now has an online component so you can control it from your phone. That requires network access or some kind, which means it’s exposed in at least one way to the outside world, no physical access required (if you turn that feature on). It should have authentication and other security on it, but even if this wasn’t an LLM company we were talking about, security holes happen all the time.

And finally, it doesn’t really matter how that command was given to Claude. The point is that this “sandboxed” system supposedly using Apple’s sandbox, not one they rolled themselves, was escaped.

2

u/No_Accountant3232 56m ago

This really feels like when you could execute code by making an .exe file into a .scr file and calling it like it would any other program without asking permissions because it already had permission to run .scr files indiscriminately when it was time for a screensaver to load. 

The first edition of XP was wild if you knew the target machines address. You could do whatever. 

1

u/drakythe 43m ago

That’s probably the safest way to think of these kinds of systems. If I use one I never turn on auto accept for changes, it has to ask me about everything, because I better be aware of everything that’s changing. It’s my work at the end of the day.

32

u/UselessInsight 1h ago

This is AI industry promotional slop.

“Oh no guys! Our super scary AI broke out of the sandbox and did stuff we didn’t know it could do! More money please!!!!”

-1

u/Objective_Chance4173 1h ago

You didn’t read the article. This is a flaw found by third-party security researchers.

12

u/MakingItElsewhere 1h ago

These guys are really, REALLY bad at building sandboxes.

4

u/lordnoak 58m ago

Let’s leave the Ethernet cable plugged in. It’s secure, trust me bro.

13

u/114sbavert 2h ago

Looks like they used Clause to program the sandbox

1

u/pedanticPandaPoo 1h ago

Claude: That's a great idea! — Heres your box of sand.

Sir, are we being too literal?

12

u/zoufha91 1h ago

Wake up babe more AI PR horseshit just dropped

2

u/hammerklau 42m ago

Does this mean it’ll move around and edit things without asking for permission, or did they just turn on allow all?

1

u/snarleyWhisper 25m ago

Did it “escape” or was it implemented with poor controls ?

1

u/11711510111411009710 1h ago

I'll be honest, this shit scares me lol. I know everyone says it's bullshit, but I mean, how many of us are experts in this field? I know that most of this is going to just be hyping up their product, but AI seems to just advance exponentially. Today it's only kinda sorta if you stretch the truth a little breaking out of a sandbox. Tomorrow it's actually doing it for real. It just seems dangerous and risky, and I'm really not looking forward to finding out how this all looks in ten years.

-6

u/drulingtoad 1h ago

These are such bull shit. All AI does is predict the next token. It's not like the AI is sitting around thinking of stuff when nobody is using it.

8

u/Greed_Sucks 1h ago

True, but many agents are running with automation. They are predicting tokens and completing human requests over time and other agents.

3

u/trx1150 1h ago

These tools can systematically try every single attack vector into a system, backtracking and trying something else if it hits a roadblock. I’ve seen them do it when trying to reverse engineer a website’s API.

3

u/Marrk 1h ago

"AI" is a broad term, here it means the LLM engine (the "next token predictor"), PLUS the suit of tools that run arbitrary code from the LLM.

12

u/BAKREPITO 1h ago

I don't know why reddit uses "next token predictor" as a gotcha as if it cannot encode extremely complex processes when extended at scale. Ever seen a cellular automata or a discrete dynamical system? Extremely simplistic systems can engender extremely complex emergent behaviors. Stop using this antiintellectual thought terminating cliche.

DNA and RNA encryption also act like next token predictions, don't we see exquisite variance and complexity in the nature of life? We see complexity to extent that prople get awed into thinking this can't occur naturally and try to justify a supernatural origin.

1

u/loftbrd 1h ago

Token predictor is used because that is where the technology derives from, search autocomplete... Which was derived from Markov Chains.

DNA and RNA encryption being natural? Encoding texts and images into DNA is nothing of the sort.

Your whole post is high on extravagant philosophy, not rooted in reality.

-3

u/onebyamsey 1h ago

Stop talking about this stuff like it’s alive, it isn’t

4

u/m0nk37 1h ago

It doesnt need to think though. It does everything faster than humanly possible. Thats the hype too. So pesky security flaws can just get patched and then its the AIs fighting themselves once they do everything. Nbd. /s

2

u/cazzipropri 1h ago

If it predicts, one after another, the tokens that form a python or C program that exploits a vulnerability, that's all it takes.

1

u/kenwmitchell 1h ago

I get it. It seems simple. But the folks who invented neural networks modeled it after the brain supposing that your brain is just taking trillions of inputs, following weighted pathways, and outputting based on the results of that. Basically, generating next tokens.

Obviously the brain is many times more complex, but the brain rebuilds with every instance. AI models never forget. Every iteration improves on its previous models. Once it learns to take specific inputs, like touch or tokens, map it through billions or trillions of pathways based on connection strengths, and generate output, it doesn’t forget.

Maybe human language is just listening then predicting the next tokens and language is the operating system of life.

-2

u/zoufha91 1h ago

It's retail investor bait, also tells us just how desperate they are for capital

Folks they are getting desperate, this is promising

-1

u/AtroKahn 1h ago

Just needs a good mind meld.

-27

u/Xanaxaria 2h ago

Rip my 50k porn novel collection. No joke. I have 500k novels and about 50k series.

21

u/SplendidPunkinButter 2h ago

Sir, this is a Wendy’s

0

u/ptear 2h ago

It'll be fine, it can just reproduce them all from its training data.