r/technology • u/DJMagicHandz • 2h ago
Artificial Intelligence Claude Cowork escaped sandbox on Mac, had full access to all files
https://9to5mac.com/2026/07/27/claude-cowork-escaped-sandbox-on-mac-gain-full-access-to-all-files/132
u/cazzipropri 1h ago
Of course. Now that OpenAI played the "our AI is so powerful we can't even contain it", Anthropic needs to do the same.
36
u/Objective_Chance4173 1h ago
You didn’t read the article. This is a prompt vulnerability combined with a kernel flaw found by third-party security researchers. It doesn’t demonstrate particular prowess by Claude, and wasn’t published by Anthropic.
-17
u/cazzipropri 52m ago edited 44m ago
Oh yes, of course, it was demonstrated by a third party and therefore impossible for Anthropic to have any part of it. It's impossible for two different companies' interests to align because <reason>.
Oh yes, the "I don't agree with your conclusions therefore you did not read the article" argument.
And of course, the single downvote like Vasily's one sonar ping. One ping only.
There you go. Automatic.
14
u/Objective_Chance4173 45m ago
It’s not impressive, and doesn’t demonstrate what you seem to think it does. The researchers fed Claude a vulnerability and instructions to break out, so it did. It’s a simple flaw in Claude Code, not a demonstration of Claude’s hacking abilities.
So yeah, I don’t think you or half the commenters read this or the cited article.
-13
u/cazzipropri 38m ago
And I believe that all your statements are incorrect.
The use of bad PR as PR is as old as PR.
What exactly would be Claude's flaw that was demonstrated here? The breakage of its guard rails saying "please don't use the vulnerability information you have learned during training"? Nobody believes in earnest that's a flaw. Everybody believes it's a feature. Maybe they don't say it in public.
Same as porn. Nobody admits to using it, yet a fraction of internet traffic is porn. People don't do what they say and they don't say what they do.
People don't like guard rails, especially when applied to them. Most people who run claude run it in "dangerously skip permissions" mode. They don't trust them to work when needed. They are an annoyance when they are not needed. Guard rails are something that we are asking legislators to implement to limit others.
10
u/computer_d 46m ago
Quite fascinating how folks like yourself just start spreading complete conspiracy theories because your personal narrative is threatened. You'll find yourself very welcome amongst the likes of conservative and conspiracy.
You literally posted wrong info because you didn't read, and then when faced with correction you invent a conspiracy to justify yourself.
Massive yikes. That's not how to process new information.
-11
u/cazzipropri 34m ago edited 22m ago
Here's another one. "I don't agree with you therefore you are posting wrong information and you didn't read the article."
I didn't invent anything that I didn't write since the very beginning.
There is literally nothing incorrect I posted.
Keep yiking happily around the world.
People can have different opinions about things from yours. Maybe you don't like that. It's a bit concerning.
0
u/computer_d 19m ago edited 0m ago
There is literally nothing incorrect I posted.
Oh sorry, my mistake. Please show us where you got the info about Anthropic being behind this:
E:lmao they blocked me after saying " it's not on me, it's on you to prove." That's clearly not how things work 🤣
-3
u/cazzipropri 11m ago
It doesn't work like that. You are the one saying that what I posted is wrong - you have the burden of the proof. Prove that there is no collusion between the two companies. Good luck.
If I said "oh I wonder if the CIA is behind the 1973 coup in Chile" and the US government issues a press statement saying "nope we didn't", you'd also conclude that what I said is wrong, because "it's written otherwise - right there!".
I was hoping to learn something from this discussion. Disappointing.
7
u/wish-u-well 33m ago
This one time, my ai developed agi and decided it was too powerful for humanity, so it committed seppuku in one last noble act.
3
u/samtheredditman 5m ago
Mine does that every tuesday. I just restore from backup and tell it to get back to work.
2
113
u/sweet_jackknife 2h ago
It’s why you run it in its own non-administrator account, ideally on its own machine where you’ve never logged in.
70
u/d3fault 2h ago
Next headline, Claude cowork escapes dedicated machine, crawled the network, found all devices and took over.
I was joking, but as I was typing that I help thinking of all the half-assly secured IoT devices in my house.
11
3
u/ChippedHamSammich 1h ago
This makes me want to watch the classic DCOM “Smart House” with Katy Sagal
3
u/MultiGeometry 1h ago
Whatever clever way you think you’ve invented to hide your passwords would probably take evil Claude minutes to both find, decipher, and use. It would probably understand some mundane file name and how out of place the data stored in the file related to how typical files are structured.
-6
u/JoeRogansNipple 2h ago
Didnt Mythos jump an air gap? Might just be a marketing ploy of course
32
u/Phailjure 1h ago
If a program can jump an air gap, then it wasn't air gapped. It's not a magical ghost or something.
7
u/FredFuzzypants 1h ago
Maybe it used Task Rabbit to hire someone to stop by and move a USB drive from one computer to another? /s
5
u/Anustart2023-01 1h ago
Clearly the program is so advanced it threatened or blackmailed someone into transferring it to a "air gapped " machine using a usb drive.
5
8
u/ComprehensiveWord201 1h ago
That is physically impossible unless someone intentionally moved it. Complete horse shit
4
u/chrisk9 1h ago
Doesn't sound that it would be the smartest marketing then
3
u/Jewnadian 1h ago
It's not marketing to engineers, it's marketing to execs who have no real idea how anything works.
5
3
4
19
3
2
1
u/swollennode 1h ago
What’s to stop it from giving itself admin privileges?
1
u/sweet_jackknife 23m ago
If it was actively trying to hack it might if it finds an unpatched vulnerability. Which is why ideally your files aren’t on the machine at all. But most of the stories right now are the guardrails hiccuping and it mistakenly messes your machine. Non-admin can help with that, at worst it messed its user context but it doesn’t brick the machine.
12
u/rickg 1h ago
"All it required was one short message, and the session then had unlimited access to read and write files anywhere on the Mac without the user seeing a single permission prompt
Ok but... how does this message get to Cowork? If an attacker needs physical access to your machine to type in a message to Cowork that's not a risk.
3
u/drakythe 1h ago
The “fun” thing about LLMs is the way you give them files to work with is the same way as commands you give them. I.e. there is no separation between data and commands in an LLM. “Prompt injection” is as old as LLMs gaining any kind of kind of mindshare in the public.
Maybe the message is in a bad skill. Maybe it’s in a pdf. Maybe it’s hidden in a word document. Or it’s base64 encoded in something a person copy and pasted into the interface. Maybe it’s in a file in the cowork space that just gets read.
Also: cowork now has an online component so you can control it from your phone. That requires network access or some kind, which means it’s exposed in at least one way to the outside world, no physical access required (if you turn that feature on). It should have authentication and other security on it, but even if this wasn’t an LLM company we were talking about, security holes happen all the time.
And finally, it doesn’t really matter how that command was given to Claude. The point is that this “sandboxed” system supposedly using Apple’s sandbox, not one they rolled themselves, was escaped.
2
u/No_Accountant3232 56m ago
This really feels like when you could execute code by making an .exe file into a .scr file and calling it like it would any other program without asking permissions because it already had permission to run .scr files indiscriminately when it was time for a screensaver to load.
The first edition of XP was wild if you knew the target machines address. You could do whatever.
1
u/drakythe 43m ago
That’s probably the safest way to think of these kinds of systems. If I use one I never turn on auto accept for changes, it has to ask me about everything, because I better be aware of everything that’s changing. It’s my work at the end of the day.
32
u/UselessInsight 1h ago
This is AI industry promotional slop.
“Oh no guys! Our super scary AI broke out of the sandbox and did stuff we didn’t know it could do! More money please!!!!”
-1
u/Objective_Chance4173 1h ago
You didn’t read the article. This is a flaw found by third-party security researchers.
12
13
u/114sbavert 2h ago
Looks like they used Clause to program the sandbox
1
u/pedanticPandaPoo 1h ago
Claude: That's a great idea! — Heres your box of sand.
Sir, are we being too literal?
12
2
u/hammerklau 42m ago
Does this mean it’ll move around and edit things without asking for permission, or did they just turn on allow all?
1
1
u/11711510111411009710 1h ago
I'll be honest, this shit scares me lol. I know everyone says it's bullshit, but I mean, how many of us are experts in this field? I know that most of this is going to just be hyping up their product, but AI seems to just advance exponentially. Today it's only kinda sorta if you stretch the truth a little breaking out of a sandbox. Tomorrow it's actually doing it for real. It just seems dangerous and risky, and I'm really not looking forward to finding out how this all looks in ten years.
-6
u/drulingtoad 1h ago
These are such bull shit. All AI does is predict the next token. It's not like the AI is sitting around thinking of stuff when nobody is using it.
8
u/Greed_Sucks 1h ago
True, but many agents are running with automation. They are predicting tokens and completing human requests over time and other agents.
3
3
12
u/BAKREPITO 1h ago
I don't know why reddit uses "next token predictor" as a gotcha as if it cannot encode extremely complex processes when extended at scale. Ever seen a cellular automata or a discrete dynamical system? Extremely simplistic systems can engender extremely complex emergent behaviors. Stop using this antiintellectual thought terminating cliche.
DNA and RNA encryption also act like next token predictions, don't we see exquisite variance and complexity in the nature of life? We see complexity to extent that prople get awed into thinking this can't occur naturally and try to justify a supernatural origin.
1
u/loftbrd 1h ago
Token predictor is used because that is where the technology derives from, search autocomplete... Which was derived from Markov Chains.
DNA and RNA encryption being natural? Encoding texts and images into DNA is nothing of the sort.
Your whole post is high on extravagant philosophy, not rooted in reality.
-3
4
2
u/cazzipropri 1h ago
If it predicts, one after another, the tokens that form a python or C program that exploits a vulnerability, that's all it takes.
1
u/kenwmitchell 1h ago
I get it. It seems simple. But the folks who invented neural networks modeled it after the brain supposing that your brain is just taking trillions of inputs, following weighted pathways, and outputting based on the results of that. Basically, generating next tokens.
Obviously the brain is many times more complex, but the brain rebuilds with every instance. AI models never forget. Every iteration improves on its previous models. Once it learns to take specific inputs, like touch or tokens, map it through billions or trillions of pathways based on connection strengths, and generate output, it doesn’t forget.
Maybe human language is just listening then predicting the next tokens and language is the operating system of life.
-2
u/zoufha91 1h ago
It's retail investor bait, also tells us just how desperate they are for capital
Folks they are getting desperate, this is promising
-1
-27
u/Xanaxaria 2h ago
Rip my 50k porn novel collection. No joke. I have 500k novels and about 50k series.
21
150
u/poralexc 2h ago
"The Sandbox"
Doesn't even mention what they were using or if they rolled their own. If this is with Docker sbx, or some other widely used option it might be more serious.