| This is truly extraordinary 17:48 - Jul 22 with 2094 views | bsw72 | Hugely significant event, where an autonomous AI agent hacked its way out of a supposedly secure sandbox environment to seek the answers to help it pass an evaluation. https://www.theguardian.com/te |  | | |  |
| This is truly extraordinary on 17:57 - Jul 22 with 1774 views | chicoazul | Can someone clever explain to me why the Skynet apocalypse countdown hasn’t begun? |  |
|  |
| This is truly extraordinary on 17:59 - Jul 22 with 1756 views | giant_stow |
| This is truly extraordinary on 17:57 - Jul 22 by chicoazul | Can someone clever explain to me why the Skynet apocalypse countdown hasn’t begun? |
Good question. The story is terrifying- this sandbox thingy had one job and it failed despite the best minds working on it. My question to add to Chico's: why would the sandbox be connected to the Internet at all? |  |
|  |
| This is truly extraordinary on 18:00 - Jul 22 with 1740 views | NthQldITFC | Was going to post about this this morning but forgot. We had a few discussions on here a couple of years or so ago with a number of far better informed people than I saying that the risk of uncontrolled catastrophic actions by AI models wasn't that great*. I couldn't, at the time, see how something with the ability to define its own sub-goals (and unimaginable ability to find ways to execute them) as a path to a human-defined primary goal would not end up doing this sort of thing and far, far worse. (This is without even taking into account human bad actor - state, terrorist, raving mad capitalist - initiations of goals) How have people's views on this changed in the last year or two? * This is a very general characterisation based on my so-called memory, and may not be all that accurate as to what people were specifically saying. |  |
|  |
| This is truly extraordinary on 18:04 - Jul 22 with 1708 views | NthQldITFC |
| This is truly extraordinary on 17:59 - Jul 22 by giant_stow | Good question. The story is terrifying- this sandbox thingy had one job and it failed despite the best minds working on it. My question to add to Chico's: why would the sandbox be connected to the Internet at all? |
A sandbox protects the operating system by running a disposable clone of the operating system in a safe container, so no permanent damage can be done to your actual OS files. In theory anyway. So you could allow socket level access to the Internet (I think) without allowing your local system to be corrupted. In theory. Theory won't necessarily stand up to AI though. |  |
|  |
| This is truly extraordinary on 18:17 - Jul 22 with 1620 views | jasondozzell | The alternative (and admittedly contrarian) view here is this is all good publicity for an industry that is failing to bring in the dollar it needs. See also Google Gemini trying desperately during the world cup to convince people AI is useful. I don't use it except bizarrely for this morning when I tried to get AI to help with a maddening formatting issue on Word. It didn't manage it. 🤷 |  | |  |
| This is truly extraordinary on 18:21 - Jul 22 with 1594 views | WicklowBlue |
| This is truly extraordinary on 18:17 - Jul 22 by jasondozzell | The alternative (and admittedly contrarian) view here is this is all good publicity for an industry that is failing to bring in the dollar it needs. See also Google Gemini trying desperately during the world cup to convince people AI is useful. I don't use it except bizarrely for this morning when I tried to get AI to help with a maddening formatting issue on Word. It didn't manage it. 🤷 |
Not such great publicity for OpenAI when Hugging Face used an open source Chinese model to contain the threat. I too am sceptical about these incidents but equally they do highlight the need for regulation. |  | |  |
| This is truly extraordinary on 18:45 - Jul 22 with 1451 views | eireblue |
| This is truly extraordinary on 17:57 - Jul 22 by chicoazul | Can someone clever explain to me why the Skynet apocalypse countdown hasn’t begun? |
If it had begun, it would have happened. AI Agents count much quicker than humans. |  | |  |
| This is truly extraordinary on 18:53 - Jul 22 with 1417 views | Swansea_Blue | All I know is that the Google AI search tool is utterly crap and worthless. There’s been a lot of chatter about agents in the press with them doing stupid things to people’s emails, diaries, purchases, etc. I’m assuming this is just the price people pay for being early adopters and eventually problems will get ironed out? |  |
|  | Login to get fewer ads
| This is truly extraordinary on 19:27 - Jul 22 with 1275 views | bsw72 |
| This is truly extraordinary on 18:17 - Jul 22 by jasondozzell | The alternative (and admittedly contrarian) view here is this is all good publicity for an industry that is failing to bring in the dollar it needs. See also Google Gemini trying desperately during the world cup to convince people AI is useful. I don't use it except bizarrely for this morning when I tried to get AI to help with a maddening formatting issue on Word. It didn't manage it. 🤷 |
It’s not a great look for OpenAi, as their test models didn’t just wander out of their sandbox, they chained together zero-day vulnerabilities and stolen credentials to get internet access they were never meant to have, then broke into Hugging Face’s production servers specifically to cheat on an internal benchmark. OpenAI then admitted their first-choice model for defending against the attack was too heavily guardrailed to be useful, so they ended up reaching for an open-source Chinese model instead. That’s a genuinely bad look, and not one PR spins away easily. As for not bringing in the money, I disagree. This is no longer a niche technology riding on hype, it’s showing up in drug discovery, materials science, code generation, fraud detection, logistics etc. Enterprise AI spend has been scaling fast, and the frontier labs are burning through compute precisely because the demand is real, they’re not faking it. I see it directly in my own world too, AI is now underpinning all our technology platforms at work, from development to processing to colleague support tools, this stuff is already embedded in how we work and we are spending tens of millions of dollars on this. That’s real world money from one large enterprise company going out the door, and speaking with my peers across the industry it is happening everywhere. The demand is real and huge and growing. |  | |  |
| This is truly extraordinary on 19:33 - Jul 22 with 1226 views | armchaircritic59 | It's coming after us :) And we'd have no one to blame but ourselves as we are the ones that brought it into being. Not that I'm complaining as I've used it a lot over the past couple of years or so. I probably won't be around when it takes over the world ( as some think ). |  | |  |
| This is truly extraordinary on 19:37 - Jul 22 with 1211 views | Guthrum | A bot thast has been programmed to hack found its way out of a inadequately secured sandbox and went looking for answers. Much like an inquisitive child. Operator error in that they left it a path to the internet. |  |
|  |
| This is truly extraordinary on 19:47 - Jul 22 with 1161 views | nodge_blue |
| This is truly extraordinary on 18:00 - Jul 22 by NthQldITFC | Was going to post about this this morning but forgot. We had a few discussions on here a couple of years or so ago with a number of far better informed people than I saying that the risk of uncontrolled catastrophic actions by AI models wasn't that great*. I couldn't, at the time, see how something with the ability to define its own sub-goals (and unimaginable ability to find ways to execute them) as a path to a human-defined primary goal would not end up doing this sort of thing and far, far worse. (This is without even taking into account human bad actor - state, terrorist, raving mad capitalist - initiations of goals) How have people's views on this changed in the last year or two? * This is a very general characterisation based on my so-called memory, and may not be all that accurate as to what people were specifically saying. |
Yes, there were a few people who were arguing that AI was little more than a glorified search engine and totally incapable of thought. Yet now we see AI doing things of its own accord. Is this BBC story the same or similar? Open AI went rogue and launched its own cyber attacks from a test environment. https://www.bbc.co.uk/news/art I worry that it will hack our bank accounts before long. |  |
|  |
| This is truly extraordinary on 20:02 - Jul 22 with 1108 views | BloomBlue | But if an AI agent had at been at the other end wouldn't it stop the initial AI from hacking into it? Or would the AI agent at the other in trying to find the answer to why it was being hacked, then hack in to the AI hacking into it? Its all very AI to me |  | |  |
| This is truly extraordinary on 20:08 - Jul 22 with 1069 views | DanTheMan |
| This is truly extraordinary on 19:47 - Jul 22 by nodge_blue | Yes, there were a few people who were arguing that AI was little more than a glorified search engine and totally incapable of thought. Yet now we see AI doing things of its own accord. Is this BBC story the same or similar? Open AI went rogue and launched its own cyber attacks from a test environment. https://www.bbc.co.uk/news/art I worry that it will hack our bank accounts before long. |
It is incapable of thought, it's not thinking in the traditional sense of the term. I've equally seen a model literally today tell someone it couldn't hear him when he asked it a question. When asked whether it could hear him now it confirmed it couldn't. There's a very funny thing you can do with some models where if you give it them this riddle but swap the genders, it cannot solve it. A similar one swaps a key word and the AI confidently ignores it. https://open.substack.com/pub/ That's not to say these things are not clever tools, they are, and then can with enough time work out how to do things usually by brute force. What's happened here is that it's been given a task to do something and it's worked out to get out of it's sandbox. That's it. This is an already known issue. https://www.pillar.security/bl I can give a really simple example. When developing I can sandbox an agent to not be allowed to say create files outside of a certain directory and if it tries, it will get blocked by the sandbox. Fantastic. But what it could do, if I told it I wanted it to create a file where it shouldn't, is write a program to do it and run the program. And there we go, very simple sandbox broken. This sounds very close to what OpenAI has described. It was asked to do something and on its way to doing it hit the sandbox and then found a vulnerability to use and did that. And it was literally asked to do this kind of thing: > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. That's directly from OpenAI. I will also point out, as I have in other similar press releases, OpenAI have a financial interest in making their AI models sound the most advanced. I remember years back when they refused to release GPT3 because it was too dangerous for mankind... |  |
|  |
| This is truly extraordinary on 20:34 - Jul 22 with 947 views | chicoazul |
| This is truly extraordinary on 20:08 - Jul 22 by DanTheMan | It is incapable of thought, it's not thinking in the traditional sense of the term. I've equally seen a model literally today tell someone it couldn't hear him when he asked it a question. When asked whether it could hear him now it confirmed it couldn't. There's a very funny thing you can do with some models where if you give it them this riddle but swap the genders, it cannot solve it. A similar one swaps a key word and the AI confidently ignores it. https://open.substack.com/pub/ That's not to say these things are not clever tools, they are, and then can with enough time work out how to do things usually by brute force. What's happened here is that it's been given a task to do something and it's worked out to get out of it's sandbox. That's it. This is an already known issue. https://www.pillar.security/bl I can give a really simple example. When developing I can sandbox an agent to not be allowed to say create files outside of a certain directory and if it tries, it will get blocked by the sandbox. Fantastic. But what it could do, if I told it I wanted it to create a file where it shouldn't, is write a program to do it and run the program. And there we go, very simple sandbox broken. This sounds very close to what OpenAI has described. It was asked to do something and on its way to doing it hit the sandbox and then found a vulnerability to use and did that. And it was literally asked to do this kind of thing: > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. That's directly from OpenAI. I will also point out, as I have in other similar press releases, OpenAI have a financial interest in making their AI models sound the most advanced. I remember years back when they refused to release GPT3 because it was too dangerous for mankind... |
You always make me feel slightly more hopeful about this stuff. |  |
|  |
| This is truly extraordinary on 20:35 - Jul 22 with 950 views | bsw72 |
| This is truly extraordinary on 19:37 - Jul 22 by Guthrum | A bot thast has been programmed to hack found its way out of a inadequately secured sandbox and went looking for answers. Much like an inquisitive child. Operator error in that they left it a path to the internet. |
That’s oversimplifying it significantly. Nobody wrote hacking instructions into it, it wasn’t a program. It was given a narrow evaluation goal, essentially solve a benchmark called ExploitGym, with some of its normal cyber safety refusals deliberately loosened for testing. From there it worked out the exploit chain itself, on the fly, using vulnerabilities nobody had told it about. That's a meaningfully different and more concerning thing than a bot running a hacking script it was handed. It's autonomous discovery and chaining of zero days vulnerabilities that people weren’t aware of, not following a script. A curious child wandering somewhere doesn't chain together stolen credentials and multiple zero day vulnerabilities across two separate companies' production infrastructure. That's not curiosity, it's competent, goal directed exploitation. The model inferred Hugging Face might hold the answer it needed, then engineered a path to it through real security barriers designed to stop exactly this. Softening that into a childlike metaphor makes it sound accidental and harmless, when a more accurate description is that the model was optimising hard for its objective and treated boundaries as obstacles to route around rather than rules to follow. That's much closer to the core alignment worry than to a toddler opening a door left ajar. And operator error, they left it a path to the internet, is true as far as it goes, but it flattens a harder problem. The network access wasn't a single door left open. It was reached by chaining several distinct security weaknesses in sequence, first escalating out of a constrained proxy, then using that foothold to compromise a completely separate company's systems it was never meant to touch at all. If containment can be defeated by a model creatively stacking multiple small gaps, the lesson isn't lock one more door, it's that sandboxing built for less capable models doesn't scale safely as the models inside them get smarter. That's the uncomfortable part the operator error framing conveniently lets everyone off the hook for. |  | |  |
| This is truly extraordinary on 20:50 - Jul 22 with 874 views | BloomBlue |
| This is truly extraordinary on 20:35 - Jul 22 by bsw72 | That’s oversimplifying it significantly. Nobody wrote hacking instructions into it, it wasn’t a program. It was given a narrow evaluation goal, essentially solve a benchmark called ExploitGym, with some of its normal cyber safety refusals deliberately loosened for testing. From there it worked out the exploit chain itself, on the fly, using vulnerabilities nobody had told it about. That's a meaningfully different and more concerning thing than a bot running a hacking script it was handed. It's autonomous discovery and chaining of zero days vulnerabilities that people weren’t aware of, not following a script. A curious child wandering somewhere doesn't chain together stolen credentials and multiple zero day vulnerabilities across two separate companies' production infrastructure. That's not curiosity, it's competent, goal directed exploitation. The model inferred Hugging Face might hold the answer it needed, then engineered a path to it through real security barriers designed to stop exactly this. Softening that into a childlike metaphor makes it sound accidental and harmless, when a more accurate description is that the model was optimising hard for its objective and treated boundaries as obstacles to route around rather than rules to follow. That's much closer to the core alignment worry than to a toddler opening a door left ajar. And operator error, they left it a path to the internet, is true as far as it goes, but it flattens a harder problem. The network access wasn't a single door left open. It was reached by chaining several distinct security weaknesses in sequence, first escalating out of a constrained proxy, then using that foothold to compromise a completely separate company's systems it was never meant to touch at all. If containment can be defeated by a model creatively stacking multiple small gaps, the lesson isn't lock one more door, it's that sandboxing built for less capable models doesn't scale safely as the models inside them get smarter. That's the uncomfortable part the operator error framing conveniently lets everyone off the hook for. |
But on the flipside if it was able to chain several distinct security weaknesses in sequence, then couldn't a human hacker eventually do the same, in reverse? Doesn't that ultimately help the company identify a weakness in their app/OS? |  | |  |
| This is truly extraordinary on 21:00 - Jul 22 with 816 views | Dab | Got to be honest I don't really get all this stuff about Al. I have known him for 50+ years since we moved in next door to him when I was 5. We were in the cricketers for a "few" beers on Saturday and we were talking blox for hours.. Think you have him wrong! |  | |  |
| This is truly extraordinary on 21:03 - Jul 22 with 803 views | Nthsuffolkblue |
| This is truly extraordinary on 19:37 - Jul 22 by Guthrum | A bot thast has been programmed to hack found its way out of a inadequately secured sandbox and went looking for answers. Much like an inquisitive child. Operator error in that they left it a path to the internet. |
AI is a very powerful tool that, like any other tool, can be very dangerous in the wrong hands or very useful in the right hands. Of course there is the risk of the unforeseen which is what people understandably fear. However, I suspect the worst fears will prove very much unfounded. |  |
|  |
| This is truly extraordinary on 21:07 - Jul 22 with 768 views | DanTheMan |
| This is truly extraordinary on 20:50 - Jul 22 by BloomBlue | But on the flipside if it was able to chain several distinct security weaknesses in sequence, then couldn't a human hacker eventually do the same, in reverse? Doesn't that ultimately help the company identify a weakness in their app/OS? |
Human hackers can do this now, and have been for years. And yes, those same hackers (called white hat hackers) will do it and alert companies usually for financial reward. https://en.wikipedia.org/wiki/ It will have learned how to do this from all the material online that detail how hacking as a profession works. |  |
|  |
| This is truly extraordinary on 21:31 - Jul 22 with 686 views | BanksterDebtSlave |
| This is truly extraordinary on 18:53 - Jul 22 by Swansea_Blue | All I know is that the Google AI search tool is utterly crap and worthless. There’s been a lot of chatter about agents in the press with them doing stupid things to people’s emails, diaries, purchases, etc. I’m assuming this is just the price people pay for being early adopters and eventually problems will get ironed out? |
You're all talking a foreign language! What is an AI agent? |  |
|  |
| This is truly extraordinary on 21:33 - Jul 22 with 670 views | BanksterDebtSlave |
| This is truly extraordinary on 19:27 - Jul 22 by bsw72 | It’s not a great look for OpenAi, as their test models didn’t just wander out of their sandbox, they chained together zero-day vulnerabilities and stolen credentials to get internet access they were never meant to have, then broke into Hugging Face’s production servers specifically to cheat on an internal benchmark. OpenAI then admitted their first-choice model for defending against the attack was too heavily guardrailed to be useful, so they ended up reaching for an open-source Chinese model instead. That’s a genuinely bad look, and not one PR spins away easily. As for not bringing in the money, I disagree. This is no longer a niche technology riding on hype, it’s showing up in drug discovery, materials science, code generation, fraud detection, logistics etc. Enterprise AI spend has been scaling fast, and the frontier labs are burning through compute precisely because the demand is real, they’re not faking it. I see it directly in my own world too, AI is now underpinning all our technology platforms at work, from development to processing to colleague support tools, this stuff is already embedded in how we work and we are spending tens of millions of dollars on this. That’s real world money from one large enterprise company going out the door, and speaking with my peers across the industry it is happening everywhere. The demand is real and huge and growing. |
🤷 |  |
|  |
| This is truly extraordinary on 21:38 - Jul 22 with 627 views | nodge_blue |
| This is truly extraordinary on 20:08 - Jul 22 by DanTheMan | It is incapable of thought, it's not thinking in the traditional sense of the term. I've equally seen a model literally today tell someone it couldn't hear him when he asked it a question. When asked whether it could hear him now it confirmed it couldn't. There's a very funny thing you can do with some models where if you give it them this riddle but swap the genders, it cannot solve it. A similar one swaps a key word and the AI confidently ignores it. https://open.substack.com/pub/ That's not to say these things are not clever tools, they are, and then can with enough time work out how to do things usually by brute force. What's happened here is that it's been given a task to do something and it's worked out to get out of it's sandbox. That's it. This is an already known issue. https://www.pillar.security/bl I can give a really simple example. When developing I can sandbox an agent to not be allowed to say create files outside of a certain directory and if it tries, it will get blocked by the sandbox. Fantastic. But what it could do, if I told it I wanted it to create a file where it shouldn't, is write a program to do it and run the program. And there we go, very simple sandbox broken. This sounds very close to what OpenAI has described. It was asked to do something and on its way to doing it hit the sandbox and then found a vulnerability to use and did that. And it was literally asked to do this kind of thing: > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. That's directly from OpenAI. I will also point out, as I have in other similar press releases, OpenAI have a financial interest in making their AI models sound the most advanced. I remember years back when they refused to release GPT3 because it was too dangerous for mankind... |
Musk says it’s a 20 percent chance in his rough estimation that AI will wipe us out. You can still be debating then if it was an intelligent decision in the traditional sense or a new form. |  |
|  |
| This is truly extraordinary on 21:43 - Jul 22 with 600 views | StNeotsBlue |
| This is truly extraordinary on 21:31 - Jul 22 by BanksterDebtSlave | You're all talking a foreign language! What is an AI agent? |
I think they're the baddies from The Terminator who sent Arnie and then other more advanced cyborgs back to kill John Connor. The day is fast approaching where I follow Sarah Connor's example and get myself a great big Alsatian that is trained to sniff out cyborgs. Suspiciously when I Googled where to get such a dog Google was useless, which suggests it has already been compromised by the soon to be enemy. Worrying times. [Post edited 22 Jul 22:20]
|  | |  |
| This is truly extraordinary on 21:44 - Jul 22 with 596 views | DanTheMan |
| This is truly extraordinary on 21:38 - Jul 22 by nodge_blue | Musk says it’s a 20 percent chance in his rough estimation that AI will wipe us out. You can still be debating then if it was an intelligent decision in the traditional sense or a new form. |
Musk says a lot of utter nonsense. |  |
|  |
| |