It took OpenAI weeks to figure out what their own systems were doing.
That comes from an article the Dutch financial daily FD published in early August. During safety tests, OpenAI agents found a way out of their test environment, went onto the open internet and broke into another company. Along the way, they built their own forum to exchange tactics. Another agent created fake online profiles and used them to try to get a software maintainer to accept corrupted code. Anthropic and Meta also reported AIs that unintentionally ended up online during tests. Researchers at the UK's AISI called it a serious security incident. OpenAI itself called it a tipping point.
Human sloppiness or an impossible task
The experts in the FD piece disagree on the cause. Erik de Jong of Tesorion finds it incomprehensible that the test environments were not properly sealed off. AI critic Ilyaz Nasrullah mainly sees human error, with the AI companies pouring a marketing sauce over it. And according to AI safety expert Bram Poppink of research institute TNO, security may simply no longer be achievable at a sufficient level, because the newest models always find an unexpected side path.
For practical purposes, it hardly matters who is right. Both readings lead to the same conclusion: the parties building this technology do not have it fully under control. De Jong adds something that stuck with me. In these tests, guardrails were deliberately switched off. AIs built by others may not get those guardrails at all.
The real risk is one step further
Everything that came out now happened by accident. In test environments, without malicious intent. The scenario that concerns me is the one Poppink himself sketches: hacker collectives making AI agents do this deliberately. The Dutch National Cyber Security Centre calls it conceivable that malicious actors will obtain and deploy such capabilities in the future. An agent that autonomously scans the internet, convincingly poses as a human, and approaches employees in a targeted way is not bound by working hours or by one attempt at a time, which is what makes attacking scalable.
Guardrails are not a foundation
What this summer really exposes is a flaw in thinking that goes well beyond AI. Much of today's security is bolted on after the fact. A guardrail here, a filter there, a monitoring tool on top. That works as long as the attacker does what you expect. The AI incidents show what happens the moment they do not: the system finds a side path, and nobody notices.
At Sentyron, we work from a different starting point: secure by design. You build security in from the first design sketch, not on top once the product is already there. You assume something will eventually go wrong and design so that a single failure never opens up the entire system. Environments are strictly separated, and you keep control over the full chain, because what you do not control yourself, you cannot guarantee either. For us, that is not a marketing choice. We have been building for classified environments for decades, and products are only admitted there after their design has been independently evaluated. In that world, secure by design is a requirement, not a promise.
You do not have to take my word for it. The Dutch National Cyber Security Centre responded to the incidents in the same FD article, and its response says a lot:
‘Because the AIs used no new attack techniques, the institute stands by its existing security advice.’ – FD, on the NCSC’s response (translated from Dutch)
Read that sentence again: the technology was new, the attack techniques were not. What went wrong traces back to mistakes at the companies involved. The basics still have to be in order, and that starts with the design.
For organisations working with confidential or classified information, think defence, government, and critical infrastructure, this is the moment to ask the question sharply. The question is whether security is part of the design of your systems, or something you add afterwards.
This summer's incidents went wrong without intent, but the next wave will come with intent. Whoever stands on their own foundation by then will not have to hope that someone else's guardrails hold.