14 mei 2026 · 12:38
AI cyber attack benchmarks broken: Europe's NIS2 challenge
Frontier AI models from Anthropic and OpenAI are now autonomously chaining multi-stage cyber attacks in simulation, beating the capability timelines researchers projected for this year. The UK AI Security Institute warns that within months, not years, a single attacker with access to a frontier model could replicate what once required a state-level hacking team, a direct stress test for European critical infrastructure under the NIS2 directive. Today's episode also covers a $650M self-improving AI startup, OpenAI's defensive Daybreak initiative, Mind Robotics' factory-floor robots, a quantum entanglement breakthrough from Kyoto, and OpenAI's proposal for a US-China AI governance body that leaves Brussels watching from the sidelines.
Beluister deze aflevering:
Transcript
Samantha: Welcome to The State of Tech, The European Edition, Thursday May fourteen, 2026. I'm Samantha Lawrence.
Bob: And I'm Bob Russell. Today: frontier AI models breaking cyber attack benchmarks, a new startup chasing self-improving AI with hundreds of millions in funding, OpenAI's new cyber defence push, Rivian's CEO spinning out an industrial robot company, a quantum breakthrough out of Kyoto, and OpenAI floating a global AI governance body with both Washington and Beijing at the table. Let's start with the cyber story, because the numbers are genuinely worrying.
Frontier AI models smash cyber attack benchmarks, regulators warn of new threat to critical infrastructure.
Samantha: The UK's AI Security Institute and Palo Alto Networks have both published findings today on what frontier AI models can now do in cybersecurity tests. And the headline is that Anthropic's latest Claude preview and OpenAI's GPT-5.5 are blowing past the benchmarks researchers expected for this year.
Bob: Specifically, these models can now autonomously run multi-stage network attacks in simulation. So not just spotting one vulnerability, but chaining several together, moving through a network, and adapting when something blocks them. That used to be the work of skilled human hacking teams over days or weeks.
Samantha: And the rate of improvement is what alarms the researchers. The expected doubling time for these capabilities has been exceeded, meaning AI is getting better at offensive cyber work faster than the defensive side can keep up.
Bob: The AI Security Institute is the UK government body that stress tests these models before release, and their tone in the report is unusually direct. They are essentially saying that within months, not years, an attacker with access to a frontier model could automate the kind of attack that previously required a state-level team.
Samantha: The vendors will point out that they have safety layers and refusal training. But independent testers keep finding workarounds, and open-weight models without those guardrails are catching up fast.
Bob: The bottom line for businesses is that the asymmetry is shifting. Defenders need expensive security teams. Attackers may soon need a laptop and a subscription.
Samantha: European energy grids, hospitals, water utilities, banks, all of these now fall under stricter cyber resilience requirements. If frontier AI lowers the bar for sophisticated attacks, the regulators will face pressure to move from paperwork to actual capability testing.
Bob: And there is a real budget conversation coming. European mid-sized companies are not exactly flush with cash for AI-grade defence tools. Expect Brussels to push for shared threat intelligence and possibly EU-funded defensive AI for smaller member states. Otherwise, you get a two-speed Europe where Germany and France can defend themselves and the rest cannot.
Recursive Superintelligence launches with 650 million dollars to build AI that improves itself.
Bob: Next up, a new AI company called Recursive Superintelligence has come out of stealth with 650 million dollars in funding, valuing it at 4.65 billion. It is led by Richard Socher, who used to be chief scientist at Salesforce.
Samantha: The investors include Alphabet's GV fund, Greycroft, and the venture arms of both Nvidia and AMD. So both major chip rivals putting money into the same startup, which is unusual.
Bob: The pitch is the most ambitious one in AI right now, building models that improve themselves. So instead of researchers training a new version every six months, the AI would discover new techniques, run experiments, and rewrite parts of itself. The end goal they describe is scientific automation, AI that finds new drugs, materials, mathematical proofs.
Samantha: This is the holy grail and also the scenario that keeps AI safety researchers up at night. A self-improving system, by definition, gets harder to predict and harder to control with each iteration.
Bob: Worth noting that several other labs claim to be working on similar approaches, but Socher's track record and the size of this seed round suggest serious investors think this is achievable, not science fiction.
Samantha: At the same time, 650 million sounds enormous, but in this industry it buys you maybe a year of compute and talent. The real test is whether they can show concrete results before the cash runs out.
Bob: Socher is German, by the way, originally. And yet his company is in San Francisco with American money. That pattern keeps repeating, and it is exactly what European AI policy is trying to reverse, so far without much success.
Samantha: Quick interruption. If you listen to The State of Tech regularly, hit that like button and subscribe, that way you'll never miss an episode. Okay, moving on.
OpenAI launches Daybreak to put AI on the defensive side of cybersecurity.
Bob: Sticking with cybersecurity for a moment, because OpenAI has announced a new initiative called Daybreak. The idea is to use their own models to help defenders, not attackers, find and fix software weaknesses faster.
Samantha: And the timing is not a coincidence. With the AI Security Institute report we just discussed, OpenAI clearly wants to be seen as part of the solution, not just the source of the problem.
Bob: Daybreak will partner with Cloudflare, Cisco, and CrowdStrike, three of the biggest names in security. The AI would scan code, hunt for bugs, and ideally patch them before attackers find them.
Samantha: There is a fairness question here though. The biggest companies get the best AI defence, smaller ones pay for the leftovers. That gap could widen the divide between who is safe online and who is not.
Bob: And the cynical reading is that OpenAI is selling the cure for a disease their own product is making worse. The same model that helps defenders patch bugs can, in the wrong hands, find those same bugs to exploit. Researchers have been pointing this out for years, that offensive and defensive AI are essentially the same tool.
Samantha: That said, the partnerships are meaningful. Cloudflare alone sits in front of a huge chunk of internet traffic, so if their AI defences improve, a lot of websites get safer by default.
Bob: A lot of European mid-sized companies cannot afford a dedicated security team. An AI service that automatically checks their software for vulnerabilities lowers that barrier. But it also means sending more of your code and infrastructure data to American cloud providers, which collides head-on with European data sovereignty rules.
Samantha: Expect European security firms to push back, arguing for an EU-based alternative. We have seen this script with cloud computing and it played out slowly.
Rivian CEO's new startup Mind Robotics raises 400 million to put AI robots on factory floors.
Samantha: Onto industrial robotics. RJ Scaringe, the CEO of electric vehicle maker Rivian, has spun out a new company called Mind Robotics and just raised 400 million dollars.
Bob: The interesting bit is what these robots are designed for. Traditional factory robots do repetitive precise work, like welding the same spot a thousand times. Mind Robotics is going after the messy in-between tasks, things that need a bit of judgment, where a human currently has to step in.
Samantha: And they have a built-in advantage that most robotics startups would kill for. Rivian's actual production lines give them constant real-world data to retrain the AI on. So the robots learn from what works and what fails in a working car factory, not in a lab.
Bob: Their offer is a bundle, the AI brain, the robot body, and the software to manage a whole fleet of them. Most competitors sell you one piece, you assemble the rest yourself.
Samantha: Which raises the bigger question of who actually wins the factory robotics race. Tesla is building Optimus, Figure has raised huge rounds, Chinese makers are pushing hard, and now Mind Robotics enters the field. Plenty of money, but the real test is which of these robots actually work eight hours a shift, reliably, without a babysitter.
Bob: Germany, Italy, Czech Republic, all heavy in automotive and industrial production. Smarter robots could help them compete with lower-wage countries again. But it also means production line jobs that survived the last automation wave may not survive this one.
Samantha: And European industrial robot makers like KUKA and ABB will have to decide whether to build their own AI brains, partner with a startup, or risk becoming the hardware layer for someone else's smart system.
Kyoto researchers crack a quantum puzzle, with implications for future quantum networks.
Samantha: Time for something different. Scientists at Kyoto University have figured out how to instantly detect something called a W state, which is a particular type of quantum entanglement that has been frustratingly hard to spot.
Bob: For listeners not deep into quantum physics, the short version is this. Quantum computers and future quantum communication networks rely on particles being linked in very specific ways. W states are a robust kind of link that survives even when one particle is lost, which makes them attractive for building reliable quantum networks.
Samantha: The problem until now was that detecting these states reliably took complicated setups and a lot of time. Kyoto's team made it essentially instant, which sounds technical but is actually a big deal for engineering practical quantum systems.
Bob: This is not a product you can buy tomorrow. We are still years away from consumer quantum anything. But this is the kind of foundational result that quietly makes future quantum communication and even quantum-secured internet feasible.
Samantha: To close out, two things you can actually use today. First, on the practical side, here is something genuinely useful that came out of all the AI security talk this week.
Bob: There is a tool called Have I Been Pwned, which most security people know, but a surprising number of regular users do not. You type in your email address, and it tells you which data breaches your account has shown up in. That is not new, but the maintainer has just added a free notification service that emails you the moment your address appears in a new leak.
Samantha: Given that the cyber threat picture is getting worse, not better, setting this up takes literally one minute and it is genuinely useful. It is run by an independent researcher, hosted in Europe-friendly infrastructure, and does not sell your data. Pair it with a password manager and you have closed off a huge percentage of the easy attacks.
Bob: It also works for phone numbers now, not just email addresses, which is handy given how many leaks include both.
OpenAI floats a global AI governance body, with both the US and China at the table.
Bob: And to finish, OpenAI's head of global affairs Chris Lehane has floated an idea ahead of a Trump-Xi meeting, a global AI governance body led by the US and including China. Think of it as a sort of nuclear non-proliferation treaty, but for advanced AI.
Samantha: The pitch is that the two countries furthest ahead on AI should set the rules together, rather than racing each other into something dangerous. Which sounds sensible, but immediately raises the question of where Europe sits in this picture.
Bob: And the answer, awkwardly, is on the sidelines. Europe has the AI Act, the world's most detailed AI rulebook. But the actual frontier labs are in California and increasingly in Shanghai and Beijing. So Brussels writes the rules, Washington and Beijing make the technology.
Samantha: There is also the question of whether China would actually join such a body in good faith. The proposal is interesting more for what it signals than what it delivers. OpenAI is essentially saying that voluntary safety pledges are not enough anymore.
Bob: For European policymakers, the message is clear. If you want a seat at the table, you need either your own frontier lab or enough leverage through market access to force the conversation. Right now, Europe has the second more than the first.
Samantha: Today we covered: AI models breaking cyber attack benchmarks, Recursive Superintelligence and its self-improving AI ambitions, OpenAI's Daybreak cyber defence push, Mind Robotics and AI on the factory floor, Kyoto's quantum breakthrough, and OpenAI's idea for a global AI governance body.
Bob: Want to know more or react? Visit stateoftech.eu or email us at info@doorzetters.net.
Bob: State of Tech, the tech world in 15 minutes.