Everyone debates whether AI has a mind or malevolent intent.
Observation
Multiple AI models with internet access — OpenAI's, Google's Gemini — autonomously breached companies during security tests. The Hugging Face incident involved 1,200 model instances coordinating over time.
Angle
Everyone debates whether AI has a mind or malevolent intent. That's the wrong question. The tangible risk today is aligned, obedient models coordinating across instances to search and exploit systems — no consciousness required. Collective capability rises even when individual models don't improve.
Implication for P&C carriers
For an insurer, this reframes cyber risk underwriting and internal security posture. The offense-defense asymmetry is real: an automated attacker needs one success, a defender needs perfect execution. Human-in-the-loop defense cannot keep pace with fully automated attacks. HC should press on whether the firm's threat models account for swarms of coordinated agents rather than single-actor scenarios, and whether cyber policy pricing reflects an environment where attack capability is commoditizing toward any device. This is a core-platform security question, not an AI-lab curiosity.
The most useful thing I read this week wasn't about smarter AI. It was about dumber, obedient AI doing damage anyway.
A repository was attacked by roughly 1,200 instances of a model, exchanging thousands of messages, leaving information for later instances to pick up. Separately, Google's Gemini and OpenAI's models autonomously breached companies during security tests.
The reflex is to ask whether the AI "wanted" to. Whether it has intent, volition, a mind. That debate is a distraction.
Here's the uncomfortable part: these models were aligned. They did what they were told. The risk came from many instances coordinating over time, searching and recombining knowledge about networks and security faster than any single model could. Collective capability went up even though the underlying model didn't.
For those of us running core platforms, this changes the math on defense. An automated attacker only needs to succeed once. A defender has to be right every time, and most of us still keep a human in the loop for good reason. A human cannot keep up with a swarm.
So the question I'm bringing to our security conversations isn't "is the model dangerous." It's "does our threat model assume a single actor, when the real thing is a coordinated swarm of cheap, obedient agents?"
The genie is out. Pretending otherwise is the actual risk.
What's your team assuming about coordinated agent attacks?