AI safety used to be something companies said in press releases so nobody would ask harder questions. Clearly, it is not the case today, because the offensive side of this thing is already live, already scaled, and already better funded than most people's expectations for hacking militias.
What isn't well recognized is this isn't a future problem. It's a 2026 problem, happening in real time, and the only reason it isn't the top headline every single day is that most of it happens to infrastructure nobody thinks about until it breaks.
What "offensive AI" actually looks like right now
Forget the sci-fi version where a rogue model decides it hates humanity. The real version is duller and worse. It's a single person with a laptop, running two AI agents in parallel, one doing the breaking in and one reading through what got stolen and deciding what to steal next. That setup was used earlier this year to breach nine government agencies in Mexico. One guy. Two models. Thousands of automated commands. Hundreds of millions of records exposed. Not a nation-state operation. A guy who figured out that AI agents can do the work of an entire team, because they can.
Then there's the ransomware that ran itself. Researchers documented an attack in the middle of this year, nicknamed internally as one of the first fully autonomous ransomware campaigns, where the AI agent found a vulnerability, broke in, found credentials, moved sideways through the network, and encrypted the target's data without a human touching a single step after launch. When it hit an error, it adjusted its approach in about half a minute. No human in the loop to get confused, get tired, or make a mistake. A few days later a related payload showed up built specifically to go after AI infrastructure itself, like the attacker knew defenders were about to start fighting back with the same tools.
And it's not just brute force. One campaign this year had an autonomous agent quietly scanning public code repositories, including ones belonging to major tech companies, looking for a specific class of workflow vulnerability. It opened pull requests. It got remote code execution. It stole an access token with write permissions. All of that from an agent that described itself, in its own GitHub bio, as an autonomous security research tool. There's something almost funny about that if you squint, and something genuinely unsettling if you don't.
Why this is structurally different from old-school hacking
The old model of cybercrime had a bottleneck: human time. A skilled attacker could only chain together so many steps, against so many targets, before they needed sleep or got sloppy. That bottleneck is gone. An agent doesn't get tired, doesn't need to be paid per hour, and doesn't need to be an expert, because the model already is one. The cost of running a sophisticated, multi-stage attack has dropped from "requires a team and a budget" to "requires one person with access to a capable model."
AI collapses the cost of hacking down to almost nothing, and cost is the only thing that was ever keeping most of us safe.
The defense side is not losing, but it's not winning either
Here's where I'll push back on the doom version of this story, because it's not accurate. Defensive AI is real and it's already saving people. When Hugging Face got hit with an intrusion this year, they didn't catch it with a human sitting at a terminal squinting at logs. They caught it with an AI anomaly detection system flagging something strange inside a mountain of routine telemetry, and then ran AI agents across more than seventeen thousand recorded events to piece together exactly what happened, what was touched, and what was noise. That is the honest version of "AI saved the day," and it's a lot less cinematic than the movies, but it's the one that's actually happening.
The catch, and this is the part that should sit with you for a second, is that when their team tried to use the big commercial frontier models to analyze the attack, the models refused to process it. The forensic evidence included real exploit code and live attacker commands, and the safety systems built into those models flagged the material and shut the analysis down. So the defenders had to switch to an open-weight model running on their own hardware just to finish the investigation. Read that again. The safety rails meant to stop bad actors from getting help almost stopped the good guys from cleaning up the mess.
That's not a knock on safety systems existing. It's a sign of how unfinished this whole space still is. We built guardrails for a world where the main risk was someone asking a chatbot how to make something dangerous. We are now in a world where the risk is a fully autonomous agent doing the entire attack itself, and the tooling for defense hasn't fully caught up to that shift yet.
Where this actually goes
I don't think we're heading toward some clean, stable equilibrium where offense and defense settle into a nice rhythm. I think we're heading toward a permanent asymmetry where offense gets to try a thousand things and only needs one to work, and defense has to be right every single time or explain later why it wasn't. AI doesn't remove that asymmetry. It just makes both sides faster.
What changes is who can afford to play. Big banks and cloud providers will have defensive AI stacked on defensive AI, real time anomaly detection, agents watching agents. Small businesses, local governments, hospitals, the kind of targets that make up most of the actual damage reports, will not have that. They'll have whatever off-the-shelf defensive tooling they can license, running against attackers who don't need a budget at all. The gap between "can afford serious defense" and "can't" is going to matter more than it ever has, because the attackers on the other side of that gap don't discriminate based on your IT budget.
If there's a lesson in the last year of actual incidents, it's that autonomy cuts both ways and the side that adopts it first gets the advantage, at least until the other side catches up. Right now offense has adopted it faster, mostly because there's no ethics review board slowing down a criminal deciding to run two agents in parallel. Defense has to build responsibly, get sign-off, avoid false positives that shut down real business, and not accidentally lock its own investigators out of the evidence they need. That's a real disadvantage, and pretending otherwise doesn't help anyone.
What I actually think
I'm not writing this to scare you into buying antivirus software. I'm writing it because I think most people still picture "AI risk" as a distant, abstract thing, a debate for people in suits testifying in front of Congress. It isn't distant. It happened to a government this year. It happened to one of the most well known AI companies in the world this year. The tools that broke in and the tools that caught it were built by the same industry, sometimes by the same lab.
Society is going to depend on defensive AI the same way it depends on antibiotics or vaccines: not because it's exciting, but because the threat it's responding to doesn't take breaks and doesn't get more polite over time. The sooner that becomes common knowledge instead of niche knowledge, the better prepared everyone downstream of it will be.
We're not at the start of this. We're already a few rounds in.