AI-Driven Penetration Testing: I Watched It Break In (2026)
I asked an AI I built to find a way into a WordPress security plugin, the kind installed on millions of websites to keep hackers out. A couple of minutes later it handed me an admin login. No password. One request. That is AI-driven penetration testing, and I recorded the whole thing so you can watch it happen.
I've spent thirty years finding bugs like this by hand. What I'm going to show you is not a demo of a clever AI. It's proof that the job of a hacker is being rebuilt in front of us, and that the valuable skill is no longer the hacking. It's directing the machine that does it, and knowing when it's wrong.
Here's the demo, why it worked, and how you become the person who can do this.
TL;DR: if you've only got 30 seconds
An AI system I built found a real admin-login bypass in a security plugin used on four million sites, and logged in with no password. (Old, patched flaw; disclosed; safe to show.)
It worked because of the system, not the AI: several models from different companies checking each other, so none marks its own homework.
The grind moved to the machine. The judgment (what to attack, what's real) is still you. That's the new, valuable skill.
It's learnable, especially if you already work in IT. The window is open now, while almost nobody can do it.
What Is AI-Driven Penetration Testing?
Start with an example, because the phrase sounds bigger than it is.
Say you want to check whether a building is secure. The old way: you hire one locksmith who walks the whole building, tries every door and window himself, and writes you a report. Slow, expensive, and only as good as one tired person's attention on the day.
AI-driven penetration testing is the other way. You are the person in charge. You point a system of AI models at the target and say "find me a way in." The AIs do the door-rattling (reading the code, tracing how a login works, trying things) at a speed no human can match. You decide what to point them at, you read what they find, and you make the call on what's real. You direct; the machine does the grinding.
That's the difference from an "AI pen testing tool" you buy off a shelf. A tool runs the same fixed checklist every time. AI-driven pen testing is you designing and steering the whole system: choosing the models, deciding what to point them at, and judging what comes back. That's exactly why it's a skill, not a purchase. (More on that, because it's the whole point.)
💡 In plain English. "Penetration testing" (or "pen testing") is paying someone to try to break into your systems on purpose, so you find the holes before a criminal does. AI-driven just means you direct an AI system to do the breaking-in attempts instead of doing them all by hand.
The security industry is being rebuilt around this, and it's hitting security first for a simple reason: in security, "fast but broken" isn't a bug. It's a breach.
Watch It Break Into a Security Plugin
Here's the recorded run (it's in the video above). I've sped it up; normally it takes a couple of minutes.
I asked HAL, the AI system I built, to review a WordPress security plugin installed on around four million sites when I recorded this (over three million today). To be clear about the ethics before anything else: this was done on an old, already-patched version of the plugin. The flaw is a real, publicly documented one, CVE-2024-10924, rated 9.8 "critical", and it was fixed back in version 9.1.2 in November 2024. I'm blurring the plugin's name on screen. I'm not dropping a live zero-day on YouTube. I'm showing you a real find, safely.
Here's what happened, step by step:
Five AI models reviewed the code at once, from three different companies (Claude, OpenAI, Google). Why five, why three companies? So that no AI marks its own homework. (That one idea is the engine of the whole thing; see the next section.)
Every one of them independently found the same bug. All rated critical. Then a separate "referee" AI checked the finding so we weren't chasing a false alarm.
HAL explained it in plain English: it's a password-and-two-factor bypass. You can log in as the administrator with no password at all.
So I made it prove it. First I tried the normal way: username admin, password password123. Rejected. This is a properly locked-down site.
Then HAL fired the exploit it wrote. It doesn't guess a password. It just tells the site, in effect, "I'm user number one, the administrator." And the site believes it. We're in. Full control.
Think about what that means. The plugin whose entire job is to keep hackers out became the way in. The AI found the flaw and broke through it, on its own, using the system I built and directed. To be precise about the scale: four million is the number of sites that run this plugin, not the number that were exploitable. This particular door only opened on the vulnerable version with the two-factor setting switched on, and it's patched now. But that's the point of finding it first: on a plugin this widespread, a single missed logic flaw is a skeleton key waiting to be cut.
💡 How the flaw actually works, in one sentence. The security check runs correctly and decides you've failed it. Then every part of the code that's supposed to act on that answer ignores it and logs you in anyway. It's the security guard who checks your ID, sees it's fake, writes "REJECTED" on a sticky note… and then waves you through while the note sits on the desk.
Why the AI Isn't the Clever Part: the System Is
Here's where most people get it wrong, and it's the most important thing in this article.
The AI is not the genius. On its own, a single AI is unreliable. It misses things. It makes things up. And when you ask it to check its own work, it just agrees with itself. It's the same way you can't proofread your own writing, because your brain reads what it meant to say.
So what broke into that site wasn't one smart AI. It was a system built around several of them. Two ideas do the heavy lifting:
No AI marks its own homework. We don't ask one model to find bugs and then trust it. We run several, from different companies, because a model from OpenAI has different blind spots from a model from Google or Anthropic. When one misses something, another often catches it. That's real AI checking AI, not one AI nodding along to itself.
A referee makes the final call. After the models argue it out, a separate judge decides what's actually real and throws out the false alarms. It never gets told which model raised which finding, so it weighs the claim, not the source.
Here's a sense of why reasoning matters more than pattern-matching here. A standard code scanner (the normal automated tool a security team runs, which looks for known bad patterns like a spellchecker looks for known misspellings) found nothing on this plugin. Zero. Because this flaw wasn't a pattern you can match. It was a logic mistake, the kind you only catch by following the code's reasoning step by step. That's what the models did that the scanner couldn't. (If you want the full recipe for running a review like this, I wrote it up in our secure code review guide.)
Let me be straight about what this one run does and doesn't prove, because that honesty is the whole point. It reproduced a real, disclosed, already-patched flaw, and every model found it. That shows the workflow catches this kind of logic flaw. It isn't a benchmark that says five models always beat one, and it isn't the same as discovering a brand-new bug from scratch. The value of running several models from different companies, and a referee on top, shows up on the harder jobs: when one model misses something, or invents a bug that isn't there, and you need a way to catch that without trusting any single one of them. On a clean, known flaw like this, they happen to agree. On live, unknown code, they don't, and that's exactly where the system earns its place.
So the skill isn't the AI. It's knowing where to point it, how to build the checking around it, and how to spot the moment it's confidently, completely wrong. That takes real thinking, and it's learnable.
AI vs Human Pen Testers: What Actually Changed
Let me be precise about what's moved and what hasn't, because the fear ("AI is coming for my job") has it backwards.
The grind has moved to the machine. Running the tools by hand, remembering the exact syntax, trawling through code line by line: that's the part the AI now does, faster than any of us.
The judgment has not moved. Deciding what to attack, architecting the system that attacks it, reading a finding and knowing whether it's real or the AI hallucinating: that's still you. If anything it matters more now, because the machine produces more to judge.
This isn't a prediction. It's already on the scoreboard. In 2025, an AI system called XBOW topped HackerOne's US bug-bounty leaderboard, the first AI ever to do it, above every human hacker on the board. It got there by submitting more than a thousand vulnerability reports, and it briefly ranked number one in the world. That's the same kind of system I built for the demo above: not one AI, but a directed system of them. (This is also why the honest answer to "is AI killing the penetration tester career?" is "no. It's changing what the job is.")
So the job most at risk is the one that was only the manual grind. I can't promise you a title or a salary, nobody honestly can, but the direction of travel is hard to argue with: the person who learns to architect and direct their own AI systems, and to judge what those systems produce, is far better placed than the person who refuses to. That's a skill worth having whichever way the market moves. (It also changes the economics of the work itself, which I dug into separately in the cost of penetration testing: AI vs humans.)
This Has a Name: Agentic Engineering
What you watched isn't a trick I invented. It's a discipline with a name.
Andrej Karpathy, a founding member of OpenAI, has been the clearest voice on this: software is becoming something you direct rather than something you hand-write. Agentic engineering is the name for the discipline that grows out of that: directing systems of AI agents to do real work. I call what it makes you AI-driven: an AI-driven pen tester, an AI-driven security engineer.
Here's the plain version. You're not learning to write code faster. You're not "vibe coding." You're learning to be the director: the person who sets the goal, assembles the AI cast, and judges the results, while the AIs do the labour. You design; they build.
And almost nobody can do it yet. That's not me hyping it. Outside a handful of elite bug-bounty hunters and the people building these systems, the field simply hasn't shifted. Which is exactly why there's an opening.
For the deeper explanation of the discipline itself, where it came from and the levels of it, I've written a full breakdown in our agentic engineering article.
How to Become an AI-Driven Pen Tester
Good news first: you don't need my thirty years.
If you already work in IT, cloud, or development (or you're just a bit techie) you have the foundations. You understand how systems fit together, which is the hard part to teach. What you're adding is a new skill on top: how to architect and direct the AI system that does the grunt work, and how to catch it when it's wrong. That's a shorter bridge than it looks.
Here's the honest catch, so you don't walk in expecting magic:
It's not push-button. Buying a tool doesn't make you an AI-driven pen tester any more than buying a stethoscope makes you a doctor. (There's a growing shelf of them. I compared the main ones in AI penetration testing: hype or real? and the best AI for hacking. But the tool is the easy part to buy and the hard part to direct.)
You still need the fundamentals. You have to understand what a login bypass is to know the AI found a real one. You won't execute the old way by hand much anymore, but you have to understand it to direct it.
The skill is judgment. The valuable, hard-to-replace part is reading what the AI finds and knowing what's true. That's the part worth building.
Where you take it is your call. A top job at a company that can't function without this skill. Independent consulting or penetration testing. Bug bounty, which is where a lot of the money is. Or building something of your own. Same underlying skill, and it's the core of what we teach in the StationX Certified AI-Driven Security Engineer path.
The window is open right now, while almost nobody can do this. A year ago, breaking into a plugin on millions of sites took an expert with decades of experience. You just watched an AI do it, directed by someone who knew what he was looking at. That someone can be you.
Start here (free)
I put the whole method (what's changing, why, and exactly how to step into it) into a short, free book: Become the Cyber Security Expert the AI Era Demands. It's the fastest way to get your head around this shift and start. Read it free here.
Watching me do it changes nothing for you. Deciding you're going to be the one who does it changes everything. And do it ethically. The difference between the person in this demo and the person you should fear is what they do when they find the door open.
Frequently Asked Questions
Can AI really hack a website?
Yes, and this article shows a real example. An AI-driven system found a genuine login bypass in a widely used security plugin and used it to log in as administrator with no password. The important nuance: it wasn't one AI being clever. It was a system of several models, cross-checked and refereed, directed by a human. On its own, a single AI is unreliable enough that you couldn't trust the result.
Will AI replace penetration testers?
It replaces the manual grind: running tools by hand, remembering syntax. It does not replace the judgment: deciding what to attack, building the system that attacks it, and knowing whether a finding is real. Testers who learn to direct AI systems become more valuable, not less. Testers who only know the manual way are the ones at risk.
Is AI penetration testing legal?
Testing is legal when you stay inside written authorization: systems you own, or a signed engagement, or a bug-bounty program's published scope and rules. Permission isn't a blank cheque, though. It covers specific targets and specific techniques, so the rule is simple: test only what the written scope allows, the way it allows. Using the same techniques outside that scope, or against systems you have no permission for, is a crime. The skill is identical; the difference is authorization. The demo in this article was run on an old, already-patched, publicly disclosed flaw, on a plugin whose name is blurred, not a live target.
Do I need to know how to code to be an AI-driven pen tester?
You need to understand systems and security fundamentals, but you do not need to be a strong programmer, and this is explicitly not “vibe coding.” The skill is architecting and directing AI agents and judging their output, closer to being a director than a coder.
What's the difference between AI pen testing and agentic engineering?
Agentic engineering is the general discipline: directing systems of AI agents to do real work. AI-driven penetration testing is that discipline applied to offensive security: pointing a directed AI system at a target to find vulnerabilities. Same craft, security-specific application.
About the Author
Nathan House, Founder & CEO of StationX
Nathan House has 30 years of hands-on cybersecurity experience and is Cambridge-educated, holding CISSP, CISA, CISM, OSCP, CEH, and SABSA. He founded StationX in 1999 — one of the UK’s first cybersecurity companies — and has secured £71 billion in UK mobile banking transactions and the London 2012 Olympics, advising clients including Microsoft, Cisco, BP, Vodafone, and VISA. He authored the world’s most popular cybersecurity course — a #1 Udemy bestseller taken by over 500,000 students — and was named Cyber Security Educator of the Year 2020, AI Security Educator of the Year, and a UK Top 25 Security Influencer 2025. A DEF CON speaker and featured expert on CNN, Fox News, NBC, and the BBC, Nathan leads StationX’s training of more than half a million students worldwide.