AI Hackers Are Faster Than Your Pen Test

Eli Cohen, Snyk19:29 · Oct 2026 · 3,774 views
Thumbnail for AI Hackers Are Faster Than Your Pen Test Watch on YouTube
TL;DR
  1. 1

    AI coding tools are increasing the amount of code shipped while also increasing the amount of insecure code entering production.

  2. 2

    Static scanning, dynamic testing, and human penetration testing each miss different classes of vulnerabilities, especially business-logic flaws.

  3. 3

    AI penetration testing should run on every pull request, use several specialized agents, consume context from other scanners, and prove that reported vulnerabilities are exploitable.

Summary

Eli Cohen argues that the old security testing cycle cannot match the speed of AI-assisted development or AI-assisted attacks. He says developers are producing more code, much of it insecure, while attackers use frontier models to chain smaller vulnerabilities into serious attacks. Static analysis can find issues such as SQL injection and cross-site scripting, but it misses runtime authorization problems. Dynamic testing can probe runtime behavior, but it does not understand application business logic well enough. Human penetration testing understands that logic, yet its cost and slow reporting mean that companies may run it only once or a few times a year. Cohen proposes continuous offensive security, with AI penetration tests on every code change and agent red teaming for multi-step attacks. His model uses an orchestrator, recon agents, vulnerability hunters, an exploit-validation judge, remediation, and reporting. He also says vendors should be judged on continuity, context, data integration, and proof of exploitability.

Key ideas
01:34

AI coding tools increase both output and the security backlog

Cohen says developers are producing 218% more lines of code per developer than before, and that more than 85% of developers use coding agents. He pairs this with his claim that 62% of LLM-generated output is insecure or broken. The result is more code reaching production while security teams create issues faster than they can remediate them. He mentions developers running several sessions of Claude Code or Codex at once. His point is that security processes built for a slower development cycle now face a larger and less secure stream of changes.

03:07

AI attackers can chain small flaws into serious attacks

Attackers are also using frontier models, according to Cohen. He says they can chain low-severity vulnerabilities into critical vulnerabilities, particularly in areas that require application context. He cites an average successful AI attack time of 24 minutes and 34 seconds, with the fastest example taking four minutes. He also says 43% of MCP servers have vulnerabilities. The practical concern is that attackers can continuously test and combine weaknesses while a company's own security backlog grows.

05:31

Static and dynamic testing cover different gaps

Static application security testing runs against code changes and can look for issues such as SQL injection and cross-site scripting. Cohen says it does not catch problems that appear only at runtime. Dynamic application security testing can send payloads to API endpoints and find runtime issues such as broken object-level authorization, where user A can access user B's data. It still struggles with the business logic produced by modern coding systems. Each method finds useful classes of bugs, but neither replaces testing that understands how the application is meant to work.

07:11

Human penetration testing understands business logic but runs too slowly

Cohen calls human penetration testing the gold standard for finding business-logic vulnerabilities because a security researcher can understand the application and try to break it. The problem is cost and speed. Companies may run a test once, twice, or perhaps three times a year, leaving the application untested for much of the year while attackers keep working. The output is often a PDF report that developers must interpret, fix, and then retest. Cohen says this process no longer fits a system that changes continuously, although he still describes penetration testing as the strongest existing method for business logic.

09:16

Continuous offensive security combines AI penetration testing with agent red teaming

Cohen proposes continuous offensive security for systems that change throughout the day. The approach includes AI penetration testing on every code change and agent red teaming for multi-step attacks. He names prompt injection, data exfiltration, and goal hijacking as examples of attacks that need more than a single request. Dynamic testing remains useful for authorization problems, so the proposed approach combines it with AI penetration testing rather than discarding it. The goal is to test production systems for unknown problems at a frequency closer to the rate of development and attack.

10:46

A team of agents divides the penetration test into distinct jobs

The system Cohen describes starts with an LLM orchestrator that follows a dynamic plan and operates the other agents. A recon agent gathers network endpoints, APIs, and fingerprints, then builds an application architecture. Vulnerability-testing agents target known vulnerability classes and hunt for flaws by trying to understand the application's business context. An exploit-validation judge checks whether each finding is real and exploitable. A remediation agent turns findings into actions for developers, while a reporting agent communicates the results. Cohen presents this division as a way to reduce false positives and make findings usable.

14:51

Context from other scanners guides the AI test

Cohen says an AI penetration-testing system should not operate alone. It should receive data from static scanners, dynamic scanners, and other agent engines. That information helps guide the LLM toward relevant parts of the application. He says richer context improves results and can make the LLM operation cheaper because the model knows where to look. He also rejects using one scanner for every problem. For example, he says dynamic testing is well suited to cross-site scripting because it can send exhaustive payloads, while penetration testing is better for business logic. The systems should be combined according to the vulnerability class.

17:30

Vendors should prove continuity, context, integration, and exploitability

Cohen ends with four questions for anyone evaluating an AI penetration-testing vendor. Does the system run continuously or only as an occasional point-in-time test? Does it reason about application context? What additional context can it consume, and how does it enrich its data? Does it prove that a reported bug is exploitable, or does it only flag the bug? These questions reflect his concern that a new tool can reproduce the limits of older scanners if it lacks context, runs too infrequently, or sends developers findings without evidence.

"Our goal is not only to help you find the bugs faster, but to actually do it better and cheaper."18:26
Who should watch
  • You ship AI-generated code frequently and need security checks that run closer to each pull request.
  • Your security program relies on SAST, DAST, or annual penetration tests, and you need to understand what each approach misses.
  • You are evaluating AI penetration-testing products and want practical questions about context, integrations, continuity, and exploit validation.