Sprocket Security | How We Test

How Sprocket runs continuous penetration testing

Technology-Powered.
Human-Verified.

Our AI agent fleet runs the parts of a penetration test that reward speed and coverage. Our human penetration testers run the parts that reward judgment. Here’s exactly where the line falls — and the safety framework that keeps the agents on their side of it.

Machine scale. Human judgment.

How the loop works

Your attack surface, re-enumerated continuously. Work routed to the agent with the right capability. Every finding that requires it is driven to closure by a human penetration tester before it reaches you. Then the loop comes around again.

The problem we solve

Most testing models make you pick a side.

Traditional consultancies

Experienced penetration testers and a report — once a year. That leaves 345 days of untested exposure between engagements, while your environment ships changes weekly.

Fully automated scanners

Continuity and speed, but they find what they have signatures for. They don't chain three medium findings into a domain compromise, and they don't understand that your low-severity IDOR exposes every customer's billing record.

Not us. Sprocket's continuous penetration testing doesn't split the difference. We assign each part of the engagement to whichever is genuinely better at it.

Phase by phase

Who does what, and why

Six phases, one engagement. Select a phase to see the AI agent fleet’s role, the human penetration tester’s role, and why the work is split that way.

Phase 01 · Attack surface discovery · AI agents lead

AI agent fleet

Apex — Discovery and recon. Continuously enumerates domains, subdomains, IP ranges, cloud assets, exposed services, and third-party footprint. It re-runs on a schedule, not on a calendar invite, so new infrastructure shows up in days rather than at the next annual engagement.

Penetration testers

Your lead penetration tester validates scope, flags assets that look like yours but aren't, and decides what's in bounds.

Why this split

Enumeration is a breadth problem — machines don't get bored on page 400 of a certificate transparency log. But breadth is exactly where automated discovery goes wrong: an AI agent that follows every link outward eventually enumerates infrastructure that isn't yours. The boundary is a human decision, enforced before the AI agent runs, not corrected after.

Phase 02 · Reconnaissance and fingerprinting · AI agents lead

AI agent fleet

Apex — Fingerprinting. Identifies technologies, versions, exposed endpoints, authentication surfaces, and known-vulnerable components across everything found in Phase 1.

Penetration testers

Penetration testers read the map. Which of these looks like it was built in a hurry? Where does the tech stack change abruptly — usually a sign of an acquisition or a legacy system nobody owns?

Why this split

Identifying a version string is pattern matching at volume. Knowing that a version string is interesting is interpretation. AI agents produce the inventory; a penetration tester reads it for the things that don’t fit.

Phase 03 · Authenticated testing · AI agents and humans together

AI agent fleet

Link — Authenticated testing. This is where most automated tooling stops. Link, our authenticated web application testing AI agent, holds session state and tests behind the login — the part of the app where your actual risk lives.

Penetration testers

Penetration testers define the roles and privilege boundaries worth testing, then work the business-logic layer the AI agent surfaces: can a standard user reach an admin function, can tenant A read tenant B's data, does the multi-step workflow enforce state at every step or only at the end.

Why this split

An AI agent can test whether an endpoint responds. A penetration tester knows whether that response should have been possible.

Phase 04 · Exploitation and chaining · Humans lead

AI agent fleet

Full AI fleet, in support. AI agents keep feeding the penetration tester fresh signal from the rest of the environment while they're deep in one attack path, so nothing goes stale during a long engagement.

Penetration testers

Penetration testers confirm exploitability by hand, then chain findings the way an adversary would. A weak password policy plus an exposed admin panel plus a stale service account is not three medium findings — it’s one critical path.

Why this split

The hard question in exploitation isn't can this be exploited — it's how far do we go to prove it. Enough to demonstrate impact, not so far that we damage a production system or touch data we shouldn't. That's a proportionality call with your business on the other side of it, and it belongs to a person who is accountable for the answer.

Phase 05 · Validation and reporting · Humans lead

AI agent fleet

Clutch and Smith — Validation and writing. AI agents assemble evidence, reproduce steps, and pre-populate the report so penetration testers spend their time on analysis rather than screenshots.

Penetration testers

Humans supervise every finding. No finding reaches your report without a penetration tester confirming it, rating it against your environment, and writing remediation guidance a developer can act on without a follow-up call.

Why this split

An automated system can produce a finding that is confident, well-formatted, severity-rated, and wrong. Fluency is not evidence. The only durable check is a person who reproduced the finding and can show you the steps — which is why validation is the one phase we don’t split at all.

Phase 06 · Retest and continuous coverage · AI agents and humans together

AI agent fleet

Apex and Torque — Coverage and new test builds. The discovery and fingerprinting AI agents keep running, so a new asset that appears in month seven gets tested in month seven.

Penetration testers

You fix something, you request a retest, a penetration tester confirms it. Unlimited retests are included — no change order, no waiting for next year’s engagement.

Why this split

Continuity is a machine strength — nothing else will re-check your perimeter every week for a year. Confirming a fix is a human one. “Fixed” has to mean the attack path is closed, not that a signature stopped matching. Those are different claims, and only one of them is worth putting in front of your auditor.

The short version

The split, on one page

Activity AI agents Penetration testers
Attack surface discovery Lead Validate scope
Fingerprinting Lead Interpret
Authenticated app testing Lead Business logic
Exploitation & chaining Support Lead
Severity & business impact — Lead
Finding validation Evidence Every finding
Remediation guidance Draft Lead
Retesting Continuous monitoring Verify fixes
What the agents are not allowed to do

Four boundaries, enforced in code

Each lane above has a boundary the agent cannot cross, enforced in code rather than asked for in a prompt. These are four of the seven properties in our published safety framework, applied per phase.

01

Bounded

Scope checked before every request.

02

Proportionate

Exploitation stopped at proof.

03

Contained

Discovered data confined to your engagement's isolated volume.

04

Transparent

An audit trail the agent cannot reach.

Four outcomes Sprocket delivers

Findings in days, not at the end of a quarter

AI agents surface candidates continuously; penetration testers validate on a rolling basis.

Fewer things to triage

Human verification happens before the finding reaches you, not after your team has spent a sprint chasing it.

Coverage that tracks your environment

New subdomain, new app, new acquisition — the discovery layer sees it without you filing a scope change.

The same penetration testers over time

Continuity means they learn your environment and this context compounds with every engagement. Year two is not a re-introduction.

FAQ

Straight answers

Is this an AI pentest?

No. It's a penetration test performed by human penetration testers, with an AI agent fleet doing the discovery, enumeration, and coverage work that machines do better. Every finding in your report is supervised by a human.

Can I see which findings came from AI agents versus humans?

Yes — consider us a glass box. Findings carry their provenance, and every one is supervised by a human.

How is this different from a scanner with a services wrapper?

A scanner reports what it matched. Our human penetration testers exploit, chain, and rate findings against your business. The AI agents exist to give those penetration testers more surface to work from, not to replace their judgment.

What happens when my environment changes mid-engagement?

Discovery keeps running. New assets enter scope as they appear rather than waiting for the next cycle.

How fast do I get results?

Findings are delivered as they're validated, not held for a final report drop.

Start with what an attacker already sees

Free external attack surface assessment. No commitment, and you keep the results either way.