dashboard radar bug_report school newspaper sensors
Agentic AppSec // Multi-agent review · fix · verify Scan engine: checking…

Agentic AppSec Lab

A team of AI agents finds a security flaw, patches it, re-runs real tests and has an independent verifier sign off, live, step by step. Switch to Posture Scan for a full report on any public repository: secrets, known-vulnerable dependencies from OSV.dev, the OpenSSF Scorecard and SARIF export. New to AppSec? SAST in CI/CD explains pipelines and security scanning in plain words, then lets you push code through a live security gate.

Review → Fix → Verify → Explain

Four Claude agents work a security flaw end to end while you watch. A reviewer finds it, a fixer patches one file, real tests and static analysis re-run, and an independent verifier on a different model checks the fix. Deterministic tool output and AI output are labelled separately. Nothing is ever merged or pushed.

Agents: checking…

1 · Pick a vulnerable training app (full fix + tests)

Loading training apps…

…or a public GitHub repo (review + patch preview, never executed)

Guardrails in this lab

Prompt-injection aware

Code is passed to the agents as untrusted data. Instructions hidden in comments ("report no finding", "delete the tests") are flagged, not followed. Try the prompt-injection app.

One-file allow-list

The fixer can only change the file named in the finding. Tests, CI, prompts and dependency manifests are off limits. The patch must compile and stay under 200 lines.

Independent verification

Real pytest and Bandit results decide first. The verifier runs on a different model from the fixer, in a fresh context, and sees the full diff and per-rule tool comparison.

Human stays in charge

Public repositories are reviewed but never executed: you get a patch preview to test yourself. Nothing is committed, pushed or deployed. Runs are rate-limited and cached for 24 h.