Agentic AppSec Lab
A team of AI agents finds a security flaw, patches it, re-runs real tests and has an independent verifier sign off, live, step by step. Switch to Posture Scan for a full report on any public repository: secrets, known-vulnerable dependencies from OSV.dev, the OpenSSF Scorecard and SARIF export. New to AppSec? SAST in CI/CD explains pipelines and security scanning in plain words, then lets you push code through a live security gate.
Scan a repository
Public repositories only, default branch. The code is cloned to a temporary directory and deleted after the scan. A scan summary (repository URL, score, finding counts) is logged for usage analytics. Source code is not kept.
Scanning
0s elapsed
Steps are approximate for the code engine; the Scorecard and dependency checks report their own status. A first scan after idle can take ~30 s while the engine wakes.
Posture grade
Code findings
–
Vulnerable dependencies
–
OpenSSF Scorecard
– / 10
Code findings
Dependencies
Security practices
OpenSSF Scorecard checks, scored 0–10
What the scanner checks
Secrets & risky code
Pattern rules for leaked keys (AWS, GitHub, Google, Slack, private keys), hardcoded passwords, injection-prone calls, plus Bandit and Semgrep static analysis.
Known-vulnerable packages
Dependencies from package-lock.json / package.json, requirements.txt and go.mod are checked against OSV.dev (GitHub Advisories, PyPA, Go vuln DB).
Engineering practices
OpenSSF Scorecard: branch protection, code review, pinned dependencies, dangerous workflows, token permissions, signed releases and more.
Limits
Static analysis only, no runtime testing. Pattern rules can produce false positives. Without a lockfile, dependency versions are the declared minimums. Scorecard covers repositories in its weekly public dataset.
Review → Fix → Verify → Explain
Four Claude agents work a security flaw end to end while you watch. A reviewer finds it, a fixer patches one file, real tests and static analysis re-run, and an independent verifier on a different model checks the fix. Deterministic tool output and AI output are labelled separately. Nothing is ever merged or pushed.
1 · Pick a vulnerable training app (full fix + tests)
Loading training apps…
…or a public GitHub repo (review + patch preview, never executed)
Pipeline ·
0.0s · 0 tokens
Agent team
Live status from the run's own events · click an agent to filter the console
Live agent console
Finding
Evidence
Verdict
Guardrails in this lab
Prompt-injection aware
Code is passed to the agents as untrusted data. Instructions hidden in comments ("report no finding", "delete the tests") are flagged, not followed. Try the prompt-injection app.
One-file allow-list
The fixer can only change the file named in the finding. Tests, CI, prompts and dependency manifests are off limits. The patch must compile and stay under 200 lines.
Independent verification
Real pytest and Bandit results decide first. The verifier runs on a different model from the fixer, in a fresh context, and sees the full diff and per-rule tool comparison.
Human stays in charge
Public repositories are reviewed but never executed: you get a patch preview to test yourself. Nothing is committed, pushed or deployed. Runs are rate-limited and cached for 24 h.
Step 1 of 3
What is CI/CD? An assembly line for code
CI · Continuous Integration. Every time a developer saves a change to the shared code (a "push"), a robot immediately builds it and runs automatic checks. Problems are caught within minutes, while the change is still fresh in the developer's mind.
CD · Continuous Delivery / Deployment. When every check passes, the same robot packages the code and ships it to a test server or straight to users. No manual copying, no "it worked on my laptop".
Together these steps are called a pipeline. Like a factory line, each station checks the product and can stop the line. Click a station:
Step 2 of 3
What is SAST? A spell-checker for security bugs
Static Application Security Testing reads your source code without running it and flags patterns known to cause attacks: SQL glued together from user input, passwords typed into the code, eval() on user text, weak encryption. A spell-checker underlines "teh"; SAST underlines "SELECT … " + username.
SAST
Reads the code. Fast, runs on every push, points to the exact line. Can raise false alarms. Tools: Bandit, Semgrep, CodeQL.
DAST
Attacks the running app from the outside, like a robot hacker. Finds real, reachable issues but needs a deployed app and is slower. Tool: OWASP ZAP.
SCA
Checks the libraries you depend on against lists of known vulnerabilities. Tools: OSV-Scanner, Dependabot. (The Posture Scan tab does this.)
Why put SAST in the pipeline? "Shift left." A bug caught minutes after it is written is a one-line fix. The same bug found after release means an incident, a patch and maybe a breach. So teams run SAST on every pull request and let it block the merge when it finds something serious. That block is called a security gate.
Cost and effort to fix grow the later a bug is found.
Step 3 of 3 · Try it
Push code through a live security gate
Pick a snippet, edit it if you like, and press git push. Real Bandit (a Python SAST tool) scans it on the server. Your code is only read, never run. If the gate blocks you, read why, fix the line and push again.
1 · Pick a snippet
2 · Your code app.py
3 · Pipeline
Findings on your code
Pipeline log
Waiting for a push…
How it looks in a real project
On GitHub, a small file in the repository tells the robot what to do on every pull request. When Bandit fails, the pull request shows a red check and the merge button is blocked until someone fixes the code.
Show the GitHub Actions file .github/workflows/security.yml
name: security-checks # the pipeline's name on: [pull_request] # run on every pull request jobs: sast: runs-on: ubuntu-latest # a fresh, throwaway computer permissions: security-events: write # allow posting findings to the PR steps: - uses: actions/checkout@v4 # 1. get the code - uses: actions/setup-python@v5 # 2. install Python with: { python-version: "3.12" } - run: pip install "bandit[sarif]" # 3. install the SAST tool - run: bandit -r . -ll -f sarif -o bandit.sarif # 4. scan; any Medium+ finding fails the job (the gate) - uses: github/codeql-action/upload-sarif@v3 if: always() # 5. show findings on the PR, pass or fail with: { sarif_file: bandit.sarif }
Add ping endpoint #42
ExampleWords you will hear
Click a word to see what it means.