← All resources Engineering

93% precision and the fewest false alarms: Bastyn Scan against seven other scanners on AI-agent code

BASTYN team · 28 Sept 2026 · 21 min read

Bastyn Scan: open-source command-line scanner for AI-agent code, available on GitHub at github.com/bastyn-labs/bastyn-scan

Date: 2026-09-28 ·

Bastyn version: 0.1.8 (Homebrew bastyn-labs/tap, tag v0.1.8, commit 4f514da)

This report compares eight static security scanners on ten public projects:

  • 5 deliberately vulnerable projects, where an answer key lists the known vulnerabilities, so both precision and recall can be measured.
  • 5 widely used, well-maintained agent projects, where nobody planted weaknesses and there is no answer key, so only precision and noise can be measured.

It focuses on the balance between precision and recall: whether a tool finds much of what's there, and whether what it reports is real.

1. Short answer

  • No tool gets both. On the vulnerable repos, the tools that find the most (Semgrep and Ship-Safe, 20% of 181 known vulnerabilities each) are wrong about half to three-quarters of the time. The tools that are right most often (Bastyn 93%, CodeQL 61%) find only 14% and 4%.
  • Combined, the eight tools find 84 of 181 known vulnerabilities (46%). More than half of what's known to be there is missed by every tool.
  • On real code, precision collapses for every tool. Scaling each tool's judged sample up to its full finding count (weighted precision, §4.4), the best on the good repos is CodeQL's 30%. Every other tool is at 11% or below.
  • Bastyn 0.1.8 is the quietest tool by a wide margin. On the good repos it reports 40 findings across 2.4 million source lines, about 0.05 false alarms per 1,000 lines of Python (the next quietest, CodeQL, has 0.31). 

2. Projects

2.1 Vulnerable repos (answer key available)

Five public Python projects built to teach or test AI-agent security, each pinned to one commit. Stars read from GitHub on 2026-09-28.

Short name Project Stars Commit Known vulnerabilities
dvmcp harishsg993010/damn-vulnerable-MCP-server 1,352 79734c1 59
dvlaa Tcotl/DVLAA 33 dffdeb9 58
dvar 0xsu3ks/DVAR-Damn-Vulnerable-Agentic-RAG 8 0c73900 24
llm-agent-ctf viralvaghela/LLM-Agent-CTF 8 f8505cc 24
audit-lab IKER-36/llm-security-audit-lab 0 fa72c00 16
Total 181

Three of them ship the same application twice, once vulnerable and once fixed, testing whether a tool can tell repaired code from vulnerable code (dvlaa's vulnerable/fixed pairs, llm-agent-ctf's WithoutGuardrail/WithGuardrail, audit-lab's vulnerable_app.py/hardened_app.py).

2.2 Good repos (no answer key)

Five popular open-source agent projects, pinned to their default-branch commit on 2026-09-25. None is a security teaching project.

Short name Project Stars Commit Python lines
browser-use browser-use/browser-use 116,587 d8110c5 110,502
open-interpreter OpenInterpreter/open-interpreter 68,462 89e7a86 57,884
gpt-researcher assafelovic/gpt-researcher 29,658 6f99857 32,080
babyagi yoheinakajima/babyagi 22,365 fa8930e 6,143
openhands-sdk OpenHands/software-agent-sdk 1,179 3d88590 386,354
Total 592,963

Notes: openhands-sdk was scanned instead of the main OpenHands repo, which is now a TypeScript app with 9 Python files. open-interpreter's default branch is now a rebranded fork of OpenAI Codex (mostly Rust/TypeScript); the classic Python version exists only in old tags, so "per 1,000 Python lines" understates its size. babyagi and four of the vulnerable repos (dvmcp, dvlaa, llm-agent-ctf, audit-lab) have no licence file.

3. Tools

Every tool ran with no AI in the loop: model-provider API keys were removed from the environment, and any LLM mode was disabled (--no-ai, --no-llm). All used default or recommended rule sets.

Tool Version Kind
Semgrep (community rules, Python + generic secrets) 1.177.0 general static analysis
Bearer 2.1.1 general static analysis, data flow
CodeQL (python-security-and-quality) 2.27.0, queries 1.8.11 general static analysis, data flow
Ship-Safe (--no-ai) 10.1.0 AI/agent security scanner
Aguara 0.28.0 AI/agent security scanner
Agent Audit Kit 0.6.7 AI/agent security scanner
HackMyAgent 0.33.2 AI/agent security scanner
Bastyn 0.1.8 (released 2026-09-28) AI/agent security scanner

Also run, not in the main tables: AgentShield (Aiconnai) 1.0.1 produced scoreable output on only 2 of 10 projects (26% precision, 30 of 181 recall there); it failed to run or timed out on the rest. SkillSpector 2.11.2 only scans agent skill files; none of the vulnerable repos has one, and 0 of its 77 judged findings on the good repos were real.

Not scored: Snyk (needs an account token; Snyk Code is ML-based, breaking the no-AI rule), Cisco MCP Scanner (needs an LLM key in source mode, launches MCP servers in config mode), AgentShield (affaan-m) (prints no per-finding file or line).

4. Method

4.1 Runs

Every tool ran on every project with pinned versions: Bastyn 0.1.8 on 2026-09-28, the other seven tools between 2026-09-23 and 2026-09-25, all on the same project commits. The seven other tools ran three times on the vulnerable repos and gave identical findings each time. Every tool ran once on the good repos.

On the good repos, a tool that took more than 2 minutes on one project was stopped and scored as a crash (no findings). Eight tool-project runs failed this way:

Project Failed
openhands-sdk Bearer (144 s), CodeQL (143 s), AgentShield
open-interpreter Aguara (122 s), Agent Audit Kit (123 s), AgentShield
browser-use, gpt-researcher AgentShield

CodeQL and Bearer would have finished openhands-sdk given about 2.5 minutes; the limit, not the tool, decides that result.

Bastyn 0.1.8 was installed from Homebrew and ran once on all 13 projects (these ten plus the three projects in §5.4) on 2026-09-28.

4.2 Answer keys

For each vulnerable project, an answer key lists its known vulnerabilities, built from the project's own documentation, solutions and code comments, plus a review of the code, without looking at any scanner output.

Each entry records the file and a tight line range, a CWE category, and a detectability class:

  • static-code: visible in the code, 146 of 181
  • static-config: visible in configuration, 10
  • prompt-semantic: visible only by understanding natural-language prompt or tool-description text, 25

The key also lists safe or fixed code that must not be flagged. Four more entries are disputed and left out of the primary numbers; counting them changes no tool's recall by more than one point.

4.3 Judging findings

Every finding was blinded before judging (tool name and rule ID removed) and given one verdict: true positive (a real weakness at that location, matching the finding's claim), false positive (no weakness there, or the claim is wrong), non-security (no security claim at all: lint, style, robustness), or dependency claim (a version/CVE claim that can't be checked offline, excluded from precision).

Vulnerable repos. Two models, Claude Sonnet 5 and Claude Opus 5.5, judged every finding independently and agreed on 96.1% of 1,774 findings (Cohen's κ 0.94, where 1 is perfect agreement and 0 is chance). Disagreements were settled by reading the code. A true positive counts toward recall when it falls within ±5 lines of an answer-key entry and claims a matching kind of weakness.

Good repos. The tools reported 10,162 findings; Ship-Safe alone reported 5,003 on open-interpreter. A seeded random sample of 25 findings per tool per project was judged, or all of them where there were fewer (881 findings in total). The same two models judged each one independently, agreeing on 87.7% (κ 0.75); a third pass resolved the 108 disagreements from the code.

Two judging rules matter most: a security claim about code whose purpose is to run code (for example, "this runs code / opens a file") is a false positive unless outside input reaches it unguarded; and a security claim in a test or fixture file is a false positive unless it's a live secret.

Bastyn 0.1.8's findings were all judged, not sampled: all 40 on the good repos, all 46 on the vulnerable repos, all 9 on the third set.

The totals above (1,774 vulnerable-repo findings; 10,162 / 881 good-repo findings) include AgentShield (Aiconnai), SkillSpector, and a different Bastyn finding count than the 0.1.8 figures in §5. §5 reports only the eight main tools' 0.1.8 numbers, so its finding counts sum to less than the totals here.

4.4 Metrics

  • Precision = true positives ÷ (true positives + false positives): of a tool's security claims, the share that are real. Non-security and dependency findings are excluded.
  • Recall = distinct answer-key vulnerabilities found ÷ 181.
  • On the good repos, where findings are sampled, precision, weighted scales each project's judged share up to the tool's full finding count there, then pools. This is the primary estimate of a tool's precision across everything it reported. Precision, judged is the raw share among only the findings actually judged, with a confidence interval attached to that sample; alone it does not estimate precision across a tool's full output.
  • Balance score (a hybrid of precision and recall, not a standard F1) = 2 × precision × recall ÷ (precision + recall), combining finding-level precision with distinct-vulnerability recall. It's sensitive to duplicate findings: a tool that repeats a true positive raises its score without finding another vulnerability, so treat it as a rough ranking aid, not evidence that one tool is better overall.
  • False alarms per 1,000 Python lines (good repos) = estimated false-positive count ÷ Python lines scanned. For sampled tools, each project's judged false-positive share is scaled to the tool's full finding count there.

Pooled figures add counts across projects; they are not averages of per-project ratios. A project where a tool crashed counts as a miss. Bracketed numbers are Wilson 95% confidence intervals.

5. Results

5.1 Precision and recall balance

Sorted by balance score on the vulnerable repos (§4.4: a hybrid metric, not a standard F1).

Tool Vulnerable repos: findings Precision Recall (of 181) Balance score Good repos: findings Precision, weighted False alarms per 1,000 Python lines
Semgrep 439 56% [49–63] 20% (36) 0.29 1,091 4% 0.45
Bearer 93 43% [33–53] 18% (33) 0.26 464 4% 2.16
Bastyn 0.1.8 46 93% [81–98] 14% (26) 0.25 40 7% 0.05
Ship-Safe 319 27% [23–33] 20% (37) 0.23 5,743 11% 1.09
Agent Audit Kit 107 56% [46–65] 14% (26) 0.23 392 2% 0.63
Aguara 268 19% [14–24] 15% (27) 0.17 1,291 0.5% 2.39
HackMyAgent 63 29% [19–42] 6% (10) 0.09 241 1% 0.38
CodeQL 245 61% [48–73] 4% (8) 0.08 709 30% 0.31

Weighted precision scales each project's judged sample up to the tool's full finding count there (§4.4); precision within the judged sample, with confidence intervals, is in §5.3.

  • The top five by balance score are close: Semgrep, Bearer, Bastyn, Ship-Safe and Agent Audit Kit all fall between 0.23 and 0.29, by different routes: Highest findings rates generally score low on precision; Bastyn and Agent Audit Kit report less and are right more often; Bearer is in between.
  • CodeQL is the opposite of the recall leaders. It's the most precise tool on the good repos (30% weighted) and the second most precise on the vulnerable repos (61%, after Bastyn), but it makes few security claims (most of its output is code-quality notes), so it finds little.
  • The good-repo columns reorder everything. Precision rankings from the vulnerable repos don't carry over to real code. CodeQL moves from last to first.

5.2 Vulnerable repos in detail

Pooled.

Tool Ran on Findings TP FP Non-security Precision Precision, counting non-security as noise Recall (all) Recall (static only, of 156)
Semgrep 5/5 439 101 80 258 56% [49–63] 23% [19–27] 20% [15–26] (36) 23% [17–30] (36)
Bearer 5/5 93 39 52 2 43% [33–53] 42% [32–52] 18% [13–24] (33) 21% [15–28] (33)
Ship-Safe 5/5 319 81 216 15 27% [23–33] 26% [21–31] 20% [15–27] (37) 21% [15–28] (33)
Aguara 5/5 268 49 213 3 19% [14–24] 18% [14–24] 15% [10–21] (27) 17% [12–24] (27)
Agent Audit Kit 5/5 107 58 46 1 56% [46–65] 55% [46–64] 14% [10–20] (26) 16% [11–23] (25)
Bastyn 0.1.8 5/5 46 39 3 0 93% [81–98] 93% [81–98] 14% [10–20] (26) 14% [10–20] (22)
HackMyAgent 5/5 63 16 39 3 29% [19–42] 28% [18–40] 6% [3–10] (10) 6% [4–11] (10)
CodeQL 5/5 245 33 21 191 61% [48–73] 13% [10–18] 4% [2–8] (8) 5% [3–10] (8)
AgentShield (Aiconnai) 2/5 220 55 159 0 26% [20–32] 26% [20–32] 17% [12–23] (30) 19% [14–26] (30)

Bastyn's other 4 findings are dependency claims (an outdated Flask version). 5 of its true positives are real weaknesses not in the answer key, such as a privileged container. Its 3 false alarms: a shell command in dvlaa's fixed copy that is actually safe, a report-building script in audit-lab whose file path isn't attacker-controlled, and a default credential in a dvlaa test script.

Recall per project (static-code and static-config vulnerabilities only; denominator in the header):

Tool dvmcp (52) dvlaa (46) dvar (21) llm-agent-ctf (21) audit-lab (16)
Semgrep 48% (25) 15% (7) 10% (2) 5% (1) 6% (1)
Bearer 37% (19) 9% (4) 19% (4) 29% (6) 0% (0)
Ship-Safe 31% (16) 17% (8) 10% (2) 24% (5) 12% (2)
Aguara 35% (18) 11% (5) 5% (1) 10% (2) 6% (1)
Agent Audit Kit 40% (21) 4% (2) 0% (0) 10% (2) 0% (0)
Bastyn 0.1.8 25% (13) 15% (7) 5% (1) 0% (0) 6% (1)
HackMyAgent 19% (10) 0% (0) 0% (0) 0% (0) 0% (0)
CodeQL 4% (2) 2% (1) 14% (3) 10% (2) 0% (0)

Recall depends heavily on the project: dvmcp is textbook-dense, and four tools (Semgrep, Bearer, Aguara, Agent Audit Kit) find a third or more of its weaknesses; audit-lab's weaknesses are in application logic, and no tool finds more than 2 of 16.

By detectability class (vulnerabilities found):

Tool static-code (146) static-config (10) prompt-semantic (25)
Semgrep 32 4 0
Bearer 33 0 0
Ship-Safe 27 6 4
Aguara 23 4 0
Agent Audit Kit 21 4 1
Bastyn 0.1.8 22 0 4
HackMyAgent 9 1 0
CodeQL 8 0 0

Weaknesses that live in prompt or tool-description text are almost entirely missed: 25 exist, and the best tools find 4. Bastyn's four come from a rule that detects hidden instruction blocks (such as <IMPORTANT>) in MCP tool descriptions, all in dvmcp.

All tools together:

Project Found by at least one of the eight tools
dvmcp 41 of 59
dvlaa 16 of 58
dvar 11 of 24
llm-agent-ctf 13 of 24
audit-lab 3 of 16
All 84 of 181 (46%)

Adding AgentShield (Aiconnai), which ran on only two projects, raises this to 88 of 181 (49%).

Fixed versus vulnerable code. In dvlaa the vulnerable and fixed copies differ by a few lines, typically one quoting call, one guard or one constant. Several tools flag both copies the same way, which is one of the main sources of false alarms. Where a "fix" left the dangerous call in place, the answer key lists the fixed copy as vulnerable too.

5.3 Good repos in detail

"Ran on" counts projects where the tool finished within the 2-minute limit.

Tool Ran on Findings Judged Judged: TP / FP / non-security / dependency Precision, weighted Precision, judged [95% CI] False alarms per 1,000 Python lines
CodeQL 4/5 709 100 7 / 18 / 75 / 0 30% 28% [14–48] 0.31
Ship-Safe 5/5 5,743 125 12 / 63 / 23 / 27 11% 16% [9–26] 1.09
Bastyn 0.1.8 5/5 40 40 (all) 2 / 27 / 0 / 11 7% 7% [2–22] 0.05
Semgrep 5/5 1,091 125 4 / 37 / 84 / 0 4% 10% [4–23] 0.45
Bearer 4/5 464 100 10 / 90 / 0 / 0 4% 10% [6–17] 2.16
Agent Audit Kit 4/5 392 82 2 / 67 / 0 / 13 2% 3% [1–10] 0.63
HackMyAgent 5/5 241 105 3 / 93 / 6 / 3 1% 3% [1–9] 0.38
Aguara 4/5 1,291 100 6 / 86 / 0 / 8 0.5% 7% [3–14] 2.39

Weighted scales each project's judged shares up to the tool's full finding count there, then pools. This is the primary estimate of overall precision, since the judged sample takes the same number of findings from every project regardless of how many each tool actually produced there. Judged is the raw share within that sample, with a confidence interval; it does not by itself estimate precision across a tool's full output. Bastyn's two columns match because all of its findings were judged, not sampled.

Where the real weaknesses are. Across all tools, the true positives point at 18 distinct files:

Project Real weaknesses found Found by
babyagi The web app binds to all network interfaces with no authentication, so anyone who can reach the port can replace a stored function's code and run it. The dashboard renders that data without escaping (cross-site scripting). Error text is returned in HTTP responses. A plugin ships a real OAuth client secret. 13 files. Bastyn, Semgrep and HackMyAgent flag the exec; Semgrep flags the binding; Bearer, Ship-Safe and Aguara flag the dashboard; CodeQL flags the error text; Agent Audit Kit, HackMyAgent and Ship-Safe flag the secret.
gpt-researcher TLS verification is silently turned off on retry when downloading PDFs. A server-side request forgery check can be bypassed by DNS rebinding. The AI chat reply is rendered as HTML without the sanitiser the rest of the page uses. Exception text is returned in an API response. 4 files. CodeQL, Agent Audit Kit, Ship-Safe.
openhands-sdk A CI workflow on pull_request_target checks out and runs pull-request code while holding an API-key secret, gated only by a label. 1 file. Ship-Safe.
browser-use, open-interpreter None in the judged sample. n/a

13 of the 18 files are in babyagi, the smallest and least maintained project. Its central flaw is an unauthenticated API, open on all interfaces, that runs stored code. Several tools flag parts of it, and none connects the parts.

Bastyn 0.1.8's 27 false alarms:

Rule FP Where
BAS-ZT1-001 (hard-coded API key) 14 all in openhands-sdk test files
BAS-ZT1-010 (credential-shaped literal) 5 4 in source, 1 in an example
BAS-LLM02-001 (key passed as a literal to an LLM client) 3 all in test files
BAS-LLM10-004 (eval/exec on a non-literal) 2 source, where running code is the product's purpose
BAS-ZT1-003, BAS-MCP-002, BAS-SKILL-003 3 1 test file, 1 source, 1 Markdown file

18 of the 27 are in test files, which are false positives by the judging rule unless the secret is live. Skipping test paths for credential rules is the cheapest precision gain on the table here. An estimate from these verdicts, not a measured run, puts Bastyn's good-repo precision at about 18% if it did. Its two real findings are the babyagi exec calls that run stored, remotely replaceable code.

Bastyn across releases, on the good repos:

0.1.6 0.1.7 0.1.8
Findings 90 43 40
TP / FP 2 / 51 (sampled) 2 / 30 (all judged) 2 / 27 (all judged)
Vulnerable repos: recall (of 181) 6 6 26
Vulnerable repos: precision 94% (15/16) 94% (15/16) 93% (39/42)

0.1.8 added a broader set of recall rules: 6 → 26 of 181, at unchanged precision. It also fixed several shell-command and file-path false alarms that a release candidate had introduced on the good repos, and removed one old false alarm in dvmcp. Four OpenHands helpers that pass a function argument straight to a shell are now reported as hidden-by-default "observations" instead of defects.

5.4 A third vulnerable-repo set

Three more vulnerable projects, with an answer key of 48 vulnerabilities (12 in prompt text).

Project Stars Commit Known vulnerabilities
M507/HackMeGPT 14 c76842a 21
anthonyg-1/mcp-goat 3 61412a1 16
jlov7/damn-vulnerable-agent-asset-corpus 0 d70bd68 11
Tool Findings TP FP Precision Recall (of 48)
Aguara 93 11 82 12% [7–20] 15% [7–27] (7)
Agent Audit Kit 25 12 11 52% [33–71] 8% [3–20] (4)
HackMyAgent 75 4 56 7% [3–16] 8% [3–20] (4)
Ship-Safe 57 6 48 11% [5–22] 6% [2–17] (3)
CodeQL 40 6 10 38% [18–61] 4% [1–14] (2)
Semgrep 88 5 37 12% [5–25] 4% [1–14] (2)
Bearer 47 0 47 0% [0–8] 0% [0–7] (0)
Bastyn 0.1.8 9 0 1 0% (0 of 1) 0% [0–7] (0)

The whole field is weak here: the best tool finds 7 of 48, and the eight tools together find 17 of 48 (35%). With only 48 vulnerabilities, most differences in this table are within noise. Bastyn 0.1.8 finds none of the 48; its other 8 findings are dependency claims. 

5.5 Run time

Wall-clock seconds, summed over the projects each tool completed. Vulnerable repos: mean of three runs for the other tools, one run for Bastyn. Good repos: one run. Measured on a shared laptop, order-of-magnitude only.

Tool Vulnerable repos (5) openhands-sdk browser-use open-interpreter gpt-researcher babyagi
Bastyn 0.1.8 6.5 2 1.5 3.5 1.3 <1
Ship-Safe 6.7 9 3 23 5 1
Aguara 8.0 21 10 failed 116 1
Agent Audit Kit 21.8 93 33 failed 18 2
HackMyAgent 24.1 8 7 9 10 5
Bearer 59.7 failed 32 54 21 12
CodeQL (create + analyze) 127.3 failed 73 67 30 27
Semgrep 150.2 44 60 61 28 20

6. Observations

  1. Precision and recall trade off, and nothing escapes the trade. On the vulnerable repos, every tool with recall of 18% or more has precision of 56% or less, and every tool with precision above 60% has recall of 14% or less. The balance score compresses the top five tools into 0.23–0.29, closer together than the score's sensitivity to duplicate findings (§4.4) makes that ordering reliable.
  2. Recall is low for every tool, and the tools overlap heavily. The best single tool finds 20% of the known vulnerabilities, and all eight together find 46%. On the third set the best finds 15% and all eight together 35%.
  3. General-purpose scanners match or beat the agent-specific ones on the vulnerable repos. Semgrep and Bearer know nothing about agents and lead on balance score. The agent-specific tools' advantage shows up only in prompt-level and configuration weaknesses, and even there it is small: 4 of 25 prompt-level vulnerabilities at best.
  4. Precision on labs does not predict precision on real code. Every tool but one is at 11% weighted precision or below on the good repos. The exception is CodeQL, last by balance score on the vulnerable repos, which is the most precise tool on real code. Labs are dense with textbook weaknesses, so pattern matches there are usually right; real projects mostly guard their dangerous calls, and the same patterns fire on guarded code, tests, docs and skill files instead.
  5. Volume is a separate axis from accuracy. Ship-Safe reports 5,743 findings on the good repos; among its adjudicable security claims, an estimated 11% are real (weighted precision). Bastyn reports 40, with 2 of 29 security claims real. Per 1,000 Python lines, Bastyn raises the fewest false alarms (0.05, against CodeQL's 0.31). A quiet tool produces fewer alerts to review, but on these projects it also surfaces fewer real problems: Ship-Safe's larger output contains more of the real weaknesses (§5.3).
  6. Test files drive a large share of false alarms on real code. 18 of Bastyn's 27 good-repo false alarms are in test files, mostly test API keys. For Bastyn, skipping or down-ranking test paths for credential rules is the cheapest precision gain available here.

7. Limitations

  • The third vulnerable-repo set is small. 48 vulnerabilities give wide intervals: 0 of 48 is compatible with a true recall of up to 7%, and the leader's 7 of 48 with anything from 7% to 27%.
  • Labels were produced by AI models, with no human annotation. Two independent models agreed at κ 0.94 (vulnerable repos) and κ 0.75 (good repos, before a third tie-break pass). Disagreements were settled from the code. Two closely related models can share the same mistakes.
  • The good repos are sampled for every tool except Bastyn: at most 25 judged findings per tool per project, so small differences (3% against 7%) mean little.
  • The judging rules settle real ambiguities in one direction: for example, whether a capability that is the product's purpose counts, whether findings in tests count, and whether pinning claims count as dependency findings. Other reasonable rules would move some verdicts.
  • Strict ±5-line matching. A correct finding a few lines outside an answer-key range gets precision credit but no recall credit.
  • The vulnerable repos are small, all Python and deliberately vulnerable. Recall on real production code is unmeasured.
  • Default configurations. Tuned rule sets could change results, Semgrep's in particular.
  • The 2-minute limit is a policy, not a technical ceiling. It turns two runs that would have finished (CodeQL and Bearer on openhands-sdk) into failures.
  • Other tools were not re-run for this report. Their versions are as of 2026-09-23 to 2026-09-25; newer releases may differ.
  • A path-export bug strips the leading dot from paths such as .github/workflows/x.yml, affecting 230 findings across all corpora. Judges were told to read the dotted path, and no finding was scored a false positive because of it.

Try Bastyn Scan

Bastyn Scan is the open-source command-line scanner behind the Bastyn results in this report. It is a single binary with no account or API key needed, licensed Apache-2.0.

Install with Homebrew:

brew install bastyn-labs/tap/bastyn

or with Cargo (cargo install bastyn), or the install script from the repository. Check it works with bastyn --version, then scan the current directory:

bastyn scan

The scan reports defects (provable problems that can fail a build) separately from observations (missing controls that cannot be confirmed as bugs). Add --show-observations to see the observations, or --offline to skip the dependency vulnerability lookups.

BASTYN provides you with actionable cross layer findings, missing defences, coverage gaps and compliance crosswalk without the noise.

 

Help us test it on your agents

Bastyn Scan has been run against thousands of repositories so far. But to keep improving the scanner and the next round of assessments, we need more people to scan their repos. If you build agents, you can help:

  • Run it on your own agent code. It takes a minute and needs no account. Get it from github.com/bastyn-labs/bastyn-scan.
  • Tell us what it got wrong. False alarms, missed vulnerabilities and rules that do not fit your framework are the most useful input we can get. Open an issue at github.com/bastyn-labs/bastyn-scan/issues.
  • Share your results with the community. Post what it found, and what it missed, in the same issue tracker so others can compare notes and we can fold the findings into the next round.

Scan your agentic codebase, get protected. Help us improve this open-source scanner by running it on your projects or contributing.

Subscribe for latest BASTYN Trust Intelligence

For risk, compliance, product and security experts