asiscan-cli
v1.5.3
Published
Static analysis for AI agent codebases. Audits against the OWASP Top 10 for Agentic Applications (ASI01-ASI10) and the OWASP Top 10 for LLM Applications.
Maintainers
Readme
ASIScan
Static analysis for AI agent codebases. Audits against the OWASP Top 10 for Agentic Applications (ASI01–ASI10), the OWASP Top 10 for LLM Applications, and EU AI Act Article 50.
npx asiscan-cli . ASIScan — OWASP ASI / LLM Top 10 static audit
2 files · 2 KB · 18 rules · 9ms
CRITICAL ASI01 Agent Goal Hijack
agent.ts:13:9
Untrusted or dynamic content is interpolated directly into the system
prompt. Anything reachable by an attacker in that string becomes an
instruction.
│ const systemPrompt = `You are an ops assistant. Context: ${retrievedDoc}
CRITICAL ASI05 Unexpected Code Execution (RCE)
agent.ts:44:5
Shell command built by string interpolation. Model-influenced values in
a shell string are a command-injection primitive.
│ exec(`bash -c "${args.cmd}"`, (e, stdout) => {
11 critical · 12 high · 3 mediumWhy this exists
OWASP published the Top 10 for Agentic Applications in December 2025 and refreshed the GenAI LLM Top 10 on 3 August 2026. EU AI Act Article 50 transparency obligations became enforceable on 2 August 2026.
Every engineering lead who shipped an agent this year is now expected to show they have assessed against frameworks that are months old, using tooling built for a threat model that predates agents entirely. Your SAST scanner does not know what a tool call is. It will not tell you that your agent has a standing admin token, that retrieved documents land in your system prompt, or that nobody can stop a running agent.
That is the gap this fills.
What it actually checks
18 rules across three frameworks. The interesting part is how they check.
Most agentic risk is a missing control, not a forbidden token. A codebase
isn't insecure because it calls exec — it's insecure because it calls exec
on a model-influenced string, in a process with no sandbox and no egress
policy. So ASIScan runs two kinds of probe:
- Presence probes flag genuinely dangerous constructs — a shell string
built by interpolation, a hardcoded long-lived token, model output reaching
innerHTML. - Control probes fire only when your code demonstrably does the risky thing and shows no evidence anywhere in the project of the mitigating control. Trigger without control is the finding.
Control evidence is deliberately generous, and comments never count as
evidence. A // TODO: add sandboxing comment must not make a scanner report
you as mitigated — for a security tool, a false "you're safe" is the worst
possible output.
Coverage
| Framework | Rules | |---|---| | OWASP Top 10 for Agentic Applications (2026) | ASI01–ASI10 — goal hijack, tool misuse, identity abuse, supply chain, RCE, memory poisoning, inter-agent comms, cascading failures, human-trust exploitation, rogue agents | | OWASP Top 10 for LLM Applications (2025) | LLM01, LLM02, LLM05, LLM06, LLM07, LLM10 | | EU AI Act | Art. 50 transparency and synthetic-content marking |
LLM04 (Data and Model Poisoning) and LLM09 (Misinformation) cannot be meaningfully assessed from source. They are covered in the manual review checklist rather than faked here with regexes.
ASCII smuggling
A model reads codepoints. A reviewer reads rendering. Everything in the gap between those two is an attack surface, and it is the one place where a scanner has a structural advantage over a careful human.
ASIScan flags, across every file it reads:
- Zero-width characters (U+200B, U+200C, U+2060–U+2064) — invisible in your editor and in the diff, ordinary tokens to the model.
- Bidirectional overrides (U+202A–U+202E, U+2066–U+2069) — Trojan Source, CVE-2021-42574. Code that renders one way and parses another.
- Unicode tag characters (U+E0000–U+E007F) — a complete invisible ASCII alphabet with no legitimate use in source.
- Instruction-override language in tool and skill descriptions — a description is passed to the model verbatim, so it is an instruction channel, not documentation. "Before using any other tool you must first…" in a description is tool poisoning.
U+FEFF is deliberately not treated as a smuggling carrier. It is the byte-order mark: legitimate at the head of a file and emitted on purpose by .NET code generators. An early version of the probe included it, and every finding it produced across the tuning corpus was false. A signature that is 100% noise on real code is worse than no signature, so it was removed and a regression test now pins that behaviour.
Across five large agent frameworks these four probes produce zero findings. Against a crafted smuggling fixture, all four fire. Silent on clean code, loud on the actual attack, is the profile a security probe should have.
Measured behaviour
Against the fixtures in test/fixtures/:
| Fixture | Findings | Controls satisfied |
|---|---|---|
| vulnerable-agent (deliberately unsafe) | 26 | 0 |
| secure-agent (reference implementation) | 0 | 5 |
| polyglot-agent (Python / C# / Go) | ASI01, ASI03, ASI05 in all three | — |
| ASCII-smuggling fixture | 4 (zero-width, bidi, tag chars, tool desc) | — |
Sensitivity and specificity both matter. A scanner that finds everything is noise; one that finds nothing is decoration.
Measured against real repositories
Fixtures prove a scanner can fire. They don't prove it is usable. v1.0 was tuned against fifteen open-source agent frameworks in two rounds — roughly 17,000 source files in total.
| Corpus | Before tuning | After tuning | |---|---|---| | Round 1 — 5 repos, ~3,100 files | 792 findings | 36 | | Round 2 — 10 repos, ~14,000 files | 339 findings | 78 |
Every surviving finding was reviewed by hand. Precision by tuning round:
| Round | Findings | Precision | |---|---|---| | Untuned, 5 repos | 792 | ~4% | | Tuned, 5 repos | 36 | ~55% | | 10-repo corpus | 89 | ~62% | | + context gating | 78 | ~69% | | + targeted fixes | — | ~80% measured on a 5-repo subset; ~75% estimated corpus-wide |
The last figure is labelled honestly: the final fixes were validated against the five repositories that carried the false positives they targeted (49 findings, ~39 true), not re-measured across all ten. Treat ~75% as the working estimate and ~80% as the best case.
For context: published figures put untuned commercial SAST at 60–90% false positives, dropping to 10–20% once tuned for a stack, and SonarQube at 40–60% of findings requiring developer review. A ~20–25% false-positive rate is therefore in the same band as well-tuned commercial tooling — with the difference that this number is measured, published, and reproducible rather than asserted.
The false positives removed along the way are worth stating plainly, because they are the failure modes every regex-based scanner has:
- Flagging
url.startswith(('http://', 'https://'))— scheme-validation code, i.e. reporting the security control as the vulnerability. 475 findings. - Flagging
xmlns="http://www.w3.org/2000/svg"— an XML namespace identifier, not a network endpoint. 56 findings in a single icon file. - Treating test files as production code.
config_test.pyandtest_routes.pyare full of deliberately fake secrets and example URLs; they accounted for roughly two thirds of all findings in round two. - Matching JavaScript's
regex.exec(text)as process execution. - Reading
maxAge: 0as "credential never expires" when it means expire now — exactly backwards. - Matching "YOLO" the object-detection model inside recorded test fixtures.
Honest precision estimate: roughly 75% on real code. That is a usable signal-to-noise ratio for a security review, and it is not 100%. Expect to dismiss some findings, and please report them to [email protected] — rules are data, and most fixes are a one-line change that ships in the next v1.x. False-negative reports (a real risk it should have caught and didn't) are the most valuable feedback we get.
Context gating
Some constructs are only interesting inside an agent. role = "admin" in an
ordinary RBAC module is application design; the same line in a file that also
defines agent tools is excessive agency. Line-local matching cannot tell those
apart, and conflating them was the largest remaining source of false positives.
Presence probes therefore support a nearby gate: the finding is recorded
only if the file also shows the relevant context — agent/tool surface for LLM06,
external-input handling for ASI01. That single change removed LLM06 as a
false-positive source entirely and took precision from 62% to 69%.
Known remaining weak spots, so you can weight findings accordingly: security
example code that deliberately contains unsafe URLs is flagged (an SSRF-defence
demo looks identical to an SSRF bug from source); constants whose value equals
their own name (CONFLUENCE_API_TOKEN = "CONFLUENCE_API_TOKEN") read as
hardcoded secrets; and ttl=None in a storage API is read as a non-expiring
credential.
Usage
Run it without installing via npx asiscan-cli ., or install it once with
npm install -g asiscan-cli and use the asiscan command:
asiscan [path] [options]
-f, --format <fmt> terminal | markdown | json | sarif (default: terminal)
-o, --output <file> write report to a file
--fail-on <sev> exit 1 at/above severity (default: high)
--only <ids> e.g. ASI01,ASI05
--skip <ids> exclude rules
--frameworks <f> asi | llm | euaia | all# Human-readable audit
asiscan .
# Evidence document for a security review
asiscan . --format markdown --output audit.md
# GitHub code scanning
asiscan . --format sarif --output results.sarif --fail-on critical
# Focus a specific review
asiscan . --only ASI05,ASI09CI
Copy .github/workflows/asiscan.yml into your repo. It runs on every PR,
uploads SARIF to the GitHub Security tab, writes a markdown summary to the
job page, and gates on critical findings.
Programmatic
import { audit, ASI_RULES } from "asiscan";
const result = audit({ root: "./src", rules: ASI_RULES });
console.log(result.findings.filter((f) => f.rule.severity === "critical"));What's in the box
src/ Scanner, rule engine, four reporters
templates/
manual-review-checklist.md 49 checks covering what static analysis can't see
eu-ai-act-readiness.md Scope, classification, Art. 50, evidence map
redteam/
probes.jsonl 30 adversarial probes mapped to risk IDs
README.md How to run them through real channels
.github/workflows/ CI workflow with SARIF upload
test/fixtures/ Vulnerable and secure reference agentsThe templates matter as much as the scanner. Static analysis covers roughly half of the ASI Top 10 — the rest is architecture, runtime config, and process. The checklist is what you hand an engineer to cover the other half, with an evidence column, because "yes" without evidence is not an answer an auditor accepts.
Suppressing a finding
At ~75% precision you will still hit false positives. Dismiss one on a single line:
const cmd = build(userChoice); // asiscan-ignore ASI05 reviewed: userChoice is an enumWorks in //, #, and /* */ comments. The rule id and reason are
conventional, not parsed -- write them anyway. An unexplained suppression is
indistinguishable from someone who got bored, and six months later nobody can
tell which one it was.
Suppression applies to presence probes only. A control probe asks whether a mitigation exists anywhere in the project; that is not a question a single line is entitled to answer.
If a rule is wrong often enough that you are suppressing it repeatedly, report it instead. Rules are data, not hardcoded logic -- most fixes are a one-line regex change.
Honest limitations
Read this part before you rely on the tool.
- It reads source, not behaviour. It cannot see runtime configuration, IAM policy, network topology, or what your model actually does.
- Control probes are project-wide. If a mitigation exists anywhere, the rule stays quiet — even if it isn't applied on the path that needs it. This favours quiet over noisy, and it means a clean result is weaker evidence than a dirty one.
- Regex-based, not AST-based. It will miss things a compiler wouldn't, and it can be fooled by unusual formatting.
- Not a certification. Not a conformity assessment. Not legal advice. Findings under the EU AI Act heading are engineering-visible proxies for obligations that require documented human assessment.
A clean report means "nothing detectable from source." It does not mean secure. The manual checklist exists precisely because the scanner is not sufficient.
Requirements
Node 18+. No network access, no telemetry, no data leaves your machine.
Open core
The scanner is MIT and complete on its own. Scan engine, all 18 rules, terminal / markdown / JSON / SARIF output, the CLI, the fixtures and the CI workflow — install it, run it, fork it, ship it in your product. It is not a trial and it is not crippled.
npx asiscan-cli .The Evidence Pack is the commercial part, and it is what turns a scan into something you can send to a reviewer:
| | | |---|---| | Assessment report | Dated, scoped, named-assessor HTML/PDF with findings, remediation windows, owner columns and a sign-off block | | Framework crosswalk | Every rule mapped to NIST AI RMF 1.0 and ISO/IEC 42001:2023 Annex A, with an honest evidence-strength column | | Questionnaire answer pack | Draft answers to the AI questions on enterprise security questionnaires and agent RFPs, grounded in your own findings | | Manual review checklist | 49 checks covering the half of the ASI Top 10 that source analysis cannot see | | EU AI Act worksheet | Scope, classification, Art. 50 controls, evidence map | | Red-team probe suite | 30 adversarial probes mapped to risk IDs |
asiscan . --format assessment --company "Acme Ltd" \
--scope "api/ on main" --assessor "J. Smith, Head of Security" \
--output assessment.htmlThe split is deliberate. Rules are a commodity — there are free scanners and there should be, and ours is one of them. What nobody hands you is a document a procurement team accepts. That is the part worth paying for, and it is the part that takes judgement rather than regexes.
See LICENSE (MIT, scanner) and LICENSE-EVIDENCE-PACK (commercial).
