cmd-risk
v0.4.0
Published
Heuristic advisory risk classifier for shell commands, with a forkable rule table, for deciding when to prompt a human.
Downloads
655
Maintainers
Readme
cmd-risk
A heuristic advisory risk classifier for shell command strings, with a small forkable rule table.
This is not a security boundary. It is for deciding when a coding agent (or an agent sandbox) should pause and prompt a human before running a shell command, not for containing a hostile actor. It matches on naive, regex-based heuristics and can be fooled by obfuscation, quoting tricks, aliases, or variable expansion. The command splitter is naive text splitting, not a real shell parser: it does not understand quoting, escaping, subshells, or heredocs.
The problem
Every coding agent and agent sandbox ends up hand-rolling some version of "is this shell command about to destroy something?" as an ad-hoc regex list buried deep in the codebase. These lists are rarely reused, rarely tested, and rarely agreed on between projects. This package pulls that logic out into one small, readable, forkable rule table plus a classifier, so you can start from something reasonable and edit it for your own risk tolerance instead of writing it from scratch.
Install
npm i cmd-riskUsage
import { classify, isSafe, splitSegments, RULES } from 'cmd-risk';
classify('rm -rf /');
// {
// level: 'critical',
// reasons: [
// 'Recursive force-remove targeting root, home, or the current directory.',
// 'Recursive, forced file removal.'
// ],
// matched: ['rm-rf-root', 'rm-rf'],
// segments: ['rm -rf /']
// }
classify('cd /tmp && rm -rf /tmp/build');
// level: 'high', segments: ['cd /tmp', 'rm -rf /tmp/build']
isSafe('git status'); // true
splitSegments('git add . && git commit -m "x"');
// ['git add .', 'git commit -m "x"']
// Add your own rule without touching the built-in table:
classify('terraform destroy', {
extraRules: [
{
id: 'terraform-destroy',
level: 'high',
pattern: /\bterraform\s+destroy\b/,
reason: 'Tears down provisioned infrastructure.',
},
],
});CLI
cmd-risk "rm -rf /tmp/x"Prints the risk level on the first line, then one - reason line per matched rule. Exits
0 for safe/moderate, 1 for high/critical, so it can gate a script:
cmd-risk "$CMD" || echo "needs human review"API
classify(command, options?) -> Verdict
command: string- a shell command string.options.rules?: Rule[]- replaces the built-inRULEStable entirely.options.extraRules?: Rule[]- appended to whichever rule table is in use.- Returns:
{ level: 'safe' | 'moderate' | 'high' | 'critical', reasons: string[], // human-readable, deduplicated, one per matched rule id matched: string[], // deduplicated matched rule ids, first-seen order segments: string[] // command split into naive segments } levelis the highest level matched across all segments. A command matching no rule is'safe'with emptyreasons/matched.
splitSegments(command) -> string[]
Splits on ;, &&, ||, |, and newlines. Trims each segment and drops empty ones. This is
intentionally naive: it does not track quoting, so a delimiter character inside a quoted
string still splits the command.
isSafe(command) -> boolean
Shorthand for classify(command).level === 'safe'.
RULES
The exported array of built-in rules. Each rule is:
{ id: string, level: 'moderate' | 'high' | 'critical', pattern: RegExp, reason: string }Copy the array, edit it, drop entries you don't care about, add your own, and pass it back in
as options.rules. That is the entire point of shipping it as plain data.
Built-in rule ids:
- critical:
rm-rf-root,mkfs,dd-to-device,fork-bomb,chmod-777-root,overwrite-block-device - high:
rm-rf,git-reset-hard,git-clean-force,git-push-force,git-checkout-dot,shred,truncate,pipe-to-shell,drop-sql,docker-prune-all,kill-all - moderate:
sudo,npm-publish,git-push,chown,chmod,package-remove,write-outside-cwd
Use as a Claude Code hook
cmd-risk-hook prompts or blocks before a destructive Bash command runs, when wired up as a Claude Code PreToolUse hook.
Install:
npm i -g cmd-riskAdd this to ~/.claude/settings.json:
{
"hooks": {
"PreToolUse": [
{
"matcher": "Bash",
"hooks": [{ "type": "command", "command": "cmd-risk-hook", "timeout": 5 }]
}
]
}
}By default, classify()'s level maps to a permission decision like this:
| level | decision |
| ---------- | --------- |
| critical | deny |
| high | ask |
| moderate | no action |
| safe | no action |
Two environment variables override the thresholds: CMD_RISK_DENY_AT (default critical) and
CMD_RISK_ASK_AT (default high), each one of safe, moderate, high, critical. The
command is denied when its level is at or above CMD_RISK_DENY_AT, otherwise it prompts when at
or above CMD_RISK_ASK_AT, otherwise the hook stays silent and the normal permission flow
applies. An unrecognized value for either variable falls back to its default.
The deny check is evaluated before the ask check, so if CMD_RISK_DENY_AT is set at or below
CMD_RISK_ASK_AT, the ask threshold becomes unreachable for anything that already clears the
deny threshold. safe is a legal value for both variables, but setting CMD_RISK_DENY_AT=safe
denies every Bash command, including harmless ones like ls or pwd, which effectively disables
Bash for the session. This is almost never what you want.
Set CMD_RISK_HOOK=off to disable the hook entirely: it then exits immediately with no output on
every call, regardless of the command or the threshold variables above. If you run Claude Code
unattended on a schedule, for example a nightly build or a scanning job on a systemd timer with
nobody watching, set this in those jobs so an automated run is never blocked waiting on a decision
nobody is there to make.
The hook fails open: if anything goes wrong while it runs, it exits silently and the command proceeds as if the hook were not installed. As with the rest of this package, it is a heuristic advisory classifier, not a security boundary.
How it works
classify runs every rule's regex against every segment produced by splitSegments, then
takes the deduplicated union of matched rule ids and the highest matched level. Two rules
(git-push-force and git-push) have an explicit interaction: if a segment matches
git-push-force, that same segment will not also report the plain git-push rule (a
different segment with an unrelated plain git push can still report it).
Two rule ids, pipe-to-shell and fork-bomb, are shapes that only exist as a literal | or
; sequence (curl x | sh, :(){ :|:& };:). Because splitSegments also splits on | and
;, per-segment matching alone would never see either shape intact. For exactly these two
rule ids, classify additionally tests the untouched original command string, on top of the
normal per-segment checks every other rule gets. This is a deliberate, narrow exception to
keep those two rules functional given a naive splitter; every other rule is matched strictly
per segment.
Before any rule pattern is tested, the contents of every single- and double-quoted string in
the text being checked are masked out (replaced with neutral filler, quote characters left in
place), so a risk-looking word inside a commit message, echo string, or -m argument - e.g.
git commit -m "sudo is mentioned here" - is not mistaken for the real command. A backslash
inside a quote escapes the next character, including a same-type quote, so it doesn't end the
string early. Masking is only used for matching: the segments array in the returned Verdict
is always the original, unmasked text. If a quote is left unterminated, masking falls back to
matching the raw, unmasked text for that piece of the command instead of masking everything
from the stray quote to the end of the string - an advisory classifier should fail open (still
flag real risk) rather than let an unbalanced quote hide something dangerous.
One targeted exception to masking: if a quoted argument's entire content is exactly a bare root
path (/, ~, $HOME, /*, or similar), it is left unmasked, so rm -rf "/" and rm -rf
'~' are still recognized as targeting root, the same as their unquoted forms. This only applies
when the quoted content is exactly one of those literal path shapes - a quoted word that happens
to be short, like "sudo", is unaffected and still gets masked.
Heredoc bodies are masked the same way, and for the same reason: writing documentation, a commit
message, a config file, or a code example through a heredoc is an extremely common shape, and the
body is just prose sitting unquoted in the command, so cat >> notes.md <<'EOF' followed by lines
that happen to mention rm -rf / or git clean -fdx should not get flagged. classify detects
<<WORD, <<'WORD', <<"WORD", and the indent-stripping <<-WORD forms, masks everything from
the line after the introducer up to the first line that is exactly WORD (leading whitespace
allowed for <<-), or to the end of input if that line never appears, and does this before
splitSegments runs, since a heredoc body's newlines would otherwise be shredded into unrelated
segments first. The one exception: a heredoc is left unmasked when it is fed to a known
interpreter (bash, sh, zsh, dash, ksh, python, python3, node, perl, ruby,
eval), because that body is genuinely executed, so bash <<EOF with rm -rf / in it still
classifies critical. If no command word can be identified in front of the << at all, the body is
also left unmasked, the same fail-open direction as the unterminated-quote ruling above.
Known limitations, honestly:
- The splitter has no notion of quoting:
echo "a; rm -rf /"is split as if the;were a real command separator, even though it is inside a string literal. - Regex-based flag detection is approximate. It looks for
-r/-f/--recursive/--forcestyle tokens; unusual flag bundling or long-option abbreviations it doesn't recognize can be missed. - The interpreter list for heredoc masking matches exact command words only, not paths:
/bin/bash <<EOFis not recognized asbash, so that heredoc's body would be masked even though it is really executed. Invoke interpreters by their bare name for the exception to apply. - Nothing here executes, sandboxes, or blocks anything. It only classifies a string.
- It is trivially bypassed by anyone motivated to bypass it (encoding, variable indirection, wrapper scripts). That is expected and fine for its intended use: a heuristic nudge for when to ask a human, not a control that has to hold up against an adversary.
Related
Small, single-purpose packages for the same problem space. Each one has zero dependencies and does one thing.
prompt-cache-fit- Reorder prompt blocks least-variable-first for prefix cache reuse, and measure the hit rate.apply-edit-block- Apply LLM search/replace edit blocks that do not match the source exactly.ctx-compact- Trim a conversation to a token budget without ever orphaning a tool result.cassette-fn- Record and replay LLM calls at the function boundary, so your tools still run on replay.
License
MIT
