style-doctor
v0.4.2
Published
Diagnose prose health the way react-doctor diagnoses a React codebase: grouped findings, a 0-100 score, CI-ready exit codes, JSON for agents.
Maintainers
Readme
style-doctor
react-doctor for prose. Point it at a folder; it scans Markdown/text and the
prose inside .astro / JSX / .html / .vue / .svelte components, groups
findings by rule, prints a 0–100 score, and exits non-zero when blocking
issues are present, so it drops straight into CI. Single file, zero
dependencies, Node ≥18.
It looks for LLM tells (delve, rich tapestry, not just X, but Y, em-dash
overuse, it's important to note, copula avoidance like serves as),
leftover AI-generation artifacts (chatbot sign-offs, oaicite/utm_source=
chatgpt.com citation scraps), LLM formatting defaults (Title Case headings,
**Label:** text bullets, bold overuse), and filler/grammar problems (weasel
words, wordy phrases, passive voice).
Use
npx style-doctor # scan ./
npx style-doctor docs/ # scan a folder
npx style-doctor CHANGELOG.md # scan specific files
extract-prose web/ | npx style-doctor - # score prose piped on stdinTemplates & components
style-doctor scans .astro, .jsx, .tsx, .html, .htm, .vue, and
.svelte by default, alongside .md/.markdown/.mdx/.txt. A naive (no-AST)
extractor pulls the linted prose out of each file:
- text nodes between tags,
alt,aria-label,title,placeholderattribute values,<meta name="description" content="…">and the like.
It skips frontmatter, <script> / <style> blocks, {…} expressions, and any
line that reads as code, not a sentence (identifiers, paths, class lists).
--no-templates turns it off and scans Markdown/text only.
It is a regex, not a compiler: a stray < in prose or an exotic JSX shape can
throw it off. For full control, pipe your own extraction in:
# a commit message
git log -1 --format=%B | npx style-doctor -
# gate a PR's added lines
git diff origin/main... | grep '^+' | cut -c2- | npx style-doctor - --stdin-name diff- reads prose from stdin; --stdin-name <label> sets the filePath reported
for those findings (default <stdin>). Combine stdin with file/dir args and
everything merges into one score.
Output mirrors react-doctor: an agent-guidance header, Scanned N files, a
score line, a per-category issue count, then grouped findings
(rule, then file:line:col per occurrence).
✔ Scanned 12 files in 21ms
Style Doctor — my-docs
Score: 78 / 100 Needs work
6 issues
LLM Tells: 6 errors
✖ Filler word "delve" ×4
style-doctor/delve
guide.md:5:21
...CI
- run: npx style-doctor --scope changed --quiet- Exit 0/1.
1when blocking issues exist.--blocking error(default) blocks on error-severity findings;--blocking warningblocks on any;--blocking noneis advisory-only. --scope changedlimits the scan to files changed vsgit HEAD(--base <ref>to compare against a PR base). Fast pre-commit / PR checks.--min <n>also fails if the score drops belown.--quietdrops the guidance header and footer.--scoreprints just the number.
For an AI
npx style-doctor --json # full report, suppresses other output
npx style-doctor --json-compact # same, one line{
"schemaVersion": 1, "tool": "style-doctor", "version": "0.4.0",
"ok": false, "score": 78, "label": "Needs work", "words": 640,
"summary": { "issues": 6, "errors": 6, "warnings": 0, "filesWithIssues": 1,
"byCategory": { "LLM Tells": { "errors": 6, "warnings": 0 } } },
"diagnostics": [
{ "filePath": "guide.md", "rule": "delve", "category": "LLM Tells",
"severity": "error", "title": "Filler word \"delve\"",
"message": "...", "help": "Use \"look at\", \"go into\", or cut.",
"line": 5, "column": 21, "match": "delve",
"id": "guide.md::5:21::style-doctor/delve" }
]
}Finding ids are stable (<file>::<line>:<col>::style-doctor/<rule>). Loop:
scan → apply help at each line:col → re-scan until ok / your --min.
Options
--category "LLM Tells"|AI Artifacts|Filler|Formatting|Grammar (repeatable,
display filter; score stays global) · --no-warnings · --only a,b ·
--ignore a,b · --exclude a,b
(skip paths, globs) · --no-templates · - (read stdin) · --stdin-name
<label> · --no-color / NO_COLOR · --rules · --selftest.
Persistent config: a .style-doctor.json file, or a "style-doctor" key in
package.json:
{ "exclude": ["vendor/**", "CHANGELOG.md"], "ignore": ["passive-voice"] }exclude skips paths (globs, matched at any depth: ** spans directories, *
stays in one segment); ignore skips rule ids. CLI --exclude / --ignore add
to the config lists.
Use it as a library too: import { buildReport } from "style-doctor", then
buildReport(dir, opts) returns the same object as --json.
Scoring
score = 100 − 4 × (weighted findings per 100 words), clamped 0–100.
Weights: error 3, warning 1. Label: ≥90 Excellent, ≥80 Healthy, ≥50 Needs work,
else Critical.
Prior art
Vale and proselint
are the established prose linters, with richer engines, but no single score, no
react-doctor-style grouped report, and setup to do. style-doctor trades
breadth for one zero-config command tuned for LLM tells: scan → score → gate.
Passive-voice detection is a naive regex (no POS tagger); pass --ignore
passive-voice if it's noisy for reference docs.
Credits
Simon Willison's llm-cliche-highlighter (tools.simonwillison.net) inspired the LLM-tell rules. It is a browser tool that highlights the words and constructions LLMs overuse; this project turns that idea into a scored, CI-friendly CLI. Thanks, Simon.
The output format and CI ergonomics follow react-doctor.
The AI Artifacts and Formatting rules (chatbot sign-offs, citation scraps,
Title Case headings, **Label:** bullets, bold overuse), and several LLM
Tells / Filler rules (copula avoidance, vague attribution, legacy praise,
hedge stacking), come from Wikipedia's
Signs of AI writing
and Conor Bronsdon's
avoid-ai-writing pattern
list.
Changelog
See CHANGELOG.md.
License
MIT © Andrea Bruno
