@qaguardian/playwright-score
v2.1.1
Published
Deterministic, AI-free Playwright spec quality score — the weighted share of your tests that are clean.
Maintainers
Readme
@qaguardian/playwright-score
Deterministic, AI-free quality score for Playwright specs — the weighted share of your tests that are clean.
Website: qaguardian.com/open-source/playwright-score
Built by: QA Guardian — managed Playwright E2E (AI drafts, engineers verify, you own the code)
Lint + score Playwright tests against community best practices via
eslint-plugin-playwright
plus suite-level metrics (locator ratio, assertion-delegation tracing,
Page Object Model import resolution, ...).
The score never calls an LLM. AI may generate or repair code using findings; rules grade code.
Real-world results
100 public Playwright suites, scored with the published package against each project's actual source — chosen to include both well-known, heavily-engineered platforms and smaller, less mature projects, plus three named testing-infra competitors scored under the exact same rules as everyone else. Not curated to look good: 35 of 100 fail the default threshold, for real, verifiable reasons (see VALIDATION.md, which also has the full before/after against the last published version and every one of the 100 entries, not just the 15 shown below).
100 suites · 7,281 files · 39,059 tests · 65/100 pass (80% threshold) · tested on 2026-09-24
Top scorers (everyone at 97+; the next tier down is a 6-way tie at 96 — see VALIDATION.md for the full 100):
| Repo (source scanned) | Score | Grade | Result | |---|---:|:-:|:-:| | pinterest/gestalt | 100/100 | A | PASS | | TryGhost/Ghost | 99/100 | A | PASS | | hashicorp/vault | 98/100 | A | PASS | | microsoft/playwright (TodoMVC example) | 98/100 | A | PASS | | plausible/analytics | 98/100 | A | PASS | | segmentio/analytics-next | 98/100 | A | PASS | | cloudflare/templates | 97/100 | A | PASS | | glpi-project/glpi | 97/100 | A | PASS | | shopify/hydrogen | 97/100 | A | PASS |
Bottom 5 scorers:
| Repo (source scanned) | Score | Grade | Result | |---|---:|:-:|:-:| | Flagsmith/flagsmith | 62/100 | D | FAIL | | activepieces/activepieces | 58/100 | F | FAIL | | th3cyb3rhub/TheCyberHub | 56/100 | F | FAIL | | DataDog/documentation | 50/100 | F | FAIL | | getsentry/spotlight | 33/100 | F | FAIL |
Repo names link to the exact commit scanned. Selected by a mix of
mechanical GitHub search (real @playwright/test usage, hand-verified)
and, for the most recent 16, hand-picked brand recognition — three of the
100 are named competitors' own example suites (Checkly, Currents,
LambdaTest), scored with no separate bar. Full methodology, the complete
100-repo table, findings breakdown, and
scripts/validate-corpus.sh to reproduce every number here yourself live
in
VALIDATION.md.
A machine-readable
validation.json
carries the same per-repo data for tooling. See
CHANGELOG.md
for what changed release to release. A self-contained visual
report of an earlier five-suite audit lives at
docs/scorecard.html — host it wherever
(GitHub Pages, qaguardian.com, ...).
See METHODOLOGY.md for the full scoring model, or the product write-up on the landing page.
Install
npm install -D @qaguardian/playwright-scorePackage: @qaguardian/playwright-score on npm.
CLI
# One-shot (scoped package — use -p so the playwright-score binary is resolved)
npx -p @qaguardian/playwright-score playwright-score ./tests --profile standard --threshold 80
# After local install
npx playwright-score ./flow.spec.ts --format json --out report.json| Flag | Description |
|---|---|
| --profile standard | Scoring profile (default, and only, profile) |
| --threshold <n> | Pass bar 0–100 |
| --format text\|json\|markdown\|sarif | Output format |
| --out <file> | Write report to file |
| --version | Print version |
Exit codes: 0 pass · 1 below threshold or no files matched · 2 tool error
Library
import { scorePaths } from '@qaguardian/playwright-score';
const result = await scorePaths({
paths: ['tests/login.spec.ts'],
profile: 'standard',
threshold: 80,
});
console.log(result.score, result.grade, result.pass, result.findings);GitHub Action
Drop this into a workflow to gate PRs on the score, post a job summary, and keep a sticky PR comment with the full findings up to date:
- uses: qa-guardian/playwright-score@v1
with:
paths: tests e2e
# threshold: 80 # defaults to 80
mode: gate # gate: fail CI below threshold · warn: report only| Input | Description | Default |
|---|---|---|
| paths | Space-separated paths/globs to score | tests |
| profile | standard (default, and only, profile) | standard |
| threshold | Pass bar 0–100 | 80 |
| mode | gate (fail CI below threshold) | warn (report only) | gate |
| comment | Post/update a sticky PR comment | true |
| version | @qaguardian/playwright-score version to run | latest |
| github-token | Token for the PR comment | ${{ github.token }} |
Outputs: score, grade, pass — usable by downstream steps (e.g. a
custom badge, a Slack notification on regression, etc).
Prefer a raw CLI call, or a non-GitHub CI system? See the CLI section above — same score, same exit codes:
- name: Playwright Spec Score
run: npx -p @qaguardian/playwright-score playwright-score ./tests --profile standard --threshold 80 --format textQA Guardian integration
QA Guardian's own codegen pipeline (playwright_runner) dogfoods this
package for the standard profile, layering its own private house-rules
gate on top internally — that layer isn't part of this package (it's
product-specific, e.g. "timeouts must be exactly 2000 or 20000ms", not
Playwright best practice) and isn't published here. playwright_runner
sets:
| Env | Values | Default |
|---|---|---|
| SPEC_SCORE_MODE | off | warn | gate | warn |
| SPEC_SCORE_THRESHOLD | 0–100 | 80 |
warn: log score; never fail the run; findings still inject into heal/generate repair promptsgate: fail the run when score < threshold- After AI generate, every spec is scored; a spec below threshold is repaired with its score findings and re-scored, up to 3 scored attempts (the initial generation plus up to two repairs). A spec still below threshold, or one that could not be scored, is flagged for Guardian review and never silently accepted.
- Heal and repair prompts always include score findings when the scorer reports issues.
# Local from monorepo
cd playwright-score && npm run build
cd ../playwright_runner && npm install
SPEC_SCORE_MODE=warn node ...Development
npm install
npm run build
npm test
npm run score -- ./fixtures/good-standard.spec.tsLicense
MIT
