@mstange/perfherder-cli
v0.3.1
Published
Query Firefox performance data from the command line: search signatures, detect steps, compare pushes, and attribute a regression to what landed
Maintainers
Readme
perfherder-cli
Query Firefox performance data from the command line. Find a benchmark, see where it stepped, work out what landed there, and say whether a regression was the same work getting slower or a slow path being taken more often.
It talks to treeherder.mozilla.org directly. No account, no configuration,
nothing to run alongside it.
npm install -g @mstange/perfherder-cli
perfherder-cli --helpWhat it is
The command-line half of perfherder2, a client-side reimplementation of Perfherder's graphs view — and built from that app's own modules rather than beside them, so an answer here and the same answer read off the graph cannot disagree. Every command prints a link that opens the graph it is describing, because a text answer nobody can check against the picture is one nobody should act on.
| Command | Answers |
| --- | --- |
| search <term...> | which signature do I mean? |
| series <ref...> | what level is it at, and how do two of them compare? |
| series <ref...> --drift | it never stepped — how far has it slid since February? |
| series <ref...> --runs | one row per job: which worker ran it, and how it read |
| machines <ref...> | is that scatter the test, or is it one worker in the pool? |
| noise <ref...> | what is the scatter made of, and what can this series resolve? |
| changes <ref...> | where did it move, and what landed there? |
| changes <ref...> --cluster | which landings moved these twenty series, not which series moved? |
| step <ref...> --at | how big is the move here, on each of these series? |
| locate <ref> --at | which push is the step actually on? |
| compare <a> <b> | statistics, distributions, and whether the modes moved |
| commits <repo> <from> <to> | what landed between two revisions |
| url <ref...> | a shareable link, without fetching anything |
A series reference is <repo>,<signatureId>[,<frameworkId>] — the first
column of search, and the same thing a series= parameter in a shared app link
contains, so references paste both ways.
A worked investigation
# 1. Find the signature.
perfherder-cli search speedometer3 android --repo mozilla-central
# 2. Levels over a window — the "A vs B" answer.
perfherder-cli series mozilla-central,270490 mozilla-central,230167 --range 60d
# 3. When did it move, and what landed?
perfherder-cli changes autoland,5350953 --range 6mo --commits
# 4. What kind of change was it? Statistics, distributions, and whether the
# modes moved or only their weights.
perfherder-cli compare autoland,5350953@<beforeRev> <afterRev> --pool 24
# 5. Did the other platforms see it, even where no bar was drawn?
perfherder-cli step autoland,5350953 --across platform --at <rev> --range 60d
# 6. Before trusting any of it: can this series resolve a change that size,
# and is the scatter the test or the pool?
perfherder-cli noise autoland,5350953 --range 30d
perfherder-cli machines autoland,5350953 --range 90d --sort name
# …and over a whole suite at once: which landings moved it, six months back?
perfherder-cli changes <refs...> --range 6mo --cluster --briefFour things here have no counterpart in the app. Two exist because silence is not evidence:
stepmeasures a change at a point you name and says which of the detector's two bars a real-but-unmarked move failed. A platform running a benchmark once per push, beside one running it twelve times, has several times the per-push noise — so the same real step is certified on one graph and invisible on the other. Reading that as "it didn't happen here" is the mistake this prevents.series --driftprints the first window of a range against the last, for a series that slid 8% over three months without ever stepping. Segmentation looks for steps and there is no step in that shape, so no bar is drawn and nothing is wrong.
And two because a point estimate is not an interval, and a series is not a cause:
locateranks every push a step could be on, by the criterion the detector itself uses to place a bar, and marks the one Perfherder alerted on.changes --clustergroups the events of a whole ref list into the landings behind them, joining on the interval each event brackets rather than on the push it was placed on. Nine events across three platforms become one row — and the intersection of their brackets is often narrower than any single series carries, sometimes a single push.
Before you trust a comparison
noise splits a series' scatter into the three levels it has — a replicate
inside its run, a run inside its push, a push mean inside the series — because
the third is the one everything else reports and the second is the one that
matters. A series can print cv 1.5% and be made of jobs that scatter by 3.2%,
which is not a smaller version of the same finding: the test is noisy and the
retriggers are hiding it. It ends on what a comparison of two single pushes could
resolve at all, which is the number to know before pushing to try rather than
after squinting at two dots.
Where that scatter turns out to be the pool rather than the test, machines
names the workers behind it: each one's level against the closest thing to a
simultaneous measurement, its standard error, and how much its own runs scatter.
--sort name is worth reaching for — a pool's device families are contiguous
under name order and scattered under every other, which is how a 53-machine
Android pool turned out to be two batches 4.3% apart.
Notes that save a round trip
--jsonprints the same object the text was rendered from, for piping.- Responses are cached under
$XDG_CACHE_HOME/perfherder-clifor ten minutes to a day depending on how fast the thing behind them changes, so narrowing a search is cheap.--no-cachewhen you need it fresh. - Ranges:
--range 90d/6mo/1y/36h, or--from/--towithYYYY-MM-DD. Every command prints the range it resolved. All times are UTC. - Links are built against https://perfherder2.netlify.app/. Point them
somewhere else with
--base <url>orPERFHERDER2_BASE_URL.
Design
docs/cli.md has the reasoning: why the CLI lives inside the app's repository, why a pooled comparison is tested over push means, what the mode analysis is for, and what four fresh sessions found when they were handed the tool with no context.
MPL-2.0.
