grok-autoresearch
v1.0.0
Published
Durable autonomous experiment loops for Grok Build.
Maintainers
Readme
grok-autoresearch
Durable autonomous experiment loops for Grok Build.
Measure a baseline, change one variable, keep improvements, discard regressions, and continue in the same Grok session.
Requirements
- Grok Build CLI
0.2.106or newer - Node.js 22 or newer
- Git, unless a disposable session explicitly sets
allowNoGit: true
The plugin has no npm dependencies. It does not require a browser, a separate API key, or additional account configuration. The optional hint tool invokes the already-installed Grok CLI and uses its existing login; it never reads or bundles Grok credential files.
Install
Recommended, through Grok's native plugin manager:
grok plugin install aa2246740/grok-autoresearch --trust
grok mcp doctor autoresearchThe npm package is a zero-dependency installer for the matching tagged release:
npx --yes [email protected] install --trust
npx --yes [email protected] doctorRestart any Grok TUI that was already running when the plugin was installed.
Autocomplete should show one flat /autoresearch command.
To update later:
grok plugin update research-labQuick start
Run Grok inside a Git repository, then explicitly start a bounded loop:
/autoresearch Optimize the parser for 20 runs; minimize latency_ms and keep tests passing.The agent prepares a stable benchmark, records the baseline, and changes one coherent variable per experiment. Each accepted result becomes a Git commit; rejected or invalid candidates are restored automatically.
Commands
| Command | Behavior |
| --- | --- |
| /autoresearch <goal> | Start or continue an experiment loop. Never runs implicitly. |
| /autoresearch resume | Resume persisted .auto state. |
| /autoresearch status | Inspect state without activating the loop. |
| /autoresearch export | Open the live local dashboard. |
| /autoresearch off | Cancel pending continuation and persist manual-off. |
| /autoresearch clear | Stop and remove current and legacy experiment logs. |
| /autoresearch finalize | Turn selected results into reviewable branches. |
How it works
controlvalidates Git state and establishes explicit activation.run_experimentexecutes the stable benchmark and optional correctness checks.log_experimentappends JSONL evidence and owns keep, commit, discard, and rollback.- A one-shot background token wakes the same Grok session for the next iteration.
- A
PreToolUseguard blocks edits between logging a result and consuming that token. - The loop stops at its configured limit, on
/autoresearch off, on interruption, or when the task is genuinely blocked.
The durable record lives under .auto/:
| File | Purpose |
| --- | --- |
| .auto/prompt.md | Objective, constraints, files in scope, and current guidance. |
| .auto/measure.sh | Stable benchmark emitting METRIC name=number lines. |
| .auto/checks.sh | Optional correctness gate; failures cannot be kept. |
| .auto/ideas.md | Candidate ideas and experiment notes. |
| .auto/log.jsonl | Append-only result ledger. |
| .auto/hooks/ | Optional before/after shell hooks. |
Configuration
.auto/config.json stays at the Grok workspace root even when workingDir
points at a nested project:
{
"workingDir": ".",
"maxIterations": 20,
"maxAutoResumeTurns": 20,
"allowNoGit": false,
"hints": {
"enabled": false,
"provider": "xai",
"model": "grok-4.5",
"thinkingLevel": "high",
"maxRecentRuns": 8,
"maxCallsPerSession": 5,
"timeoutSeconds": 120
}
}maxIterations bounds durable experiments. maxAutoResumeTurns separately
bounds automatic same-session wakeups. Hints are disabled by default and are
advisory only; the active Grok executor retains file, shell, benchmark, and
keep/discard control.
Included capabilities
- Current
.auto/and legacyautoresearch.*path recovery - JSONL segments, secondary metrics, confidence estimates, and ASI diagnostics
- Stable benchmark guard,
METRICparsing, timeout/cancel, checks, and output spill files - One-shot same-session continuation with a pending-work tool guard
- Before/after shell hooks with JSON stdin, an 8 KB stdout cap, and observability entries
- Deterministic compaction summaries
- Live local dashboard with JSONL and SSE updates
autoresearch-create,autoresearch-hooks, andautoresearch-finalizesupport skills
Pi's fullscreen terminal widget is host-specific and is not included. Grok uses concise tool status plus the same live web dashboard.
Safety model
Autoresearch deliberately edits files and creates commits. Use a clean Git
worktree or disposable branch, define correctness checks, choose a finite run
limit first, and review kept commits before merging. /autoresearch off cancels
the next pending continuation without deleting the experiment ledger.
The dashboard listens on loopback only. Benchmark and hook commands are project code and run with the same operating-system permissions as Grok Build.
Development
npm test
grok plugin validate .
grok plugin install "$PWD" --trust
grok mcp doctor autoresearchThe automated suite covers command registration, explicit activation, Git keep/discard behavior, correctness gates, timeouts, large output, one-shot continuation, pending-work guards, hooks, compaction, the dashboard, the MCP handshake, working directories, bounded hints, and the npm installer argument boundary.
Provenance
This is a Grok Build port of the MIT-licensed
pi-autoresearch implementation
at commit 5c13dd07e3d6c52d7a6da538198ce965ad8a5123. The upstream copyright and
license are preserved in LICENSE.
License
MIT. See LICENSE.
