npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

pi-observable-coding-agent

v0.0.2

Published

A local-first observability and workflow review extension for Pi coding agents

Readme

Pi Observable Coding Agent

A local-first Pi extension that turns coding-agent sessions into structured task runs, deterministic workflow reviews, and evidence-backed Skill improvement candidates.

Version 0.3 observes the agent without changing its decisions. After Pi fully settles a task, it:

  1. Correlates tool calls and results, including duration and failure fingerprints.
  2. Normalizes built-in file reads and uses workspace snapshots to resolve opaque shell mutations.
  3. Separates completed tasks from tasks with passing verification evidence.
  4. Classifies task type, user corrections, validation retries, and repeated errors.
  5. Records Skill version hashes, activation level, and optional deterministic checkpoints.
  6. Calculates both per-run findings and cross-run workflow insights.
  7. Persists the complete structured run as JSONL and a durable Pi session entry.
  8. Exposes /trace, /review, and a local Web dashboard for inspection.
  9. Turns deterministic lessons into previewable, project-local Skills on explicit request.
  10. Groups multiple TaskRuns by their originating Pi session and assigns chronological run positions.

Install for local development

npm install
npm run verify
pi -e /absolute/path/to/observable-coding-agent

To install it as a Pi package after building:

pi install /absolute/path/to/observable-coding-agent

Commands

  • /trace: show the active trace or the most recently completed task.
  • /review: show deterministic bottlenecks and Skill candidates for the last task.
  • /skill-draft [run-id]: preview generated Skills without writing files.
  • /skill-create <candidate-number> [run-id]: create one project-local Skill.
  • /observe: start the local dashboard and print its URL.
  • /observe stop: stop the dashboard server.

The completed task also appears as a durable [observe] card in the Pi transcript. Expanding the card shows the review.

Local dashboard

Run /observe, then open the returned http://127.0.0.1:<port> address. The dashboard opens on cross-run Insights: session and run counts, verification rate, average duration and tool usage, user-correction rate, recurring bottlenecks, task categories, and Skill-version evidence. The sidebar groups TaskRuns by Pi session and labels each run with its chronological position, for example Session 019fd654 · Run 3/4. Every aggregate can be traced back to the individual runs, timelines, files, validations, findings, and Skill candidates. It refreshes every five seconds. A run-level Skill lesson includes a Generate Skill action that writes only to that task project's .agents/skills/ directory.

Failed tool, validation, and assistant events display their redacted reason in the run Overview, inline in the Timeline, and beside failed file operations. When a validation tool returns structured blockers or changed-scope diagnostics, those concrete reasons take precedence over a truncated JSON preview. A single failure remains visible even when it does not meet the threshold for a workflow bottleneck finding.

The server binds only to 127.0.0.1, serves a fixed asset allowlist, and applies a restrictive Content Security Policy. GET is used for observation; the sole write endpoint accepts same-origin POST requests for project-local Skill creation. It stops when the Pi session shuts down or when /observe stop is called.

Storage

The full TaskRun is appended to observability/task-runs.jsonl next to the current Pi session file. For an in-memory session it falls back to <cwd>/.pi/observability/task-runs.jsonl.

Override the directory when needed:

PI_OBSERVABILITY_DIR=/absolute/private/path pi

Files are created with user-only permissions where supported. The extension stores paths, bounded/redacted tool inputs, and short redacted result previews. It does not intentionally store full file contents. Keys matching token, password, secret, cookie, authorization, and API-key patterns are removed; bearer tokens and common uppercase credential assignments are also redacted.

Existing v0.1 JSONL records are normalized in memory when read. New records use schemaVersion: 2; the original append-only data is never rewritten.

Observable Skill checkpoints

Loaded Skills are distinguished from explicitly or implicitly activated Skills. Only activated Skills participate in checkpoint evaluation and cross-run Skill metrics.

A Skill can optionally place observability.json beside its SKILL.md:

{
  "checkpoints": [
    {
      "id": "read-failing-source",
      "description": "Read the exact failing source before editing",
      "tool": "read"
    },
    {
      "id": "run-targeted-check",
      "commandPattern": "typecheck|tsc|test"
    },
    {
      "id": "touch-test-file",
      "filePattern": "(?:^|/)test/|\\.test\\."
    }
  ]
}

Checkpoint evidence is deterministic: successful tool calls, matching commands, or matching file activity. Invalid patterns simply produce no evidence; they never break the agent run.

Per-run reviews expose provisional “Skill lessons” so useful evidence is visible immediately. They are recommendations, not generated Skill files. The Insights view promotes a lesson to an evidence-backed Skill candidate only after the same bottleneck appears in at least three distinct runs, and lists the contributing run IDs.

Generate a Skill

Preview lessons from the latest run or a run ID prefix:

/skill-draft
/skill-draft da50b44b

Create candidate 1 from that run:

/skill-create 1 da50b44b

The explicit create command and Dashboard button generate:

.agents/skills/<generated-name>/
├── SKILL.md
├── observability.json
└── agents/openai.yaml

One or two evidence runs produce a provisional Skill; three distinct runs produce an evidence-backed Skill. Creation uses an atomic temporary directory and never overwrites an existing target. Generated Skills remain project-local; global installation and automatic updates are intentionally outside this version.

Run /reload in Pi after creation so the current session discovers the new project Skill. A newly started Pi session discovers it automatically after the project trust check.

Current deterministic findings

  • Excessive exploration before the first successful mutation.
  • Multiple failed tool calls.
  • Repeated identical errors.
  • Multiple validation → repair → validation loops, even if the final validation passes.
  • High-frequency repeated mutation of the same file.
  • Repair mutations made before the affected file is re-read after a failed validation (a heuristic blind-edit signal).
  • Mutations without recognized verification evidence.
  • Failed verification with no later passing evidence.
  • User corrections that indicate goal misunderstanding.
  • Activated Skill checkpoints without evidence.

These rules are deliberately explainable. A later analyzer can add model-assisted interpretation without replacing the deterministic evidence layer.

Current limitations

  • Workspace snapshots use file size and modification time and exclude large generated directories. Extremely large workspaces fall back to an unknown workspace mutation.
  • A verified outcome means the final recognized validation passed; it does not erase earlier rework findings, independently prove product correctness, or distinguish every possible pre-existing failure.
  • Implicit Skill activation is conservative and only inferred when the prompt names the Skill. Loaded Skills are not treated as used.
  • Skill creation requires an explicit terminal command or Dashboard click. Existing Skill directories are never overwritten or updated automatically.
  • Cross-run aggregation reads up to 200 records per workspace and keeps the latest 500 globally. SQLite indexing and long-term retention controls are not included yet.

npm release

Run the complete test, build, smoke, package-manifest, and sensitive-content checks:

npm run release:check

Preview exactly what npm would publish without changing the registry:

npm run release:dry-run

Bump the package version without creating a Git tag, then publish explicitly:

npm run release:version:patch
npm run release:publish

release:version:minor and release:version:major are also available. The publish command always targets the public npm registry and reruns the release gate before uploading.