@popoverai/browser-automation
v0.15.4
Published
Narrated demo videos of your web app, recorded by your coding agent.
Downloads
1,221
Maintainers
Readme
agentic-demo
Narrated demo videos of your web app, recorded by your coding agent.
Ask Claude Code, Codex, Cursor, or any agent that can run shell commands for a walkthrough of a flow. It explores the app with agent-browser, writes a short script (the browser commands for each step and a sentence to say over it), and agentic-demo records the script as an mp4 with a voiceover. No screen recorder, no microphone, no editing.
The script stays behind as a file. Nothing improvises during the take, so the same script gives the same video, and when the app changes you re-record instead of redoing the demo.
Quick Start
Paste this into Claude Code, or any coding agent:
Use `agentic-demo` to create a demo of this feature. Start from `npx agentic-demo --help`.The agent does the rest. It will tell you if it needs anything from you, such as an OpenAI API key for the voiceover. You get a video file, plus the script it was made from, so you can ask for a fresh recording whenever the feature changes.
Before You Record
- The voiceover costs money. It is generated by a voice provider (OpenAI by default) and billed to your account with them.
- Your pages reach your agent's model. Your coding agent reads the pages it explores, and like anything else it reads, that content goes to the model behind it. agentic-demo itself sends only the narration text, to the voice provider (or the voice endpoint you point it at); the recording and the video stay local. A cloud browser's provider sees the pages too.
- The video shows whatever is on screen. Record against demo data, not real customers.
- Passwords stay out of the script. The guide tells your agent to sign in through agent-browser's encrypted credential store, so no password is written into the file.
AGENTS.md / CLAUDE.md
Add this to your project's instructions so your agent reaches for agentic-demo whenever you ask for a demo:
## Demo Videos
Use `agentic-demo` to record narrated demo videos of web flows. Run `npx agentic-demo guide` for the workflow before writing a steps file.Reference
Installation
npm install -g agentic-demo # Global
npm install -D agentic-demo # Pinned in a project
npx agentic-demo --help # Without installing
npm install -g agentic-demo@latest # UpdateRequires Node.js 22+. Installing downloads an ffmpeg binary (about 77 MB); pnpm 10 blocks that download until you run pnpm approve-builds, or point --ffmpeg / FFMPEG_BIN at another ffmpeg. agent-browser runs as npx agent-browser, which uses your installed copy if there is one. The same CLI is also published as @popoverai/browser-automation.
Versions are 0.x, so a minor version can break things. The changelog lists every change.
Commands
agentic-demo <steps.json> # Record to an mp4 (same as `record <steps.json>`)
agentic-demo validate <steps.json> # Check a steps file without opening a browser
agentic-demo guide # Print the workflow guide for agents
agentic-demo schema # Print the steps file's JSON Schema
agentic-demo example # Print a starter steps fileFor the step-by-step workflow, run agentic-demo guide.
Options
| Option | Description |
| -------------------------- | ---------------------------------------------------------------- |
| -o, --out <dir> | Output directory (default: a new temporary directory) |
| --tts <provider[:model]> | Narration provider and model (default: openai:gpt-4o-mini-tts) |
| --tts <url> | Get narration from a voice endpoint instead |
| --voice <voice> | Voice id for the narration provider |
| --silent | Silent audio track sized to the narration; no key needed |
| --keep | Keep each segment's audio, video and frames next to final.mp4 |
| --session <name> | agent-browser session to drive |
| --agent-browser "<cmd>" | How to run agent-browser (default: npx agent-browser) |
| --trailing-delay <ms> | Default trailingDelay for every step (default: 1000) |
| --max-fps <n> | Cap the capture frame rate (default: uncapped) |
| --ffmpeg <path> | ffmpeg binary (default: ffmpeg-static's, or FFMPEG_BIN) |
| --json | Print the summary in final.json on stdout as well |
Without --json, stdout is the path to final.mp4 and progress goes to stderr.
Output
A recording writes final.mp4 and, beside it, final.json: the video's length, and for each step where it begins in the video and how long it took. Every run writes it, narrated or --silent; --json prints the same summary.
{
"videoPath": "/home/me/demo/final.mp4",
"outputDir": "/home/me/demo",
"durationSeconds": 16.16, // Length of final.mp4
"segments": [
// One per step, in the steps file's order
{
"index": 1, // The step's position in the steps file, from 0
"instruction": "click #red", // The step's commands
"narrative": "Turn the page red.", // The step's narration
"startSeconds": 2.323, // Where the step begins in final.mp4
"captureSeconds": 1.01, // How long its commands took to run
"narrationSeconds": 1.6, // Length of its narration
"renderedSeconds": 1.88, // Its length in final.mp4
"frameCount": 2, // Frames captured; 1 means the step changed nothing on screen
},
],
}startSeconds is the time of the step's first frame in final.mp4, rounded up to the millisecond, so a player that seeks there shows that step rather than the end of the one before. Use it to open the video at a step. Don't add up renderedSeconds instead: joining the steps starts each one a few hundredths of a second later than that sum.
Steps File
{
"url": "https://app.example.com/login", // Optional: opened before recording; omit to record the current page
"speech": { "provider": "openai", "voice": "alloy" }, // Optional: narration settings
"steps": [
{
"narrate": "Sign in with the demo account.", // Said over this step
"commands": [
["auth", "login", "demo", "--no-navigate"], // "demo": credentials saved with `agent-browser auth save`
["wait", "--load", "networkidle"],
], // agent-browser commands, one argv array each
"trailingDelay": 1000, // Optional: ms to keep capturing after the last command
},
],
}Also optional: openArgs (extra arguments for open), speech.endpoint (a voice endpoint URL, in place of speech.provider), speech.model, speech.instructions, speech.speed, speech.language, speech.providerOptions, and a per-step speech block. agentic-demo schema prints the full schema.
A step can change the browser's size, with ["set", "viewport", "375", "667"], to show a flow on a phone after a desktop. The video keeps one size, the widest and tallest the recording reached, and fits the smaller view inside it, centred on black.
If agent-browser stops sending the browser's picture partway through, the recording stops with an error that names the step, instead of finishing a video that freezes there.
Narration
Five voice providers ship with the CLI. Each reads its key from its own environment variable:
| Provider | --tts | Key | Models |
| ----------------- | ----------------------------------------------------------- | ------------------------------------------- | -------------------------------------------------------- |
| OpenAI | openai[:model] (default gpt-4o-mini-tts, voice alloy) | OPENAI_API_KEY | gpt-4o-mini-tts, tts-1-hd |
| ElevenLabs | elevenlabs[:model] (default eleven_multilingual_v2) | ELEVENLABS_API_KEY | eleven_v3, eleven_flash_v2_5 |
| Hume | hume | HUME_API_KEY | (single model) |
| Deepgram | deepgram[:model] (default aura-2) | DEEPGRAM_API_KEY | aura, aura-2 |
| Vercel AI Gateway | gateway[:model] (default openai/tts-1-hd) | AI_GATEWAY_API_KEY or VERCEL_OIDC_TOKEN | openai/tts-1, fish-audio/s2-pro, spacexai/grok-tts |
Choose the provider and voice in the steps file's speech block, or for one run with --tts and --voice.
The AI Gateway reaches several providers' voices with one key. Its model ids name the provider too, as in --tts gateway:openai/tts-1; the AI Gateway model list has the current set. Without --voice, the model's own default voice speaks.
Voice endpoint
Instead of calling a voice provider itself, agentic-demo can ask an HTTP endpoint for each narration line. Use this when the machine that records should not hold a provider key: your app keeps the key, decides whether to narrate, and returns the audio. Name the endpoint with --tts https://… or "speech": { "endpoint": "https://…" }, and put the credential your endpoint expects in AGENTIC_DEMO_TTS_TOKEN.
Before it opens the browser, agentic-demo narrates the first step, so a refused token or a used-up quota stops the run before any recording. The rest of the lines are requested while the video is rendered, in parallel, one request per step. The first line's audio is reused, so a run makes one request per step in all.
Request
POST <endpoint URL>
Authorization: Bearer <AGENTIC_DEMO_TTS_TOKEN>
Content-Type: application/json
Accept: audio/*
{
"text": "Sign in with the demo account.",
"model": "gpt-4o-mini-tts",
"voice": "alloy",
"instructions": "Friendly product walkthrough; unhurried.",
"speed": 1.0,
"language": "en",
"outputFormat": "mp3",
"providerOptions": { "openai": { } }
}Authorizationis sent only whenAGENTIC_DEMO_TTS_TOKENis set.textis always present.outputFormatis"mp3"unless the steps file sets another. Every other field is present only when the steps file (or--voice) sets it, and is passed through without checking, so your endpoint decides which to honour and which models and voices to allow.modelandvoicecome from the steps file only when itsspeechblock names this same endpoint; a voice chosen for a provider is not sent.--voiceis always sent. A step's ownspeechblock overridesvoice,instructions,speedandlanguagefor that line.
Success: any 2xx status with the audio bytes as the body and an audio/* Content-Type (audio/mpeg, audio/wav, …). The format is read from the bytes, so any format ffmpeg can decode works. A 2xx answer without an audio/* Content-Type is treated as an error.
Refusal: any non-2xx status. Put the message for the person in the body, as JSON or as plain text:
HTTP/1.1 402 Payment Required
Content-Type: application/json
{ "error": "Your team has used this month's narration quota. Upgrade at https://example.com/billing." }agentic-demo stops, exits with status 1, and prints the error string exactly as sent, on its own line after the status. When the body is not JSON with a string error, it prints the whole body. It does not retry. A 401 or 403 when AGENTIC_DEMO_TTS_TOKEN is unset also says that no credential was sent.
[agentic-demo] the voice endpoint https://app.example.com/api/voice refused to narrate (HTTP 402):
Your team has used this month's narration quota. Upgrade at https://example.com/billing.License
Apache-2.0. Originally forked from @browserbasehq/mcp-server-browserbase.
