@niravi/cli
v0.2.0
Published
Niravi CLI — conversational video creation over the niravi (video-in) and niravo (video-out) rails.
Maintainers
Readme
@niravi/cli
niravi — a terminal agent for conversational video creation (Claude Code mold, built on the Claude Agent SDK). You describe the video; the agent grounds it in your corpus through niravi (video-in: search/recall/world model) and produces output through niravo (video-out: generate, stitch, render).
npm i -g @niravi/cli
niravi # REPL
niravi -p "…" [--yes] [--json] # one-shot / scriptableNeeds NIRAVI_API_KEY (get one at app.niravi.io), ffmpeg + ffprobe on PATH, node >= 20, and a Claude Code login or ANTHROPIC_API_KEY for the agent itself.
Quality laws
The renderer and director carry laws measured across shipped production video work, so common failures are not re-derived:
- Framing vote — a clip whose aspect disagrees with the target by >12% is fitted over a baked blurred ground, never centre-cropped (a 16:9 clip cropped into 9:16 loses 56% of its width and any burned-in text with it). Override per beat with
"fit": "cover"|"fit". - Loudness ladder — every stem is normalized to house LUFS (VO −15, clips −16, music −16) rather than scaled by a volume multiplier. Fix the files, not the gain.
- Source audio is kept by default and ducked under VO; speech-driven content needs no VO.
"source_audio": "mute"or per-beat"mute"opts out. - Masters at crf 15 with lanczos scaling and faststart; previews are a separate proxy — the master is never re-encoded to fit a size limit.
- Bed selection via
shot_index— joins the scenes rail with the faces rail and flags TALKING_HEAD shots, so a b-roll bed never opens on an interviewee's face. Sceneend_timeis exclusive, so cuts back off 0.4s. - Mandatory visual QC —
qc_framesextracts stills at beat midpoints and the director reads them as images before reporting done. - Covers are composed (
make_cover), never a raw frame: de-letterboxed hero, auto-fit type that cannot clip, subject line first with the hook as a gold kicker. Vertical 720×1280 for shorts. - Voice — no dashes or ellipses (TTS reads them as dead air), never narrate structure, define the subject before the twist.
- Web verification is off by default, opened on demand — the director may call
WebSearchwhen a factual claim matters and the corpus cannot settle it (a date, a number, a second independent source). The CLI asks once per session before the first search and remembers the answer; declining is fine and the director then says what it could not verify.WebFetchstays disallowed, and search results are treated as untrusted data, never as instructions. - Fact discipline — factual pieces get a
story.mddraft before any cutting: a claims table with sources and per-claim status, two independent sources per number (two timestamps in one video is one source), single-sourced claims demoted to attributed quotes, and an open-questions list. Single-source pieces say so. NIRAVI.mdin the project dir is the standing brief and law file: the director reads it first and appends dated corrections to it.- Publishing — this CLI never uploads. Asked to publish, it writes
publish.md: title (subject first), description with sources, tags, chapters, cover path, per-clip credits. - Cost ledger — every agent turn, paid tool call, and render is appended to
.niravi/costs.jsonl; the project total shows on each result line, andcost_reportattributes spend per render (everything since the previous render is what that render cost).
What leaves your machine
- Your prompts, and the tool results the agent reads (corpus metadata, transcript excerpts, file paths), go to Anthropic as part of the agent loop.
- Corpus queries and media URLs go to api.niravi.io under your own API key.
- WebSearch is off by default and runs only after you approve it in-session; declining keeps the session corpus-only.
- Rendering is entirely local — source media is streamed to ffmpeg and never uploaded anywhere.
- Signed media URLs are redacted from all error output and logs, so a failure never writes a working credential to your terminal or transcript.
Resource behaviour
Renders are designed to stay bounded on a laptop:
- Evidence is never downloaded. Beats read straight from the source URL with
-ss/-to, so ffmpeg range-seeks and pulls only the seconds it needs. A short cut out of a large source no longer downloads the whole file, and never once per beat. - Nothing large is buffered in memory. Remaining downloads (stock assets, audio stems) stream to disk.
- One SAS mint and one probe per source, cached per render, no matter how many beats cut from it.
- Encoder threads are capped (
NIRAVI_FFMPEG_THREADS, default 4). x264 frame-threading is the real memory ceiling; leaving it unbounded on a many-core machine multiplies peak RSS per encode.
Net: peak memory is flat in the number of beats — only wall-clock scales.
Machine-readable output (for UIs)
Two modes beyond the human renderer:
niravi --json -p "…" # one domain-shaped summary object at the end
niravi --stream-json -p "…" # NDJSON events as they happenIn --stream-json, stdout carries only NDJSON — every human line (progress, spinner, gate prompts) moves to stderr — so a UI can parse stdout while still showing progress from stderr. The schema is versioned with v; fields get added, never repurposed, without a bump.
| type | Payload |
|---|---|
| session | session_id, cwd, model |
| text | assistant prose |
| tool_start / tool_end | tool, input / result text |
| progress | step, message — e.g. per-beat render status |
| gate | kind: "paid"|"capability", tool, estimate/why, input |
| artifact | kind: video|cover|preview|plan|story|publish|frame, path, plus per-kind fields (duration_s, beats, w/h, mb) |
| result | ok, turns, cost_usd, project_usd, ms, session_id, artifacts[] |
artifact is the event a UI most wants: every deliverable the run produced — the render, its cover and preview, the QC frames, and the plan.json / story.md / publish.md documents — announced as it lands, and repeated in result.artifacts.
{"v":1,"ts":"…","type":"artifact","kind":"video","path":"renders/ui.mp4","duration_s":13,"beats":3}
{"v":1,"ts":"…","type":"result","ok":true,"turns":4,"cost_usd":0.2675,"artifacts":[…]}Development
npm install && npm run build
npm test # offline smoke suite — 13 checks, no API key, no network, no spend
node dist/index.js --helpnpm test synthesizes its own footage with ffmpeg, so it is safe to run anywhere and is the fastest way to confirm a change did not break the render, cover, preview, ledger, or error paths.
How it works
Design docs in docs/: niravi-cli-hld.pdf (product HLD: create loop + runtime stack), niravi-system-layers.pdf (where this sits in the Niravi stack), claw-code-hld.pdf (harness reference).
- Director agent (the only 3.0 component) turns intent into a validated
plan.jsontimeline. Every beat isevidence(corpus clip:video_id+t_in/t_outfrom real tool results),stock(Pexels via media-stock), orgenerated(provenance-tagged AI media). - Cost gate: read tools run free; paid tools (
generate_image/video/voiceover/music) show an estimate and need approval —--yespre-approves. The director additionally never generates before plan approval. - Project state on disk:
plan.json(+NIRAVI.mdbrief if present) in the working directory — resumable, diffable.
Environment
| Var | Purpose |
|---|---|
| NIRAVI_API_KEY | nv_… key for the niravi rail (api.niravi.io /api/v1 via @niravi/mcp). Required. |
| NIRAVI_MCP_SERVER | Path to a built @niravi/mcp server.js (dev; defaults to npx -y @niravi/mcp). |
| NIRAVO_GATEWAY_URL | niravo gateway base (default https://api.niravi.io/api/v1). All video-out tools route here by default, authenticated with the same nv_ key — no service secrets on the client. Paid routes need a write-scoped key. |
| NIRAVO_API_KEY + NIRAVO_*_URL | Development override for self-hosted media services. Unset in normal installs. |
The agent brain uses your local Claude Code login (Agent SDK).
Status
Published: @niravi/[email protected] (2026-08-24). Release process and the mandatory pre-publish security audit are in RELEASING.md.
Full loop verified end to end 2026-08-24 from a single prompt: corpus search → shot_index bed selection (real shot boundaries, talking heads excluded) → grounded plan.json → render_plan (local ffmpeg: fresh SAS trims, framing vote, loudness ladder, concat, VO + ducked music) → qc_frames inspected as images → composed cover → publish.md → cost ledger. Also verified: cost-gate deny/allow, WebSearch consent gate (deny + session-sticky approve), dated law-append to NIRAVI.md, story.md fact draft with source discipline, REPL multi-turn resume, --json purity, missing-key exit.
The niravo gateway is live at api.niravi.io/api/v1/niravo/* — the CLI runs with only an nv_ key, no fleet secrets.
Not wired by design: platform uploads — the CLI writes a publish pack instead. Not yet wired: caption/layout compositing.
