@amirhosseinnouri/momgen
v2.3.0
Published
Turn a meeting recording into a short Minutes of Meeting markdown file.
Downloads
154
Maintainers
Readme
momgen
Turn a meeting recording (audio or video) into a short, to-the-point Minutes of Meeting markdown file: ElevenLabs Scribe writes the transcript, and a chat model of your choosing writes the minutes.
Setup
Run it straight from the registry with whichever package manager you already have — there is nothing to clone and nothing to install globally:
npx @amirhosseinnouri/momgen ./video.mp4
pnpx @amirhosseinnouri/momgen ./video.mp4
bunx @amirhosseinnouri/momgen ./video.mp4All three are equivalent; they differ only in which registry client fetches the package.
The published bin is plain JavaScript built for Node >=20, so no Bun is required to
run it — Bun is only the development runtime for this repository.
One thing does have to be on PATH first: ffmpeg and ffprobe. Every local audio
step shells out to them, and neither is bundled.
brew install ffmpeg # or apt install ffmpeg, winget install Gyan.FFmpegCloning is for changing the code — the CLI itself needs no checkout. If that is what you are here for, see Development.
Then save your two keys once, with the CLI:
npx @amirhosseinnouri/momgen config set --elevenlabs_api_key sk-... --llm_api_key sk-...That writes ~/.momgen/config.json — owner-readable only — and every run picks it up no
matter which directory you start it from.
The same settings can come from the environment instead, which is what a .env in the
directory you run momgen from gives you:
LLM_API_KEY=... # minutes
ELEVENLABS_API_KEY=... # transcriptionEnvironment variables win over the saved file, so a one-off
LLM_MODEL=... momgen ./video.mp4 overrides a saved default without erasing it. A blank
variable does not: LLM_MODEL= means unset, not "forget what I saved".
Configuration
momgen config set --llm_model qwen3:30b # save one or more settings
momgen config list # everything saved, API keys masked
momgen config get llm_model # one value, unmasked, for scripts
momgen config unset llm_model # forget a setting
momgen config path # where the file lives
momgen config # the list of settingsEvery setting has a saved name and an environment variable, the same string in different
cases: --llm_api_key is LLM_API_KEY, --segment_seconds is SEGMENT_SECONDS. Values
are checked when you set them, so a price that is not a number or a segment length of 0
is refused there and then rather than at the start of the next run. MOMGEN_CONFIG_PATH
points momgen at a different file if you keep more than one profile.
Give the ElevenLabs key the user_read permission as well as speech-to-text. It is not
required to transcribe, but it is what lets the cost report show the credits left on the
account — and an empty account is otherwise invisible: ElevenLabs drops the connection
part-way through an upload rather than refusing it, and the run fails on a socket error
minutes in. With user_read, momgen stops before sending anything and says the balance is
spent.
The minutes model is reached through the AI SDK over a plain
OpenAI-compatible /chat/completions endpoint, so any provider speaking that dialect works
— point LLM_BASE_URL at it and name the model in LLM_MODEL:
# OpenCode Zen (the default)
LLM_BASE_URL=https://opencode.ai/zen/go/v1
LLM_MODEL=deepseek-v4-flash
# OpenRouter
LLM_BASE_URL=https://openrouter.ai/api/v1
LLM_MODEL=anthropic/claude-sonnet-5
# a local vLLM / Ollama
LLM_BASE_URL=http://localhost:11434/v1
LLM_MODEL=qwen3:30bAgent setup
The setup above is also packaged as a skill, momgen-setup, so a coding agent can do it
for you: check the prerequisites, walk you through both keys, work out where the .env
has to live for the way you invoke momgen, and finish with a verification run.
A second skill, momgen-add-template, covers Templates — ask an agent for
minutes shaped around your standup, retro or client call, or for a warmer or more formal
voice, and it writes the sections and the style, saves them, selects them, and shows you
the result.
Install them with the skills CLI, which reads the skills straight out of this repository — no clone, and nothing to install from the registry:
npx skills add amirhosseinNouri/momgenIt finds both skills under skills/, asks which of your agents to install into (Claude
Code, Cursor, Codex, …) and links them into each one's skills directory. Useful flags:
npx skills add amirhosseinNouri/momgen -g # user-level instead of this project
npx skills add amirhosseinNouri/momgen --all # every agent, no prompts
npx skills add amirhosseinNouri/momgen -l # just list what the repo offersThen ask the agent to "set up momgen" and it takes over from there.
If you would rather not add a tool: the skills are plain markdown with no tooling
assumptions beyond a shell, so point any agent at skills/momgen-setup/SKILL.md or
skills/momgen-add-template/SKILL.md directly, or link them yourself — Claude Code
auto-discovers skills under .claude/skills/:
mkdir -p .claude/skills
ln -s ../../skills/momgen-setup .claude/skills/momgen-setup
ln -s ../../skills/momgen-add-template .claude/skills/momgen-add-templateskills/momgen-setup/references/env-vars.md holds the full variable reference, which the
agent loads only when a specific variable comes up.
Both skills ship in the npm package as well, so npm i @amirhosseinnouri/momgen puts them
under node_modules/@amirhosseinnouri/momgen/skills/ — also valid symlink targets.
Usage
npx @amirhosseinnouri/momgen ./meeting.mp3 # or pnpx / bunx
npx @amirhosseinnouri/momgen ./standup.mp3 --template standupIt takes an audio or video file, or no argument at all, in which case it prompts for one.
--template picks the section list the minutes are built from — see Templates.
The .env is read from the directory you run in, not from wherever the package was
installed. That is the usual reason an npx run reports a missing key while a .env sits
one directory over. Export the variables from your shell profile if you run momgen from
arbitrary directories.
Each run writes to a fresh output/<source-name>-<timestamp>/ directory containing
transcript.md and mom.md. output/ is git-ignored.
Templates
Out of the box the minutes have a fixed shape — Attendees, Agenda, Decisions, Action Items, Discussion Summary — written tersely. A template replaces both:
momgen template list # every template, its sections and its style
momgen template add standup --file ./standup.md
momgen template show standup # one template
momgen template remove standup
momgen template path # where the files liveA template is a Markdown file with an optional --- block on top:
---
description: Quarterly board meeting
style: |
Formal and impersonal. Third person, no contractions.
Attribute every decision to the person who made it.
---
# Board Meeting
## Present
## Resolutions
## Action Items (markdown checklist)It controls exactly two things — the sections in the body, and the style the prose
inside them is written in. The --- block takes those two fields and description, which
is for you and never reaches the model; a misspelt field is an error rather than a setting
that silently never applies. A file with no block at all is a valid template written in
momgen's default style.
What a template cannot change is momgen's own instruction: that the model is an expert
meeting minutes writer, that it writes in the configured language, and that it uses
exactly the sections it was given. The last one is the reason the structure holds — the
minutes land in a file nobody reviews before reading, so a run that invents its own headings
is worse than one that leaves a section empty. Style is voice, not the job: formality,
person, sentence length, how much detail survives. The language is mom_language, not a
style.
Changing only the tone does not mean restating the sections — with no body, add reuses the
default template's:
momgen template add warm --style 'Warm and conversational. Full sentences, first person plural.'Templates are plain files at ~/.momgen/templates/<name>.md, beside config.json — so
MOMGEN_CONFIG_PATH moves a profile's templates along with its settings, and sharing one
is copying a file. momgen template add also reads the body from stdin, and --style wins
over a style: already in the file, which is what makes both scriptable:
cat <<'MD' | momgen template add retro --style 'Bullet fragments, not sentences.'
# Retro
## Went well
## Did not go well
## Action Items (markdown checklist)
MDPick one for a single run with --template <name>, or for every run with
momgen config set --mom_template <name>. MOM_TEMPLATE in the environment overrides the
saved setting, and --template overrides both. With none of them the built-in default is
used — it has no file, so a fresh install always has a working template, and saving one
named default shadows it.
The name is resolved before anything is uploaded, so --template standupp fails
immediately rather than after the transcription has been paid for, and the run prints what
the minutes will be in a Minutes panel above the cost confirmation:
◇ Minutes ─────────────────────────────────────────╮
│ │
│ Template retro │
│ Language English │
│ Style Bullet fragments, not sentences. │
│ │
│ # Retro │
│ ## Went well │
│ ## Did not go well │How it works
- Audio extraction — for video inputs, ffmpeg strips the audio to mp3. Cached by the
video's SHA-256, so re-running the same file skips the decode. The audio is re-timed
(
aresample) because call recordings often carry jittery timestamps that the mp3 muxer would otherwise drop speech over. If the extracted audio is materially shorter than the video's declared duration — the signature of an interrupted download, whose MP4 header still claims the full length — the run warns rather than quietly transcribing a fraction of the meeting. - Silence removal — transcription is billed per second of audio uploaded, and silence
transcribes to nothing, so gaps longer than a second are cut out before upload. A short
pause is left in place of each gap (
stop_silence), otherwise the words on either side run together and come back as one mangled token. If the filter eats almost the whole recording — a wrong threshold for an unusually quiet source — the original audio is used instead, since paying to transcribe a mangled file is worse than paying for the silence. - Cost estimate — the speech-only audio is split into segments, segments already in the cache are subtracted, and the run prints what is left to upload with its price before asking to continue. Nothing reaches the provider until that prompt is answered.
- Transcription — each uncached segment goes to Scribe and the full response is written to the cache before anything else runs, so an interrupted run resumes where it stopped and never pays for the same audio twice. Scribe is left to detect the language itself, which it does well on mixed Persian/English speech — forcing a code makes the loanwords worse.
- Minutes — the full transcript is streamed to the configured chat model, which writes the minutes under the selected template's sections and in its style (by default: Attendees / Agenda / Decisions / Action Items / Discussion Summary, written tersely — see Templates). Streaming is not cosmetic: a reasoning model can think for minutes before its first token, and a non-streamed request sits idle long enough to be killed in transit. The run aborts only after five minutes with no delta at all — reasoning counts as progress.
Configuration
| Env var | Default | Purpose |
|---|---|---|
| LLM_API_KEY | — | required |
| ELEVENLABS_API_KEY | — | required |
| LLM_BASE_URL | https://opencode.ai/zen/go/v1 | OpenAI-compatible endpoint, without /chat/completions |
| LLM_MODEL | deepseek-v4-flash | minutes model |
| ASR_LANGUAGE | unset (auto-detect) | language hint for Scribe |
| ASR_PRICE_PER_HOUR | 0.22 | rate used for the estimate |
| ELEVENLABS_STT_MODEL | scribe_v1 | Scribe model |
| SILENCE_THRESHOLD | -35dB | below this counts as silence |
| SILENCE_MIN_SECONDS | 1 | shorter gaps are left alone |
| SEGMENT_SECONDS | 900 | audio segment length |
| MOM_LANGUAGE | English | language of the minutes |
| MOM_TEMPLATE | default | name of the template — sections and style — see Templates |
Nothing reads process.env outside src/lib/config/. The table above is that module's
zod schema restated: a variable that is unset — or set to an empty string, which a
half-filled .env line produces — falls back to the default, and one that is set to
nonsense (SEGMENT_SECONDS=nine, SEGMENT_SECONDS=-5, a malformed LLM_BASE_URL) fails
at startup naming the variable, rather than surfacing as a $NaN cost estimate several
minutes later. LLM_API_KEY is checked at startup too, even though the minutes are the
last step — a run that dies there has already paid for the transcription.
Caches live in $TMPDIR/momgen-cache: extracted audio, speech-only audio, segments, and
Scribe responses (keyed by model and language hint, so changing either re-transcribes).
Everything ffmpeg produces is written to a temporary name and moved into place only once
it exits cleanly, so an interrupted run never leaves a short file behind for the next run
to read back as a complete one.
Development
Only needed if you are changing momgen itself — to use it, see Setup.
git clone https://github.com/amirhosseinNouri/momgen
cd momgen
bun install
bun start ./video.mp4 # runs src/index.ts directly, no buildBun is the development runtime: it runs the TypeScript sources as-is and provides the test
runner. What gets published is different — bun run build compiles src/ into a single
ESM file at dist/index.js with rollup (rollup.config.mjs), prefixed with
#!/usr/bin/env node, and that file is the package's bin. Runtime dependencies stay
external and are resolved from node_modules as normal; only the project's own modules are
bundled.
bun run build # dist/index.js + sourcemap
node dist/index.js ./video.mp4
bun test
bun run typecheckprepublishOnly runs typecheck, tests, and the build, so a publish cannot ship an artifact
that does not compile. dist/ is git-ignored and rebuilt on every publish.
The one thing Bun does that Node does not is load .env automatically, so src/env.ts
does it explicitly with dotenv. It is imported by the entry point only — never by the
library modules, which would put the repository's own .env underneath the tests.
Layout
Each module is a directory holding the implementation, its types, its tests, and a barrel that is the only thing other modules import from:
src/lib/video/
video.ts implementation
video.types.ts exported interfaces
video.test.ts unit tests
index.ts barrel — `import { extractAudio } from '../video'`| Module | Role |
|---|---|
| src/index.ts | entry point — nothing but run() |
| src/env.ts | loads .env for the Node build; imported by the entry point only |
| src/lib/config/ | every env var and every default, in one schema |
| src/lib/proc/ | the only place a child process is spawned — ffmpeg/ffprobe go through it |
| src/lib/cli/ | orchestration and the terminal UI: input → audio → cost prompt → transcript → minutes |
| src/lib/video/ | video detection, audio extraction, the truncated-download check |
| src/lib/audio/ | ffmpeg/ffprobe work: hashing, duration, silence removal, segmenting |
| src/lib/asr/ | transcription planning, cost estimate, per-segment cache, retries |
| src/lib/elevenlabs/ | Scribe HTTP client and the cost report it bills by |
| src/lib/llm/ | streamed chat via the AI SDK + the minutes prompt |
| src/lib/template/ | the sections and style the minutes are written to — read, write, resolve |
| src/lib/output/ | run directory naming and creation |
| output/ | generated runs, git-ignored |
| dist/ | the built bin, produced by bun run build, git-ignored |
| skills/momgen-setup/ | the setup skill an agent follows — see Agent setup |
| skills/momgen-add-template/ | the template skill an agent follows — see Templates |
| web/ | the landing page (Next.js + shadcn/ui) — see web/README.md |
The tests never call a provider — the ASR and chat clients are driven through a stubbed
fetch, and planTranscription takes an injected provider. The ffmpeg-backed tests build
their own two-second fixtures with lavfi and skip themselves when ffmpeg is absent.
