markdrilldown
v0.2.0
Published
Converts a markdown document into a self-contained, drill-down HTML page with LLM-generated summaries at every heading level.
Maintainers
Readme
markdrilldown
Converts a single Markdown document into a self-contained, drill-down HTML page: an LLM-generated summary at every heading level down to H3, expanding in place until the reader reaches the original source text.
Usage
export GOOGLE_GENERATIVE_AI_API_KEY=...
npx markdrilldown README.mdThat writes README.html next to the input — a single file with no external assets, ready to
open in a browser or drop on a static host.
markdrilldown <input.md> [-o out.html] [--model provider:model-id] [--summary-threshold n]
[--skip-export] [--rebuild-cache]-o,--output— output path. Defaults to the input's basename with a.htmlextension, in the same directory.--model— the model to summarize with, asprovider:model-id. Defaults to the model of whichever provider you have credentials for; see Providers and API keys.--summary-threshold— how short a section's own text has to be, in characters, for it to be shown as written rather than summarized. Defaults to 200;0summarizes every section.--skip-export— leave the document's markdown source, and the Export Markdown button that downloads it, out of the page.--rebuild-cache— ignore every cached summary and generate the whole document again. Use it after switching to a better model, or whenever you want the summaries rewritten.
Settings you would otherwise pass every time, and your API key, can live in a file instead —
see Settings in ~/.markdrilldown.
A section with less text of its own than the threshold gets no summary: its own text is shown in place of one, with its subsections already visible beneath it. A section that is nothing but subheadings shows those subsections directly. Only the section's own text counts, not its subsections', and the document as a whole is always summarized.
Providers and API keys
| Provider | --model prefix | Environment variables | Default model |
| --------- | ---------------- | ---------------------------------------------------------- | --------------------------------------------- |
| OpenAI | openai: | OPENAI_API_KEY | gpt-5-nano |
| Groq | groq: | GROQ_API_KEY | openai/gpt-oss-20b |
| Mistral | mistral: | MISTRAL_API_KEY | mistral-small-latest |
| Cohere | cohere: | COHERE_API_KEY | command-r-08-2024 |
| Google | google: | GOOGLE_GENERATIVE_AI_API_KEY | gemini-3.1-flash-lite |
| Anthropic | anthropic: | ANTHROPIC_API_KEY | claude-haiku-4-5 |
| Bedrock | bedrock: | AWS_REGION, AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY | us.anthropic.claude-haiku-4-5-20251001-v1:0 |
| xAI | xai: | XAI_API_KEY | grok-4.3 |
Credentials are read from the environment, or from ~/.markdrilldown if you keep your key there
(see Settings in ~/.markdrilldown). A build fails immediately
— before any network call — if the provider is unrecognized, or if anything it needs is
missing; the error lists every missing variable at once, not one per attempt.
Bedrock runs Anthropic's models through an AWS account you already have, with no direct Anthropic key. It reads its credentials from those three variables, like every other provider reads its key, rather than from an AWS profile, an SSO session or an instance role — a build that cannot authenticate has to fail before it writes anything, and looking those up means asking a file or a token endpoint. If you use one, export it first:
eval "$(aws configure export-credentials --profile my-profile --format env)"
export AWS_REGION=us-east-1AWS_SESSION_TOKEN is picked up when it is set, which temporary credentials need. The default
model is a US cross-region inference profile; from elsewhere, name your own — --model
bedrock:eu.anthropic.claude-haiku-4-5-20251001-v1:0. The colon inside a Bedrock model id is
fine: --model splits on the first colon only.
Configure any one provider and markdrilldown has a provider: with no --model, it uses the
first provider in the table above that has everything it needs, at that provider's default
model. A provider with only some of its variables set is skipped. That order is cheapest default
model first, by published input price per million tokens — summarizing sends a lot of text and
gets a little back — and it is fixed, so the same variables pick the same provider on every run
and on every machine. Bedrock and Anthropic reach the same models at the same price, so
Anthropic wins that tie: AWS credentials are often in your environment for something else
entirely. Configure several providers and the highest one in the table wins; name a --model
and it wins over all of them. Configure none and the build stops before it generates anything,
listing every provider and the variables it wants.
The provider wiring is tested, but the default model ids come from each vendor's published
documentation — a published id can be withdrawn without the documentation saying so. Five have
been run against a live API: OpenAI's, Anthropic's, Google's, Groq's and Cohere's, OpenAI's being
the one you get by default. The other three — Mistral, xAI and Bedrock — are wired but unproven.
If a build fails because its provider's default model no longer exists, pass
--model provider:model-id naming one that does.
Adding configuration can change which provider writes your summaries, silently: a provider ranking above the one you have been using becomes the new default. Cached summaries are reused whichever model wrote them, so the change costs nothing — but the next summary generated is written by a different model than the last.
Settings in ~/.markdrilldown
Write the settings you would otherwise pass every time — and your API key — into
~/.markdrilldown, and every build afterwards uses them with no flag and no environment
variable:
model: anthropic:claude-haiku-4-5
api-key: sk-ant-...
summary-threshold: 300
skip-export: trueThose four are the whole list: model, api-key, summary-threshold and skip-export, each
named for the flag it stands in for. Anything else in the file is an error, as is a file that
does not parse — a typo you cannot see is worse than a build that stops and tells you. Having
no file at all is perfectly normal.
A flag beats the file, and the file beats the environment. A model here means markdrilldown
does not go looking for a provider at all, however many other keys you have set, and the
api-key is used ahead of that provider's own environment variable. The key is never written
into your environment, and never appears in the page, in anything the tool prints, or in any
error it raises.
The model and the key go together: give both or neither. They are one way of reaching one
provider, and half of it answers nothing — a key with no model would sit unused while
markdrilldown picked a provider by price, and a model with no key still sends you to the
environment. The consequence is worth knowing: --model naming a different provider does not
get this key. That provider reads its own environment variable, and the build stops immediately
if it is not set. Bedrock cannot be configured here at all, since three variables do not fit in
one key; export them as the Providers section describes.
The file holds a secret, so keep it to yourself — chmod 600 ~/.markdrilldown. A build warns
when a file holding a key is readable by anyone else, and then carries on: the file is yours to
manage, and a permission your system chose is not a reason to refuse to build your document.
Caching
Summaries are cached in a sibling file, <input.md>.summary-cache.json, keyed by the content
hash of each section. Rebuilding a document you have only partly edited re-summarizes only the
changed sections, and a fully cached rebuild needs no API key at all.
The model is no part of the key, so switching models costs nothing: the summaries you already
have describe the same text, and every one of them is reused. The trade is that a better model
does not improve a document you have already built — pass --rebuild-cache to generate every
summary again at the model you are now using. Those results are cached in turn, so the build
after it is free again.
Summaries are written to the cache file as they are generated, not only when a build succeeds. A build that fails partway — a provider rate limit, an outage, an interrupted run — keeps everything it generated before the failure, so running it again asks the provider only for the summaries that are still missing. A document that trips a per-minute rate limit finishes after enough attempts rather than failing in the same place every time. The failed build still writes no HTML file.
Nothing is ever removed from a cache file. Editing the document leaves the entries for the text you replaced behind, so a cache file grows with every edit you make, and that growth is doing nothing for you.
It is safe to delete a .summary-cache.json file at any time. The next build simply regenerates
what it needs, at the cost of the API calls. Delete one to reclaim the space, or when the
document has changed enough that little of the cache still applies.
Commit the cache file alongside the document if you want reproducible builds in CI without paying for summaries twice.
Requirements
Node.js 22 or newer. Several of the Vercel AI SDK provider packages this depends on declare
node: >=22 themselves, so an older runtime will not install cleanly.
