@markdownee/markdownee
v0.8.6
Published
Standalone CLI and npm library for Markdownee. Built on Trafilatura Core and Crawlee.
Maintainers
Readme
Markdownee
Markdownee is a web scraper that crawls websites and saves their content as Markdown, HTML or plain text for LLMs, retrieval pipelines, and research datasets.
Choose which links to follow, set page and depth limits, and select how much page content to keep. Control tables, links, images, and user comments separately.
Boilerplate removal is powered by Trafilatura Core, our open-source fork of Trafilatura.
The Core in its name means it is reduced to one task: extracting main content by removing boilerplate. Other packages handle output conversion, including Markdown.
Trafilatura Core is ported from the original Python Trafilatura, with go-trafilatura as a DOM translation aid.
Crawlee handles crawling and uses Playwright for browser rendering.
Two language versions — TypeScript and Python: self-host with the npm CLI, npm library, or Python library, or use the hosted Apify Actor. Source code is on GitHub.
Save image files with the optional image downloading mode.
This package provides the TypeScript library and CLI.
Use fetch to return one page directly without opening a dataset or key-value
store. Use crawl to collect records and export to write stored results to an
output directory. purge removes the selected local storage. Inspect the first
result before expanding a collection; the
CLI guide explains each command's input
and output.
Install
Use Node.js 22.22.2+ on 22.x, 24.15.0+ on 24.x, or 26+. In a new project directory:
npm init -y
npm install @markdownee/markdowneeThe HTTP examples below need no browser. For adaptive or Chromium crawling, also install Chromium; install Firefox instead when selecting the Firefox crawler:
npx playwright install chromiumUsage: library
Save this as extract.mjs:
import { fetch } from '@markdownee/markdownee';
const { markdown } = await fetch(
'https://en.wikipedia.org/wiki/Web_scraping',
{ crawlerType: 'cheerio' },
);
console.log(markdown);node extract.mjsThe program prints extracted Markdown. fetch() follows no links, returns selected
formats, and throws on request failure. It rejects image saving; use a file-producing
CLI fetch or a crawl for downloaded images. The
library guide covers createCrawler,
options, output layouts, storage, and export APIs.
Usage: CLI
After the local installation above:
npx markdownee fetch https://en.wikipedia.org/wiki/Web_scraping \
--crawler-type cheerio --save markdown-file -o page.mdThis writes page.md; diagnostics go to stderr. Omit the file route to print
Markdown to stdout. The other commands are crawl for stored collection, export
for files and a manifest, and purge for permanent removal of selected local buckets.
See the CLI guide for configuration,
limits, image files, and the full flag reference.
Why Markdownee
Choose what to fetch, which content to retain, and where to save it. Page limits, link filters, and separate controls for images, tables, links, and user comments keep those choices explicit. Optional image downloading keeps local assets with extracted content.
Trafilatura Core ports original Python Trafilatura, with go-trafilatura as a DOM translation aid. Crawlee drives Playwright for browser rendering.
Use the npm CLI, npm library, native Python library, or hosted Apify Actor.
Support
Report problems through the issue tracker.
License
Licensed under Apache-2.0.
