npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

rag-extract

v0.1.1

Published

Document extraction, chunking and retrieval for RAG

Downloads

3,035

Readme

rag-extract

Turn documents into retrievable text: extract, chunk, index, retrieve. Runs on Node and Bare.

Installation

npm i rag-extract

Usage

import { RagExtract, MemoryVectorIndex } from 'rag-extract'

const doc = await RagExtract.fromBytes({ bytes, fileName: 'report.pdf' })
const chunks = doc.getChunks()

const index = new MemoryVectorIndex()
await index.upsert(
  doc.id,
  chunks.map((chunk, i) => ({
    id: chunk.id,
    content: chunk.content,
    metadataJson: chunk.metadataJson,
    embedding: vectors[i]
  }))
)

const hits = await doc.search({ query: 'fuel pressure', limit: 5 }, { index, embed })

embed is yours: (texts: string[]) => Promise<number[][]>. This library never talks to a model, and vectors above come from embedding chunks.map((chunk) => chunk.content).

The same flow as plain functions, for a pipeline whose stages run in different places:

import {
  extractRagDocumentFromBytes,
  chunkRagDocuments,
  MemoryVectorIndex,
  searchRagChunks
} from 'rag-extract'

const document = await extractRagDocumentFromBytes({ bytes, fileName: 'report.pdf' })
const chunks = chunkRagDocuments([document], {})
// ...embed and upsert as above...
const hits = await searchRagChunks({ query: 'fuel pressure', limit: 5 }, { index, embed })

Each area is also its own entry point: rag-extract/extract, rag-extract/chunk, rag-extract/vector, rag-extract/retrieval, rag-extract/artifact. The root and rag-extract/extract load the PDF engine; the other four do not.

API

RagExtract

  • RagExtract.fromBytes(input, opts?) — extract a document; new RagExtract(document) wraps one you already have.
  • doc.id, doc.document — the document id its chunks carry, and the extracted document.
  • doc.getChunks(opts?) — the chunks for this document.
  • doc.search(query, { index, embed }) — retrieval scoped to this document.

rag-extract/extract — bytes to text

  • extractRagDocumentFromBytes(input, opts?) dispatches on the classified kind. opts.extractors injects handlers for formats the runtime cannot read alone.
  • classifyRagDocument, isSupportedRagDocument, RAG_ACCEPT_EXTENSIONS.
  • extractCsvText, extractHtmlText, extractRtfText, decodeTextBytes, extractPdfText, extractPdfHead — the individual decoders.

PDF, HTML, CSV/TSV, RTF and plain text work out of the box. Office documents (.docx, .xlsx, .pptx, .doc, .odt, …) are classified but not decoded — pass your own converter as opts.extractors.office or extraction throws RAG_UNSUPPORTED_FILE:

await extractRagDocumentFromBytes(input, {
  extractors: { office: async (bytes, fileName) => convertToMarkdown(bytes, fileName) }
})

The converter's optional third argument carries signal and the input's mimeType, so it can select a format when the file name has no extension. Legacy .xls and .ppt remain unsupported.

rag-extract/chunk — text to chunks

  • chunkRagDocuments(docs, opts?) — sliding token windows snapped to paragraph boundaries.
  • ragDocumentId(document) — the id every chunk of a document carries.
  • buildRagChunkMetadata, parseRagMetadata, hitMatchesFilters — the chunk metadata envelope.
  • tokenizeForRag, estimateRagTokens.

rag-extract/vector — the index contract

  • VectorIndex — the interface. Implement it over whatever store you have.
  • MemoryVectorIndex — an exhaustive in-memory cosine index. A correct reference implementation, not a production store.

rag-extract/retrieval — chunks to an answerable context

  • searchRagChunks(query, deps) — embed the query, scan the index once, shape the result.
  • retrieveRagKnowledge(query, deps) — the two-tier flow with named-document scoping and a similarity floor.
  • selectCoverage(hits, { budgetTokens }) — spread a budget across whole documents, sampling oversized ones along their length.
  • isWholeDocumentQuery(query) — a bare summarize ask has no subject to retrieve on, so the caller reads the document whole instead.
  • formatRagContext — render selected hits into a prompt block.
  • sourcesFromHits / webSourcesFromToolResult — hits to citation records. Both take an optional bounds bag (maxSources, excerptChars, maxUrlChars) so a host with a wire format can impose its own limits.
  • chooseRagNotice and the notice constants, createRagSourceTagFilter, citedByPhraseOverlap, resolveTargetDocument, RAG_MIN_SIMILARITY.

deps is { index, embed, debug? }; retrieval is silent unless you pass debug.

rag-extract/artifact — a portable index

  • encodeRagArtifact / decodeRagArtifact — the QVRAGIDX1 container (text, metadata and packed float32 embeddings).
  • documentIndexArtifactId, pipelineVersionFor, contentHash — deterministic, pipeline-version-sensitive artifact ids.

The QVRAGIDX1 byte layout and the id composition are frozen. Changing the chunker, the extractor or the hash invalidates every artifact produced by an earlier version, so those are major-version changes.

License

Apache-2.0