rag-extract
v0.1.1
Published
Document extraction, chunking and retrieval for RAG
Downloads
3,035
Readme
rag-extract
Turn documents into retrievable text: extract, chunk, index, retrieve. Runs on Node and Bare.
Installation
npm i rag-extractUsage
import { RagExtract, MemoryVectorIndex } from 'rag-extract'
const doc = await RagExtract.fromBytes({ bytes, fileName: 'report.pdf' })
const chunks = doc.getChunks()
const index = new MemoryVectorIndex()
await index.upsert(
doc.id,
chunks.map((chunk, i) => ({
id: chunk.id,
content: chunk.content,
metadataJson: chunk.metadataJson,
embedding: vectors[i]
}))
)
const hits = await doc.search({ query: 'fuel pressure', limit: 5 }, { index, embed })embed is yours: (texts: string[]) => Promise<number[][]>. This library never talks to
a model, and vectors above come from embedding chunks.map((chunk) => chunk.content).
The same flow as plain functions, for a pipeline whose stages run in different places:
import {
extractRagDocumentFromBytes,
chunkRagDocuments,
MemoryVectorIndex,
searchRagChunks
} from 'rag-extract'
const document = await extractRagDocumentFromBytes({ bytes, fileName: 'report.pdf' })
const chunks = chunkRagDocuments([document], {})
// ...embed and upsert as above...
const hits = await searchRagChunks({ query: 'fuel pressure', limit: 5 }, { index, embed })Each area is also its own entry point: rag-extract/extract, rag-extract/chunk,
rag-extract/vector, rag-extract/retrieval, rag-extract/artifact. The root and
rag-extract/extract load the PDF engine; the other four do not.
API
RagExtract
RagExtract.fromBytes(input, opts?)— extract a document;new RagExtract(document)wraps one you already have.doc.id,doc.document— the document id its chunks carry, and the extracted document.doc.getChunks(opts?)— the chunks for this document.doc.search(query, { index, embed })— retrieval scoped to this document.
rag-extract/extract — bytes to text
extractRagDocumentFromBytes(input, opts?)dispatches on the classified kind.opts.extractorsinjects handlers for formats the runtime cannot read alone.classifyRagDocument,isSupportedRagDocument,RAG_ACCEPT_EXTENSIONS.extractCsvText,extractHtmlText,extractRtfText,decodeTextBytes,extractPdfText,extractPdfHead— the individual decoders.
PDF, HTML, CSV/TSV, RTF and plain text work out of the box. Office documents (.docx,
.xlsx, .pptx, .doc, .odt, …) are classified but not decoded — pass your own converter as
opts.extractors.office or extraction throws RAG_UNSUPPORTED_FILE:
await extractRagDocumentFromBytes(input, {
extractors: { office: async (bytes, fileName) => convertToMarkdown(bytes, fileName) }
})The converter's optional third argument carries signal and the input's mimeType, so it
can select a format when the file name has no extension. Legacy .xls and .ppt remain
unsupported.
rag-extract/chunk — text to chunks
chunkRagDocuments(docs, opts?)— sliding token windows snapped to paragraph boundaries.ragDocumentId(document)— the id every chunk of a document carries.buildRagChunkMetadata,parseRagMetadata,hitMatchesFilters— the chunk metadata envelope.tokenizeForRag,estimateRagTokens.
rag-extract/vector — the index contract
VectorIndex— the interface. Implement it over whatever store you have.MemoryVectorIndex— an exhaustive in-memory cosine index. A correct reference implementation, not a production store.
rag-extract/retrieval — chunks to an answerable context
searchRagChunks(query, deps)— embed the query, scan the index once, shape the result.retrieveRagKnowledge(query, deps)— the two-tier flow with named-document scoping and a similarity floor.selectCoverage(hits, { budgetTokens })— spread a budget across whole documents, sampling oversized ones along their length.isWholeDocumentQuery(query)— a bare summarize ask has no subject to retrieve on, so the caller reads the document whole instead.formatRagContext— render selected hits into a prompt block.sourcesFromHits/webSourcesFromToolResult— hits to citation records. Both take an optional bounds bag (maxSources,excerptChars,maxUrlChars) so a host with a wire format can impose its own limits.chooseRagNoticeand the notice constants,createRagSourceTagFilter,citedByPhraseOverlap,resolveTargetDocument,RAG_MIN_SIMILARITY.
deps is { index, embed, debug? }; retrieval is silent unless you pass debug.
rag-extract/artifact — a portable index
encodeRagArtifact/decodeRagArtifact— theQVRAGIDX1container (text, metadata and packed float32 embeddings).documentIndexArtifactId,pipelineVersionFor,contentHash— deterministic, pipeline-version-sensitive artifact ids.
The QVRAGIDX1 byte layout and the id composition are frozen. Changing the chunker, the
extractor or the hash invalidates every artifact produced by an earlier version, so those are
major-version changes.
License
Apache-2.0
