npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@docture/llm

v0.2.0

Published

Model strategies for Docture extraction, classification, splitting, vision, and constrained layout assistance.

Readme

@docture/llm

Structured extraction and constrained layout assistance through any AI SDK model.

pnpm add @docture/llm ai

ai is a peer dependency, so there is exactly one copy of the model types in your tree and the model you construct is the model that runs.

import { LLM } from "@docture/llm";
import { openai } from "@ai-sdk/openai";

new LLM(openai("gpt-5"))             // a provider instance
new LLM(anthropic("claude-opus-5"))  // a different provider
new LLM("openai/gpt-5")              // an AI Gateway model string

LLM is the ModelBackend that extractor.loadLlm(...) takes. @docture/core does not depend on the AI SDK, because a deterministic pipeline should not install a model client. So a model reaches core as something that knows how to turn itself into strategies.

const extractor = new Extractor()
  .loadDocumentLoader(new DocumentLoaderPdfJs())
  .loadLlm(new LLM(openai("gpt-5"), { pricing }), { vision: true, classify: true });

Configure it once and every strategy it produces inherits the cache, the pricing table and the instructions.

Your schema reaches generateObject in its original form. Zod stays Zod, so refinements and descriptions get to the provider instead of surviving a round trip through JSON Schema.

What is here

| | | |---|---| | LLM | a model, and everything it can be turned into, which is what loadLlm takes | | llm(…) / LlmStrategy | extract from the document's text | | vision(…) / VisionStrategy | extract from page images, which is what reads a scan | | llmClassifier(…) / LlmClassifier | decide the document kind from a cheap text excerpt | | llmLayout(…) / LlmLayoutAnalyzer | classify and order existing conversion elements without returning text | | TextSplitter | cut a bundle by reading the page text | | ImageSplitter | cut a bundle by looking at the pages, so it needs a rasterizer | | createLlmExtractor(…) | the terse wiring for the common case |

const extractor = createLlmExtractor({
  loaders: [new DocumentLoaderPdfJs()],
  model: openai("gpt-5"),
  before: [myDeterministicParser],   // tried first, since a parser that works costs nothing
});

Long documents

completion says how to handle a document too large for one request:

| | | |---|---| | "forbidden" (default) | one request, whole document | | "paginate" | one request per page, then merge | | "concatenate" | pack pages into requests under maxCharsPerRequest, then merge |

Multi-request runs send a schema with required stripped and validate only the merged result. Without that, a page that legitimately has no invoice number makes the model's honest partial answer a hard failure. See src/relax.ts. Merging concatenates arrays and keeps the first present scalar, so a later chunk cannot overwrite a good value with a null.

Cost and caching

pricing is USD per million tokens, keyed by model id. Without it, costUsd is honestly undefined rather than 0, which would read as "this was free".

cache (memoryCache() or fileCache(dir)) keys on model id, prompts and the JSON Schema, so a schema change can never be served an old shape.