npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@scanmate/extract

v0.24.6

Published

PDF pages as rasters with their metadata: text layer with positions, scanned-or-born-digital, real resolution, an original paired with its scan - and where its fields are, found from the labels it prints.

Readme

scanmate extract

@scanmate/extract

PDF pages as rasters with their metadata: what kind of page it is, what resolution it really holds, and the text layer with every run's place, face and size — and an original paired with its returned scan.

the original's text layer, every run boxed where it sits

Each box is one run of the text layer, with the text, face and size it carries: exact, and free. Made from the IRS Form W-9 (a work of the United States government, in the public domain), filled in as a generator would.

Install

npm install @scanmate/extract
import { extractPages, extractPair, inspectDocument } from '@scanmate/extract'

// What is in this file, without rendering anything:
const inspection = await inspectDocument('fw9-returned.pdf')
inspection.pages[0].metadata?.kind          // 'vector' | 'scanned' | 'scanned-with-text-layer' | 'empty'
inspection.pages[0].metadata?.effectiveDpi  // what the scan really holds

// One document - the boxes in the picture above:
const pages = await extractPages('fw9-issued.pdf', { dpi: 150 })
pages[0].image.raster                        // decoded RGBA
pages[0].metadata.textItems                  // every run, in points from the top-left

// Or a pair, ready for @scanmate/align:
const { pages: pairs, unpaired } = await extractPair({ original: 'fw9-issued.pdf', scanned: 'fw9-returned.pdf' })

What it decides

  • Whether a page is a scan, from two signals: how much of the page the image operators actually cover (80% is the line), and whether there is any text. A generated form's letterhead covers 1.2% of the page; every page of three real scans of it covered 100%. A scan with an OCR text layer over it is its own kind, because its text must never be trusted.
  • What resolution to render at. 'match' — the default for a pair — renders both sides at the scan's own resolution. Measured on real scans at 93, 120 and 144 dpi, that beat every fixed choice from 150 to 300 on alignment confidence, on overlap, on false "added ink", and on time. Rendering the original finer than the scan adds no information; it just makes the two disagree at every stroke edge.
  • What a scan holds when its page is not its paper. A photo stored at one pixel per point says 72 dpi of itself. Under 'match' the scan is measured against the original's page size, so a 3024 × 4032 picture of an A4 sheet is read as the ~345 dpi it is, and its own pixels are never resampled.

The text layer

Each run carries its text, its box in points from the page's top-left as displayed (after /Rotate), its baseline, its font size, pdf.js's name for the face, its angle, and whether it ends a line. The box comes from the font's ascent and descent rather than a guess, which is what lets @scanmate/ocr claim read words by position and file glyph templates by face and size.

Only the original's text layer is meant to be trusted. A returned document's is reported and never used: it can be stale, or planted, and it is not what the person signing the paper saw.

Where the fields are

A generated document's fields move with its content: one more line in an address pushes the signature block down, sometimes onto the next page. What does not move is a field's place beside its label. locateFields finds each label in the text layer and places the fields from it:

import { locateFields } from '@scanmate/extract'

const { regions, anchors, problems } = await locateFields('fw9.pdf', [
  {
    anchor: 'Signature of U.S. person',
    fields: {
      signature: { dx: 44, dy: -3.8, width: 262, height: 22 },
      date:      { dx: 328, dy: -3.8, width: 171, height: 22 },
    },
  },
])
regions    // [{ page: 1, id: 'signature', x: 120, y: 577, width: 262, height: 22 }, { page: 1, id: 'date', ... }]
problems   // [] - or why a field could not be placed, or should not be trusted

the signer fields, resolved from the anchor the original prints

The boxes were not measured by hand: they are offsets from a label found in the IRS Form W-9's own text layer (a work of the United States government, in the public domain).

  • A label is found as it is printed. The W-9 prints "Signature of" and "U.S. person" as two runs on two lines; the anchor is the whole label, and wraps with it. Runs are joined along a line, and onto the next line under it.
  • Whole words only, compared after the suite's normalisation with the punctuation at their edges set aside: 'Signature' finds Signature:, and 'Date' never finds Update.
  • The whole document is searched, and each region carries its page. A label printed twice is refused as anchor-ambiguous rather than guessed at: name the occurrence (counted in page order) or the page.
  • from measures the offsets from any corner of the label. A field to the right of its label is best measured from 'top-right', so it does not move when the label's wording does.
  • The regions are checked: off the page, overlapping another on its page, zero-sized, or an id used twice all come back in problems.
  • Turned pages and turned text work. Everything is in points from the top-left of the page as displayed - the frame the pixel comparison measures in and markPages draws in - so the regions go straight to either.

It reads the text layer and renders nothing: about 200 ms on the W-9. resolveFields does the same from text already in hand, and locateAnchor only finds a label. A label inside a longer run has its edges placed in proportion to its characters, and says so with estimated: true.

Pairing

Scans lose pages and gain cover sheets, so extractPair reports what it could pair and what it could not, rather than throwing or truncating silently:

const { pages, unpaired, pageCount } = await extractPair({ original, scanned }, { pairing: 'index' })
unpaired.original   // [7]  - no scanned page for these
unpaired.scanned    // [1]  - a cover sheet, say

pairing is 'index', 'page-number', or an explicit list of [originalPage, scannedPage].

Options

| option | default | | |---|---|---| | dpi | 'match' for pairs, 'native' for one document | Or a number. | | fallbackDpi | 200 | For a page with no native resolution. | | minDpi / maxDpi | 72 / 400 | Bounds on a resolution read from a file. | | pages | all | Numbers, or a range string like '1-3,5'. | | output | 'png' | 'jpeg', or 'none' to keep rasters only. | | quality | 92 | For lossy output. | | background | white | PDF pages are transparent where nothing is drawn. | | includeText | true | Read the text layer. | | pairing | 'index' | How scanned pages match original ones. | | onProgress | — | Called before and after each page. |

extractPairStream yields pages as they are rendered, for documents too large to hold at once.

Rendering

Pages are rendered with pdfjs-dist onto @napi-rs/canvas. The standard-14 fonts ship inside pdfjs-dist, so nothing needs system fonts or fontconfig in a container. Both are prebuilt, with no system package to install.

pdf.js detaches the buffer it is given. This package passes a copy; any caller doing its own getDocument should too, or it will find its own bytes empty afterwards.

How it decides

documentation/algorithms.md has the algorithms in full: what each step measures, the decision flows, every constant with the measurement behind it, and what the package deliberately does not do.