npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

saturn-migrator

v0.1.1

Published

Saturn Migrator

Readme

Saturn Migrator (smigrate)

CLI for migrating data into Saturn. Two subsystems:

  1. Structured migrations (extract / migrate) — migration-definition .js files map CSV/JSON/API sources to ./output/{destination}.json.
  2. Document collection (discover / collect / investigate / status / ui) — extracts structured record data from folders of PDFs and images using OCR + a small LLM, with per-field confidence tracking, then investigates each record as a case: a per-record agent verifies values against the documents, learns what kinds of paperwork exist (and how they changed over time), links folders that belong to the same person, and reports contradictions and each record's story.

Both roads end in the same ./output/<Destination>.json shape, and both are driven from the UI (smigrate ui) or the CLI.

npm install -g saturn-migrator

Requires Node ≥ 18. For document collection you'll also want:

brew install tesseract      # local OCR (optional — falls back to Claude vision)
export ANTHROPIC_API_KEY=…  # LLM extraction (Claude Haiku 4.5 by default)
# or use another small model — pick provider + model from the model badge in the UI's menu bar:
export OPENAI_API_KEY=…     # gpt-5.4-mini / gpt-5.4-nano / gpt-5.6-luna / legacy minis
export GEMINI_API_KEY=…     # gemini-3.8-flash (default) / 3.5-flash-lite / 2.5-flash-lite / 3.1-pro-preview

Document collection workflow

Run everything from a project directory (state lives in ./.smigrate/: small hand-editable JSON for config/schema/scan, and one SQLite database, collection.sqlite, for everything that grows with the collection — records, files, results, and the OCR text with a full-text index). Projects created before the database existed are imported automatically the first time they are opened; the old JSON is parked in .smigrate/legacy/.

The UI opens on ▶ Start, which shows the same seven steps with the project's live status on each. In order:

| # | Step | What it does | CLI | |---|---|---|---| | 1 | Schema | Fields to capture, identity keys, the use-case description and the document guide (what each kind of paperwork is for, which fields it provides, how to recognise it from filenames) | smigrate ui | | 2 | Model (menu-bar badge) | Default provider/model and API keys; route labelling and planning to the cheapest tier, investigation to the strongest, with the reasoning lever per task | edit .smigrate/config.json | | 3 | Discovery | Point at the root folder, confirm how one record is stored, build the collection map | smigrate discover --root <dir> --confirm <n> | | 4 | Seed (optional) | Import a CSV/JSON you already have as a low-confidence starting layer | smigrate seed --file data.csv | | 5 | Extraction — the sweep | Over every pending record: plan the few files that cover the fields, OCR them once, one structured extraction call, escalate weak required fields | smigrate collect | | 6 | Investigate — the case review | On linked, incomplete or flagged records: an agent verifies values, cross-checks documents, searches the collection, reads related records, reports corrections, contradictions and the story | smigrate investigate --linked-only --incomplete-only | | 7 | Records | Review needs review and incomplete, correct in the drawer, export to ./output | smigrate status, UI → Export |

Extraction vs. investigation. Extraction is the bulk pass: cheap, fast, runs on everything, and produces a value with confidence, source file and evidence for every field. Investigation is the second look: it runs on the records that deserve it, costs several model calls each, and produces corrections with evidence, verified fields, unresolved inconsistencies, verdicts on related records (same person? switched agents?), the record's story and a needs review flag. Run extraction first over the whole collection, then investigate where it matters, then review; repeat as documents arrive or the guide improves.

If you change the root directory after the map was built, the next run checks whether the same files exist under the new root: if so the map is re-pointed automatically (the collection moved); if not the run refuses, and every screen shows a banner with Rebuild map for the new root — one click re-scans the new folder with the record pattern you already confirmed. Records are keyed by their root-relative path, so any that also exist under the new root keep their results; records only found under the old root are marked removed; new folders start pending. (CLI: smigrate discover --root <dir> --confirm <n>.)

How it works

  • Discovery walks the root directory, infers how records are stored (folder-per-record, file-per-record, nested categories), you confirm the pattern, and it materializes a collection map (.smigrate/map.json) — the single artifact progress is tracked against. Pick the root folder with the 📁 Browse… button next to either root-directory field — it opens a local folder browser with a Drives & volumes bar so you can jump straight to a mounted removable/external drive (macOS /Volumes, Windows drive letters, Linux /media & /mnt) instead of typing the path. You can also describe the organization in your own words (UI textarea, or smigrate discover --hint "each client has a folder named by file number…") — the model treats your description as authoritative when ranking candidates, and proposes a pattern directly if the heuristics missed it. The description is saved on the schema and reused. The scan streams live per-folder progress (which folder, running document counts) in the UI and the CLI, and the walk result is cached to .smigrate/scan.json — reopening the app shows the last scan's candidates instantly (with a rescan link) instead of re-walking a large tree. Re-running discovery merges: extracted records keep their results.
  • Planning — deciding what to open without opening anything: the filename is a hint about what a file is, never the truth, and on a large collection most files are irrelevant to most fields. So before any file is read the planner predicts each file's document type from three sources, in order: the label of a file already read earlier (content-based, cached by hash), the document guide's filename patterns (the most specific matching pattern wins, so *approval* beats app*), and filename stems the catalog has learned (a stem must have been seen planning.min_stem_evidence times). From "file → probable type" and "field → document types that provide it" (guide rows, plus catalog evidence above planning.min_provider_share) it picks the smallest set of files covering the wanted fields — a single primary form usually covers most of them — and skips files whose type is not known to hold anything wanted (a bank statement, say). Only fields with no candidate go to one cheap LLM planning call over filenames plus the same knowledge. If a field still has no candidate, at most planning.max_unknown_files_per_round files of unknown type are opened, smallest first, so something is learned about them — never the whole folder. After extraction, weak required fields trigger up to two escalation rounds where the same logic runs again with the new labels ("approval_date is weak → open the Approval Letter we skipped"). Every file gets a written reason (why opened, why skipped, what it was taken to be), shown in the record drawer, and what would the planner open? dry-runs the plan for any record without opening anything.
  • Document guide (Schema screen → Documents & use case): the place to tell the system what the paperwork is. A free-text use-case description ("applications filed by agents 2014–2023; the Application Form holds almost everything; the Approval Letter is issued on the approval date…") plus one row per document type: purpose, filename patterns, which fields it provides, whether its own date is a field (an approval letter's date is approval_date — once such a file is opened and labelled, the value follows from the label with no second look), whether it is the primary form, and the period it was in use. Suggest from filenames drafts the guide with the model from your description and a sample of the collection's filenames (no file is opened); Check patterns shows how many real files each row matches and which filenames nothing explains. The guide is also handed to the planner's model call and to the investigator.
  • OCR chain (pluggable, .smigrate/config.json → ocr_chain): pdftext (PDF text layer, free) → tesseract (scanned pages + images; PDFs rasterized in pure JS) → claude_vision (last resort). Results are cached by content hash — no file is ever OCR'd twice, even if renamed or shared between records.
  • Reading order (pdftext): a PDF's text layer is stored in content-stream order, which is not visual order. Filled forms and certificates typically draw the static template in one pass and the filled-in values in a later pass, so a value can be emitted hundreds of items away from its own label — a certificate reading on …………, ……… has applied would extract as on , has applied with 1ST / JANUARY / 1984 orphaned at the end of the page, and the model then attributes those stray values to the wrong fields. pdftext therefore reads each text item's coordinates and rebuilds true reading order (top-to-bottom, left-to-right, with dotted leader lines collapsed), yielding on 1ST JANUARY , 1984 has applied in the prescribed manner. This assumes a single text column — true for the forms, certificates and IDs this tool ingests; set ocr_pdftext_order: "stream" in .smigrate/config.json to revert to raw stream order for genuinely multi-column documents.
  • Parallelism (Extraction screen → Performance, or .smigrate/config.json): two axes, plus a safety ceiling. concurrency is how many records run at once; file_concurrency is how many of one record's files are OCR'd at once (a record with 30 scans used to be read strictly one file at a time, so a single large record serialized the run). Because records × files could spawn far more tesseract processes than the machine has cores, ocr_concurrency is a run-wide ceiling on CPU-bound OCR that holds however the other two are set. All three accept "auto" (the default) and size themselves from the detected core count — os.availableParallelism(), so container CPU limits are respected. The UI shows "N CPU cores detected · suggested records: M" and the resolved numbers in effect; the CLI takes --concurrency <n> and prints the plan at the start of a run. Tuning: keep OCR near your core count (it's CPU-bound); records can go higher since they mostly wait on the LLM, but too high will hit provider rate limits (429s are retried with backoff, so you'd see latency rather than failures).
  • OCR cache invalidation: cache entries are stamped with a cache_version. Because the cache is keyed by file content, it cannot tell that the same bytes are now extracted differently — so whenever extraction behaviour changes the version is bumped and stale entries are re-read automatically. You never have to remember --force-ocr after upgrading. To force a re-read anyway (e.g. you replaced a file in place, or want to re-run a different OCR chain): smigrate collect --force-ocr, the re-read documents checkbox on the Extraction screen, or re-read documents (ignore OCR cache) next to Enrich from documents in the record drawer. Note that plain Enrich from documents re-runs the model but reuses cached document text — tick the box when you want the files themselves re-read.
  • Extraction: compacted, provenance-marked text ([file: x.pdf, page 3]) goes to the LLM with a structured-output schema. Every field returns {value, confidence, source_file, page, evidence}; a stable cached system prompt keeps per-record token cost low.
  • Folder names are evidence: every extraction call carries a record context block — the record's folder path, the folder categories above it (agent, year…), and the list of its files with subfolders — and the provenance markers name each file by its record-relative path ([file: Dependents/Jane Doe/passport.pdf, page 1]), which source_file copies. The organisation note, the use-case description and the document guide are part of the extraction system prompt (stable per run, so they cache), so the model knows what the folder names mean — a passport in a dependent's subfolder is the dependent's, not the applicant's.
  • Confidence: LLM self-report is clamped by type-validation and OCR quality into an effective confidence, bucketed good / review / low / missing. The UI grid and smigrate status surface gaps; manual corrections in the UI set confidence to 1.0.
  • Prompt inspector (debugging): on the Records screen, open any record and click 🔍 View prompt to see the exact prompt last sent to the model for that record — the system prompt (including the field schema as the model sees it: name, type, description, required), the compacted document text with its [file: …, page N] provenance markers, and the structured-output JSON schema. Escalation rounds appear as separate tabs. This is the fastest way to diagnose values landing in the wrong field: usually two field descriptions read ambiguously to the model, so tighten the descriptions (or add a per-field source hint) and re-run. Captures are written on every extraction to .smigrate/tmp/prompts/<record>.json (git-ignored, since they embed raw document text).
  • Seed layer (layered sources): import a CSV/TSV or JSON array (UI → Seed screen, or smigrate seed) keyed by a reference number column. Rows start unbound and join records three ways: by the reference number extracted from the documents (authoritative — runs automatically after every extraction), by fuzzy name matching against folder names (unambiguous matches auto-bind, the rest queue for one-click review), or manually. Seeded values land at csv tier (default confidence 0.6 — amber in the grid): they fill gaps and corroborate, but documents always outrank them and manual corrections outrank everything. When a document disagrees with the seed, the seed value is kept as a visible conflict; when they agree, confidence gets a corroboration boost. A bind-only pass ("Find references in documents") extracts just the reference field per record so binding can happen before paying for full enrichment.
  • Export (UI → "Export to ./output" or POST /api/export) writes records as {type, id, …fields} to ./output/{destination}.json — the same shape smigrate migrate produces, so collected data flows into Saturn unchanged. includeUnboundSeeds optionally exports seed rows that never matched a folder (flagged _source: "seed_only").

Records are cases, not folders

Three layers turn the sweep into an investigation. All three are learned from the collection itself, not configured up front, because the paperwork was never uniform: forms were required for a few years and then dropped or revised, files were named by whoever scanned them, and a person can have folders under two different agents.

  • Document catalog (.smigrate/catalog.json, UI → Investigate → Document catalog, CLI smigrate investigate --catalog). Every file that gets read is labelled from its content — document type, its own date, issuer, subject, form version — with one small model call, cached by content hash so a file is labelled once for the life of the project (config classify_documents, default on). Because every extracted field carries source_file provenance, the catalog learns which document types actually supply which fields ("date_of_birth: Passport 62%, Application Form 30%") and, from document dates, which years each form was in use. This knowledge block is handed to the model on every planning and investigation call, so file selection reasons about "the passport" or "the 2019 application form" rather than trusting filenames. It also tallies filename stems per document type, so a never-opened file can be predicted from its name once the same naming has been seen often enough. The catalog is a set of GROUP BY queries over the database, so it stays cheap on very large collections. Files read before this feature existed can be labelled with Label files & rebuild.
  • Identity links (.smigrate/links.json, UI → Investigate → Related records, CLI smigrate investigate --links). Tick Identity on the schema fields that identify a person (passport number, national id, file reference — fields with such names are used automatically if none are ticked). After every run, records whose identity values match are linked; name + date of birth is always a secondary key and name alone is a weak lead. Links are grouped into cases with a hypothesis — "appears under 2 different top-level folders (Agent A / Agent B) — likely started with one agent and switched" — and shown with the folder categories. Give a verdict yourself (same person / same case / family member / different person) or let the investigator read both folders and decide. Verdicts are kept and respected on rebuild.
  • Investigation (UI → Investigate, or the 🕵 button in any record's drawer; CLI smigrate investigate). An agent loop per record: it sees the current values with their evidence and confidence, the record's files with their learned labels, the catalog knowledge, related records and any seed row, then decides one action at a time — open_files (OCR + extraction of exactly those files, shown alongside what they said and whether it agrees or disagrees with the current value), search_collection (full-text search across every OCR'd document, to follow a reference number or an unusual name into other folders), read_related_record, or finalize. Steps are capped (config investigation.max_steps, default 8) and a max-cost cap applies. The final report carries corrections (accepted only with verbatim evidence that passes the field's type check — its tier agent beats plain extraction, never a manual value), verified fields, inconsistencies (plus deterministic checks: identity fields that disagree with a record judged the same person, and document-vs-document conflicts left unreconciled), verdicts on related records, and the record's story. Records are marked needs review when a high-severity contradiction or a missing required field remains; the Records grid filters on it and the drawer shows the story, the contradictions and every step the agent took.

Every decision is a structured-output call (complete() on each LLM provider), so the investigator runs identically on Claude, OpenAI and Gemini.

Scale

The design target is collections of millions of files (terabytes of scans). What makes that work: discovery streams the tree and keeps only per-directory aggregates; records, files, results and per-field values live in SQLite (better-sqlite3, WAL mode), so updating one record is one row write and the Records grid is one indexed, paged query; OCR text is stored per page with an FTS5 index, so the investigator's collection-wide search is an index lookup instead of a corpus scan; catalog and link statistics are SQL aggregates. What does not scale by itself is compute: OCR is CPU-bound (budget roughly 2 s per page per core with tesseract), and SQLite is single-machine — for several OCR machines, shard by top-level folder into separate projects and merge the exports, or move the same schema to Postgres. Keep spend bounded with the planner (few files per record), --max-cost, and by investigating only linked and incomplete records.

Useful flags

smigrate collect --limit 5          # first 5 pending records
smigrate collect --record <id>      # one record
smigrate collect --force            # re-extract everything
smigrate collect --ocr-only         # pre-warm the OCR cache, no LLM calls
smigrate collect --max-cost 2.50    # stop when estimated spend hits the cap
smigrate collect --batch            # Message Batches API (50% cheaper, slower)
smigrate collect --batch-resume <batchId>
smigrate investigate --limit 10      # investigate the first 10 not-yet-investigated records
smigrate investigate --record <id>   # one record
smigrate investigate --linked-only   # only records that share identity keys with another
smigrate investigate --incomplete-only --max-cost 5
smigrate investigate --steps 12      # allow more decisions per record
smigrate investigate --links         # rebuild + print who appears in more than one folder
smigrate investigate --catalog       # label unlabelled files, print the document catalog
smigrate ui --port 4600 --no-open

Interrupting a run (Ctrl-C) is safe: in-flight records finish, and the next collect resumes from the map.

Swapping engines

Three LLM providers ship built in. One model serves every screen (discovery suggestions, planning, extraction, document labelling, guide drafting, investigation), so its configuration lives in the UI's menu bar: the provider · model badge at the top right shows the active model and whether its key is set, and opens the model settings (provider, model, API keys, Vertex AI) from any screen. The CLI reads the same .smigrate/config.json → llm/model:

| Provider | Env var | Models | |---|---|---| | claude (default) | ANTHROPIC_API_KEY | claude-haiku-4-5 (default), claude-sonnet-5, claude-sonnet-4-6, claude-opus-5, claude-opus-4-8, claude-opus-4-7, claude-opus-4-6, claude-fable-5-1, claude-fable-5 | | openai | OPENAI_API_KEY | gpt-5.4-mini (default), gpt-5.6-luna (cheapest current), gpt-5.6-terra, gpt-5.6-sol (flagship), gpt-5.5, gpt-5.4, gpt-5.4-nano, plus legacy gpt-5-mini / gpt-5-nano / gpt-4.1-mini / gpt-4o-mini | | google | GEMINI_API_KEY / GOOGLE_API_KEY | gemini-3.8-flash (default, newest Flash), gemini-3.7-flash, gemini-3.6-flash, gemini-3.5-flash, gemini-3.5-flash-lite, gemini-3.1-flash-lite, gemini-3.1-pro-preview, gemini-2.5-pro, gemini-2.5-flash, gemini-2.5-flash-lite (cheapest), plus legacy gemini-2.0-flash. Gemini 3.x Flash prices are promotional until 2026-12-31 and double after; Pro models double their rates above 200k input tokens, which the cost tracking applies. |

The model picker shows each model's price per million tokens and a one-line note; the lists and prices come from the provider modules (MODELS in providers/llm/*.js, served at /api/models), so adding a model is one line. Spend is computed from each response's actual usage at that model's rate.

The reasoning lever (model settings → Effort, per task in the routing table, or config.effort / config.models.<task>.effort). Every provider exposes reasoning depth differently and every model accepts a different subset of levels, so the tool keeps one ladder — off · none · minimal · low · medium · high · xhigh · max — and each model lists the rungs it accepts. Whatever you pick is mapped to the nearest rung that model supports at call time, so one setting works however a task is routed (xhigh on Gemini becomes high; off on a Claude model becomes low). The picker shows each model's own levels and default. What is sent:

| Provider | Parameter | Levels by model (default in bold) | |---|---|---| | Claude | output_config.effort | Sonnet 5, Opus 5 / 4.8 / 4.7, Fable 5 / 5.1: low, medium, high, xhigh, max · Sonnet 4.6, Opus 4.6: low, medium, high, max · Haiku 4.5: not accepted (never sent) | | OpenAI (Chat Completions) | reasoning_effort | GPT-5.6 Sol / Terra / Luna: none, low, medium, high, xhigh, max · GPT-5.5: none, low, medium, high, xhigh · GPT-5.4 / 5.4-mini / 5.4-nano: none, low, medium, high, xhigh · GPT-5 mini / nano: minimal, low, medium, high · GPT-4.1 mini, GPT-4o mini: not accepted | | Gemini (generateContent) | generationConfig.thinkingConfig.thinkingLevel, or thinkingBudget: 0 for off | 3.8 / 3.7 / 3.6 Flash: low, medium, high · 3.5 Flash: minimal, low, medium, high · 3.5 / 3.1 Flash-Lite: minimal, low, medium, high · 3.1 Pro, 2.5 Pro: low, medium, high (cannot be turned off) · 2.5 Flash: off, low, medium, high · 2.5 Flash-Lite: off, low, medium, high · 2.0 Flash: none |

Reasoning tokens bill as output on every provider, so the lever is a direct cost control: extraction and labelling are routine — none / off / minimal / low keep them cheap — and the investigator earns medium or high. Long-context surcharges are applied to spend tracking: OpenAI's 5.4+ flagships and the 5.6 family bill 2× input and 1.5× output above 272k input tokens; Gemini Pro models double above 200k.

Per-task routing (model settings → Per task, or config.models): the pipeline makes five kinds of call with very different shapes, and one model for all of them overpays somewhere. Each task — classify (labelling a file: a few hundred tokens, once per distinct file, the highest count by far), plan (filenames plus knowledge → which files to open), extract (the bulk of the tokens), investigate (a few reasoning-heavy calls on linked or incomplete records) and suggest (guide drafting, discovery pattern) — can name its own {llm, model, effort}; anything left blank uses the default. A typical split: labelling and planning on the cheapest tier (Haiku 4.5, gpt-5.6-luna, gemini-2.5-flash-lite), extraction on Haiku or Sonnet 5, investigation on Opus 5 at medium or high effort. Cross-provider mixes work, and splitting never hurts prompt caching (each task has its own stable system prompt; caches are per model anyway). If a routed provider has no key, that task falls back to the default with a warning. Run logs record the model used per task and spend per task (usage_by_task), shown in the CLI summary and in the routing table's Spent column, so you can see where a cheaper or stronger model would matter. Batch mode submits the extraction task's requests and therefore needs extraction routed to claude.

Refusals and fallbacks: Claude Opus 5 and Fable 5.1 run safety classifiers that can decline a request (stop_reason: "refusal"). On those models the tool opts into Anthropic's server-side fallback (fallbacks: "default"), so a declined extraction is re-run on another Claude model inside the same call and priced at that model's rate (the run log's usage records served_by). Batch mode does not support fallbacks, so a refusal there fails the record. Fable models additionally require 30-day data retention on the Anthropic account; a zero-data-retention org gets a 400.

Set the key for your chosen provider either as an environment variable (export ANTHROPIC_API_KEY=… before launching — an exported key always takes precedence), or enter it in the UI via the model badge in the menu bar → API keys. Keys entered in the UI are written to .smigrate/config.json, which is git-ignored so they aren't committed. The picker's status line shows whether the selected provider's key is set.

Gemini via Vertex AI

The google provider can also talk to Vertex AI instead of the AI Studio API — same models, but authed with your Google Cloud project and Application Default Credentials rather than an API key.

From the UI: open the model badge in the menu bar, pick the google provider and a Google Vertex AI section appears. Tick Use Vertex AI, enter your project (and optionally a location, default global), Save, then click Sign in with gcloud — this runs gcloud auth application-default login, which opens a Google consent page in your browser and stores credentials locally. The panel shows whether gcloud is installed and whether credentials are active. Settings persist in .smigrate/config.json.

From the environment (same convention as the official google-genai SDK; an exported value always takes precedence over the UI):

export GOOGLE_GENAI_USE_VERTEXAI=true
export GOOGLE_CLOUD_PROJECT=my-gcp-project
export GOOGLE_CLOUD_LOCATION=global        # optional; default "global", or e.g. us-central1
gcloud auth application-default login       # or set GOOGLE_APPLICATION_CREDENTIALS to a service-account key

Credentials resolve through standard ADC: gcloud auth application-default login, a service-account key file pointed to by GOOGLE_APPLICATION_CREDENTIALS, or the GCE/Cloud Run metadata server. No GEMINI_API_KEY is needed in this mode. Requires the google-auth-library package (installed as a dependency).

All providers share identical prompts and structured-output schemas (providers/llm/common.js), so switching engines never changes what the model is asked to do. Notes: --batch uses the Anthropic Message Batches API and therefore requires the claude provider; the claude_vision OCR fallback also always runs through Anthropic (drop it from ocr_chain if you don't set ANTHROPIC_API_KEY).

OCR providers live in providers/ocr/, LLM providers in providers/llm/ — each is one file implementing a small contract (available/supports/extractText or planFiles/extractFields). Add a file, register it in the provider index.js, and reference it from .smigrate/config.json.

Structured migrations (original subsystem)

For data that is already tabular — a CSV export, an Asana project — the original flow still applies, and it is now visible in the UI under ⚙ Migrations (list of definitions, their field mappings and transformer source, source/output file status, preview of rows, and Extract / Migrate buttons that run exactly what the CLI runs, with a live log).

Each migration definition is a .js file in the project directory:

module.exports = {
  enabled: true,
  source: "Members",           // reads ./source/Members.csv
  source_format: "csv",        // csv | json
  destination: "Member",       // writes ./output/Member.json, records get type "member"
  destination_format: "json",
  id_field: "account_number",  // id becomes "member-<account_number>"
  id_field_transformer: (v) => (v ? v.replace("/", "-") : ""),
  default_field_map_function: (col) => col, // destination for columns with none
  return_all_source_fields: false,          // pass unmapped columns through as-is
  extractor: "assana",                       // optional: "assana" | "csv" | (config, migration) => {...}
  extractor_config: { token: process.env.ASANA_TOKEN, project: "123", fields: "name,notes" },
  map_fields: [
    { source: "Customer Account No", destination: "account_number" },
    { source: "PaymentAmount", destination: "payment_amount",
      transformer: (row, col, moment) => Number(row[col].replace(",", "")) },
    { source: "Branch", destination: "branch", mapped_resource: "Branch" }, // → "branch-<value>"
  ],
};
smigrate extract   # runs each enabled definition's extractor → ./source/<source>.<format>
smigrate migrate   # maps ./source rows through map_fields/transformers → ./output/<Destination>.json

Extractors live in providers/ (assana.js pages an Asana project into ./source, csv.js copies a CSV from extractor_config.file). Transformers receive (row, columnName, moment) and return the destination value; mapped_resource turns a value into a <resource>-<value> reference. Output records are { type, id, ...fields } — the same shape the document-collection Export writes.