ndi-sdk
v0.23.0
Published
TypeScript SDK for NDI (Nace Document Intelligence)
Maintainers
Readme
NDI TypeScript SDK
TypeScript client for NDI (Nace Document Intelligence): parse, split, classify, extract and ground documents, and build searchable workspaces over a corpus.
npm install ndi-sdkRequires Node 20+ (or any runtime with native fetch, FormData, and Blob: browsers, Deno,
Bun, edge). Zero runtime dependencies. Dual ESM + CommonJS.
The Python SDK is a separate package (pip install ndi-sdk). This package covers the same
curated /v1 + legacy surface.
Quickstart
import { NdiClient } from "ndi-sdk";
const client = new NdiClient({ api_key: "ndi_sk_…" });
const upload = await client.documents.createUpload("report.pdf");
let job = await client.documents.parse(upload, { wait_seconds: 60 });
job = await client.jobs.wait(job.job_id);
console.log(job.result && "markdown" in job.result ? job.result.markdown : job.result);api_key defaults to $NDI_API_KEY and the host to $NDI_BASE_URL, so a configured
environment needs only new NdiClient(). There is one async client — fetch is async-only.
close() is a no-op for the default fetch; if you pass your own fetch (proxies, mTLS, or a
mock in tests), you own its lifetime.
The two APIs
NDI has two surfaces and this SDK covers both.
| | Platform /v1 | Legacy /api/v1 |
|---|---|---|
| Where | client.workspaces, client.files, client.ingestion, client.documents, client.tools, client.search, client.jobs, client.domains | client.legacy |
| State | Workspaces persist; files are ingested once and searched many times | Stateless, one-shot, nothing retained past the job |
| Use it for | Anything new | Maintaining an integration already written against it |
Everything slow is a job
Ingestion, parse, extract, search and workspace deletion all return a Job rather than a
result, because any of them can outlast a request. There are two ways to wait, and they
compose:
// Ask the server to hold the response open, up to its ceiling.
let job = await client.ingestion.ingest(workspace_id, { path_prefix: "reports/", wait_seconds: 120 });
// Poll from the client. Returns as soon as the job is terminal.
job = await client.jobs.wait(job.job_id, { timeout: 600 });wait_seconds saves a round trip for work that finishes quickly; jobs.wait covers the rest.
It raises JobFailedError if the job failed (pass raise_on_failure: false to get the failed
job back instead) and JobTimeoutError if your budget runs out — the job keeps running
server-side either way.
job.result is a discriminated union keyed on result_type. A parse is a ParseResult class
(with markdown / text getters); an unrecognised result_type arrives as UnknownResult
with its payload intact.
Job progress is also available as Server-Sent Events (client.jobs.events(job_id)), an
AsyncIterable. Treat it as a latency convenience: the stream can end early.
Workspace flow, end to end
A workspace is a durable corpus: upload files, ingest them once, then search and query them repeatedly.
import { NdiClient } from "ndi-sdk";
const client = new NdiClient();
const workspace = await client.workspaces.create({ name: "fy25-audit" });
await client.files.upload(workspace.workspace_id, "local/balance-sheet.xlsx", {
path: "reports/balance-sheet.xlsx",
labels: { engagement: "fy25" },
});
await client.files.uploadFromUrl(workspace.workspace_id, {
path: "reports/minutes.pdf",
url: presignedUrl,
file_name: "minutes.pdf",
});
let ingestion = await client.ingestion.ingest(workspace.workspace_id, { path_prefix: "reports/" });
ingestion = await client.jobs.wait(ingestion.job_id, { timeout: 1800 });
const search = await client.tools.hybridSearch(workspace.workspace_id, {
query: "total liabilities",
k: 10,
});
for (const hit of search.hits) {
console.log(hit.path, hit.locator, hit.snippet);
}
const answer = await client.jobs.wait(
(
await client.search.deep(workspace.workspace_id, {
query: "What were total liabilities at year end, and where is that stated?",
})
).job_id,
);client.search.automatic lets the service answer one query with a one-pass fact search or an agent that reads
documents, queries tables, and selects documents from the metadata catalog.
For a single file, client.uploadAndIngest(...) collapses the two calls — it uploads, then
ingests the returned file_id. Prefer the explicit form when batching several uploads into
one ingestion run.
Direct end-client uploads
Mint a short-lived upload grant and hand its token to the browser:
const grant = await client.files.createUploadGrant(workspace.workspace_id, {
path: "inbox/statement.pdf",
max_bytes: 10 * 1024 * 1024,
ttl_seconds: 3600,
});Ledger understanding measures the ingested journal package and stores a report
on the workspace. create / update accept a ledger_understanding policy
block (trigger, gl_parsing):
await client.workspaces.update(workspace.workspace_id, {
ledger_understanding: { trigger: "auto", gl_parsing: true },
});
let understanding = await client.workspaces.startLedgerUnderstanding(workspace.workspace_id);
understanding = await client.jobs.wait(understanding.job_id, { timeout: 3600 });
const report = await client.workspaces.getLedgerUnderstanding(workspace.workspace_id);
console.log(report.stale, report.blocking, report.parsing);Re-ingest after files change with stale_only: true. Deletion is irreversible and needs the
name back as confirmation:
const job = await client.workspaces.delete(workspace_id, { confirm_name: "fy25-audit" });
await client.jobs.wait(job.job_id);One-shot document operations
client.documents is stateless. Four kinds of source are accepted:
import { UrlSource } from "ndi-sdk";
UrlSource({ url: presignedUrl, file_name: "invoice.pdf" }); // NDI fetches it
{ type: "upload", upload_id: upload.upload_id }; // staged bytes
{ type: "workspace_file", workspace_id, file_id }; // a file already in a workspace
{ type: "parse_result", job_id: parseJob.job_id }; // reuse a parseFor local bytes, stage them once. The handle from createUpload is accepted directly:
const upload = await client.documents.createUpload("local/invoice.pdf");
const parse = await client.jobs.wait((await client.documents.parse(upload)).job_id);Pass parse_mode: "medium" to route to the best parser, or "high" to add page
verification and correction. The default is "low"; reused parse_result sources
support only that default.
Extract
const schema = {
type: "object",
properties: {
invoice_number: { type: "string" },
total: { type: "number" },
},
required: ["invoice_number", "total"],
};
const validation = await client.documents.validateExtractSchema({ json_schema: schema });
const job = await client.jobs.wait((await client.documents.extract(upload, { json_schema: schema })).job_id);Ground, split, classify
await client.documents.ground(upload, { targets: [{ id: "total", text: "1,200.50" }] });
await client.documents.split(upload, {
classes: [{ id: "invoice", label: "Invoice", description: "A supplier invoice" }],
});
await client.documents.classify(upload, {
classes: [{ id: "invoice", label: "Invoice", description: "A supplier invoice" }],
});classify rejects a parse-result source.
Workspace tools
The tools read a workspace's structure and content directly. They answer inline — no jobs.
await client.tools.folderMetadata(workspace_id, { directory: "reports/" });
await client.tools.fileMetadata(workspace_id, { path: "reports/model.xlsx" });
await client.tools.readFile(workspace_id, { path: "reports/minutes.pdf", pages: [3, 4] });
await client.tools.qaFile(workspace_id, { path: "reports/report.pdf", query: "What is the invoice total?" });
await client.tools.queryTables(workspace_id, {
paths: ["reports/q1.xlsx", "reports/q2.xlsx"],
query: "Compare quarterly revenue",
});
await client.tools.kgInfo(workspace_id);
await client.tools.kgSearch(workspace_id, { query: "the parent holding company" });
await client.tools.kgWalk(workspace_id, { start_node_ids: ["entity:acme"], hops: 2 });Pagination
Every listing is cursor-paginated. Take a page at a time, or let the SDK follow the cursor:
import { JobKind } from "ndi-sdk";
const page = await client.files.list(workspace_id, { path_prefix: "reports/" });
for await (const file of client.files.iterAll(workspace_id, { path_prefix: "reports/" })) {
console.log(file.path, file.ingestion_status);
}
for await (const job of client.jobs.iterAll({ workspace_id, kind: [JobKind.INGESTION] })) {
console.log(job.job_id, job.status);
}Errors
Every failure is an NdiError. HTTP failures carry the server's typed error code.
import { ConflictError, ErrorCode, NdiError, RateLimitError } from "ndi-sdk";
try {
await client.files.upload(workspace_id, "local/report.pdf", { path: "reports/report.pdf" });
} catch (exc) {
if (exc instanceof ConflictError && exc.code === ErrorCode.PATH_CONFLICT) {
await client.files.upload(workspace_id, "local/report.pdf", {
path: "reports/report.pdf",
on_conflict: "new_version",
});
} else if (exc instanceof RateLimitError) {
console.log(exc.retryable, exc.request_id);
} else if (exc instanceof NdiError) {
throw exc;
} else {
throw exc;
}
}Retries and idempotency
Transient failures (429, 5xx, dropped connections) are retried automatically with
full-jitter exponential backoff, honouring Retry-After. A 408 is not retried: on NDI it
means a legacy synchronous endpoint's wait window closed, so it raises SyncWaitTimeoutError.
Every job-creating call carries an Idempotency-Key. Pass your own idempotency_key to
extend that guarantee across process restarts. An empty string throws rather than being
replaced with a fresh key.
import { RetryPolicy } from "ndi-sdk";
const client = new NdiClient({ retry_policy: RetryPolicy({ max_attempts: 5, initial_backoff: 1.0 }) });
await client.ingestion.ingest(workspace_id, { idempotency_key: `nightly-ingest-${new Date().toISOString().slice(0, 10)}` });Legacy /api/v1
Kept whole under client.legacy. Each action has two forms. The plain form waits inline
(timeout_seconds); the *Async form starts the job and hands back its id.
const handle = await client.legacy.upload("local/invoice.pdf");
const result = await client.legacy.parse(handle, { timeout_seconds: 60 });
const accepted = await client.legacy.extractAsync(handle, { json_schema: schema });
const job = await client.legacy.getJob(accepted.job_id);A 202 is JobAccepted; any other 2xx is JobStatus. A replayed idempotency key can answer
an async start with the existing job. The same key is sent in the JSON body and the
Idempotency-Key header.
Forward compatibility
- Unknown fields on a response are kept, not rejected.
- Unknown enum members (a new
JobKind, a newErrorCode) arrive as their string value and still compare equal to it (JobStatus("archived") == "archived"). - Unknown job results arrive as
UnknownResultwith the payload intact.
Versioning
Semantic versioning, with the usual pre-1.0 caveat:
while the major version is 0, a minor bump may change the API surface. Anything that
breaks a caller is listed under a Breaking heading in CHANGELOG.md.
