@sequenceholdings/pipeline-spec
v0.1.0
Published
Data Pipelines spec SDK — typed stage specs (ingestion/transformation/serving), directory loader, target resolution with deterministic hashing, and offline graph validation
Readme
@sequenceholdings/pipeline-spec
Schema and offline validation support for Sequence Managed Pipeline stage specifications.
This package is published so @sequenceholdings/studio-cli can load its
pipeline commands in standalone repositories. Its JavaScript module export is
an internal implementation detail, not a supported public SDK contract. Author
and validate pipeline repositories through seq-studio pipeline.
Install
Install the CLI and its pipeline-spec peer dependency in the repository that owns your pipeline:
pnpm add --save-dev \
@sequenceholdings/studio-cli \
@sequenceholdings/pipeline-specValidate a pipeline repository
Run validation from the repository root:
pnpm exec seq-studio pipeline validate .The validator discovers *.stage.yml files recursively. A stage specification
describes one deployable unit and its inputs, outputs, runtime, ownership, and
environment policy. For example:
schema_version: 1
stage: ingest-events
title: Event ingestion
description: Pulls event records into the bronze landing zone.
type: ingestion
owners:
- [email protected]
criticality: standard
environments:
- staging
runtime: databricks
source:
system: events_api
credential: events_api_key
feeds:
events:
description: Event records updated since the previous run.
snapshot_mode: incremental
natural_key:
- event_id
cursor:
column: updated_at
format: json
schedule:
cron: "0 5 * * *"
tz: UTC
entrypoint: src/ingest_events.pyCredentials are symbolic references such as events_api_key; never put secret
values, tokens, connection strings, or environment-specific URLs in a stage
specification.
Table and column descriptions
Keep bronze source documentation inline with each feed:
feeds:
events:
description: Event records updated since the previous run.
columns:
event_id:
description: Identifier assigned by the source system.
updated_at:
description: Time the source last updated this event.Each optional columns entry requires a description string and does not
accept types or constraints. The source determines bronze columns and types.
Description edits change specification provenance, not the data contract.
After deployment, seq-studio pipeline schema <assetId> -e <env> --json
returns descriptions alongside observed columns when supported by the server.
Stage YAML files allow up to 5 MiB; other specification files retain a 1 MiB
limit. Existing column_descriptions and documentation_ref declarations
remain supported; inline columns.<name>.description takes precedence.
Validation coverage
Offline validation checks:
- strict schemas for ingestion, transformation, and serving stages;
- references between stages, feeds, outputs, and triggers;
- asset column compatibility and contract changes;
- symbolic credential and webhook references;
- deterministic specification and contract hashes; and
- serving-stage compatibility with declared ORM bindings.
Validation reports all discoverable issues in one pass so repository authors can correct the complete graph before deployment.
Compatibility
- Node.js 20 or newer
- package versioning follows the release consumed by
seq-studio - the root JavaScript export may change without a standalone public SDK deprecation cycle
This package is distributed under the license declared in package.json.
