npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

ease-design

v0.6.0

Published

Multi-runtime design-cli — describe what you want, get pro-taste UI.

Readme

DESIGN:OS is a multi-runtime design CLI. You drive it through the agent CLI you already use (Claude Code, Codex CLI, or Antigravity) with plain-language /ui:* commands. Describe the interface; DESIGN:OS routes the work to the right execution arm, loads the surface-specific craft and design-system context, and keeps delivery claims honest.

| Surface | Start with | What you get | |---|---|---| | Web marketing | /ui:generate <intent> | A qualified route to production HTML, bounded by typed delivery contracts. | | Figma | /ui:to-figma <intent> | Idiomatic canvas authoring through the opt-in Figma hand. | | Native macOS | /ui:native-macos <intent> | A first-class, SwiftUI-first workflow — available now with provisional assurance. | | Native iOS | /ui:native-ios <intent> | Compact-first SwiftUI craft for iPhone — available with provisional assurance. | | Native iPadOS | /ui:native-ipados <intent> | Resizable SwiftUI craft for iPad workspaces and mixed input — available with provisional assurance. |

Your host agent writes the implementation. The deterministic ui kernel handles routing, contract validation, and fail-closed claim boundaries; it does not call a model or generate Swift. No API keys, no design tokens to hand-edit, no taste vocabulary to learn.

Apple-native arms are available, not overclaimed. macOS, iOS, and iPadOS have separate routes, artifacts, evidence, and assurance. They share SwiftUI fundamentals without flattening platform behavior. PROVISIONAL means an arm can route and produce work today while live accessibility, hardware/device evidence, and owner acceptance remain separate qualification gates.

Native mobile proof, retained

The content-addressed proof board retains the current evidence boundary. TocChien is the iOS proof: its three image-first screens — champion catalogue, champion detail, and game dictionary — carry 16 controller-replayed simulator tests, paired light/dark captures, independent visual review, and owner acceptance bound to exact source and capture hashes. iPadOS retains 12 simulator tests, for 28 simulator tests across the board, but its visual review and owner acceptance remain pending.

This proof remains PROVISIONAL. iOS Tier 2 provenance is unqualified: the final paths have mixed Terra/Sol authorship without an immutable intermediate checkpoint. It does not authorize physical-device, live assistive-technology, iPad visual, release-qualification, assurance-upgrade, or qualified-delivery claims.


Generated by DESIGN:OS

Public sites shipped end to end with this toolchain. Every repo documents how to reproduce it with DESIGN:OS — OPAH ONE walks it command by command. The grid grows as new demos ship.

In the controlled three-way study, orchestration beat raw prompting and prompt enhancement in all three categories — a 1.77-point mean lift in independent blind review — and the follow-up repeatability study scored nine fresh orchestrated runs at 8.96 / 10 with 0.25 population standard deviation.

Interaction boundary benchmark

Three treatments test different surfaces with the GSAP motion-direction layer active:

The architecture treatment has since shipped as its own public repo — design-os-rill-architecture — and leads the showcase grid above. These are runnable treatments, not static mockups: clone the repo and open them from showcase/018-improving-proof-benchmark/runs/, or watch the recordings — DESIGN:OS, Nouri, Rill Architecture — and the rendered critique. The benchmark proves differentiated delivery across product marketing, native-mobile direction (a browser prototype, deliberately not counted as native implementation evidence), and premium service; it does not claim world-class superiority until matched controls receive blind scores.

The prompts behind the evidence

The user prompts were intentionally short. DESIGN:OS did not receive a polished creative brief:

Architecture: Build a premium site-specific residential architecture landing page.
Nutrition: Build an AI nutrition experience that helps someone understand a meal.
Planning: Build a planning SaaS landing page for connected decision context.

Before implementation, DESIGN:OS compiled each sentence into the generation packet the host model actually received. The product-specific audience, outcome, and action changed per case; the orchestration contract below stayed fixed:

directions: 3
selectBy: topic fit, evidence strength, execution risk, convergence risk
regions: [hero, proof, context, process-or-connection, conclusion]
eachRegionMustDeclare:
  - purpose and narrative role
  - distinct layout family
  - composition anchor and hierarchy event
  - visual type and rationale
  - responsive transformation
  - craft investment
imagery:
  - plan section-specific image prompts before implementation
  - use purposeful evidence where CSS would become a placeholder
  - keep labels, controls, data, and product state as live HTML
hero:
  - fit message, action, and primary demonstration in the initial desktop viewport
  - show hierarchy, state, and consequence
depth: preserve topic specificity and design investment through the conclusion
composition: test content-led and golden-ratio candidates; release the ratio when content fails
preflight:
  - no placeholder primary visual
  - no generated fake product screenshot
  - no repeated section shell without rationale
  - no quality decline toward the footer

Nothing is hidden behind a marketing paraphrase. Read the exact historical packets for Architecture, Nutrition, and Planning.

The current interaction-boundary benchmark deliberately changed two surfaces. Its raw requests were:

DESIGN:OS: Build a promotional page for DESIGN:OS and push interaction and animation craft.
Nutrition: Build a native mobile nutrition interface, not a responsive website.
Architecture: Build a premium site-specific residential architecture service page.

The resulting benchmark contract keeps the web and native claims separate: the Nouri browser prototype communicates native-mobile direction but does not count as native implementation evidence.

Why this is our best-performing method

A longer prompt is not the advantage. The advantage is converting vague intent into decisions the builder and curator can verify:

  1. Product truth before styling. Audience, outcome, action, and prohibited claims constrain what the page must communicate.
  2. Divergence before convergence. Three structural directions are compared before one visual language becomes expensive to change.
  3. Every section has a job. Region contracts prevent the familiar strong-hero, weak-footer quality collapse.
  4. Visual evidence is planned. Image prompts, aspect ratios, crop behavior, and narrative jobs are decided before imagery is generated or sourced.
  5. Product truth stays editable. Controls, labels, data, and states remain live HTML rather than being flattened into a convincing but unusable generated screenshot.
  6. Qualification requires evidence. Desktop/mobile renders and deterministic gates can reject work the model would otherwise describe as finished.

That makes orchestration the best-performing method in the workflows tested here, not a claim that DESIGN:OS is universally best. The same study exposed binary-rule misses and two mobile overflow failures. Those failures are public because the next product improvement is a deterministic repair gate, not a stronger adjective.

The nine repeatability runs live in-repo under showcase/world-class-benchmark/evidence/repeatability-study/ (clone to open the HTML). Full evidence: controlled comparison, repeatability result, and blind curator report.


The agent learns in the work

DESIGN:OS now separates four claims that are often blurred together:

  1. ALIVE — the project has a learning loop.
  2. LEARNING — evidence became a bounded lesson.
  3. APPLIED — later work cites and uses that lesson.
  4. IMPROVING — a preregistered suite of at least three holdouts across at least two categories clears the declared mean-quality, aggregate-repair, recurrence, and safety thresholds — the deterministic engine computes the suite verdict itself.

The proof is causal rather than volumetric. A thousand memory events do not prove learning. An APPLIED claim needs the original evidence, a later retrieval/application receipt, the affected decision, the artifact, and its outcome. Escaped defects need a gate, a failing negative fixture, and a later run.

The first dogfood study used three real project histories:

| Project | Supported level | What it proves | |---|---|---| | ease-design | APPLIED | User feedback changed orchestration and a later Nutrition Planning delivery | | client-portal (confidential) | LEARNING | Persistent project identity and harvested project lessons | | platform-design-system | APPLIED | Escaped defects became executable gates with negative fixtures and later detections |

The controlled Nutrition Planning run compared a memory-disabled control at 89/100 against a learning-enabled treatment at 94/100. Static comparison images are intentionally omitted from this README.

The treatment eliminated prohibited text/Unicode interface glyphs 27 → 0, added motivated scroll reveals, preserved reduced-motion handling, and won the independent blind review by five points. After the first blind findings were reapplied, it improved 89 → 94 without the user repeating the feedback.

The deterministic verdict is still APPLIED, not IMPROVING. The study did not clear the predeclared +10 comparison delta, used one treatment repair round, and has not completed contradiction or cross-project isolation runs. That boundary is intentional: the final design crossed the world-class visual threshold, but one successful case is not proof of an everyday trend — the engine now enforces a suite-level gate, so no single comparison can graduate on its own.

Run the evidence report:

design-os evolution \
  --dir . \
  --proof showcase/017-living-agent-proof/evidence/ease-design-proof.json

Read the controlled result, forensic project comparison, and machine-readable proof.


Built by the studio behind DESIGN:OS

Five live products, five different worlds — an AI dev-tool, a farm's direct-to-Hanoi fruit brand, a plugin marketplace, a villa-care service, a deal CRM for real-estate brokers. Same studio, same design discipline this toolchain distills. Real scrolling recordings of the live sites — every demo links through.

Under the hood the same mechanism scales from one screen to a whole product: 26 personas compile the same 27-component, paired-token starter kit into any system (ui ds init), and every qualified surface must provide evidence for the declared machine gates before it ships. The toolchain is also dogfooded on a production internal developer platform — a 129-component Figma library scanned, hygiene-audited, contrast-proven, and VR-baselined end to end.


Six daily verbs

These are the six common moves. The full adapter exposes additional workflows (plus the internal critique gate) when the task needs more.

| Verb | What it does | | --- | --- | | /ui:generate <intent> | Weak intent compiled into a typed brief and generation contract; one candidate is rendered, repaired, and delivered by qualification status. | | /ui:learn | Brownfield onboarding — compile the DS from your project's own evidence (code, a URL, or Figma) instead of a persona default. | | /ui:iterate · /ui:refine | Tweak in plain words; surgical line-diffs, re-scored; the DS hash-seal stays intact. | | /ui:from-url <url> | Extract a live site's design system into a portable folder (spec + tokens + audit). | | /ui:to-figma <intent> | Author idiomatic Figma on the canvas — auto-layout, real instances, token-bound variables. | | /ui:why <question> | Ask why a past design decision was made — answers with provenance from the project's design memory. |

| Command | What it does | |---|---| | /ui:generate <intent> | Start a fresh design from a plain-language description. Token-bound variants across diverse personas. | | /ui:native-macos <intent> | Build a SwiftUI-first native macOS surface through an available, provisional arm; it never claims qualified platform delivery. | | /ui:native-ios <intent> | Build a compact-first SwiftUI iPhone app with native navigation, touch, keyboard-safe layout, and provisional evidence. | | /ui:native-ipados <intent> | Build a resizable SwiftUI iPad workspace with split navigation, scene state, mixed input, and provisional evidence. | | /ui:iterate <change> | Tweak the current design in plain words; applied as a surgical line-diff, re-scored by the gate. | | /ui:refine | Run the full critique→refine polish loop on the current design. | | /ui:redesign <intent> | Reimagine an existing page in a different persona/direction. | | /ui:from-url <url> | Extract a live site's design system into a self-contained ./<slug>/ folder (spec + tokens + audit). | | /ui:from-ref <path> | Generate from a reference (image/markup), matching its look on your design system. | | /ui:figma | Reproduce a Figma source 1:1 as HTML (keeps source colors intentionally). | | /ui:to-figma <intent> | Author idiomatic Figma on the canvas from intent — needs the figma-agent hand. | | /ui:extract | Pull a design system out of existing HTML. | | /ui:slides <intent> | Generate a token-bound slide deck. | | /ui:learn | Compile the DS from the project's own evidence (code, URL, or Figma). | | /ui:design <brief> | The AI-designer flow — scope-aware facet planning + curator scoring on a full brief. | | /ui:diagram <intent> | Author an accessible offline diagram as inspectable SVG across 19 grammars; product-flow views preserve source IDs and disclose every fidelity trade-off. | | /ui:chart <intent> | Author a standalone chart of quantities across 9 grammars — zero baselines, declared truncation, no fabricated data. | | /ui:audit <target> | Run the deterministic audit families against a produced design. | | /ui:evidence | Intake user evidence (interviews, tickets, analytics) into the anti-fabrication ledger. | | /ui:why <question> | Trace picks, edits, verdicts, and token changes from the design memory, with provenance. | | (internal) /ui:critique | The gate — runs inside every HTML-emitting flow. |

Native diagrams and charts without a DSL

/ui:diagram covers 19 grammars — architecture, sequence, product-flow, swimlane, data-flow, process, ER, state, flowchart, tree, org-chart, layers, nested, loop, and the platform family. /ui:chart is a sibling capability covering 9 — bar, line, scatter, radar, gantt, timeline, quadrant, venn, pyramid.

Each invocation produces one self-contained HTML file with hand-authored inline SVG. No Mermaid, no PlantUML, no charting library, no headless browser, no network request at view time. Colours resolve to your project's design tokens and fall back to a documented neutral palette when there is no design system yet.

See all 28 → design:os example gallery

What the gate actually proves

ui diagram lint and ui chart lint are deterministic, and deliberately narrow — they check the facts a static reader can prove, never whether the picture is any good:

| Check | Catches | | --- | --- | | hardcoded-svg-color | A colour in an SVG presentation attribute. ds-usage-lint reads CSS declarations only, so fill="#eb6c36" would otherwise bypass the design system while showing green. | | diagonal-line | A connector running off-axis in a grammar whose layout reads by alignment. Grammar-gated — radial and hierarchical shapes are exempt by construction, not by exception. | | svg-labelledby · no-script · no-external-ref | An artifact that is unreadable to a screen reader, or that reaches outside itself. | | zero-baseline-required · baseline-declared | A bar or pyramid whose length encoding starts anywhere but zero, and any chart that lets a reader assume zero without saying so. | | series-label · no-dual-axis | Series separated by colour alone, and two scales overlaid on one frame. |

A clean lint means the contract holds. Whether the diagram is right is the taste rubric's job, and it runs afterwards at the same ≥ 7/10 gate every other surface meets.

Routing across 28 grammars stays decidable

Nineteen diagram grammars include deliberate near-neighbours — a swimlane, a data-flow, and a process diagram are the same grid at three levels of detail. Three mechanisms keep selection from collapsing into a coin flip: trigger tokens are pairwise disjoint across both capabilities, diagram-craft.md states precedence once per collision family, and a golden corpus of briefs asserts each resolves to exactly one grammar. The third is the only one that tests behaviour rather than shape.

Diagrams can also be redrawn from an existing source: the drawio and Mermaid extractors emit a node/edge digest that a grammar is authored from, never converted into. Imported geometry never reaches the artifact, and the redraw carries a ledger of what it simplified.


Quick start

npm install -g ease-design     # installs the `ui` kernel (zero runtime deps)
ui doctor                      # verify the install is healthy

Wire it into the project you want to design for:

cd ~/code/your-app
ui init --runtime claude       # or: --runtime codex | --runtime antigravity

Then open your agent CLI in that project and type:

/ui:generate a pricing page for a developer-tools SaaS — 3 tiers, dark theme

That's the whole loop: describe → activate → compile → qualify. Unsupported surfaces stop before compilation instead of silently falling back to HTML. Have an existing app? Run /ui:learn first so the DS is compiled from your product's own evidence.

ease-design is DESIGN:OS's npm name — the package ships the ui kernel. v0.3.0 (published from CI with sigstore provenance) carries all 45 commands, including the tenant scroll engine and the asset gates behind the showcase grid.

Full studio (clone)

git clone https://github.com/jangtrinh/design-os.git && cd design-os && ./setup.sh

One idempotent script builds and links the whole studio, not just the kernel: semantic recall, the living-agent evolution/heartbeat/harvest loop, rendered accessibility audits, and the tenant + gflow scroll-cinema toolchain — repo-only hands the npm kernel doesn't ship. The Figma plugin (and its 1:1 mirror) installs from its own repo. Needs Node ≥ 22 (the recall hand) and uv for the design-os umbrella; ./setup.sh --check verifies prerequisites without changing anything.

One hand is opt-in, never silent: gflow (Google Flow — the scroll-cinema asset generator). An interactive run explains what it is, that it needs a paid AI Ultra/Pro subscription, and that it automates a real browser session on your Google account — then asks. A non-interactive run skips it. --with-gflow / --no-gflow answer ahead of time. Skipping costs nothing but the ability to generate new footage, and design-os doctor reports the gap up front instead of failing mid-generation.

| Get it | How | What you get | Update | |---|---|---|---| | Kernel | npm i -g ease-design | the 42-command ui binary + /ui:* adapters — generate, gates, tokens, DS, tenant scroll engine | npm i -g ease-design@latest | | Full studio | git clone + ./setup.sh | everything above at HEAD + recall, rendered a11y, heartbeat/evolution, and the opt-in gflow browser hand | git pull + design-os update | | Figma plugin | design-os-figma-plugin | the plugin + figma-agent CLI + the 1:1 mirror — versioned in its own repo | its own repo's releases |

The ui kernel never phones home — it is deterministic and network-free by design, so it will not nag about new versions; updating is always your move, with the one-liner above.


Qualified Delivery

A better prompt improves the first draft. It does not prove the draft is ready. /ui:generate now treats generation as a typed, evidence-bearing release process:

raw request
  → capability-activation request
  → ui knowledge activate
      ↳ unavailable surface → CAPABILITY_UNQUALIFIED (route: null)
      ↳ qualified web marketing → capability-activation.json
      ↳ provisional native macOS/iOS/iPadOS → exact native route with qualified-delivery claim forbidden
  → design-brief.json v2 (activationRef)
  → generation-contract.json
  → candidate + 1440/768/390 renders
  → qualification-record.json
  → QUALIFIED | DRAFT_WITH_CONCERNS | BLOCKED_BY_EVIDENCE

The host model owns design reasoning. The zero-dependency ui kernel only validates contracts and rejects false-green verdicts:

ui knowledge activate capability-request.json --json > capability-activation.json
ui delivery validate design-brief.json
ui delivery validate generation-contract.json
ui delivery validate qualification-record.json
ui delivery validate learning-record.json

QUALIFIED is deliberately difficult to emit. It requires:

  • all six declared static gates passing;
  • rendered evidence at desktop, tablet, and mobile widths;
  • every Must acceptance criterion covered;
  • zero unsupported claims;
  • zero unresolved findings.

The repair loop is capped at three attempts. Missing rendered evidence never silently passes; the result remains DRAFT_WITH_CONCERNS. Automated checks narrow known failure classes—they do not claim WCAG conformance, desirability, or business performance.

Qualified Delivery v2 adds an implementation-grade craft contract:

  • Phosphor icons by default for greenfield UI, with evidenced brownfield exceptions;
  • third-party logos resolved from SVGL and cached with provenance;
  • purpose-fit original imagery through Codex image generation backed by GPT Image 2;
  • intentional 390 / 768 / 1440 responsive transformations;
  • hero, scroll, loading, interaction, reduced-motion, and JavaScript-failure evidence;
  • custom-styled semantic controls with keyboard and pointer proof for custom behavior;
  • a named composition thesis, signature spatial move, and whitespace strategy.

Version 1 artifacts remain valid historical records. Only version 2 proves the current craft contract. A v2 design brief binds activationRef to a ROUTED + QUALIFIED + QUALIFIED_DELIVERY_ALLOWED web-marketing → generate → html receipt; ui delivery validate recomputes its request and installed-catalog digests. A native receipt can route with PROVISIONAL assurance but has QUALIFIED_DELIVERY_FORBIDDEN, so it cannot satisfy a marketing brief. The validator also resolves a v2 qualification record's contractRef relative to the record and rejects missing or contradictory evidence.

The design prompt orchestrator compiles weak intent before generation: provenance-bound facets, product truth, three structurally divergent directions, complete region briefs, actionable visual DNA, and controlled content-led versus golden-ratio candidates. ui prompt-plan validate and ui prompt-plan preflight block incomplete, generic, forced-ratio, or over-budget builder plans.

The world-class learning loop adds controlled orchestrated and art-directed variants after qualification. It keeps critical floor failures separate from ceiling judgments, compares raw / enhanced / qualified / orchestrated content-led / orchestrated golden / selected / art-directed artifacts under blinded evaluation, and stores lessons as hypotheses, candidates, promoted knowledge, or rejected counterevidence. Promotion requires explicit expert approval or repeated winning evidence across at least three cases and two categories.

The schemas are public contracts: capability profiles · capability activation · design brief · prompt plan · generation contract · qualification record.


Advanced motion direction

ui init now installs design-os-gsap-motion, a runtime-neutral skill adapted from the official GreenSock GSAP Skills. It teaches the host agent to direct one story-bearing interaction, choose the smallest capable motion layer, and implement GSAP timelines, ScrollTrigger, responsive branches, framework cleanup, and plugin effects without sacrificing accessibility or runtime stability.

The skill is contextual. CSS remains the default for simple transitions; GSAP activates for coordinated web choreography, pinned or scrubbed scenes, spatial continuity, SVG transformation, or gesture physics. Native mobile interfaces keep native animation and haptic APIs.

trigger -> changed property -> user meaning -> reduced-motion equivalent

Every advanced scene must declare that contract and provide normal-motion, reduced-motion, cleanup, INP, CLS, and frame-stability evidence. Read the motion direction and runtime skill.


Gradient fields

ui init installs design-os-shader-gradient, a T6 capability for animated 3D gradient fields built on ShaderGradient (MIT). Ten named presets, plus twelve hand-configurable shader-by-mesh surfaces no preset reaches.

Rendered from the published renderer at the pinned version, each preset in its own palette. In a real page a field's colours come from the active design system instead.

It is gated, not offered. A field is reachable only after the motion ladder selects T6 and the persona's motion cap allows it — and it shares one per-viewport budget with the Canvas UI effect capability, so a page gets one T6 surface, never one of each.

A field fails two independent ways, so it carries two fallbacks and shipping one as though it were both is a reject:

| Trigger | Fallback | Why this one | |---|---|---| | prefers-reduced-motion: reduce | the same field, frozen | WebGL still works; the visitor asked for less motion, not a different design | | no WebGL, or context loss | a token-derived CSS gradient | there is no canvas to freeze, and something must still light the surface |

No preset ships a static default, so the frozen state does not exist until it is wired — ui knowledge check fails any matrix row whose fallback never names it.

The field's three colours are derived from design-system tokens, never carried over from the preset. A canvas is invisible to ui ds-usage-lint, so nothing would catch a field that quietly became the page's colour authority — which is exactly why it is a rule and not a lint. Read the gradient direction and runtime skill.


Your project gets a staff — the agents

One command turns the design system into a small team of Claude Code subagents, each bound to the project's declared identity and scoped to exactly one job:

| Agent | Name pattern | Does | Never does | |---|---|---|---| | Designer | designer-<studio>-<project> | generation & iteration on the project DS | scores its own work | | Curator | curator-<studio>-<project> | critique, audits, verdicts with evidence | edits an artifact | | Figma hand | figma-<studio>-<project> | canvas operations via figma-agent, drift-asserted | simulates when the plugin is down |

The prefix is generic on purpose — every DESIGN:OS project has a designer, a curator, a figma hand — while the suffix carries the identity: a studio soul named MERIDIAN on the project meridian-store yields designer-meridian-meridian-store.

Set it up

ui ds soul init --studio    # once per MACHINE — your studio's stance; its name: names the agents
ui ds soul init             # once per PROJECT — the project's stance (or let /ui:learn draft it from evidence)
ui agents init              # writes .claude/agents/<role>-<studio>-<project>.md

Both souls start as status: draft scaffolds — fill Never / Always / Voice, then set status: ratified. ui ds soul check (and --studio) lints the structure. No studio soul yet? Agents still generate as designer-<project> and the command tells you what you're missing.

Use them

In Claude Code, delegate in plain words — "ask the designer agent to build the empty state for API Reviews" — or just describe the task and let Claude Code auto-route: each agent's description declares its scope, so build requests find the designer and review requests find the curator.

Identity is live, not baked. Every agent's first action is ui ds context, which carries the project soul (and the studio soul beneath it). Edit a soul — the whole staff changes its taste at the next task, no regeneration needed. Only the names are stamped at generate time.

The roster is yours — one agent or all three: ui agents init --roster designer,curator. ui agents list shows what exists; ui agents check fails (agent-stale, exit 1) on hand-edited or outdated files — heal with ui agents init --force. Claude Code runtime only for now.


The machine floor

Most design guidance is prose the model can talk itself past. The DESIGN:OS floor is code: deterministic linters run on every delivery, specialist gates cover design-system usage and flow structure, and the rendered tier checks what static analysis cannot see. A detected blocking breach cannot receive a QUALIFIED verdict.

The floor is no longer only for the HTML the model just wrote. Point it at a directory and it reads the app you already have — HTML, React, Vue, Svelte, plain CSS, SwiftUI, Flutter — because the rules are written against design facts, not against CSS syntax. One rule, every language. Adding a platform is one extractor, and no rule is edited.

| Layer | What it proves | |---|---| | ui taste-lint — 34 absolute checks | The generated-UI tells: transition: all, layout-property keyframes, missing reduced-motion, overshoot easing, italic display headings, uppercase line-height < 1, focus rings that fade in, z-index inflation, off-grid spacing, mixed icon sets… | | ui validate-layout — 20 checks | Structural + overflow safety: unclosed tags, fixed-width overflow, 100vw traps, root overflow-x: hidden breaking sticky. | | ui content-lint — 16 checks | Honest copy: lorem-ipsum, placeholder copy, placeholder names (Jane Doe / Acme), click-here links, all-caps shouting. | | ui a11y-lint + ui ds a11y | Tier-1 static WCAG checks + token-pair contrast (every {role}/{role}-foreground pair ≥ AA, hover/active states included). Never claims "compliant" — says exactly what it checked. | | ui tell-lint — 43 checks, any language | The design tells: the side-tab accent, the stock purple palette, nested cards, a kicker above every heading, a pulsing dot that reports nothing, an overused face (including Flutter's default Roboto — SF is not a tell). Mostly advisory: a tell prints, it never fails a build. | | ui a11y-lint — computed contrast | The real WCAG ratio against the nearest opaque ancestor, from a resolved cascade. It refuses where the answer would be a fiction — a gradient background, a translucent veil — and says the run was partial. | | a11y-audit + page-shot + ui vr | The rendered tier: axe-core over live Chrome, deterministic PNG renders, pixel-level visual-regression gates per component (design-os vr-matrix). |

The pairing is structural: every standard ships an emitter and a linter in the same commit — prose-only rules drift; enforced rules hold.

So is the honesty. Every run reports what it could not do as loudly as what it found: an UNDERCOUNT tier, an unresolvable cn() call, a rule NOT-EVALUATED for want of facts, a contrast pair that could not be computed, a walk that hit its budget. A low finding count is never allowed to read as a clean page.


Workflow map — current generation and learning loop

The host agent performs the design reasoning. The ui kernel validates artifacts and gates; it never generates assets, calls a model, or accesses the network.

flowchart TD
    U["1 · User intent or reference"] --> CA{"2 · ui knowledge activate<br/>typed surface + visible evidence"}
    CA -- unavailable --> STOP["Stop · CAPABILITY_UNQUALIFIED<br/>route: null"]
    CA -- web-marketing qualified --> B["3 · Preserve activation receipt<br/>compile design-brief.json v2"]
    CA -- native-macos provisional --> N["Route native-macos<br/>PROVISIONAL · qualified delivery forbidden"]
    CA -- native-ios provisional --> NI["Route native-ios<br/>PROVISIONAL · qualified delivery forbidden"]
    CA -- native-ipados provisional --> NP["Route native-ipados<br/>PROVISIONAL · qualified delivery forbidden"]
    B --> PP["4 · Compile prompt-plan.json<br/>3 structural directions · region jobs · visual DNA"]
    PP --> PF{"ui prompt-plan<br/>validate + preflight"}
    PF -- fail --> B
    PF -- pass --> G["5 · Ground the project<br/>scan · DS context · soul · memory prior"]
    G --> S["6 · Select direction<br/>content-led vs golden candidate"]
    S --> C["7 · generation-contract.json v2<br/>sections · assets · states · viewports · motion"]
    C --> A["8 · Resolve assets<br/>project → approved source → image generation → no image"]
    A --> I["9 · Implement one candidate<br/>semantic responsive UI · Phosphor · GSAP only when justified"]
    I --> SG{"10 · Six static gates"}
    SG -- fail --> R["Targeted repair<br/>worst evidence-backed finding only"]
    SG -- pass --> RE["11 · Render evidence<br/>1440 · 768 · 390 · states · reduced motion · no-JS"]
    RE --> Q["12 · Independent curator<br/>qualification-record.json"]
    Q -- DRAFT_WITH_CONCERNS --> R
    R --> SG
    Q -- BLOCKED_BY_EVIDENCE --> ASK["Ask for missing product truth"]
    Q -- QUALIFIED --> D["13 · Deliver with evidence<br/>then integrate through code-surface intake"]
    D --> EXP{"Quality-learning task?"}
    EXP -- no --> RCPT["14 · Record evidence receipt"]
    EXP -- yes --> AD["Art-direct qualified control<br/>section reference boards · max 3 ceiling revisions"]
    AD --> LR["Blinded comparison + learning-record.json"]
    LR --> PR{"Expert approval or<br/>3 wins across 2 categories?"}
    PR -- no --> HYP["Keep as hypothesis or counterevidence"]
    PR -- yes --> PROM["Promote bounded lesson with causal receipt"]
    RCPT -. "relevant weak prior on a later job" .-> G
    PROM -. "relevant weak prior on a later job" .-> G

The repair loop is capped at three attempts. Missing rendered evidence cannot become QUALIFIED. Memory can influence a decision only as a cited weak prior; project truth, current evidence, and the ratified soul outrank it.

Entry surfaces and outputs

flowchart LR
    subgraph IN["Entry surfaces"]
      E1["Plain intent"]
      E2["Existing codebase"]
      E3["Live URL"]
      E4["Figma file"]
      E5["Tokens / existing DS"]
      E6["Image reference"]
    end
    E1 --> W["DESIGN:OS workflows"]
    E2 --> W
    E3 --> W
    E4 --> W
    E5 --> W
    E6 --> W
    W --> H["Host model + DESIGN:OS knowledge and installed skills"]
    H --> K["42-command deterministic ui kernel"]
    H --> O1["Responsive production UI"]
    H --> O2["Idiomatic Figma canvas"]
    H --> O3["Design-system artifacts"]
    H --> O4["Evidence, audit, and learning records"]
    K --> DS[("design/ store<br/>tokens · registry · manifest · memory")]

Brownfield work routes through /ui:learn before generation unless the user explicitly chooses a fresh direction. Figma, rendered accessibility, screenshots, image generation, SVGL, GSAP, and semantic recall remain optional hands used by the host workflow, not hidden dependencies of ui.


The plugin — an agent and a designer in the same Figma file

1,675 tests green in its own repo · four adversarial review rounds · a 200-target apply into a 10,000-record registry in ~1 ms

Every AI-on-Figma setup hits the same three walls. The agent's work collapses into one giant undo step — press ⌘Z after twenty minutes of automated work and twenty minutes disappear. The designer's own edits are invisible to the codebase — a human moves a component, and the registry the agent builds from quietly goes stale. And when both act at once, they trample each other — half-applied scripts, overwritten changes, no way to tell what happened. This plugin exists to remove exactly those three walls.

Your edits reach the codebase — the right codebase

Every change a designer makes is captured live and, after the file goes quiet, offered back as one prompt: "N changes ready — Sync now." A click runs the deterministic kernel over the ledger — no model sits in that path, so a sync is reproducible and auditable. And a sync cannot guess where it goes: each Figma file is bound to its project once —

figma-agent bind --file "Design System v4" --dir ~/code/your-app

— and from then on that file's changes land in that project's registry, never in whatever directory a process happened to start from. Unbound files stage safely and migrate on bind. The confirmation never flatters itself: "Synced — 3 added, 1 updated" only when records were actually written, "Nothing synced" when the run landed nothing, and a failure keeps the prompt alive so the retry is still there.

The agent cannot wreck your file

Every mutating operation seals its own undo step — ⌘Z rolls back one thing, not the whole session. Arbitrary scripts run inside a bracket that undoes itself on error, and reports rolledBack: true only when the rollback actually completed. Mutations are jobs: one runs per file at a time, the rest queue in order, reads bypass the queue. A timeout tells you the work was not cancelled and hands you a job id to poll; cancel actually cancels; a reply that arrives after a job was killed is discarded, never served. Nothing is ever lost silently — every eviction, prune, and rotation leaves a counter, an archive, or an audit record, and if your project already owns a file the plugin wants (its own component-registry.json), the kernel yields rather than co-opts. Every error lands in design/figma-errors.jsonl with its full untruncated reason — a log written for the agent that caused it, so it can read and fix.

Why this beats a write-only bridge

Most integrations write into Figma and stop. This one closes the loop in both directions — the designer's hand edits flow back into the same registry the agent generates from, so drift is detected and reconciled instead of accumulating — and it holds under load: per-file cursors and log rotation keep sync fast on files with thousands of components (a 200-target apply into a 10,000-record registry measures about a millisecond), and the panel reports only what actually happened — a name appears in the activity feed only when the reply carried one, a count only when one parsed. The whole thing shipped through four adversarial review rounds in which every blocker was reproduced before it was believed, and each fix carries a test that fails against the previous code.

This plugin now ships as its own repo — design-os-figma-plugin — so it can version, release, and take contributions independently of this kernel. Install, build, and bind instructions live there; this section stays as the introduction to why it exists.


The Figma hand

figma-agent drives a Figma plugin over a local self-healing broker: reconnect back-off, heartbeats, a park queue that holds commands through a broker respawn, and a multi-file registry — full detail in the plugin repo's own README.

Four moves, both directions:

  • Readdesign-os figma scan exports components/variables/styles; ui ingest-figma-ds turns them into tokens + registry + DESIGN.md.
  • Auditdesign-os figma audit runs ten deterministic DS-hygiene detectors over the open file's component library — a raw one-pass scan that survives 160k-instance files, judged entirely in fixture-tested CLI code.
  • Write/ui:to-figma authors idiomatic canvas: auto-layout, real instances, token-bound variables, drift-asserted geometry.
  • Mirrordesign-os figma reconcile --apply keeps each component's registry record a 1:1, rebuildable reflection of its Figma node — the same buildable representation the write path dr