npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

apple-llm

v0.2.0

Published

Apple's on-device (Foundation Models) and Private Cloud Compute LLMs from Node: tools, streaming, structured output, an OpenAI-compatible server and a Vercel AI SDK provider. No API key, no account.

Readme

apple-llm

Apple's built-in LLMs from Node — the on-device model and Private Cloud Compute. No API key, no account, no model download, zero runtime dependencies.

npm install apple-llm
import { AppleLLM } from 'apple-llm';

const llm = new AppleLLM();
await llm.text('Summarize this in one line: …');

That is the whole setup. A small Swift helper compiles on first use (a few seconds, once) and stays warm between calls.

  • Everything an LLM SDK should have — streaming (for await), structured output typed by Zod, tool calling with your own JavaScript functions, multi-turn chat, cancellation, timeouts, token counting, usage.
  • Three ways in — this library, an OpenAI-compatible server for any tool that speaks OpenAI, and a Vercel AI SDK provider.
  • Two tiers — on-device (private, fast, free) and Apple's big server model on Private Cloud Compute — on macOS 27 Golden Gate, the one Siri AI is built on: reasoning, images, live web knowledge, 32k context. Free, no key, behind the same API.

Read this first

  • macOS 26+ on Apple Silicon, with Apple Intelligence on. Nothing works elsewhere, but importing always succeeds and probe() says exactly why.
  • The on-device model is small (~20B sparse, 1–4B active; 8,192-token context). Good at extraction, classification, tagging, rewriting, summaries and short answers. Bad at code and long reasoning. Use it for the jobs it suits.
  • The cloud tier is not local. tier: 'cloud' sends your prompt to Apple's Private Cloud Compute. The default, auto, uses the device when it can.

Quick tour

Structured output, typed by your schema

On device, the shape is guaranteed by constrained decoding — the model cannot produce anything else. Pass a Zod 4 (or ArkType, or Valibot) schema and the result is typed and validated; plain JSON Schema works too.

import { z } from 'zod';

const Contact = z.object({
  name: z.string(),
  email: z.string().email(),
  phone: z.string().nullable(),
});

const contact = await llm.json('Reach Grace Hopper at [email protected]', { schema: Contact });
//    ^? { name: string; email: string; phone: string | null }

Refinements the decoder cannot enforce (.email(), .min()) are checked afterwards; a reply that fails gets one automatic retry with the problems pointed out, then a SchemaValidationError carrying the issues and the raw text.

Streaming

for await (const delta of llm.stream('Write a haiku about bridges')) {
  process.stdout.write(delta);
}

const text = await llm.stream('…');              // or just await the whole reply
const { usage } = await llm.stream('…').result;   // or everything about it

Breaking out of the loop cancels generation. Structured output streams too — partial objects as the model fills them in, for rendering a card or form while it is still being written:

for await (const partial of llm.streamJson(prompt, { schema: Contact })) {
  render(partial); // { name: 'Grace' } → { name: 'Grace', email: '…' } → …
}

Tools: your JavaScript, called by the model

import { tool } from 'apple-llm';

const getWeather = tool({
  name: 'getWeather',
  description: 'Current weather for a city',
  parameters: z.object({ city: z.string() }),
  execute: async ({ city }) => fetchWeather(city), // city: string
});

const result = await llm.generate('Do I need an umbrella in Paris?', {
  tools: [getWeather],
});
result.text;      // "No — it's 21°C and sunny in Paris."
result.toolCalls; // [{ name: 'getWeather', arguments: { city: 'Paris' }, output: '…' }]

execute runs in your process, mid-generation, with your credentials and database. Arguments are schema-guaranteed like json(). A tool that throws ends the call with a ToolExecutionError whose cause is your error. Tools mix freely with Apple's built-ins: tools: ['ocr', getWeather].

Leave out execute and generation stops at the call instead (finishReason: 'tool-calls'), so you can run it yourself — send the result back as a tool message to continue.

Conversations

Pass a message list anywhere a prompt goes:

import type { ChatMessage } from 'apple-llm';

const messages: ChatMessage[] = [
  { role: 'system', content: 'You are terse.' },
  { role: 'user', content: 'My cat is called Biscuit.' },
];
const reply = await llm.generate(messages);
messages.push(reply.message, { role: 'user', content: 'What is my cat called?' });
await llm.text(messages); // "Biscuit."

History becomes Apple's native transcript, not a text preamble. When a long conversation outgrows the context window, the oldest turns are dropped to fit and result.trimmedTurns says how many (trimHistory: false to get a ContextLengthError instead).

Or let the helper keep the thread — it survives restarts:

const chat = llm.conversation('support-ticket-42', { system: 'You are terse.' });
await chat.text('My cat is called Biscuit.');
await chat.text('What is my cat called?'); // "Biscuit."

Cancellation and timeouts

await llm.text(prompt, { signal: controller.signal }); // AbortError
await llm.text(prompt, { timeoutMs: 5_000 });          // TimeoutError
new AppleLLM({ timeoutMs: 30_000 });                   // default for every call

Cancelling stops the model itself, not just the promise: the next call does not wait behind a reply nobody wants.

Images, documents, long text

await llm.text('What is in this photo?', { images: ['./photo.jpg'] });
await llm.text('Read the total.', { images: ['./receipt.png'], tools: ['ocr'] });
await llm.text('Compare these quotes.', { documents: ['./a.md', './b.md'] });
await llm.summarize(fiftyPageReport); // longer than the window? summarised in parts

Knowing what a call costs

const { tokens, contextSize } = await llm.countTokens(messages, { schema, tools });
const { usage, finishReason } = await llm.generate(prompt);
// usage: { inputTokens, outputTokens, totalTokens, cachedInputTokens }
// finishReason: 'stop' | 'length' | 'tool-calls' — 'length' means maxTokens cut it off

The big model: Private Cloud Compute

const cloud = new AppleLLM({ tier: 'cloud' }); // after `npx apple-llm setup-cloud`, once

await cloud.text(puzzle);                                   // reasons it through
await cloud.text('What is in this photo?', { images: ['./photo.jpg'] });
await cloud.text('Who won the last World Cup?', { webSearch: true }); // live web knowledge
await cloud.json(prompt, { schema: Contact });              // typed, validated

On macOS 27 Golden Gate this reaches the next-generation server model Apple built Siri AI on. probe() reads what it can do straight from the framework — reasoning, vision, tool calling and a 32,768-token window — without calling it. Measured against the on-device model: it solves multi-step logic puzzles the small model loses track of, finds a fact buried in 25,000 tokens, reads images, and with webSearch answers questions about last month's news correctly. It costs no API key and no money, only Apple's per-device quota (probe().cloud.quota), and your prompt leaves the Mac.

OpenAI-compatible server

npx apple-llm serve
# apple-llm serving an OpenAI-compatible API at http://127.0.0.1:11436/v1

Point any OpenAI client at it — the official SDKs, LangChain, LlamaIndex, editor plugins, Open WebUI — with any API key:

import OpenAI from 'openai';
const openai = new OpenAI({ baseURL: 'http://127.0.0.1:11436/v1', apiKey: 'unused' });
await openai.chat.completions.create({ model: 'apple-on-device', messages });

/v1/chat/completions supports streaming, tool calls, response_format with a JSON Schema (guaranteed on device), images as data or http URLs, seed and top_p. It binds to localhost, sends no CORS headers unless you pass --cors <origin>, can require --api-key, and serves the on-device model unless a client asks for apple-private-cloud by name. Or from code: import { serve } from 'apple-llm/server'.

Vercel AI SDK

import { generateText, streamText, stepCountIs, tool } from 'ai';
import { apple } from 'apple-llm/ai-sdk';

const { text } = await generateText({
  model: apple(), // on device; apple('cloud') for Private Cloud Compute
  prompt: 'What is the weather in Paris?',
  tools: { getWeather: tool({ inputSchema: z.object({ city: z.string() }), execute: … }) },
  stopWhen: stepCountIs(3),
});

AI SDK 6 and 7. generateText, streamText, generateObject / Output.object (constrained on device), multi-step tools, images and abort signals. Apple-only settings go in providerOptions: { apple: { useCase, guardrails, builtInTools } }.

CLI

apple-llm chat                                  # interactive, streamed; Ctrl+C stops a reply
apple-llm run "Summarize: …"                    # one completion; "-" reads stdin
apple-llm run --json "…"                        # text + usage + tool calls as JSON
apple-llm run --schema person.json "Ada, 1815"  # guaranteed JSON
apple-llm run --schema person.json --stream "…" # partial objects, one per line
apple-llm run --image photo.png "What is this?"
apple-llm run --tool ocr --image receipt.png "Read the total."
apple-llm serve [--port 11436] [--api-key k]    # OpenAI-compatible API
apple-llm count --system "…" -                  # tokens before sending
apple-llm probe                                 # what this machine can do
apple-llm setup-cloud                           # one-time, for the cloud tier

apple-llm --help for everything, including the Write-with-Siri presets (rewrite, proofread, summarize, draft) and ask-screen.

Two traps worth knowing

Never set temperature: 0. Constrained decoding already guarantees the schema, so greedy decoding buys nothing and reliably degenerates — it padded an unbounded array forever, then ran away inside a single string, turning a 2s call into 20s. The default is 0.4. For reproducible output use a seed instead: sampling: { mode: 'topK', k: 50, seed: 42 } returns byte-identical text across runs, without the degeneration.

Bound your arrays. Apple honours maxItems (.max(5) in Zod) but ignores maxLength on strings.

The two tiers

| | on-device | cloud (Private Cloud Compute) | |---|---|---| | Model | AFM 3 Core (~1–4B active) | Apple's next-gen server model (Siri AI's, on macOS 27) | | Privacy | nothing leaves the Mac | your prompt leaves the Mac | | Context | 8,192 tokens (4,096 on macOS 26) | 32,768 tokens | | Reasoning | no | yes | | Latency | ~0.3–2s warm | ~2s, ~10s at 25k tokens | | JSON | guaranteed by constrained decoding | requested, recovered, validated | | Streaming | yes, text and objects | one chunk (Shortcuts is not incremental) | | Images | yes (macOS 27) | yes (macOS 27 Golden Gate) | | Web knowledge | no | yes, with webSearch | | Your tools, sessions | yes | no | | Setup | none | apple-llm setup-cloud, once | | Limits | none | Apple's quota (read before calling: probe().cloud.quota) |

Errors

Every error is an AppleLLMError with a stable code, and a class to branch on:

| class | when | |---|---| | ModelUnavailableError | Apple Intelligence off, model downloading, unsupported Mac/OS (.reason says which) | | ContextLengthError | the request does not fit (.contextSize, .tokenCount) | | SchemaRejectedError | Apple's decoder refused a schema | | SchemaValidationError | the reply failed your schema's validation (.issues, .text) | | ToolExecutionError | your tool threw (.toolName, .cause) | | AbortError / TimeoutError | your signal fired / the deadline passed | | QuotaError | rate limited (.resetDate when Apple gives one) | | RefusalError | the model's guardrails declined | | ModelBusyError | other processes are using the model; already retried with backoff | | SetupRequiredError, UnsupportedError | a setup step is missing / the tier cannot do that |

Resource use

One AppleLLM keeps one helper process warm — construct it once. The helper does not keep your process alive, so a script exits without close(); call close() (or use await using llm = new AppleLLM()) to release it early in a long-running program. Apple serialises inference, so requests queue rather than run in parallel — concurrency buys nothing here.

More

The repository README has the measurements behind every default, the eight rules of Apple's schema dialect, and why the cloud tier goes through Shortcuts.

MIT.