npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@openmycel/expo-llama

v0.4.1

Published

Run GGUF language models on iPhone with llama.cpp — an Expo native module in Swift.

Downloads

1,100

Readme

expo-llama

Run GGUF language models on iPhone, fully on the device. An Expo native module in Swift over llama.cpp: load a model, stream an answer token by token, stop it, unload it. Plus the file-side helpers a model download needs: SHA-256 of a multi-gigabyte file, a reachability check that tells "offline" from "blocked", and exclusion from iCloud backups.

Built for OpenMycel, a private assistant that works offline.

Status: early. iOS only. Sessions (0.3.0) are measured in the iOS Simulator only, not yet on an iPhone — see What it does not do.

Install

Needs Expo SDK 57+ with React Native 0.86+, iOS 16.4+, and a development build (native modules do not run in Expo Go).

npx expo install @openmycel/expo-llama
npx expo prebuild   # or: cd ios && pod install

The llama.cpp engine ships inside the package as a ready llama.xcframework (8.7 MB download, 26 MB unpacked): nothing is compiled or downloaded during pod install.

Usage

import { generate, loadModel, stop, unload } from '@openmycel/expo-llama'

const info = await loadModel(fileUri, { contextSize: 2048 })
// { description: "qwen3 0.6B Q8_0", parameters: 596049920, sizeBytes: 633495552, contextSize: 2048 }

const result = await generate(
	[
		{ role: 'system', content: 'Answer briefly.' },
		{ role: 'user', content: 'Two ideas for a quick dinner?' },
	],
	{ temperature: 0.7, topP: 0.8, topK: 20, minP: 0, presencePenalty: 1.5 },
	(piece) => console.log(piece), // streamed as it is generated
)
// { text, promptTokens, cachedTokens, tokens, stopped, promptMs, generationMs, tokensPerSecond }

stop() // from anywhere: the running generate resolves with what it has so far
await unload()

The prompt is built with the model's own chat template (llama_chat_apply_template), so the app passes plain { role, content } messages. The module keeps no conversation: pass the whole history each time; the session's cache saves reading it again (Sessions).

API

| Function | What it does | | ----------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | loadModel(path, options?) | Loads a GGUF file (path or file:// URI), replacing the model loaded before. contextSize (default 4096), gpuLayers (default -1 = all on Metal). | | generate(messages, options?, onText?) | Answers the conversation; onText gets each piece of text. Options below. | | stop() | Ends a running generate; it resolves with stopped: true. | | unload() / isLoaded() | Frees the model and every session / reports whether a model is loaded. | | loadEmbedder(path, options?) | Loads an embedding model next to the chat model (Embeddings). | | embed(texts) | One vector per text, each of unit length: a dot product is the cosine. | | unloadEmbedder() / isEmbedderLoaded() | Frees the embedding model / reports whether one is loaded. | | sha256File(path) | SHA-256 of a file as lowercase hex, read in 4 MB chunks with CryptoKit — a 3 GB model never sits in memory. | | cancelSha256File() | Stops the running sha256File; it rejects with HashCancelledException within one chunk. | | checkSource(url) | One HEAD request (ephemeral session, no cookies) and why it failed: offline, unreachable (DNS, TLS, timeout, HTTP 403/451 — what a block looks like) or http. | | onMemoryWarning(listener) | iOS says memory is low: unload() now, before the app is killed. Returns the unsubscribe. | | excludeFromBackup(path) | Keeps a re-downloadable model out of iCloud and device backups. |

generate options; defaults are llama.cpp's own:

| Option | Default | What it does | | ----------------- | ------- | ------------------------------------------------------------------------------------ | | maxTokens | 1024 | Stops after this many tokens. | | temperature | 0.8 | 0 = greedy. | | topK | 40 | | | topP | 0.95 | | | minP | 0.05 | | | presencePenalty | 0 | 0 = off; looks at the last 64 tokens. | | seed | random | | | grammar | — | GBNF with a root rule: the answer can only be what it allows, e.g. one JSON shape. | | session | "" | The context to run in, made on first use (Sessions). | | contextSize | load's | The session's context size in tokens (Sessions). |

The result: text, promptTokens, cachedTokens (of them, taken from the session's cache), tokens, stopped, promptMs, generationMs, tokensPerSecond.

Errors are thrown as typed exceptions with a readable message: ModelNotFoundException, ModelLoadException, ContextException, ModelNotLoadedException, ChatTemplateException, PromptTooLongException, DecodeException, HashCancelledException, EmbedderNotLoadedException, PoolingException.

Sessions

A session is a context of its own on the loaded model: the weights are shared, the KV cache is its own. It keeps what it last read, so a prompt that starts the same way as the one before — the same system prompt, a chat one message longer — is read only from where it differs. cachedTokens in the result says how much was skipped.

Use one session per kind of prompt, so they do not overwrite each other's cache: the default "" for the chat, "router" for a fixed instruction asked again and again. Calls in different sessions still run one at a time, on the same queue.

Each session costs memory: its KV cache is allocated for its whole context when the session is first used — for an f16 cache, 2 × layers × KV heads × head size × 2 bytes per token. By default that is the contextSize of loadModel; a session that only reads short prompts can ask for less with the contextSize option of generate. Asked with another size than it has, the session is made again and its cache is lost. unload() frees all sessions; there is no way to free one.

Embeddings

An embedding model turns a text into a vector; texts close in meaning get close vectors. Use it to find which of many items a message is about before the chat model reads them — for example, the few skills of hundreds that a message may be for.

import { embed, loadEmbedder } from '@openmycel/expo-llama'

await loadEmbedder(fileUri, { pooling: 'last' })
const { vectors } = await embed([
	'Instruct: Given a message to a phone assistant, retrieve messages that ask the assistant for the same action\nQuery:Set a timer for 10 minutes',
	'Start a countdown timer for a length of time.',
])
const cosine = vectors[0].reduce((s, x, i) => s + x * vectors[1][i], 0)

It is a second model with its own weights, loaded next to the chat model: both take memory at once. Options: contextSize — the tokens one text may have, a longer text is cut (default 512); gpuLayers; threads — CPU threads, default 2, so the two models do not fight for the cores; pooling — mean, cls or last, as the model's card says (default: what the GGUF file says). The result: vectors, tokens, ms. A prefix or an instruction the model expects (e5's query: , Qwen3-Embedding's Instruct: …\nQuery:) goes in the text. Calls run on the same queue as generate.

What it does not do

  • Android. iOS only; Android is planned in the same module.
  • Freeing one session. Only unload(), which frees the model and all of them.
  • Two generations at once. One queue: a second generate waits for the first.
  • Keeping the cache across launches. A session's cache lives in memory until unload().
  • Measured on an iPhone (0.3.0). Sessions and the prompt cache are checked in the iOS Simulator, on the CPU. Speed and memory with Metal on a device are not measured yet.
  • Images, audio, LoRA adapters, tool calls. Text in, text or vectors out; grammar is the only way to shape the answer.
  • Embeddings measured on an iPhone. embed is checked in the iOS Simulator; speed and memory of two models at once on a device are not measured yet.
  • Downloading models. The app downloads the file; the module checks it (sha256File, checkSource) and keeps it out of backups.

Architecture

 JavaScript (the app)
   src/index.ts ── thin wrappers; subscribes to "onToken" for the length of one generate
        │  Expo Modules (JSI)
 Swift ─┼──────────────────────────────────────────────────────────────────────────
   ExpoLlamaModule.swift ── definitions, argument records, typed exceptions
        │                    stop() is a plain Function: it must reach a running generate
        ▼
   LlamaEngine.swift ── one model + a context per session; every call on one serial queue
        │   load:     llama_model_load_from_file → llama_init_from_model (session "")
        │   generate: chat template → tokenize → keep the cached prefix (llama_memory_seq_rm)
        │             → decode the rest of the prompt in n_batch chunks
        │             → sample (penalties → top-k → top-p → min-p → temperature → dist)
        │             → emit complete UTF-8 only → decode the token → repeat
        │   stop:     a lock-guarded flag checked before every token
        ▼
   llama.xcframework (llama.cpp b11146, built from source) ── Metal on device, CPU in the simulator

   ModelFiles.swift ── sha256 (CryptoKit), excludeFromBackup, checkSource (URLSession HEAD)

Decisions worth knowing:

  • One serial queue. Load, generate and unload never overlap, so the engine needs no locks around llama.cpp state. Only stop crosses threads.
  • UTF-8 at token boundaries. A multi-byte character can be split across tokens; bytes are held until they form valid UTF-8, so the app never renders a broken character.
  • CPU in the simulator. With Metal in the iOS Simulator llama.cpp produced garbage (only ! tokens, Xcode 27, b11146); the CPU gives correct output. gpuLayers is ignored there. Real speed can only be measured on an iPhone.
  • Quiet logs. llama.cpp logs every graph node at load; the module keeps warnings and errors only.
  • No network in the engine. The only request the module can make is checkSource, and only when the app calls it. No telemetry.

Building the engine

llama.cpp publishes a prebuilt llama-bNNNNN-xcframework.zip with every build, but it holds iPhone and macOS only — no iOS Simulator slice. So the module builds its own from source:

npm run build:llama   # scripts/build-llama.sh

The script downloads the pinned tag (b11146, the build behind the v0.5.0 release), refuses the archive unless git get-tar-commit-id matches the pinned commit, and runs llama.cpp's own build-xcframework.sh ios-sim ios-device. The arm64 code is compiled for armv8.2-a+dotprod+fp16 — the dot-product and half-precision instructions every iPhone since the A13 has, which ggml picks at compile time; the script refuses a binary without them. Needs Xcode and CMake; on an M3 Max it takes about 90 seconds.

Output: ios/Frameworks/llama.xcframework (iPhone arm64 + simulator arm64/x86_64) — in the npm package, not in git. Debug symbols make up ~170 of ~195 MB, so they are split off into dist/llama-<tag>-dSYMs.zip for the GitHub release, where crash reports can be symbolicated.

Developing the module

Link it next to an app:

npm install ../expo-llama        # "file:" dependency, a symlink
cd ios && pod install            # autolinking picks up ExpoLlama.podspec

Metro has to watch the module's folder and resolve expo and react-native from the app, so the bundle holds one copy of each; TypeScript needs preserveSymlinks. See apps/mobile/metro.config.js and tsconfig.json in OpenMycel. New Swift files need another pod install.

Credits

llama.cpp by the ggml authors, MIT; the package ships its license next to the framework, ios/Frameworks/LICENSE-llama.cpp. Module layout from Expo's expo-module-template, MIT.

Author and license

Created by Nikita MRCS — @nmrcs. Contributors are credited in the git history.

MIT.