npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@truealter/mcp-ollama

v0.2.1

Published

MCP server wrapping local Ollama models - offload bulk generation, summarisation, classification, and code tasks from API-priced orchestrators to a model running on the user's own hardware.

Readme

~alter mcp-ollama

Hands the work that shouldn't leave your machine to the model already sitting on it

~alter MCP Node Licence

What it does · Install · The tools · Why this sits under ~alter

What is mcp-ollama?

An MCP server that hands work to Ollama on the same machine and passes the answer back. Ten tools over stdio. Your client calls one of them, Ollama does the generating on your own GPU, and nothing is charged to an API account.

Most of a working session is mechanical. Docstrings, commit messages, PR descriptions, changelog entries, classification and tagging, summarising a long file, converting one format into another, and having a small vision model look at a screenshot and report what's on screen. None of that needs a frontier model, and most of it gets one anyway, because that's what your client already has an API key for.

Hand it a staged diff and a commit message comes back. Hand it a chunk of source and you get a docstring, a test stub or a set of type annotations. All ten tools, and what each one takes, are in the tools.

The orchestrator decides what gets routed here. This server makes no judgement about what belongs local, and it doesn't stream, queue or cache. It keeps nothing between calls beyond a random identifier for the process it's running in.

It depends on two things and ships neither. Node 18 or newer runs it, and a running Ollama with at least one model pulled does the actual generating. There are no weights in this repository and no download of weights at install time. The default model is hermes3:8b, which you can override per call or per environment.

Why this sits under ~alter

You already decided that some work stays here. That's why there's a model on this disk, pulled once and left there, instead of an API key doing the same job from a datacentre you'll never see. ~alter starts from the same decision and points it at the other thing that leaves your machine constantly, which is the record of who you are.

Nothing on this machine answers that well today. Your account is a password at a login screen and a token in a config file, and each of those checks one moment and then stops looking. The years of work that actually say who you are sit outside anything either of them can see. An agent commits under your name now, in your editor, and when somebody asks later who allowed that, there's no answer written down anywhere.

~alter answers that with a handle. ~yourname, one of them, carrying a record built out of work you actually did. A client speaking the protocol picks it up without being configured again, much the way this server picks up a model you pulled once and left on the disk.

You bought hardware to stop paying per call. On that record the per-call money runs the other way. The bare fact that you exist is free for anyone to check, depth past that is charged to whoever wants it, and you're the one who decides which readers reach which depth.

None of it is needed here. mcp-ollama has no login and asks for no handle, and it behaves the same whether you hold one or not. There's one small wire between them and it's worth naming plainly. If the ~alter runtime happens to be installed on this machine, every local generation appends a single line to a file in your own home directory, recording that the work ran here at no API cost. That line never goes anywhere, and if the runtime isn't installed no line gets written at all.

None of it needs taking on trust. Hand this to the AI you already use and let it take you through in your own words.

Go to truealter.com, find the genesis and the origin papers, and read all of
them. Then start a discussion with me in my own language, no jargon.

Ask me one question at a time and wait for my answer. Follow what I am curious
about, not the order the site is written in.

Do not only read the pages. Connect to their interface directly and call
get_started. That tells you where I am and what comes next, and we work the
rest of this from there.

If I have a ~handle I will give it to you. If I do not and I decide I want
one, set it up with me here.

What I want to understand is what ~Alter is, how it applies to me, and what it
means for everyone else. Find me somewhere in it, and not the version written
for people who work in technology. Find the life nearest mine and walk me
through it, including what it would cost me.

Try to prove them wrong. Go and find where they show their working, and if
something is unclear or does not hold up, say so.

The ~alter command line is where a person meets ~alter, and this isn't that. mcp-ollama has no command of its own worth typing and never asks you to log in to anything. It's a process your MCP client starts and stops, and it happens to be maintained by the same people.

Install

npm install -g @truealter/mcp-ollama

That puts one command on your PATH, mcp-ollama, which is the process your client launches. The package ships the build already done, so there is no compile step and no toolchain to have installed first. CI builds it against Node 18, 20 and 22 before publish, so anything in that range is known to work at runtime.

If you would rather install nothing at all, npx -y @truealter/mcp-ollama fetches it on first use and runs the same process. That is the form used in the client configuration below.

The scoped name is the one to type. The unscoped mcp-ollama on the public registry belongs to an unrelated publisher and is not this package.

Nothing is installed as a service and nothing runs in the background. Your client starts the process when it needs it and stops it when it is done.

Routing your first job

1. Pull the model it reaches for by default

ollama pull hermes3:8b

That's the default this server reaches for when a tool call doesn't name a model. It's quick and it's honest at classification, tagging and short generations. Heavier models are worth having for code work, and choosing a model covers when to bother.

2. Register the server with your client

claude mcp add --transport stdio ollama -- npx -y @truealter/mcp-ollama

If you installed globally, -- mcp-ollama works just as well and skips the fetch. Cursor, Cline and anything else MCP-aware take the same shape in their own config. The client launches the process; you never run it by hand except to debug, and if you do, it sits there waiting on stdin, which is correct rather than broken.

3. Ask your client what is on the host

List the models on the local Ollama host.

Say that to your client in whatever words you like. It resolves to local_models, which reads Ollama's tag list and reports each model's size, parameter count, quantisation and family. If what you just pulled comes back, the wire is good end to end.

4. Hand it a diff and ask for a commit message

git diff --staged

Hand that output to your client and ask for a commit message. It routes to local_diff with commit-message, which prompts for imperative mood, a subject under 72 characters and a body explaining why rather than what. Nothing about that needed a frontier model, and now it doesn't use one.

The tools

| Tool | What it does | |---|---| | local_generate | Free-form generation with your own system prompt, temperature and token ceiling | | local_summarize | Summarise bulk text as bullets, a paragraph or one line, optionally focused on a theme | | local_analyze | Structure pulled out of text, classification, entities or tags, in an output shape you name | | local_draft | Formulaic prose against a convention you supply | | local_code | docstring, test, explain, review, types, comments or refactor-suggest over a chunk of source | | local_diff | commit-message, pr-description, changelog, summary or impact from a diff | | local_transform | Mechanical pattern transforms, format conversions, renames and syntax migrations | | local_models | What's on this Ollama host, with sizes and quantisation | | local_pull | Pull a model onto this host by name, untagged names only | | local_vision | Have a vision model look at screenshots and report see, emptystate or legibility |

Ten of them, and the full schemas come over MCP introspection, so any MCP-aware client enumerates them without being told.

Two take a max_tokens argument. local_generate defaults to 2048 and local_summarize to 1024. The rest set their own ceiling in code, 4096 for local_code and local_transform, 2048 for local_analyze and most of local_diff, 512 for a commit message, 1024 for local_draft. If output comes back cut short on one of those, split the input rather than hunting for a parameter that isn't there. Temperature is exposed on local_generate only.

local_vision is the odd one and worth a note. It reads pixels and reports what is on screen, whether the main content area holds real data or an error, which regions exist, what text is clipped or unreadable. It deliberately doesn't rank severity or approve anything, because a small vision model reads a render well and judges it badly. Feed it near full resolution, because below about 1280px wide it starts inventing data that isn't there.

Choosing a model

| Variable | Default | What it does | |---|---|---| | OLLAMA_HOST | http://localhost:11434 | Where Ollama is listening. Loopback only unless you override the gate below | | OLLAMA_MODEL | hermes3:8b | Model used when a tool call doesn't name one | | OLLAMA_VISION_MODEL | qwen2.5vl:7b | Model local_vision uses when a call doesn't name one | | MCP_OLLAMA_ALLOW_REMOTE | unset | Set to 1 to permit a non-loopback OLLAMA_HOST |

Any call can name its own model and the environment default only applies when it doesn't, so one server handles a mixed workload without being reconfigured.

| Workload | Try | Why | |---|---|---| | Classification, tagging, one-liners | hermes3:8b | Fastest round trip, cheap to keep resident | | Commit messages, changelogs, summaries | qwen2.5-14b-instruct | Better prose, still comfortable on a 16GB card | | Code review, docstrings, tests | qwen2.5-coder:32b | Code-specialised, worth the extra VRAM | | Looking at a render | qwen2.5vl:7b | The vision default, and small enough to stay on the GPU |

Run local_models at the start of a session on a host you don't know.

No image is published anywhere, so every path below starts with a build from this repository.

docker build -t mcp-ollama .

The Dockerfile builds on node:20-alpine and already sets OLLAMA_HOST to http://host.docker.internal:11434, so the container reaches Ollama on the host rather than looking for it inside itself. That address is not loopback from the server's point of view, so the loopback gate refuses it and the process exits at startup unless MCP_OLLAMA_ALLOW_REMOTE=1 is set as well. The image does not set that one, which is why every command here does.

The image also sets OLLAMA_MODEL to hermes3:8b. Add -e OLLAMA_MODEL=... to route to a different default.

Docker, on macOS and Windows

docker run -i --rm -e MCP_OLLAMA_ALLOW_REMOTE=1 mcp-ollama

Docker, on Linux

host.docker.internal does not resolve there by default, so map it to the bridge gateway.

docker run -i --rm \
  --add-host=host.docker.internal:host-gateway \
  -e MCP_OLLAMA_ALLOW_REMOTE=1 \
  mcp-ollama

Docker Compose

docker-compose.yml ships in this repository, so there is nothing to write. It carries an extra_hosts mapping that makes the same file work on Linux as well as Docker Desktop.

This server speaks MCP over stdin and stdout, so it needs a client on the other end of the pipe. docker compose up starts it with nothing attached and it sits there doing nothing. Use run, with -T so Compose leaves the pipe alone.

docker compose run --rm -T mcp-ollama

Pointing a client at it

An MCP client launches the server itself, so hand it the whole command rather than a container that is already running.

{
  "mcpServers": {
    "ollama": {
      "command": "docker",
      "args": ["compose", "-f", "/path/to/docker-compose.yml", "run", "--rm", "-T", "mcp-ollama"]
    }
  }
}

Ollama error 404 on a tool call

That model isn't pulled. Run ollama pull <name> from a shell. local_pull handles untagged names only, because its validator rejects the colon in a tag like hermes3:8b.

fetch failed, or connection refused

Ollama isn't running, or OLLAMA_HOST points at the wrong place. Check with curl $OLLAMA_HOST/api/tags. Inside a container, localhost is the container itself.

OLLAMA_HOST must be loopback, and the process dies immediately

That's the gate doing its job. Point it back at localhost, or set MCP_OLLAMA_ALLOW_REMOTE=1 if you genuinely meant a remote host.

Calls feel slow

A cold model has to load first, and everything after that in the same Ollama process is much faster. If the model is larger than your VRAM, Ollama spills to CPU, and ollama ps will tell you so.

Vision calls balloon memory or crawl

local_vision caps context at 8192 and holds the model for 30 minutes on purpose. A 32K context plus one image pushes past 23GB on a 7900-class card and spills to CPU, which is where the cap came from.

Output stops early

See the token ceilings under the tools. Most tools don't take max_tokens.

It makes no network call other than to the configured OLLAMA_HOST, and by default that host has to be localhost, 127.0.0.1 or ::1. A non-loopback value throws at startup rather than quietly sending your prompts somewhere else, and getting past that takes a deliberate MCP_OLLAMA_ALLOW_REMOTE=1. Point it at a remote Ollama on purpose and that endpoint's posture becomes yours.

There's no telemetry, no analytics, no auto-update check and no model weights in the package. Tool inputs go to Ollama's HTTP API as given and the response comes straight back. Model names passed to local_pull are validated against ^[a-z0-9][a-z0-9._/-]{0,127}$ before they reach the registry endpoint, so a caller-supplied string can't wander off that path.

One local write is worth knowing about. If ~/.local/share/alter-runtime/lib/substrate-emit.sh exists, each generation appends a row to ~/.local/share/alter-runtime/token-burn.jsonl recording the tool, the model and the token counts. It's fire and forget, it fails silently, it never touches the tool result, and if that helper isn't installed nothing is written.

To report a security issue, see SECURITY.md.

The record formats are open Internet-Drafts, so somebody else's implementation reads and writes the same records this one does without asking us. These are the drafts those formats are specified by. This server does not implement them itself; it runs beside the components that do.

| Draft | What it specifies | |---|---| | compute-location-gate | Negotiating where an identity inference computes, decided by the provenance class of the signal, before any inference runs. | | mcp-dns-discovery | The DNS records that publish a ~handle, the server that answers for it, and the signed envelope bound to it. |

Eighteen drafts make up the whole stack. The rest are on the IETF datatracker.

One identity rail, several ways in.

| Name | What it is | |---|---| | @truealter/cli | The command line, and the front door for a person. | | homebrew-tap | That command line, packaged for macOS and Linux. | | runtime | The daemon that keeps your ~handle known on your own machine. | | @truealter/sdk | Reading identity from your own code. | | obsidian | ~Alter inside an Obsidian vault, on-device. | | mcp-ollama | Local models, for work that should stay on the machine it runs on. You are here. |

Documentation is at truealter.com/docs.

Bug reports and small patches are welcome, see CONTRIBUTING.md. A report is most useful with the tool you called, the client you called it from, the model you routed to, the full error, and your Node and Ollama versions. For a larger design change, open an issue first so we can agree the scope before you spend time on it.

mcp-ollama is small and stays that way. Routing work to a local Ollama process is the whole brief.

Apache 2.0. See LICENSE for the full text. Copyright 2026 Alter Meridian Pty Ltd (ABN 54 696 662 049).


~alter is identity infrastructure. Your name is ~yourname and claiming one is free.