npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@sidx1scr-apps/prefrontal

v1.2.20

Published

Prefrontal — Local AI Chatbot · 100% Private · Ollama + Llama.cpp + OpenRouter + Groq + Together AI + OpenAI

Readme

🧠 Prefrontal — Local AI Chatbot

Node.js Platform Release License Offline Auto Release + GitHub Packages Prefrontal CI (Stable + Cross Platform)

100% Offline · No Ads · Your Data Stays on Your Device

Prefrontal is an open-source, privacy-first chat interface for local AI models. It works with Ollama (desktop), Llama.cpp (any platform, including Android via Termux), and optionally OpenRouter if you'd rather skip local setup entirely. No required cloud dependency, no telemetry, no subscriptions.

🔗 Live demo — see what the UI looks like before installing.

[!NOTE] The hosted live demo only supports the OpenRouter runtime (it can't reach a local Ollama or Llama.cpp server from the web). It also stores your API key in a browser cookie, so opening the demo in a private/incognito window will prompt you to set one up again. Run Prefrontal locally if you want the full privacy picture — see Privacy & Data.


Table of Contents


✅ Requirements

  • Node.js v18+ and npm — to run the Prefrontal server itself
  • Git — only if you choose the Clone & Run or GitHub Packages install method below; the Release download method needs no git at all
  • One AI backend, set up below:

⚡ Quick Start

Pick whichever backend matches your needs — cloud convenience or fully offline privacy. (Want Llama.cpp or Android instead? Jump to Setting Up Your AI Backend.)

[!TIP] Don't want to install git? Skip the git clone step below and use Download a Release instead — everything else is identical.

🌐 Easiest: OpenRouter (no local AI install needed)

OpenRouter gives you free access to powerful cloud models — no GPU, no Ollama, nothing extra to install. The fastest way to get running.

1. Get a free API key Sign up at openrouter.ai and copy your API key from the dashboard.

2. Install and run Prefrontal

git clone https://github.com/sidx1-scratch/prefrontal
cd prefrontal
npm install && npm start

3. Configure and go Open http://localhost:3000, go to Settings, and set:

  • Runtime → OpenRouter
  • API Key → paste your key
  • Model → pick any free model (e.g. mistralai/mistral-7b-instruct:free)

Done — no GPU or model download required.

Optional: turn on Web Search Still in Settings, flip the Web Search toggle when using any OpenAI-compatible runtime. The model can then pull in live DuckDuckGo results before answering, with any sources it used shown as clickable chips under its reply. See Web Search below.

[!NOTE] the duckduckgo instant answer api only knows more known things so if you for example tell it to search up the prefrontal repo it wont find it because instant answer doesn't know about it. also when it searches up something it looks like this:


🖥️ Fully Offline: Ollama (no API key, your data never leaves your machine)

1. Install Ollama and pull a model

# Download Ollama from https://ollama.com, then:
ollama serve &
ollama pull gemma3:4b

2. Install and run Prefrontal

git clone https://github.com/sidx1-scratch/prefrontal
cd prefrontal
npm install && npm start

3. Configure and go Open http://localhost:3000, go to Settings, and set:

  • Runtime → Ollama
  • Server URL → http://localhost:11434
  • Model → gemma3:4b

📦 Installation Options

Four ways to get Prefrontal running locally — pick whichever fits how you work.

| Method | Best if you... | Git needed? | |---|---|---| | 📥 Download a Release | don't want to install or touch git at all | ❌ No | | ⚡ Quick Install Script | want a one-line install with no git and no extra repo clutter | ❌ No | | 📦 Install from npm | want a simple global install from npm | ❌ No | | 🧬 Clone & Run | are comfortable with git and want to stay on the latest commit | ✅ Yes | | 🌍 GitHub Packages (Recommended) | want to install once and run from anywhere without keeping a source folder around | ✅ Yes (login only) |

Option 1: Download a Release (no Git required)

If you don't want to mess around with git, you don't have to — every release ships with a ready-made source archive you can just download.

1. Grab the latest release Open the Releases page and, under Assets, download Source code (zip) (Windows/macOS) or Source code (tar.gz) (Linux/macOS).

2. Extract it Unzip the archive anywhere you like — Desktop, Documents, wherever's convenient.

3. Install and run Open a terminal inside the extracted folder — on Windows, Shift + Right-click the folder and choose "Open in Terminal" or "Open PowerShell window here" — and run:

npm install
npm start

That's the whole process. No git, no GitHub account, no cloning.

[!TIP] To update later, just download the newest release the same way, extract it to a fresh folder, and run npm install again there.

Option 2: Quick Install Script (no Git required, fetches only the files the app needs)

Prefer not to clone the whole repo (including docs, CI workflows, and tests) just to run the app? This one-liner downloads only the runtime-required files — app.js, index.html, style.css, manifest.json, server.js, .env.example, package.json, package-lock.json, and the vendor/ libraries — straight from GitHub, no git involved.

curl -fsSL https://raw.githubusercontent.com/sidx1-scratch/prefrontal/refs/heads/main/install.sh | bash
cd prefrontal
npm install && npm start

By default it installs into a prefrontal folder in your current directory. To install into a different folder name instead:

curl -fsSL https://raw.githubusercontent.com/sidx1-scratch/prefrontal/refs/heads/main/install.sh | bash -s -- my-folder-name

[!TIP] To update later, just re-run the same command — it'll re-fetch the latest versions of the required files into the same folder.

[!NOTE] Windows users: the install script requires a shell that supports bash — Git Bash (installed alongside Git for Windows) works well. If you'd rather not install that, the Download a Release, Clone & Run, and GitHub Packages options all work natively on Windows without it.

Option 3: Install from npm

Install Prefrontal directly from npm.

npm i -g @sidx1scr-apps/prefrontal

[!TIP] Update later with:

npm update -g @sidx1scr-apps/prefrontal

and run with:

npm explore @sidx1scr-apps/prefrontal -- npm start

Option 4: Clone & Run (quickest if you're comfortable with git)

No authentication needed — just clone and go. Best if you want to just get done with git for the day.

git clone --depth 1 https://github.com/sidx1-scratch/prefrontal
cd prefrontal
npm install
npm start

A local web server starts and the app opens automatically at http://localhost:3000.

Or you can just do

git clone https://github.com/sidx1-scratch/prefrontal.git
cd prefrontal
npm install
npm link
npm start

If you want to contribute or just want the full Git history.

npm link registers a global prefrontal command (the server is the bin) so you can start it from anywhere with prefrontal — useful if you run the Prefrontal Agent in the background, which auto-pairs with a server launched this way.

[!TIP] To update later, run git pull from inside the prefrontal folder. If the dependencies changed, run npm install afterward just in case.

Option 5: Install via GitHub Packages

More setup upfront, but once installed you can run Prefrontal from anywhere on your machine without keeping the source folder around.

1. Authenticate with GitHub Packages

npm login --scope=@sidx1-scratch --auth-type=legacy --registry=https://npm.pkg.github.com

Use your GitHub username and a Personal Access Token as the password.

2. Install globally

npm install -g @sidx1-scratch/prefrontal

3. Run from anywhere

npm explore @sidx1-scratch/prefrontal -- npm start

[!TIP] To update later, just run npm install -g @sidx1-scratch/prefrontal again.


🧠 Setting Up Your AI Backend

You need one of the following backends. Pick whichever suits your setup.

Option A: OpenRouter (easiest — no local install, free models available)

OpenRouter is an API that routes to dozens of AI models, many of which are free to use. No GPU or model download required — just an API key.

1. Create a free account Sign up at openrouter.ai and grab your API key from openrouter.ai/keys.

2. Configure Prefrontal

  • Runtime → OpenRouter
  • API Key → paste your key
  • Model → enter any model ID from openrouter.ai/models

Recommended free models to start with:

| Model | ID | |---|---| | Mistral 7B Instruct | mistralai/mistral-7b-instruct:free |
| Llama 3.2 3B | meta-llama/llama-3.2-3b-instruct:free | | Gemma 4 31B | google/gemma-4-31b-it:free |

[!NOTE] Free model IDs and rate limits change as providers rotate promotions. If one of the IDs above 404s, check the current list at openrouter.ai/models?q=free. Also the mistral model recommended is hard to find so use this link: Mistral 7B Instruct Free


Option B: Ollama (best for fully offline use — Windows, macOS, Linux)

1. Install Ollama Download from ollama.com and run the installer.

2. Start the server

ollama serve

This starts Ollama on http://localhost:11434. Keep this terminal open.

3. Pull a model

# Recommended — small, fast, great quality:
ollama pull gemma3:4b

# More capable (needs more RAM):
ollama pull gemma3:12b
ollama pull llama3.2
ollama pull mistral

4. Configure Prefrontal

  • Open Settings (⚙️ icon or Ctrl+,)
  • Runtime → Ollama
  • Server URL → http://localhost:11434
  • Model → type gemma3:4b or click Refresh to pick from installed models

Option C: Llama.cpp (cross-platform — Windows, macOS, Linux, Android)

Llama.cpp exposes an OpenAI-compatible REST API.

🖥️ Desktop (Windows / macOS / Linux)

Download a pre-built release: github.com/ggml-org/llama.cpp/releases

Or build from source:

git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build -DGGML_BLAS=ON   # optional: swap in a GPU flag, e.g. -DGGML_CUDA=ON
cmake --build build --config Release

Start the server (run from the llama.cpp folder):

# Option 1: auto-download a model from Hugging Face by name
./build/bin/llama-server -hf ggml-org/gemma-3-4b-it-GGUF --port 8080 --host 0.0.0.0 -c 8192

# Option 2: point -m at a .gguf file you already have
./build/bin/llama-server -m models/your_model.gguf --port 8080 --host 0.0.0.0 -c 8192

Configure Prefrontal:

  • Runtime → Llama.cpp / OpenAI
  • Server URL → http://localhost:8080/v1
  • Model → the model name/filename you started the server with (e.g. gemma-3-4b-it-GGUF or gemma-2-2b.gguf)

📱 Android / Mobile (via Termux)

1. Install Termux Get it from F-Droid — not the Play Store version, which is outdated.

2. Install dependencies

pkg update && pkg upgrade
pkg install clang git cmake make python3 libcurl

3. Build Llama.cpp

git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build
cmake --build build --config Release -j4

4. Start the server (auto-downloads a small model on first run)

./build/bin/llama-server -hf ggml-org/gemma-3-1b-it-GGUF \
  --port 8080 --host 0.0.0.0 -c 4096

[!TIP] Prefer to manage the file yourself instead of auto-downloading? Grab any compact GGUF model — something in the 1–4B parameter range, quantized to Q4, fits comfortably on most phones — and pass it with -m path/to/model.gguf instead of -hf.

5. Open Prefrontal In a second Termux session, start the Node.js server:

cd prefrontal && npm start

Then open http://127.0.0.1:3000 in your Android browser.

Configure Prefrontal:

  • Runtime → Llama.cpp / OpenAI
  • Server URL → http://127.0.0.1:8080/v1
  • ⚠️ Use 127.0.0.1 instead of localhost on Android — they may not resolve the same way

External or LAN Servers

If your AI backend runs on a different machine (home server, another PC), bind it to 0.0.0.0 and enter its LAN IP in Settings — e.g. http://192.168.1.100:11434 for Ollama or http://192.168.1.100:8080/v1 for Llama.cpp.


🎭 Personality Presets

Prefrontal includes 4 built-in personality presets that change the system prompt and temperature simultaneously. Switch between them on the welcome screen or in Settings.

| Preset | Temp | Best For | |---|---|---| | ⚖️ Balanced | 0.7 | General tasks, Q&A, learning | | 🎨 Creative | 1.1 | Writing, brainstorming, storytelling | | 🎯 Precise | 0.2 | Facts, summaries, concise answers | | 💻 Developer | 0.3 | Code review, debugging, technical docs |

You can also write a fully custom system prompt — edit the System Prompt box in Settings, and the preset selector switches to "Custom" automatically.


🌡️ Temperature Control

Temperature controls how random or creative the AI's outputs are.

| Value | Behavior | |---|---| | 0.0 | Fully deterministic — same input produces the same output | | 0.2–0.4 | Precise and focused | | 0.7 | Balanced (default) | | 1.0–1.2 | Creative and varied | | 1.5–2.0 | Wild, experimental, sometimes incoherent |

[!TIP] Temperature is sent with every message — just adjust the slider, save, and it applies to the next generation. No restart needed.


🔎 Web Search (OpenAI-compatible runtimes)

When you're running on an OpenAI-compatible runtime—Llama.cpp, OpenRouter, OpenAI, Groq, or Together AI—Prefrontal can let the model search the live web before it answers. This isn't a provider plugin; Prefrontal implements it itself, on top of the free DuckDuckGo Instant Answer API, so it works with any compatible model.

How to enable it

  1. Settings → set Runtime to Llama.cpp / OpenAI-compatible, OpenRouter, OpenAI, Groq, or Together AI.
  2. A new Web Search toggle appears (it is hidden for Ollama). Flip it on.
  3. Save. That's it — every message you send from then on can trigger a search when the model decides it's useful.

How it works Turning the toggle on quietly appends a short instruction to the system prompt, telling the model that if it wants to search, it should reply with only a small JSON object — {"search_query": "..."} — instead of a normal answer. Prefrontal watches for that JSON: if a reply matches it, nothing is shown to you yet (you'll see a "🔎 Searching the web for…" status instead of raw JSON), Prefrontal queries DuckDuckGo directly from your browser, and feeds the results back to the model as a follow-up message so it can write the real answer. Any pages the model drew on come back as clickable source chips underneath the final reply. This whole exchange (search request → DuckDuckGo → real answer) happens automatically within a single one of your messages — capped at two search rounds so a stubborn model can't loop forever — and only the final answer is saved to your chat history.

[!NOTE] Web search applies to OpenAI-compatible runtimes only and only while the toggle is on. Ollama is unaffected because the feature is not enabled for that runtime. Because this uses DuckDuckGo's free Instant Answer API rather than a full search index, it's strongest for facts, definitions, and well-known topics, and can come back empty for very narrow or breaking-news queries — the model is told to just answer from its own knowledge when that happens. Since this is prompt-based rather than a model-native tool-call feature, it depends on the model actually following the instruction; most capable instruction-tuned models handle it reliably, but very small/free models occasionally ignore it.


🎨 Features

| Feature | Details | |---|---| | 💬 Chat | Full conversation history with any local AI model | | 🔄 Streaming | Real-time, token-by-token generation | | 🎭 Personality Presets | 4 built-in modes with one click | | 🌡️ Temperature Control | Live-sent with every request | | 🔎 Web Search | OpenAI-compatible runtimes: prompt-driven DuckDuckGo search with clickable source citations | | 📚 Multi-Backend | Ollama, Llama.cpp, OpenRouter, OpenAI, Groq, and Together AI | | 🌐 LAN & External | Connect to dedicated AI servers on your network | | 💾 Multi-Chat | Unlimited saved local conversations | | 🔍 Search | Instant conversation search | | 📱 PWA & Mobile Ready | Install as an app on iOS/Android; safe-area aware | | 🎨 Refined UI | 4 themes: Dark, Midnight, Emerald, Light | | 🧑 Local Profile | Device identity stored locally — no accounts | | 📤 Export | Export chats as Markdown, profile as JSON | | 🤖 Agent Chat (/agent) | drive the paired Prefrontal Agent from the chatbox; progress + permissions stream inline | | ✅ Interactive Questions | the agent's ask_user tool renders selectable options in chat (mouse or ↑/↓ + Enter) | | 🔗 Auto-connect | agent auto-pairs in the background via a localhost shared secret — no pairing token | | 🔒 100% Private | Zero network calls except to your own model server |


🤖 Prefrontal Agent (coding agent)

Prefrontal also ships a coding-agent panel: pair with the Prefrontal Agent — a separate, zero-dependency local agent that runs commands inside an isolated Podman sandbox and keeps filesystem access scoped to an approved workspace — and drive it from this UI. Output streams live, permission prompts (e.g. network) appear inline, and the agent dials out to this server, so no ports need forwarding.

# 1. Install the agent (separate repo, zero npm deps)
git clone https://github.com/sidx1-scratch/prefrontal-agent
cd prefrontal-agent
./setup.sh
npm link                   # registers the `prefrontal-agent` command globally
#    (optional) build the sandbox image: prefrontal-agent sandbox build

# 2. Run it — it auto-connects to this server in the background
prefrontal-agent           
#    Auto-pair happens automatically because this server writes a local
#    shared secret to ~/.prefrontal-agent/shared-secret when it starts.
#    Original manual flow still works:
#      Open the Agent panel → Pair new agent → copy the token,
#      then:  prefrontal> pair <token>

# 3. Drive it from the panel, e.g.
#    run npm test | write src/app.js⏎content | list src
#    ...or from the chatbox:  /agent create a blog project and publish it

The integration adds no dependencies to this project — the relay in server.js uses only Node built-ins and agent.js is plain vanilla JS.

On the same machine the agent auto-pairs in the background: this server writes a random shared secret into ~/.prefrontal-agent/shared-secret at startup (mode 0600), the agent reads it and dials the backend automatically. Auto-pair is localhost-only and can be disabled with PREFRONTAL_NO_AUTO_PAIR=1.

See docs/agent-integration.md for the full guide, security model, and relay protocol.


⌨️ Keyboard Shortcuts

| Shortcut | Action | |---|---| | Ctrl + N | New conversation | | Ctrl + , | Open Settings | | Enter | Send message (configurable) | | Shift + Enter | New line in input | | Escape | Close modal |


🔐 Privacy & Data

Running Prefrontal locally (Ollama or Llama.cpp), no chat data or API keys leave your machine — the app makes network calls only to the model server you configured, which is also on your machine or LAN.

If you use the OpenRouter runtime instead, your messages are sent to OpenRouter's API to be routed to whichever model you picked, so at that point you're trusting OpenRouter and the model provider with your prompts — the same as using any other cloud AI service.

[!WARNING] The hosted live demo is a special case: it only works with OpenRouter, and it stores your API key in a browser cookie rather than anywhere on a device you control. Treat it as a way to preview the UI, not as your daily driver — for real use, run Prefrontal locally.


🔧 Troubleshooting

❌ "Cannot connect to server"

  • Make sure your AI backend is running before opening Prefrontal.
  • Ollama: run ollama serve in a terminal.
  • Llama.cpp: confirm llama-server is running with the correct port.
  • Verify Settings → Server URL matches your backend exactly (including /v1 for Llama.cpp).
  • On Android, use http://127.0.0.1localhost may not resolve correctly.

❌ "Model not found"

  • Ollama: run ollama list to see installed models; copy the exact name including its tag.
  • Llama.cpp: the model name is the .gguf filename (or -hf repo name) you passed to llama-server.
  • Click Refresh in Settings to auto-populate the model list.

❌ Streaming stops or output is garbled

  • Lower the Context Window setting — the model may be running out of memory.
  • Try a smaller model (e.g. gemma3:4b instead of gemma3:12b).
  • Restart the backend server and try again.

❌ Temperature changes have no effect

  • Temperature applies to the next message only — not retroactively.
  • Make sure you click Save Settings after adjusting the slider.
  • Some models enforce a minimum temperature internally; very low values may behave like 0.0.

❌ App opens but is blank

  • Confirm you ran npm start, not npm run build.
  • Try npx serve . as an alternative server.
  • Open the browser console (F12) and check for errors.

npm install fails or hangs after downloading a release zip

  • Make sure you extracted the entire archive before running commands inside it — a partial unzip is a common culprit.
  • Confirm Node.js v18+ is installed: node -v.
  • Delete node_modules (if present) and run npm install again.

🤝 Contributing

Issues and PRs are welcome at github.com/sidx1-scratch/prefrontal. If you spot a bug or have a feature idea, open an issue or comment in a discussion before submitting a large PR so the approach can be discussed first.

Contribution Rule

Prefrontal is a zero‑build, zero‑dependency project. You're welcome to contribute, but all changes must keep the app lightweight and fully runnable as plain HTML/CSS/JS. No TypeScript, no frameworks, no bundlers, no compilation — ever. The project must remain simple enough that anyone can open the folder and understand it.


📄 License

Prefrontal is released under the GNU GPLv3 license.


“Built with cloud AI. Designed for local AI.”