@gotcos/glasses-server
v6.57.1
Published
COS Glasses — self-hosted AI heads-up-display server for Even G2 smart glasses, powered by Claude Code, Codex, Cursor Agent CLI, or local Ollama
Maintainers
Readme
COS Glasses Server
Self-hosted AI heads-up display for Even G2 smart glasses. Runs on your Mac, talks to your local Claude Code, Codex, or Cursor Agent CLI, and pushes answers, voice transcription, and notes to the lens. Your data never leaves your machine, and no API key is pasted into the phone for chat.
Quick start
npx --yes @gotcos/glasses-server@latestFor the optional COS Control macOS menu bar app, run the same non-mutating readiness check it uses before guided installation:
npx --yes @gotcos/glasses-server@latest --prepare-onlyCOS Control then installs the same npm package as a launchd-managed runtime. The original foreground command remains supported and unchanged.
Normal server start checks Node, finds your CLI, checks voice and image
processing, writes ~/.cos-glasses/.env, and starts the server on
0.0.0.0:3141. Optional Whisper and Kokoro models are provisioned when their
local services start; --prepare-only intentionally does not download or
install optional models, write COS configuration, or start a listener. It may
invoke an installed agent CLI's read-only version/auth probe, and that CLI may
maintain its own user cache. On boot the server prints
an API token — paste that into the COS Glasses app. Only one COS Glasses
server may run on a Mac at a time; a second npx or source runner exits before
opening ports or touching shared conversation/media state. Version 6.6.0 also
gives that server a durable identity and boot-scoped display replay, allowing
build 188+ to reconnect after a Tailscale, Wi-Fi, or process interruption
without silently losing completed replies.
Requirements
- Node.js 20.11+ — https://nodejs.org
- Claude Code CLI (Opus/Fable/Sonnet). Claude Desktop alone does not install
the terminal command. Install it on one line with
npm install -g @anthropic-ai/claude-code(never withsudo), then runclaudeand finish the browser sign-in or Codex CLI (GPT Frontier/Balanced) — https://developers.openai.com/codex/, thencodex login - Optional: Cursor Agent CLI for Composer 2.5 Fast and the newest Grok high-fast.
Ensure
agentis onPATH, runagent login, and verifyagent modelslistscomposer-2.5-fastand acursor-grok-*-high-fastid. COS mapscursor-grokto the newest high-fast it finds; it never silently substitutes Claude or Codex. - Even G2 glasses + the COS Glasses app from the Even Hub
brew install whisper-cppfor free local voice (the launcher can download the model)- Optional:
brew install [email protected] ffmpeg poppler espeak-ngfor local Kokoro spoken replies on Apple silicon (Python 3.11-3.12 is supported).ffmpegalso enables photo/video attachments;popplerenables PDF text and page previews. TXT, Markdown, CSV, and JSON attachments need no extra tool. Text chat remains available without these optional dependencies. - Optional: Tailscale so your phone reaches your Mac from anywhere
No provider API key is needed for chat when using signed-in CLIs. Usage is billed to the corresponding Claude, Codex, or Cursor subscription. Pick a provider per query, or set a default with
COS_G2_DEFAULT_MODEL(opus|fable|sonnet|codex-frontier|codex-balanced|cursor-grok|cursor-composer|ollama). Claude tier aliases and the two GPT slots resolve dynamically, so new model releases do not require a new glasses package. GPT discovery refreshes every 15 minutes and retains its last-known-good catalog through transient failures. Cursor discovery also refreshes every 15 minutes and retains its last-known-good catalog through transient failures. Cursor Agent mode can edit files and run shell commands in the selected workspace. Choose Ask mode when you want a non-editing answer; clients that omit the execution mode default to Ask. ExistingCOS_CODEX_MODEL/COS_CODEX_REASONING_EFFORTsettings remain supported on the migrated Frontier slot; leave them blank for auto-latest. Codex runs sandboxed read-only by default. SetCOS_CODEX_SANDBOX=workspace-writefor workdir writes + outbound network (sandbox_workspace_write.network_access=trueis passed by the managed server and should also be set in~/.codex/config.tomlfor interactive Codex). Local Ollama is a fourth picker, shown only whenollama serveanswersGET http://127.0.0.1:11434/api/tagswith a pulled model. DirectPOST /api/chat— not Codex--oss. Optional pin:COS_OLLAMA_MODEL.COS_CODEX_EXTRA_ARGSis a Codex CLI hatch, not the Ollama UX. See Configuration. Claude is the most permissive provider by default. It runs with--dangerously-skip-permissions, so a glasses query on the Claude/Opus path can run shell commands and read, edit, and write files on this Mac without prompting you. That is what makes the glasses useful for real work, and it has been the behavior for some time — but as of 6.18.3 the model is also correctly told it has those tools, so you will see it use them more readily than before. SetCOS_CLAUDE_TRUST_MODE=allowlistto remove Claude's permission bypass and restrict it to COS's explicit per-query tool allowlist; undeclared tools then fail closed without prompting. In allowlist mode the query keeps web search/fetch and read-only workspace access (Read, Glob, Grep) — no shell, no edits, no writes. Only the exact valueallowlistrestricts anything — any other value logs a warning and stays trusted. Servers before 6.41.0 denied ALL workspace reads in allowlist mode; if a hardened install answers "I don't have access to your workspace files", update the server.
Connect your phone (the one gotcha)
The glasses app runs on your iPhone and must reach this server on your Mac.
- The launcher binds
0.0.0.0(all interfaces) for you. - Same WiFi (simplest): find your Mac's LAN IP (System Settings > Wi-Fi > Details), and in the COS Glasses app enter
http://192.168.x.x:3141. - From anywhere: install Tailscale on the Mac + iPhone (same account), note the Mac's
100.xaddress, and enterhttp://100.x.x.x:3141. - Either way, paste the API token the server printed at boot.
To restrict the server to localhost only, set BIND_HOST=127.0.0.1 in ~/.cos-glasses/.env.
The built-in IP allowlist blocks public-internet traffic regardless. Its mesh
range is the exact Tailscale/CGNAT allocation (100.64.0.0/10), not all of
100.0.0.0/8; RFC1918 LAN ranges remain supported.
What it does
- Ask anything, get a streamed answer on the lens (
/api/query,/v1/chat/completions) - With COS Glasses build 204+, server-owned durable queries are on by default: accepted work survives phone backgrounding, WebView reloads, and network handoffs, then reattaches without duplicate work or duplicate replies
- Choose Opus, Fable, Sonnet, GPT Frontier, GPT Balanced, Composer 2.5 Fast, or the newest Grok high-fast. Cursor slots fail closed when the local CLI or concrete model is unavailable; optional redacted tool activity streams only to the authenticated query that requested it
- Message History + cross-day "reference message N" — your chats are archived by day
and every message keeps a permanent number you can recall (
/api/archive,/api/message/:num) - Recent/history responses preserve validated photo references. Recovery uses exact session + global-message + message-era identity without exposing storage paths. Ambiguous pre-version historical refs fail closed; unversioned refs are recovered only inside the active era when created and associated after its boundary.
- Send phone photos with queued prompts, and review assistant-selected generated, research, or explicitly used email images in Messages and on the G2 lens
- Recover long voice prompts after phone, network, or server interruptions. Audio chunks are saved before transcription and retained locally for 72 hours. On compatible app builds, their warm transcript also appears live while speaking; final HQ transcription remains authoritative.
- Live voice capture + transcription during meetings
- With COS Glasses build 209+ and server 6.11.0+, meetings continue recording locally through a network interruption. Reconnecting reconciles the exact chunks already stored by the Mac, uploads only missing audio, and finalizes through an idempotent save receipt without duplicating the meeting.
- Since 6.19.0, meeting audio whose save never lands is quarantined for 72 hours
(
COS_UNSAVED_AUDIO_RETENTION_HOURS) instead of being cleaned up, surfaces on/api/healthasunsaved_captures, and can be recovered into a durable meeting scribe with one authenticated call (POST /api/meeting/orphans/:sessionId/recover; list viaGET /api/meeting/orphans). - Local whisper.cpp transcription (free and local-only by default). OpenAI
Whisper fallback is optional and requires both the exact
COS_OPENAI_WHISPER_FALLBACK=1opt-in and a configured key; a key alone never uploads audio. - HQ prompt dictation is requested by default; the phone's Fast mode switch opts into turbo. Server 6.16.0 reports whether full local large-v3 actually ran, and compatible companions alert once if an HQ request used Fast or Cloud instead of silently claiming HQ.
- Local-first spoken reply playback through Kokoro on Apple silicon. The first
use creates a private Python environment and downloads its model without
blocking the API. Selecting Local fails closed;
local_firstcan fall back to OpenAI TTS only when a key and budget are available./api/healthreports the independenttts_localstate and current engine. - Tasks / calendar / people context if you run the
COS Starter Kit (
COS_SCRIPTS_DIR); otherwise it is glasses + AI only - Welcome weather on the glasses home screen via authenticated
GET /api/welcome-context?lat=&lon=. The phone supplies GPS (Even Hub location permission); this server only proxies Open-Meteo. Without phone coords the route uses last-known process coords, then optionalCOS_WEATHER_DEFAULT_*, otherwise omits weather. OptionalnextEventappears whenCOS_SCRIPTS_DIRcalendar data is available.
Session hooks (6.48.0)
Claude Code can tell the server what each session is doing (started, prompt submitted, waiting on a permission, turn stopped, ended) through its own hooks. Install them once:
npx --yes @gotcos/glasses-server@latest --hooks install --dry-run # shows the merge, writes nothing
npx --yes @gotcos/glasses-server@latest --hooks install # merges into ~/.claude/settings.json
npx --yes @gotcos/glasses-server@latest --hooks status
npx --yes @gotcos/glasses-server@latest --hooks uninstallThe install keeps every hook you already had, backs the file up, and copies a small
POSIX sh script to ~/.cos-glasses/bin/cos-session-hook. The script writes one file per
event into ~/.cos-glasses/data/hook-spool. It contacts the server in exactly two cases:
a permission request, only after the desk has been idle for 90 s (see the 6.52.0 note
below), and a Cursor composer's Stop, only when a turn is queued for it (6.51.0). Every
other event is a file write, so a server that is down or restarting never delays a
session. Sessions report state_source: hook on /api/agent-sessions and
/api/claude-sessions from their next event on (Claude Code 2.1.272 reloads its hooks
when the settings file changes, so open tabs need no restart; they may show Claude's
"hooks modified externally" notice once, which is expected: the user-level file changed
under them). A later COS Control offers the same install from its Sessions tab. Rows
change only while the server runs with COS_CLAUDE_SESSIONS_ENABLED=1
(--hooks status prints serverApplies). Turn the ingestion off with COS_SESSION_HOOKS=0
(the spool is still drained and stamped, rows are exactly as before). Removing the package
does not remove the hooks: run --hooks uninstall first.
Since 6.48.1 the hooks also drive the follow-up queue (a queued Continue lands when the
engine closes the turn: its Stop hook followed by its registry record flipping idle,
which is when every Stop hook has returned), the attach gate (a Desktop holder whose
turn just ended reads idle at once), and the live session stream (every status draft
carries agent_state and friends; ?after=<epoch>.<cursor> replays what the ring holds
after a reconnect and seeds afresh when it holds nothing, ?seed=turn opens at the
current prompt). COS_SESSION_HOOK_SSE=0 omits the extra status fields and the live
state drafts; the cursor, epoch and id: line stay.
Since 6.49.0 a Continue can land in the running session itself, so the Desktop tab or
the claude in a terminal that you left open shows the turn and answers it with its own
context, instead of a resume child writing to the transcript behind that window. It is
off until COS_CONTINUE_LIVE=1. The server writes the turn to the session's own inbox
(the socket Claude Code publishes in ~/.claude/sessions/<pid>.json, in the line format
its help text documents), pinned to the full session id, and calls it delivered only
when the session's transcript shows the message accepted; every other outcome takes the
6.48.2 path unchanged. A frame written but not seen accepted holds the turn as
live_unverified (retryable) and is never written twice; only after 60 s without an
acceptance row does the resume child run. /api/health reports continueLive
(enabled, attempts, delivered, fallbacks by reason). The live stream also carries a
tool_outcome on the status draft that follows each tool result (+14 -2,
235 lines, 41 lines, exit 1), which a client that predates it takes as a no-op.
Where to set the flag: ~/.cos-glasses/.env (read at boot; the plist wins when
both carry it). COS Control's allowlist does not carry COS_CONTINUE_LIVE through
0.5.233, so a value set only in the LaunchAgent plist is dropped by the next
Install/Repair/Update Server; Control 0.5.234 is to add it.
Since 6.49.1 a lease yields: a completed, unpinned Continue binding stops holding its
thread 90 s after its turn ends (it held it for the whole 30 min TTL before), so a
follow-up from another surface, or the queue drainer, attaches at once instead of
waiting out the lease; a binding with a turn in flight still refuses native_target_busy.
The turn ledger is read across every binding of a thread, so a draft re-sent through a
new binding (the phone attaches per send; a parked draft drains through the drainer's
own) replays rather than repeats.
GET /api/agent-sessions/:provider/:id?turns=N (6.50.0, N from 1 to 40) adds
recent_turns (oldest first, { role: 'user' | 'assistant', text, at? }, tools and
sub-agents omitted) and recent_turns_more to the detail payload, read backward from
the end of the transcript; without turns the payload is unchanged. The glasses use it
to scroll back through a running session's conversation.
Since 6.52.0 a session question (the AskUserQuestion card) or a tool approval can be
answered from the glasses or the phone while you are away from the Mac. There is nothing
to reinstall: the hook script is the 6.51.0 one, and the server now answers the URL it
already calls, POST /api/permission-requests/ask (hook-token auth, checked before the
body is read; that one method and path is exempt from the API token, nothing else is).
The hook posts only after the desk has been idle 90 s; the server holds the request only
while the desk stays idle and a client that can answer keeps polling, and answers {}
(the Mac's own dialog) to everything else at once. A client that never polls (COS
Glasses 6.9.511 and earlier, COS Control) changes nothing.
The client contract (also at /api/models capabilities.sessionQuestions):
- Poll
GET /api/session-questions?client=glasses(orclient=phone) with X-Cos-Token everypollIntervalMs(10,000). Only those twoclientvalues count as a live answerer; a poll withoutclient, or with any other value, is answered but never counts. A client quiet forliveWindowMs(30,000) hands every held request back. - The response carries
enabled,mode,protocolVersion(1),pollIntervalMs,liveWindowMs,pendinganditems(a question'squestions, or an approval'sapproval:{tool, summary, detail}, the command, path or URL exactly as it will run, cut at 160 and 4,000 characters with...(+N chars), with only secrets redacted). - Answer with
POST /api/session-questions/:id/answer:{clientAnswerId, answers}for a question (one{labels, other}per question; free text with a comma is sent quoted),{clientAnswerId, decision}for an approval (allowordeny). The first answer wins; a retry with the sameclientAnswerIdreplays it. 409already_answered,handed_to_deskorexpired, and 410hook_gone, say why an answer came too late; 404not_foundafter a server restart or once a settled item is pruned (10 minutes); 503 during a maintenance drain, when every held item has already gone back to the Mac.
Touching the Mac hands a held request back to its dialog at once; the deadline is 110 s.
Allow once or deny only: no permission rule is ever written. Rows carry
pending_question_id or pending_permission_id while a request is held, and a question
row reads waiting_kind: question whatever the switch says; /api/health reports
permissionBroker (counters and timestamps, no ids or text).
Cancel a run from the glasses (6.53.x)
The glasses' Cancel run row stops what a session is doing, the way the desk's own stop
button does. It needs server 6.53.0 or later and COS Glasses 6.9.529 or later;
6.53.4 is recommended (it owns the whole process tree, holds a desk cancel, fixes
the false fences 6.53.1 left, and never releases a fence when ps cannot answer).
- A turn COS started (
cos_turn) is stopped by the server: its whole process tree is signalled and the turn settlesturn_cancelled. - A run at the desk (
desk_run: a Desktop tab or a terminalclaude) is stopped by the session hook at its next tool call, so a reply that is pure text finishes first. This needs the hooks installed, and after every server update that changes the hook script (6.53.0, 6.53.3, 6.53.4) you must run Install hooks once (COS Control, ornpx --yes @gotcos/glasses-server@latest --hooks install). Until then--hooks statusreadsscript_outdatedand itsadviceline says so; a script from 6.53.0 to 6.53.3 can still stop desk runs meanwhile, anything older answershooks_outdated. The PreToolUse hook runs before every tool call on the Mac and never contacts the server; it writes a spool file only for AskUserQuestion, ExitPlanMode and prompts, and always drains its input. - Codex and Cursor runs at the desk answer
cancel_unsupported("Stop it there"). - Fences a cancel left release themselves once every process the cancelled turn
started is confirmed gone: only
psanswering "no such process" counts, and apsthat times out, cannot start or is missing keeps the fence. Any other fence is still released by hand (GET /api/agent-sessions/fences, thenPOST .../fences/releasewithconfirm: true). - A cancel that lands while COS is handing a turn to the open session writes the desk marker, and re-arms it once if the next prompt that session starts is exactly the delivered one. If your own prompt reaches the session first, the re-arm is dropped (and logged) rather than stop your run.
Meetings: a closed meeting's record ages out after 4 hours, after which its status
reads missing and the phone stops asking to save it. Saving a known meeting that has no
transcript answers 404 nothing_to_save (with its state); only an unknown id answers
session_not_found.
The summary behind the prompt box (6.54.0)
When you hold the ring on a finished session or message, COS Glasses 6.9.545 floats your
live words over a dimmed card of what you are replying to. From 6.54.0 the server writes
that card: Outcome (Answer for a message), So what and You asked, one lens row
each, instead of the reply's first lines. Older phones and older servers keep the
6.9.544 card.
Longer lines, and meetings (6.55.0). A phone with more room can say how long each line
may run, with "chars": {"first": 100, "soWhat": 100, "third": 50} in the request (COS
Glasses 6.9.546 asks for these, giving the first two lines two rows each). Each number must
be a whole number and is held to its bounds: 20 to 150 for the first line and So what, 20
to 80 for the third. Anything else is a 400. Without chars, or with null, the lines keep
their 6.54.0 lengths and the prompt is word for word what 6.54.0 sent, so 6.9.545 sees no
change. The budget is part of the cache key, so a card written long is never handed to a
phone that asked for a short one.
"kind": "meeting" writes a meeting's card from the meeting's own notes (title, time,
summary, decisions and action items), sent as reply with no ask: Gist (what the meeting
settled or covered, returned as outcome), So what (your next steps, left empty when
nothing is needed) and Open (one question still unresolved, returned as open and left out
when nothing is open). An ask sent with a meeting is ignored, and the notes are data the
model is told not to follow, the same as a reply.
Pick the engine the way the COS indexer does: COS_LENS_GIST_ENGINE wins, then the
choice saved with PUT /api/lens-gist/config, then the default. One engine never falls back
to another; off turns the summary off.
Or chain them, for resilience. claude:sonnet,codex:gpt-5.6-terra asks Claude first and
Codex when Claude fails, is cooling down, or says its session or usage limit is spent (a
limit, read from the provider's own error and never from an answer, skips that engine for 30
minutes at once). Running low on one plan no longer turns the summary off. A chain gets 70
seconds in all, shared in order. It is opted into, never the default: it sends a session's
reply to every provider in it, and an installed CLI is not consent (ChatGPT.app ships
codex). PUT takes either engine or chain, never both, and a saved off wins.
| Engine | Default model | Measured (2026-09-25, one call) | Tools it keeps |
|---|---|---|---|
| claude (default) | sonnet | 2 to 5 s, about 1.2k tokens | none (--tools "", hooks off) |
| codex | gpt-5.6-terra (gpt-6-luna, gpt-6-sol work as well) | 3.2 to 3.7 s, about 11.4k tokens | shell, browser, image, agent, app and plugin tools off; web search off |
| cursor | grok-4.7-low-fast | 11 to 13 s, about 11k to 17k tokens | ask mode: read-only tools in an empty workspace |
| ollama | the local model lens queries use | 2.5 s warm, 14 s cold | none |
With nothing chosen and no Claude CLI installed, the summary is off rather than sent to a
provider nobody picked. Claude haiku timed out at 30 s in testing and is not a default.
curl -s -X PUT -H "X-COS-Token: $COS_API_TOKEN" -H 'content-type: application/json' \
-d '{"chain":"claude:sonnet,codex:gpt-5.6-terra","dailyCap":150}' http://127.0.0.1:3141/api/lens-gist/config
curl -s -X PUT -H "X-COS-Token: $COS_API_TOKEN" -H 'content-type: application/json' \
-d '{"engine":"codex","model":"gpt-6-sol"}' http://127.0.0.1:3141/api/lens-gist/config
curl -s -H "X-COS-Token: $COS_API_TOKEN" http://127.0.0.1:3141/api/lens-gist/configGET /api/lens-gist/config lists every engine, whether it is installed, and the models it
can run here. The saved choice lives in <data>/lens-gist.json and survives Update
Server; a file that cannot be read turns the summary off until it is saved again.
How it stops. The phone asks once per finished reply it has open (a session that is
still running or waiting on an approval asks for nothing). Answers are cached on disk per
engine and model. A daily cap (COS_LENS_GIST_DAILY_CAP or the saved dailyCap, default
150) counts answers, and attempts stop at twice that. A reply that failed is refused for 10
minutes, then 6 hours (a spent provider limit does not count against the reply). Three
failures in a row, or one limit answer, pause that engine for 30 minutes, then one trial call
runs. At most 2 calls run at once with 4 waiting. A drain for Update Server stops queued
calls before each engine. Every call writes a g2-lens-gist row to the token audit and a line
to <data>/lens-gist-runs.jsonl (with fallbackFrom when a later engine answered).
/api/health shows lens_gist (engine, model, calls today, cap, open breakers; no error text
and no chain); GET /api/lens-gist/config shows the chain.
Configuration
Config lives at ~/.cos-glasses/.env (created on first run). Every key is
optional except an installed CLI. Highlights: BIND_HOST, PORT,
COS_API_TOKEN (auto if unset), COS_OPENAI_WHISPER_FALLBACK=1 plus
OPENAI_API_KEY (explicit cloud transcription/TTS fallback),
COS_TTS_ENGINE (local_first or openai_primary),
COS_TTS_KOKORO_VOICE (local voice id),
COS_TTS_LOCAL_DISABLE=1 (disable the sidecar), and
COS_TTS_PRONUNCIATIONS_JSON (optional local/cloud pronunciation overrides),
COS_EXTRA_TOOLS (comma-separated mcp__server__tool or
mcp__server__* selectors shared by full and lightweight Claude paths),
COS_CLAUDE_MCP_CONFIG (optional absolute config path when .mcp.json is not
in the managed CLI working directory),
COS_CURSOR_AGENT_BIN (optional absolute Cursor agent binary),
COS_CURSOR_PERSIST_SESSIONS=0 (disable Cursor session resume),
COS_CODEX_SANDBOX=workspace-write (workdir writes + outbound network on GPT),
COS_OLLAMA_MODEL (optional pin; must match a name from ollama list),
COS_OLLAMA_HOST (optional loopback origin, default http://127.0.0.1:11434;
non-loopback is refused),
COS_CODEX_EXTRA_ARGS (Codex exec operator hatch — not the Ollama picker.
--oss --local-provider ollama --model qwen2.5-coder still retargets GPT
Frontier spawn. Put it in ~/.cos-glasses/.env; a plist-only value is dropped
on Update Server. G2 Codex exec only — Sessions Continue / Fork ignore it.
No inline # on the EXTRA_ARGS line.),
COS_WEATHER_DEFAULT_LAT / COS_WEATHER_DEFAULT_LON /
COS_WEATHER_DEFAULT_CITY (optional home fallback when phone GPS is denied),
COS_SCRIPTS_DIR (full pipeline), COS_DURABLE_QUERY_JOBS=0 (optional
machine-wide rollback for build 204+ server-owned query recovery),
COS_MESSAGES_TRAIL=0 (6.52.0: no Messages trail; every job stream and snapshot is the
6.51.0 one), COS_PERMISSION_BROKER (6.52.0: 0 answers every permission request with
the Mac's own dialog, questions holds questions but not tool approvals; unset holds
both), COS_PERMISSION_BROKER_DESK_IDLE_S (default 90, never below 30) and
COS_PERMISSION_BROKER_TIMEOUT_S (default 110, clamped to 5 through 120),
COS_MEDIA_ROOT (optional image/video store location; default
~/.cos-glasses/data/media), and COS_VIDEO_UPLOAD_V2=1 (private 6.27.3+
resumable-video canary, managed by COS Control 0.5.20). The V2 canary retains
accepted original chunks (1 MiB on new sessions; leftover 256 KiB drafts keep
that size) and finalize receipts across restarts; keep it off when
using an older companion. Your name + transcription vocabulary live in
~/.cos-glasses/.cos-profile.json (see .cos-profile.example.json).
Factory example values are ignored; add the real names, companies, acronyms,
and specialist terms you say often. Guided Setup writes a safe empty profile
instead of biasing Whisper toward placeholder text.
Telegram activity export is disabled by default even when a private COS
pipeline contains .telegram_config.json; enable it only with the explicit
COS_TELEGRAM_NOTIFICATIONS=1 opt-in.
Put rollback switches such as COS_MESSAGES_TRAIL=0 and COS_PERMISSION_BROKER=0 in
~/.cos-glasses/.env: that file survives Update Server. COS Control 0.5.239 also keeps
them in the LaunchAgent environment; 0.5.238 drops a value set only in the plist.
Morning brief (6.43.0)
A start-of-day brief runs on a schedule inside the server and waits in the
inbox as a numbered reply. Default: weekdays at 07:00 in the Mac's timezone,
with Calendar, recent-meeting decisions, tasks due this week, and what is
waiting on you. Turn on more sources (knowledge graph, reflection, health, an
opening reading, a metrics pulse, one of your own skills such as
/good-morning, a custom section), reorder them, and set their windows from
COS Control or the companion, or directly:
curl -H "X-COS-Token: $COS_TOKEN" http://127.0.0.1:3141/api/morning-brief
curl -H "X-COS-Token: $COS_TOKEN" -X PUT -H 'Content-Type: application/json' \
-d '{"time":"06:30","sources":[{"id":"skill","enabled":true,"options":{"name":"/good-morning"}}]}' \
http://127.0.0.1:3141/api/morning-brief
curl -H "X-COS-Token: $COS_TOKEN" -X POST http://127.0.0.1:3141/api/morning-brief/runOne provider run per local day, remembered in a ledger so a restart never doubles it; a Mac asleep at the slot still fires inside a three-hour catch-up window; "Run now" is capped at five a day. Off entirely when Background jobs are off. The brief is read-only by contract.
Since 6.43.1 the same GET /api/morning-brief response carries coverage,
one row per source saying what the server can see behind it (meetings stored,
memories and threads, whether the named skill exists) and every run reports
which sections the answer actually opened. Numbers reserved by a running
brief count toward /api/message-counter, so the phone never mints the same
#NNN twice.
Speaker diarization (opt-in)
Without a voiceprint model this server does not classify speakers at all — it
passes through whatever label the client sends (Unknown when the client sends
nothing; the COS companion sends its own wearer/Ext labels). Named
per-speaker diarization needs a ~26 MB voiceprint model that is deliberately
not shipped in the npm package, so it is a bolt-on:
npx --yes @gotcos/glasses-server@latest --setup-speaker-modelThat downloads the model to ~/.cos-glasses/models/, verifies it against a
pinned SHA-256, and refuses to install anything that does not match. Restart the
server afterwards and check /api/health — speaker_id should read active.
To place it by hand instead:
mkdir -p ~/.cos-glasses/models
# put 3dspeaker_speech_eres2net_sv_en_voxceleb_16k.onnx there, then restartThe server searches, in order: COS_SPEAKER_MODEL_PATH (explicit full path to
the .onnx), ~/.cos-glasses/models/, then a bundled server/models/ copy
(source checkouts only). Use the data home, not the installed package —
anything inside the package is destroyed by the next update, while
~/.cos-glasses/ survives.
Use an absolute path if you set COS_SPEAKER_MODEL_PATH — a relative one
resolves against the working directory, which under the managed LaunchAgent is
the installed package.
Verify with /api/health → speaker_id:
| Value | Meaning |
|---|---|
| active | model loaded, diarization running |
| unavailable | no model found — labels come from the client |
| error | a model is present but the runtime rejected it (see the startup log) |
The model is read once at startup, so restart after adding it. A corrupt or
mismatched .onnx is screened and probed in a child process first, so a bad
download disables diarization instead of taking the server down — but it does
mean a wrong file fails silently apart from that log line.
The wearer's label comes from owner_speaker_label in
~/.cos-glasses/.cos-profile.json (default Me); set it to match the profile
name you enrol under. Train voices via /api/voice/enroll?name=… (the default
name is owner_speaker_label); profiles persist in
~/.cos-glasses/data/voice-profiles.json.
Upgrading from an older server: before this release the enrollment default
was hardcoded to MU, so an existing install may hold a profile under that name
while owner_speaker_label resolves to Me. /api/voice/status will then
report enrolled: false. Set owner_speaker_label to MU to keep the existing
voiceprints rather than re-enrolling, which would split the same voice across two
profiles.
HQ dictation
Prompt dictation defaults to HQ. The phone owns the preference: Fast mode OFF requests HQ, and Fast mode ON requests turbo. The Mac performs all decoding; the phone does not run Whisper.
For the recommended Balanced setup, run:
npx --yes @gotcos/glasses-server@latest --setup-transcription --transcription-tier balancedThat keeps three jobs separate: Small.en supplies provisional prompt words on
the lens, Large-v3-Turbo commits the authoritative live transcript, and
Large-v3 polishes saved prompts and meetings. Small.en never writes the
recovery ledger and receives no decoder-bias prompt.
If its sidecar is missing or unhealthy, preview falls back to Turbo without
changing final quality. Set COS_WHISPER_PREVIEW_MODEL=turbo to keep one live
model, or off to disable provisional peeks. Existing installs that only
update the server remain on Turbo until Guided Setup opts them into Small.en.
Max is an opt-in tier for powerful Macs:
npx --yes @gotcos/glasses-server@latest --setup-transcription --transcription-tier maxMax keeps Turbo resident in the isolated preview sidecar for low-latency provisional words, while Large-v3 remains authoritative for live commit and saved-work polish. Canonical transcription has strict GPU priority: a cosmetic preview is dropped or aborted instead of competing with a committed decode. If Large-v3 is missing, health reports the downgrade and the server falls back to Turbo rather than making transcription unavailable. COS Control is the supported owner of the machine-wide tier; the per-lane environment variables remain advanced overrides.
Server 6.21.8 adds a default-off meeting-completion canary. With
COS_MEETING_PROGRESSIVE_HQ=1, sealed meeting windows can be polished ahead of
Stop/save on a CPU-only, single-flight lane; finalization reuses only matching
audio/model/context checkpoints. Balanced is capped at two background threads
for fanless M1/M2 MacBook Airs, while Max defaults to six and remains capped by
available CPUs. COS_MEETING_EARLY_SYNC=1 separately gives the Operations sync
pipeline a stable meeting identity before HQ completes. Either switch can be
disabled without changing canonical live transcription or raw meeting audio.
Server 6.21.32 adds a separate default-off cleanup canary for retained review
audio. Set COS_MEETING_AUDIO_ADAPTIVE_PLAYBACK=1 (or use COS Control 0.5.11+)
to profile a retained PCM chunk and create a cached playback-only copy when a
reviewer presses Play. The raw WAV remains byte-identical, still owns the
seven-day retention clock, and is served on every analyzer/FFmpeg failure.
Capture, live preview, canonical transcription, speaker attribution, save, HQ,
and meeting sync are unchanged. Append ?raw=1 to an authenticated playback
URL for an immediate raw-versus-cleaned A/B check.
While a meeting is actively recording, playback automatically stays raw so the
optional cleanup process cannot contend with live transcription. Cleanup uses
one global worker, serves raw while that worker is busy, and preempts within
100 ms if a meeting starts after a replay request was admitted.
Server 6.21.33 lets Review Meetings browse an existing single-library tree such
as meetings/YYYY-MM/*.md. Set COS_MEETINGS_ROOT to the folder that directly
contains the month folders. It is intentionally read-only. For the full COS
sync and enrichment pipeline, keep using COS_OPERATIONS_DIR with
<domain>/meetings/YYYY-MM/*.md; arbitrary domain names are supported. When a
direct library and an operations root are both configured, the server merges
them with standalone G2 recordings and prefers the enriched writable record
for the same session.
Server 6.21.35 adds authenticated, read-only Memory and Threads browsing for
full COS installs. With COS_SCRIPTS_DIR configured, the companion can show the
complete Bot Memory count/type split, bounded recent summaries, exact logical
memory IDs, and existing tracked/manual threads. Exact detail requests are
resolved by stable ID so a spoken follow-up can carry the selected snapshot as
context. Embeddings, vector-store point IDs, cache files, secrets, and local
paths never cross the API boundary.
Memory and Threads from plain markdown (6.22.0)
You do not need a Python bridge, a virtual environment, or a vector database.
Make a folder with a memory/ or threads/ subfolder, put markdown files in it,
and point COS Data at it (or set COS_CONTEXT_DIR):
notes/
memory/ any nesting, any filenames
2026-08-09-hiring-call.md
decisions/pricing.md folder name becomes the type
threads/
website-rebuild.mdFront matter is optional and every field degrades rather than rejecting:
---
type: decision # else the containing folder name, else "note"
date: 2026-08-09 # else a YYYY-MM-DD in the filename, else file mtime
status: resolved # threads only
---
# Held Rain POS for v25
Body text. The first heading becomes the summary.Two tiers, and the bridge always wins when it is present:
| Tier | Requires | Provides |
| --- | --- | --- |
| Files | a folder of markdown | Browse, read, and reference memories and threads |
| Bridge | + venv, cos_api_bridge.py, vector store | Adds semantic recall, dedup, type statistics, retention |
/api/context/status reports source: "bridge" or source: "files" so a client
can say which tier it is showing. Fields a file-backed record cannot have are
empty rather than invented: a file thread has velocity: "", meeting_count: 0
and stale: 0 because nothing computed them.
Root resolution, first match wins: COS_CONTEXT_DIR (exclusive when set), then
COS_OPERATIONS_DIR, COS_MEETINGS_ROOT and its parent, the parent of
COS_SCRIPTS_DIR, then ~/.cos-glasses — so mkdir ~/.cos-glasses/memory is a
complete setup. The file tier is read from the code path taken only when no
bridge is configured, so adding it cannot change the behaviour of an install that
already has one.
Since 6.44.5 the bridge tier also serves recent learning and the knowledge graph,
read-only: /api/context/learning (events, cursor-paged, with a per-store
coverage map), /api/context/learning/status, /api/context/learning/:id, and
/api/context/graph/{status,search,entity,passages}, plus POST
/api/context/graph/index, which only asks the pipeline to start a detached index
build and answers 202. Nothing in the file tier can answer these, so they return
503 cos_pipeline_not_configured there; older servers 404 them, which is how a client
tells the versions apart. /api/context/status carries learning and graph
blocks when the bridge can produce them and omits them otherwise.
The API is read-only in both tiers. Standalone installs with neither a bridge nor a notes folder report the feature as unavailable without affecting messages, meetings, transcription, or agents.
The first server start downloads the real-time turbo model. True HQ additionally
requires the full ggml-large-v3.bin model (about 3.1 GB):
mkdir -p "$HOME/.local/share/whisper-models"
curl -fL --progress-bar \
https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-large-v3.bin \
-o "$HOME/.local/share/whisper-models/ggml-large-v3.bin.partial"
mv "$HOME/.local/share/whisper-models/ggml-large-v3.bin.partial" \
"$HOME/.local/share/whisper-models/ggml-large-v3.bin"Restart the server, then confirm
capabilities.transcription.hq.hqAvailable: true at /api/health. The response
does not expose local paths. If the CLI or model is unavailable, dictation stays
usable on Fast and reports the downgrade truthfully. Set
COS_HQ_SPECULATIVE_WARM=0 to disable background HQ warm immediately; set
COS_BATCH_LARGE_V3=0 to explicitly use turbo. Interactive HQ uses beam 2 by
default (COS_HQ_BEAM_INTERACTIVE); meeting batch remains beam 5.
Run from source
git clone https://github.com/ukaoma/cos-glasses-server.git
cd cos-glasses-server
npm install
BIND_HOST=0.0.0.0 npm run start:serverTroubleshooting
- Claude Desktop is installed but COS says Claude Code is missing — Desktop
and the terminal CLI are separate. Run
npm install -g @anthropic-ai/claude-codeon one line withoutsudo, then runclaudeand complete sign-in. Verify withclaude --versionbefore starting COS again. - npm reports EACCES or a root-owned cache — never run COS or npm with
sudo, and do not recursively change system ownership. Use a private COS cache instead:npm_config_cache="$HOME/.cos-glasses/npm-cache" npx --yes @gotcos/glasses-server@latest. Version 6.12.2+ never runs a second install from inside npm's temporary cache. - Phone can't connect — check
BIND_HOST=0.0.0.0, the same Tailscale account on both devices, and the correct100.xIP + token. - Safari connects but the app does not — confirm
npx --yes @gotcos/glasses-server@latestis 6.6.0+, then use the app's server reconnect/edit control to verify the current URL and token. Do not run a second source ornpxserver alongside it. - AI queries fail — run
claude auth status,codex login status, oragent statusfor the selected provider, then authenticate withclaude auth login,codex login, oragent loginwhen signed out. - Composer or Grok is missing — update to server 6.16.1+, confirm
agentis discoverable on the servicePATH, and runagent models./api/healthmust reportfeatures.cursor: true; authenticated/api/modelsmust include bothcursor-composerandcursor-grok. Missing models fail closed instead of falling through to another provider. - Voice getting billed? — voice is local-only by default in 6.12.0+. Confirm
/api/healthreportscapabilities.transcription.mode: "local-only". RemoveCOS_OPENAI_WHISPER_FALLBACK(or set it to0) to disable an earlier opt-in. - Local voice unavailable? — install
whisper-cpp, restart the server, and confirm/api/healthreportsfeatures.whisper: true. A typed retryable 503 keeps compatible prompt/meeting audio available for retry instead of silently sending it to OpenAI. - Local spoken replies unavailable? — on Apple silicon, install
[email protected] ffmpeg espeak-ng, restart the server, and wait for the first-run Kokoro model download. Confirm/api/healthreportstts_local.ready: true. If Python lives outside the normal Homebrew paths, set its absolute 3.11 or 3.12 path asCOS_TTS_BOOTSTRAP_PYTHONin~/.cos-glasses/.env. Selecting Local never falls back to cloud; setCOS_TTS_ENGINE=openai_primaryonly when OpenAI playback is intentionally configured. - Photos unavailable? — install
ffmpeg, restart the server, and confirm/api/healthreportsfeatures.mediaProcessingReady: true. - Video or PDF attachments unavailable? — install
ffmpeg poppler, restart the server, and confirm/api/healthreportsfeatures.videoProcessingReady: trueandfeatures.pdfProcessingReady: true. Uploads are limited to five items, 64 MiB each; videos are represented by up to eight bounded still frames and PDF/text contents are quoted as untrusted reference data rather than executable instructions. - Prompt recovery unavailable? — update with
npx --yes @gotcos/glasses-server@latest, then confirm/api/healthreportsfeatures.promptRecovery: true. - Durable query recovery unavailable? — build 204+ requires server 6.10.0+.
Restart once, then confirm
/api/healthreportsfeatures.durableQueryJobs: true, protocol1, and stateready. To roll back, setCOS_DURABLE_QUERY_JOBS=0; accepted jobs still drain while new prompts use legacy streaming. - Session questions never reach the glasses? The server holds one only when every
gate passes.
/api/healthpermissionBroker.lastFastPathnames the last reason a request went straight to the Mac's dialog:no_client(no poll withclient=glasses|phonein the last 30 s: the app build has no question cards, or its timers are frozen),desk_active(the Mac saw input),approvals_off(COS_PERMISSION_BROKER=questions),broker_off,cursor,unsupported_tool(ExitPlanMode is never held).lastQuestionsPollAtsays when a client last counted.--hooks statusmust readinstalled(right after updating to 6.53.4 it readsscript_outdateduntil Install hooks is run once; itsadviceline says so). To turn it off with no reinstall, setCOS_PERMISSION_BROKER=0in~/.cos-glasses/.envand restart. For the Messages trail,/api/healthmessages_trail.readerErrorscounts lines a reader could not map. - Offline meeting recovery unavailable? — build 209+ requires server 6.11.0+.
Restart once, then confirm
/api/healthreportsfeatures.localFirstMeetings: trueandcapabilities.localFirstMeetings.protocolVersion: 1. Older app builds keep using their existing live-transcription and meeting-save paths.
License
MIT. Learn more at gotcos.com.
Optional owner-local memory runtime
The package includes a checksummed public Python runtime. It uses a unique owner, separate data root, reviewed memory states, current source snapshots, explicit document-link graph paths and saved explorations. Initial memory search is keyword based; semantic embeddings and inferred extraction are separate optional integrations. No private COS checkout or corpus is included. Existing plain-file setup remains available.
Run glasses-memory-setup --runtime-dir /absolute/path/to/cos-memory-runtime --data-root /absolute/path/to/private-memory with Python 3.11+ installed. Set COS_SCRIPTS_DIR to the runtime directory in your server configuration, then restart the server. Add an explicitly granted Markdown/text source using venv/bin/python3 manage.py grant /path/to/notes.md --title "Project" from that runtime directory. The bundled README documents captures, reviews, document links, source refresh, upgrade and staged restore. Keep the latest independent deletion checkpoint separately from restorable snapshots.
This is a local owner instance, separate from app.gotcos.com tenants and hosted Sessions. Installing the package does not upload your corpus or enable household sharing. A source build or local setup test is not proof of a published registry version; use the release ledger for actual availability.
