@awak-app/simy-cli
v0.6.22
Published
Local SIMY executor for non-coding Local Tasks and Agentic Loop runs.
Readme
SIMY CLI
Local execution agent for SIMY Local Tasks and Agentic Loop runs.
Action candidate search runs once for each authenticated live conversation context, and again when that context changes or Refresh is selected. Background binding renewal uses the lightweight binding endpoint instead of repeating candidate search. Settled ordinary conversations and terminal managed tasks do not automatically search candidates. Bound Actions still renew to reflect changes to deadlines and manual links. Candidate text expires after one minute; select Refresh to retrieve current suggestions after editing Actions in SIMY. Candidate search identities are cached in memory, so restarting the host may retrieve live contexts again.
simy --version prints the installed CLI version without starting an agent or
checking for updates. simy doctor performs read-only local connection and
provider login checks and prints a JSON diagnostic report: installed/running
versions, stale discovery records, session expiry, provider readiness, and recent
structured SQM failures, plus host_exit with the latest native desktop host's
abnormal exit time and status. failures.total_failures counts scoped SQM
failures only; zero does not mean the desktop has never crashed. host_exit is
local machine history, independent of account/organization and current daemon
reachability. A recorded exit remains visible after restart; not_recorded means
no retained exit record, and unavailable means it could not be verified. It is
not a lifetime crash count. Only the fixed header is read; stderr is never
included. Doctor does not run a model, renew login, or restart tasks.
Provider login is separate from model access; unverified model access remains
unknown. Diagnostic journals retain stage/code/HTTP status only, use private
files scoped to the Web/account/organization, rotate each writer's four 64 KiB
segments, and keep lifetime failure counts after event rotation.
The desktop host continues sampling connection and access while optional conversation metadata is pending. It retains the last completed conversation metadata only within the same verified scope, clears it after access or account changes, and merges completed metadata with the newer connection observation. This prevents a slow metadata read from making a live service appear disconnected; it does not conceal actual host exits, failed connections, or revoked access.
The organization menu's Allow Codex / Claude fallback setting defaults to off and grants only read-only SQM analysis permission to use the other signed-in local provider. It is saved per Web/account/organization. Switching requires an explicit native model rejection or unavailable executable, a server that supports actual-provider provenance, and fresh actor/settings/lease checks. Authentication, quota, malformed output, arbitrary failures and coding tasks do not switch engines. Both providers share one analysis deadline. A deterministic provider failure pauses new claims and retains failed history until environment/settings change or Retry analysis is selected.
After 6 consecutive SQM analysis failures, this machine stops claiming new SQM analyses for the current Web/account/organization. Codex and Claude share the counter; a completed analysis resets it. Empty polls and user cancellations do not change it. The pause and partial failure count survive restart, sign-in and provider changes. A review interrupted by a host exit counts once on restart; graceful cancellation clears that unfinished marker without changing the count. Reconnecting or selecting Retry cannot clear this pause. Installing and running a newer SIMY version resets the counter and resumes claims automatically, subject to normal organization access and analysis settings. The native organization menu shows when an update is required. This pause does not change server analysis consent or stop ordinary implementation tasks.
Agentic Loop preflight accepts the task's existing requirement, acceptance
criteria, required checks and charter context alongside the repository/provider.
Its startup summary lists all blockers, unverified conditions and recovery
actions together. Environment-only calls remain supported. Task checks reuse the
execution charter policy; model entitlement, external test resources and assurance
selection evidence are not guessed from a successful login or recommendation.
Local continuation
The recorded Bounded and Progress policy fields remain readable for compatibility. Independently verified acceptance criteria permit further work. Local token usage, elapsed time and attempt counters do not stop execution, including old tasks restored from bounded policies. Repeated work without new proof pauses the same task. Approval requests, explicit stops and authorization failures still pause execution. Provider text claiming success does not count as verification. Token records remain available for history and usage warnings. The Web controller supports connected local Codex/Claude engines; unsupported cloud execution is refused before dispatch.
A completed Agentic worker's structured result is saved privately before evidence collection and the final observation flush. If the controller restarts after that save, the authorized daemon recovers the same attempt and collects fresh evidence without dispatching the completed worker again. Collector read failures retry up to three times; exhausted reads retain the result for an explicit continuation. A result receipt does not prove acceptance. Stop, changed identity or objective, and an already consumed remote result prevent automatic recovery. An unfinished worker or a missing, invalid or redacted structured result still requires reconciliation; no unknown external operation is automatically replayed.
After CI finishes successfully, the authenticated daemon refreshes the independent review of the same saved candidate without dispatching its implementation worker. The reviewed HEAD, fetched base, objective, preparation evidence and minimum completion matrix must still match; changed candidates require explicit work. Pending/failing checks, a dirty checkout and missing required checks do not trigger this review. Human peer approval remains required when the charter requires it. The refreshed review is checked against newly collected evidence before readiness can pass. Stop cancels review continuation. A failed or interrupted refresh is not repeated on every monitor tick; explicitly continue the same run to recover. Older reviews without a saved CI basis can refresh only when their prior evidence still establishes the complete reviewed identity; otherwise explicitly continue.
For interrupted Agentic work, GET /v1/agentic-loop/:id/recovery provides a
read-only objective and checkout fingerprint. Reconcile an uncertain side effect
with POST to that endpoint (action: reconcile, authoritative
operation_status, observation, evidence_refs) before resuming. A paused
run retains the legacy metadata endpoint
POST /v1/agentic-loop/:id/continuation-ceilings for compatibility. These fields
no longer impose local limits and changing them never resets usage or resumes work. Both routes require the original authorized
account, organization and device throughout the request.
Provider reconnection and saved consent
Disconnecting an AI provider revokes its conversation monitoring and automatic
continuation with a durable connection epoch. Reconnecting restores discovery;
previous work remains paused until explicitly selected again. Old or damaged
conversation locks do not block this operation and are never deleted by it.
Individual locked conversations may still require support before they can be
edited or continued. Provider login is separate: for a signed-out Codex CLI,
run codex login in Terminal and refresh connections.
After the first connection change, the saved settings deliberately appear
disconnected to older CLI/plugin versions that cannot enforce the consent
epoch. Update the runtime and installed SIMY plugins together and start a fresh
provider session before using their hooks. Do not edit the saved connected
flag to bypass this fence. A rollback to older code leaves those integrations
disconnected; pause work and use a compatible version to restore the connection.
Connection failures write only provider, operation, and a bounded error code to
the private logs/desktop-host.log; they do not include credentials or conversation
content. Verify reconnect, fresh explicit consent, and that previous work remains
paused when validating a rollout. This does not prove provider model access.
Apple Silicon Mac app
The macOS installer includes a native status window and menu bar app. Open SIMY from Applications and choose Sign in. The window shows SIMY login, the connected Web address and last successful heartbeat, Codex installation and login status, tasks, and updates separately. A Codex login check does not prove that a task can execute successfully; permissions and service availability are checked when work starts.
The sidebar's Organization selector is global across Sessions, Tasks and SQM and shows active memberships of the connected SIMY account. Switching is available while work is running and opens browser authorization for the selected organization. It cancels running and queued SIMY analyses before changing organizations; ordinary implementation tasks keep their original organization. Use the same account in the browser; cancel an unfinished switch from the sidebar.
Browser authorization links expire after 20 minutes. The app shows Connection timed out and Reconnect when an active request expires. Reconnect starts a new link; old or canceled links cannot complete authorization. Installed clients with the loopback redirect transport finish on a self-contained local browser page so Safari does not need an HTTPS-to-HTTP cross-origin fetch. The session itself still lasts 48 hours; the 20-minute limit is the browser connection window.
Local SQM, above the connection status, allows this Mac's Codex / Claude Code to claim local SQM reviews. It defaults to off and is saved separately for each Web environment, account, and organization. Turning it off stops new claims; current analysis finishes and submits its result. The toggle does not grant SQM subscription or organization permissions. First-time repository history discovery still starts from the web workspace.
Quality checks shows local analysis and check history separately from ordinary Tasks. Recent records survive restart and are scoped to the Web environment, account, and organization. Failed checks retain safe diagnostic codes, stages, HTTP status, and request identifiers when available. The web SQM workspace retains the authoritative analysis history. Persistent provider login, model or access errors pause new local claims with a visible recovery hint; refreshing after reconnecting permits another probe, with bounded automatic cooldowns.
Conversation checks blocked by Sign-in required automatically recheck after a fresh Codex or Claude login probe succeeds, including after an app restart. An unchanged ready status cannot repeatedly re-run a failed check. Manual retry remains available; continuing the conversation still requires Auto-continue.
You can enable Auto-continue while a conversation is running when its original provider configuration is verified. SIMY saves the preference and waits for the current turn to finish before checking remaining work. Configuration and runtime blockers appear in the conversation view with recovery guidance; turning the setting off does not interrupt the current turn.
Claude read-only structured reviews retry once when the provider explicitly reports exhausted turns or structured-output repair attempts, within the original timeout. Other failures retain distinct codes and recovery hints. Safe native result subtype, exit status, output size and recognized model identifiers survive restart in SQM history and correlate with the task ID in diagnostic logs. Provider responses, prompts and credentials are not retained. Older generic failure records cannot acquire details that were never recorded.
Detached daemon and desktop-host diagnostics are saved to
$SIMY_HOME/logs/daemon.log and $SIMY_HOME/logs/desktop-host.log
(SIMY_HOME defaults to ~/.simy). Each log is limited to 1 MiB plus two
backups; credentials, Cookie headers and URL queries are redacted. Installer
handoff diagnostics remain in ~/.simy/installer/desktop-update.log.
The macOS window also saves the latest nonzero background-host exit in
$SIMY_HOME/logs/desktop-host-exit.log, including its timestamp, exit status and
up to 16 KiB of stderr. Credential-bearing lines are omitted; each abnormal exit
replaces the previous record. This includes failures before JavaScript starts.
simy --version or simy -v prints the package version without starting work.
Closing the window keeps the agent running. Open the app again or choose
Show SIMY from its menu bar icon to return. Quit SIMY asks before
interrupting active tasks. The normal simy command remains available.
Signed patch/minor updates wait for idle and verify the new runtime before handoff; failed startup restores the previous installed version. After a runtime update, choose Restart to update to load the updated native interface. This action asks to stop running tasks before restarting. Stop task… in a task detail asks for confirmation and terminates its local Provider process, keeping existing output and history. An explicit update can likewise stop active work after confirmation; automatic updates still wait for idle. ⌘Q exits SIMY and its background services directly, including when work is running. The window reports App and runtime versions separately while they differ, and normal app launches select the newer installed interface.
Older installers that predate native app relaunch support need a one-time reinstallation from the official download page to update their Applications entry. Updating only their background runtime does not replace that entry.
For development, npm run build:macos-installer:dev creates an unsigned app/DMG;
production builds require Developer ID signing and notarization. Never publish
unsigned development artifacts through the stable update manifest.
For normal use, install SIMY globally and keep its background agent running:
npm install -g @awak-app/simy-cli
simy --daemonWhen an interactive simy --daemon starts without managed SIMY workflows,
the CLI asks once before installing the Codex and Claude Code plugin
sources and Codex/Claude opt-in SQM skills. Existing unrelated plugins
and hooks are preserved. Use --install-workflows for explicit non-interactive
installation or --no-workflow-prompt to start the daemon without asking.
SQM checks apply when you explicitly select @simy, invoke the SQM skill, or ask to use SQM for the current task. Ordinary PR creation and git push do not opt in. Upgrading and restarting the daemon migrates the legacy global PR gate and shipped personal skills; unrelated hooks are preserved. Start a new coding session to reload skill instructions.
After installation, restart Codex and find and install SIMY in Plugins
before using @simy in a new task.
The plugin also includes a non-blocking connection reminder. At task startup,
resume, and prompt submission it reads the origin-scoped local discovery record
and checks the loopback CLI health endpoint. It warns once per disconnected
period in each Codex session, then rearms after the CLI reconnects. The hook does
not start or install the CLI, read session tokens, enable SQM, or send email.
Review and enable the plugin hooks through Codex's /hooks command.
For Claude Code, register the generated local marketplace and explicitly enable SIMY:
claude plugin marketplace add ~/.simy/integrations/claude-marketplace
claude plugin install simy@simy-localUse /simy:sqm-loop or /simy:pr-review in Claude Code. The workflows are generated
from the same maintained source as Codex; only provider terminology and hook/manifest
format differ. Installing source files does not imply that either provider has enabled
the plugin. Existing Claude settings, unrelated marketplaces, and plugins are preserved.
After that first installation, the daemon checks installed workflows at startup
and every hour, including after a CLI automatic upgrade. It updates only existing
platform installations and refreshes an already enabled SIMY Codex plugin through
codex plugin add. Enabled user-scope Claude plugins are refreshed with
claude plugin update simy@simy-local --scope user, after verifying that the
marketplace still points to SIMY’s generated local source. Plugin versions include a content digest so changed bundles
receive a fresh cache. Use a new Codex task to pick up updated skills; running
tasks are not restarted. Missing or disabled plugins are not enabled automatically.
Failed refreshes retry on the next check without blocking daemon startup.
Auto-continue ON saves consent for the selected original conversation; it does not
immediately launch a second turn. A verified desktop discovery admission or an
existing exact admission saves this preference without another provider RPC.
After the current turn settles, continuation separately rechecks context, the
live provider channel, approvals and the opt-in epoch. Required human decisions
remain pending, and OFF prevents future continuation without interrupting the
current turn. The macOS switch and label explain these boundaries through their
tooltips and accessibility help.
Conversation statistics estimate work time from activity events within the last seven days. Usage counters alone leave work time, idle time and utilization unavailable; a single observed activity event is a measured zero-duration sample. Missing usage records remain distinct from recorded zero tokens.
The desktop Sessions page loads a recent window of up to 200 conversations, ranked by last activity. Codex candidates come from CLI, VS Code and app-server conversation sources; non-interactive exec, unknown sources and delegated agents are excluded. A hook observation alone cannot restore a filtered conversation. Explicitly tracked, monitored, auto-continue and unsettled supervision records remain visible even outside that window. When a retained Codex conversation is omitted from the recent provider page, SIMY reads its exact metadata without loading turns or taking writer ownership. The same interactive source filter and a matching active rollout path are required before restoring provider capabilities and probing its latest turn. Exact reads and turn probes share a bounded observation deadline. Claude discovery ranks file metadata before opening selected transcripts. A conflicting saved workspace leaves only that conversation's Action association unavailable; it does not stop discovery or change the saved identity. The page lets users choose monitoring or supported session supervision with a goal, acceptance criteria and an autonomy policy. Local time, token and continuation counters do not stop work. Codex enrollment uses an exact app-server session and checks writer ownership before each continuation. Claude can continue through a verified live SIMY hook in its original process; sessions without that hook support monitoring only, and cold takeover remains unavailable. See session supervision.
Provider discovery skips outdated or unversioned executables while looking for a
compatible installation. Codex session supervision requires 0.158.0 or newer,
independently of the lower general execution minimum. On macOS, discovery also
checks the bundled CLI in ChatGPT and Codex apps in system and user Applications.
Claude discovery likewise falls back to a compatible native or Desktop bundle
when an older PATH installation cannot meet its version requirement.
On macOS and Linux, the desktop host reads the user's interactive login-shell PATH once at startup, so providers installed through nvm, Volta or custom prefixes can be found from GUI launches, including Finder or the Dock on macOS. Shell paths take priority; existing launcher paths remain available. If the five-second shell probe cannot read PATH, startup uses the launcher PATH. Windows keeps its existing PATH.
Installed Codex and Claude hooks use an absolute Node runtime path. When available,
the SIMY installer's current link or Homebrew's opt link replaces a versioned
runtime path, allowing hooks to follow updates without relying on the provider's PATH.
Use simy delivery activity --watch for a read-only local snapshot every minute,
or add --json for one JSON object per sample. It separates execution, waits,
stops, failures and repeated attempts without new verified outcomes, with task
identity, agent, reasons and outcome timestamps. Missing evidence is shown as
unknown; open time is never counted as productive time. See work activity.
The plugin also provides an opt-in, hook-based Delivery Loop in the user's existing
Codex or Claude Code conversation. Install the plugin, review and enable SIMY hooks
with /hooks, then request /simy:delivery-loop for the task. Installing or updating
the plugin never approves hook trust. SIMY's Connection & settings panel links to
Codex settings and the Claude setup guide and shows observed hook activity and
blocked/recovery states.
The Connection tab also checks each provider's actual SIMY plugin inventory and
shows installed/enabled, installed/disabled, missing, or unavailable separately
from app sign-in. Claude enablement is checked from the home folder; project
settings can differ, and project-scoped installations are labeled explicitly.
The legacy agentic-delivery-loop skill remains retired; the Web/CLI executor is
unchanged. See hook delivery behavior and limits.
--no-auto-update or SIMY_AUTO_UPDATE=0 disables these background updates too;
--no-workflow-prompt only suppresses the initial installation prompt.
This is the recommended installation because the CLI can safely install patch
and minor releases, and the daemon restarts only when no Agentic Loop task is active. To run
the interactive chat after installation, use simy.
For a one-time evaluation or source development, these modes remain supported:
npx @awak-app/simy-cli
npm start
simy --no-tui
simy --host https://simy.example.comThe published CLI connects to https://app.simy.one by default. Use --host
with an absolute SIMY Web origin for another deployment. Non-loopback hosts
must use HTTPS. SIMY_WEB_ORIGIN and the legacy SIMY_API_ORIGIN remain
available for automation, but the command-line flag takes precedence.
Work Pattern discovery checks the connected account's server-side personal
setting before reading local Codex or Claude Code history. An explicit opt-out
prevents both local inspection and upload. The daemon reads bounded user and
assistant conversation chunks from active and archived Codex sessions and
Claude Code projects, excluding delegated agents, tool output, reasoning, and
injected runtime context. Dated messages outside the configured lookback are
excluded even when their conversation continues today. Legacy transcripts with
no message timestamps use the file timestamp as a fallback. The local cursor stores only opaque hashes and
timestamps scoped to the connection, account, and device. Uploads require the
user_conversation.v2 Web, Crew, and Backend contract and exact per-chunk
receipts before the cursor advances. Deploy those companion changes before
enabling this CLI version; earlier receivers cannot acknowledge v2 history.
History sync runs only in daemon mode, at startup and then at most every 15
minutes, and only while a Codex or Claude Code executable passes the version
check. The endpoint is the session's api_base_url (for app-dev,
https://app-dev.simy.one/api/local-cli/). The history roots are
$CODEX_HOME/sessions, $CODEX_HOME/archived_sessions, and
$CLAUDE_CONFIG_DIR/projects, which default to ~/.codex and ~/.claude. When
the daemon skips a sync, it logs the reason once to stderr. Reasons include no
daemon mode, no valid session, and no compatible provider with each provider's
status. To test with synthetic history only, point every root at temp
directories and run the daemon in the foreground so its log stays visible.
Stop it with Ctrl-C after the uploaded line:
export SIMY_HOME="$(mktemp -d)" CODEX_HOME="$(mktemp -d)" CLAUDE_CONFIG_DIR="$(mktemp -d)"
# write synthetic *.jsonl files under $CODEX_HOME/sessions and $CLAUDE_CONFIG_DIR/projects/<project>/
SIMY_DAEMON_CHILD=1 simy --daemon --host https://app-dev.simy.one --no-auto-update --no-workflow-promptSIMY_DAEMON_CHILD=1 keeps the daemon attached to the terminal. The sync
cursor is written under $SIMY_HOME/work-pattern-history-sync/.
The CLI starts a loopback HTTP agent on a random available port, keeps a 48-hour, origin-scoped local session, and opens a native coding chat. The first chat message creates the remote audit ledger and starts the local Agentic Loop state machine; SIMY Web mirrors the lifecycle but is not required as the task entry point. The CLI also accepts backend-issued launch challenges for Web-led runs. It builds the requirement charter, runs Codex or Claude Code, audits structured completion evidence, and re-instructs the executor while verification shows useful progress.
Before Web creates an Agentic Loop, it classifies whether the request actually needs a managed PR, verification, evidence, and merge lifecycle. A bounded ordinary task is sent to CLI 0.2.2 or newer through the one-time direct executor instead. CLI independently recomputes the guardrail, restricts read-only work at the Provider command boundary, and permits exactly one Provider invocation. It does not create an Agentic Loop ledger, audit session, or retry lifecycle for that task.
Local Tasks
Purpose runs may opt into independent local evidence by freezing user-approved
verification_recipes in the server claim. capabilities.purpose_verification
advertises the device's Ed25519 public key; the server pins that key for the
attempt. Older clients without this capability cannot supply this evidence.
POST /v1/local-tasks/:id/verify-purpose with {} checks the original plan only
after the task succeeds. GET only reads the cached task.host_verification.
File recipes independently check bounded regular files and optional JSON format;
command recipes execute approved structured argv for Node tests or package
test, build, check, and lint scripts. Commands require write permission,
share a five-minute deadline, and retain only output hashes and exit status.
Optional protected_paths bind check definitions to their pre-execution hashes.
A receipt binds the actor, device, attempt, objective, recipe plan, local task, and before/after repository identity. Changed repositories, stopped or uncertain tasks, unknown prior check effects, and changed keys fail closed. Fresh POSTs are idempotent; an expired confirmed receipt may be refreshed explicitly on the same task, with at most three check rounds and no new Provider invocation. Checks never upload source files or raw command output. Exit zero proves the specified command's result; it does not prove product acceptance or that a package script has no external effects. This host evidence boundary does not isolate a malicious process running as the same OS user.
CLI API contract v9 promotes non-coding work to a first-class Local Task runtime. A Local Task is separate from Agentic Loop: it runs one bounded Codex or Claude Code session without Git branches, pull requests, audits, or automatic Provider retries. Knowledge-work deliverables use an isolated task sandbox by default, so they do not require a repository.
POST /v1/local-tasks/start accepts JSON or the same verified multipart
attachment envelope used by Agentic Loop and requires a stable
Idempotency-Key. Repeating the same request returns the original task; using
the same key for different content returns
LOCAL_TASK_IDEMPOTENCY_CONFLICT. GET /v1/local-tasks/reconcile with the same
header recovers a task when the original response was lost. Only the key hash is
persisted. Task state, verified inputs, and the event log stay under
~/.simy/tasks/<task-id>/. Tasks that do not need a Git
repository run in ~/simy/<task-id>/, with generated files in its outputs/
directory. The public snapshot exposes this user-facing workspace as
workspace_display_path without exposing internal state paths. Set
SIMY_LOCAL_TASK_WORKSPACE_ROOT to override the default workspace root.
Local Tasks provide:
GET /v1/local-tasks/<id>for the current structured snapshot;GET /v1/local-tasks/reconcilefor request-key reconciliation;GET /v1/local-tasks/<id>/streamfor live Server-Sent Events;POST /v1/local-tasks/<id>/controlwith{"action":"stop"};GET /v1/local-tasks/<id>/artifacts/<artifact-id>for a verified output;- one Provider invocation with recorded token usage; local token, elapsed-time and attempt counters do not interrupt execution, including restored tasks;
- restart-safe terminal persistence. An interrupted active task fails closed after CLI restart and is never replayed automatically.
Interactive Pipelines can use the same daemon worker for a local_task node.
The CLI still invokes the selected Provider exactly once. Before the node is
reported as successful, every output file is re-verified, uploaded through a
task-bound short-lived URL, and committed only after SIMY rechecks its byte
size and SHA-256 digest. Ordinary My AI Local Tasks keep their existing local
verified-download behavior. Scheduled or unattended Pipeline runs are not
allowed to select a desktop implicitly.
Web and Backend may opt into durable_local_task_steps for a versioned
multi-step Local Task. The daemon then claims and executes one closed-schema
step at a time (maximum eight), with an independent Provider invocation
boundary per step and aggregate step, token, and runtime budgets. The task
workspace is stable for the turn while each lease attempt has its own local
record. A step that crossed its Provider boundary is never replayed
automatically.
Plan initialization uses a two-phase claim. Web stores an authoritative
execution_spec.controlled_plan; the CLI validates that its Provider,
operation, permission, repository, approval target, and aggregate budgets
cannot exceed the enclosing execution spec, then submits the exact plan for
Backend persistence. The CLI never asks a Provider to infer or expand a plan.
Older one-shot requests receive a deterministic one-step compatibility plan.
Human approval and unknown Provider results are distinct non-executable
dispositions. Neither holds a Provider process, heartbeat lease, or update
work reservation. Only bounded summaries and verified artifact references from
completed steps may enter the next step's context. Unknown plan fields,
unscoped approvals, unsafe paths, and Provider-supplied risk metadata fail
closed. The original local_task_execution.v1 one-shot path remains available
unchanged for clients that do not negotiate this feature.
Contract 10 additionally advertises bounded_local_task_subtasks. Backend
admits at most two Provider children at once from a closed parallel_reads
plan with 2-4 source-bound candidates. When Crew declines or times out, the
same v2 contract can admit one serial_uploaded read-only child for 2-4
uploaded sources; this fallback never routes those uploads through the v1 CLI
staging materializer. The CLI never prefetches beyond an available execution slot:
each child starts only after a fenced server claim and an explicitly confirmed
invocation boundary. Children inherit the parent Provider and model, run in
separate read-only task sandboxes, receive only verified immutable SIMY uploads,
and cannot use mutation tools or nested delegation. Downloaded bytes are checked
against the claimed size and SHA-256 digest, then exposed through a mode 0400
file without revealing the storage URI to the Provider.
Cancellation uses one pool-wide abort signal and prevents new claims. Unknown
invocation or final-status boundaries are never replayed automatically. Sibling
fan-in summaries remain explicitly untrusted data, and all lease, generation,
budget, usage, and elapsed-time bigint fields cross the JSON boundary as
canonical decimal strings. The contract 9 durable_local_task_steps path
remains serial and unchanged.
External side effects such as sending messages, submitting forms, uploading, or
purchasing are outside this MVP. The router rejects them instead of silently
granting broad permissions. The legacy /v1/direct-execution/* contract remains
available for older Web clients.
CLI updates
Every command checks the stable npm release before dispatch, including simy,
simy --daemon, simy sqm min-check, and simy sqm check. Globally installed
patch releases update silently; minor releases update automatically with one
short notice. The updated executable is verified and then re-executes the exact
original argv once. A process-wide nonce/version guard prevents update loops.
Major releases are never installed at startup: simy update asks for explicit
confirmation (simy update --yes is available for non-interactive automation).
The macOS desktop app runs from /Applications/SIMY.app. Its signed updater
keeps downloaded releases in the user cache, then replaces the desktop bundle
after the app and its executor have stopped. The previous bundle is retained
for recovery. If permissions or verification prevent replacement, SIMY shows
an installer recovery action instead of reporting the desktop app as current.
Concurrent handoffs share one exclusive installation lock. Packages before
0.5.29 may require one system-authorized installer update because their bundle
root is not writable by administrators. New packages grant the admin group
write access only to that root directory so later signed updates can exchange
the app; standard users still use the authorized installer when required.
Handoff diagnostics are written to ~/.simy/installer/desktop-update.log.
Closing the window with ⌘W keeps SIMY running in the menu bar; reopening it
restores the existing window. ⌘Q quits the app and its service.
Use these commands to inspect or apply an update explicitly:
simy update --check
simy updateA running daemon also checks hourly. Before installing, its strict idle gate confirms there is no active, paused, blocked, or human-waiting task, executor child process, or pending ledger write. It stops accepting new work during installation, verifies the installed package, starts the new daemon, verifies its version and local health handoff, and only then closes the old process. Fixed-port daemons retain the manual restart path.
Only a global npm installation may be changed automatically. npx, a source
checkout, and a normal dependency installation are read-only and display an
exact recovery command instead:
npm install -g @awak-app/simy-cli@<version> && simy --daemonIf installation or startup verification fails, the updater restores the
current exact version when possible and prints both the pinned recovery command
and the target command. Use --no-auto-update or SIMY_AUTO_UPDATE=0 to disable
all update checks and writes for any command.
CLI 0.3.3 predates this startup updater and cannot use it to upgrade itself.
The last required manual bridge is:
npm install -g @awak-app/[email protected]From 0.4.0 onward, subsequent compatible minor releases update automatically.
When SQM Cloud returns structured incompatible_cli, the CLI runs the same
verified updater and re-executes the original SQM command at most once. An SQM
min-check or full check pins its CLI version and signed knowledge bundle for the
entire execution; no update is applied mid-check.
The current update state is also exposed by GET /v1/health as cli_update.
Auto-update E2E screenshots are generated locally under
.artifacts/cli-auto-update/; this ignored directory must not be committed.
Executor compatibility
SIMY checks the selected executor with --version before an Agentic Loop can
start. The currently validated compatibility floors are Codex 0.144.0 and
Claude Code 2.1.200. A missing, older, or unversioned executable fails the
preflight and returns the installed version, required version, and a recovery
command to SIMY Web. Update Codex with
npm install -g @openai/codex@latest; update Claude Code with claude update.
Your Desktop executor
SIMY CLI is the desktop execution target, shown to users as Your Desktop.
It can run either Codex or Claude Code as a child process inside the verified
local Git checkout. Device heartbeats advertise both desktop provider
capabilities and their installed/required versions. Preflight and every standard
provider dispatch fail closed unless the selected binary is compatible; there
is no fallback to an unverified executable name.
Legacy local requests without execution_target continue to resolve to
desktop. A request marked for the hosted simy target is rejected by the
local CLI boundary so hosted and desktop execution cannot be confused. The
target and bound device id are preserved in the run Charter, redacted Web
ledger snapshot, and restored run.
Blocked recovery contract
Every waiting_human, blocked, or failed lifecycle event includes a
versioned detail.recovery object. It provides a stable reason code, a
user-safe explanation, a primary action, executable action metadata, and
bounded technical details. Preflight failures use the same contract for
executor updates, repository selection, and CLI reconnection. Web clients must
validate the action type before rendering or executing it; raw local error text
is not copied into the recovery contract.
Provider token budget
Each run accepts a configurable token_budget (or the explicit
provider_token_budget alias) from 1 to 15,000,000 Provider tokens. The default
is 250,000. Provider tokens are the Codex or Claude Code usage reported for the
run; they are a safety limit and are not billed as SIMY tokens.
When a run reaches this limit, its recovery contract suggests a higher rounded
limit. SIMY Web can edit that suggestion and call
POST /v1/agentic-loop/:run_id/provider-token-budget with
provider_token_budget. The CLI persists the higher limit and resumes the same
run. SIMY token consumption remains a separate platform billing record.
An in-flight provider call can report usage above the previous limit. Recovery keeps that full usage history, including reported cached input; it never resets the counter or increases the limit automatically. An explicitly approved higher limit must exceed usage already recorded and remain within the finite cap. Retry limits and all completion gates still apply after a budget increase.
SQM local checks
SQM deliberately has two identity-bound stages:
simy sqm min-checkquickly inspects committed changes, the working tree, untracked files, and deletions. It runs only deterministic low-cost rules, reports preliminary findings, and writes the applicable state, transition, stress, and scenario plan. A passing min-check is not a complete validation.simy sqm checkverifies that repository, base, HEAD, working-tree digest, and signed bundle still match the min-check. It then evaluates applicable state transitions, invariants, stress/strength expectations, and scenarios, writesproof.json, and signs that proof with the local device identity.
The full-check summary reports applicable modules, evaluated rules, scenario
statuses and blocking/advisory scope. passed with no_applicable_rules or zero
scenario evaluations does not establish feature or historical incident coverage.
An empty system/repository bundle adds no_system_knowledge and an explicit
zero-case warning; knowledge that exists but does not match this change uses a
separate warning. Both retain the existing passed outcome.
Scenario statuses come from deterministic rule evaluation, not browser/provider
execution. Keep concrete regression and delivery acceptance evidence separately.
If knowledge is unavailable, the command returns exit 2 with the service cause
and writes no new proof; any older artifact is not evidence for that attempt.
Cloud storage is separate from the check outcome. check retains the completed
proof and local check history if its evidence POST fails, returns exit 2 when
storage is unconfirmed (exit 1 still takes precedence for a failed check), and
writes a redacted storage receipt to <proof-path>.upload.json. Only a server
acknowledgement matching the proof ID, evidence ID and proof digest counts as
stored. --no-upload intentionally keeps evidence local and does not return
exit 2 for storage. A successful upload is not evidence of applicable coverage.
After resolving a storage/authentication failure, explicitly recover the original
run with simy sqm upload-proof --proof /path/to/proof.json (add
--signed-evidence /path/to/signature.json if using a custom signature path, and
--host <original-origin> for its SIMY environment). This verifies the original
signature and current organization/account/device before sending the same proof;
it never reruns checks, resigns evidence or automatically replays an ambiguous
POST. The backend reconciles repeated proof IDs against their exact digest and
evidence ID. The receipt includes the saved record ID and evidence ID on success.
This recovers a historical run, not acceptance of the current working tree.
Dev/production record visibility, deployed gateway/schema and UI destination
permissions still require validation in the corresponding environment.
Normal runs do not require a separate sync command: the CLI uses a verified local cache and revalidates it with TTL and ETag semantics. The Agentic Loop runs min-check after each code modification and the full check before PR readiness. Findings and the redacted signed-proof summary are recorded in its audit; source text is never uploaded.
Configure the bundle service with SIMY_SQM_BUNDLE_URL. Authentication uses
the current SIMY session in the coding loop, or SIMY_SQM_TOKEN/
SIMY_ACCESS_TOKEN for the diagnostic command. On the first authenticated
HTTPS request, the CLI receives the bundle verification key from the SIMY
gateway, verifies the signed bundle, and pins that key to the endpoint,
account, organization, and repository cache identity for offline use. Explicit
trust anchors can instead be configured as a JSON key-ID map in
SIMY_SQM_PUBLIC_KEYS_JSON, or as one SIMY_SQM_KEY_ID plus
SIMY_SQM_PUBLIC_KEY. Public keys may be PEM, base64 SPKI, or a base64 raw
32-byte Ed25519 key. Optional settings are:
SIMY_SQM_CACHE_DIR: verified bundle cache location.SIMY_SQM_CACHE_TTL_MS: revalidation TTL, defaulting to five minutes.SIMY_SQM_ACCOUNT_IDandSIMY_SQM_ORGANIZATION_ID: tenant identity for diagnostic checks. Coding-loop checks receive tenant identity from the authenticated session. Cache entries are isolated by endpoint, account, organization, and repository, and the organization is verified against the signed bundle.SIMY_SQM_ORGANIZATION_IDis required for standalone diagnostic checks; without a tenant identity SQM fails closed in shadow mode.
Use either phase with --offline to prohibit network access and --refresh to force
ETag revalidation, --base <branch> to select a base branch, and --json for
machine-readable results. Full check additionally accepts --min-result,
--proof, --signed-evidence, and --no-upload. The min result and full proof
share one pinned signed bundle and one exact repository identity, so a code
change cannot silently reuse an old preflight. If Cloud is unavailable, the CLI clearly warns and
uses the last unexpired signature-verified bundle only after a network failure
or Cloud 5xx response. Authentication, authorization, compatibility,
signature, schema, tenant, and expiry failures never execute stale knowledge.
With no verified cache it skips SQM
in shadow mode. Bundles contain only the fixed declarative checker vocabulary;
the CLI never executes shell, JavaScript, Python, WASM, or another remotely
supplied implementation. A module's matcher array has explicit AND semantics:
every matcher must match before any check in that module runs; it is not an OR
list.
Interactive console
Running simy in a terminal opens the Agentic Loop chat. It discovers the
current GitHub repository and branch, selects an available executor, and keeps
the composer active so a natural-language request can start immediately. Each
run appears in a selectable list with its lifecycle state, process ID, and a
bounded live transcript. Use --no-tui for the foreground HTTP-agent behavior
or --daemon to run without the chat.
The console supports these direct controls:
Up/Downorj/k: select a run.i: enter human guidance. Guidance submitted while a provider process is active is queued for the next executor handoff; a waiting run resumes as a new attempt in the same lifecycle record.p: pause or resume the selected local process on POSIX systems.r: refresh pull request, review, and CI evidence.x, thenx: stop the selected run with confirmation.?: show the command reference;q: leave the console.
Use /new, /repo <owner/name>, /branch <name>, and
/executor <codex|claude> to compose another task. /attach <path> adds a
verified local file reference without copying or deleting the source file.
Slash commands also provide explicit human gates: /continue, /approve,
/criteria, /checks, /recheck, /pause, /resume, and /stop. Approval
notes, acceptance criteria, and required checks are recorded in the run charter
before orchestration continues. The console does not pretend that noninteractive
codex exec or claude -p accepts mid-process stdin; active-run guidance is
shown as queued until the next executor attempt.
Structured run snapshots and process events are persisted remotely. Executor instructions and the latest bounded log lines are redacted at both the CLI and Web ingestion boundaries before they are persisted with SHA-256 integrity hashes. Raw transcripts, environment values, and unbounded stdout are never stored remotely; live raw output is streamed directly from localhost to the active Web UI. The persisted snapshot records the redaction policy and whether input or output was redacted or truncated.
The local PR lifecycle is intentionally separate from deployment:
- Classify request risk and build a requirement charter.
- Require explicit acceptance criteria plus an approved design summary for high-risk work.
- Run the coding executor and verify its result against local Git evidence.
- Run a separate AI audit session without granting it an implementation role.
- Require eight revision-bound PR-preparation rows: latest base/head, current CI and disclosed skips, exact companion revisions, E2E outcomes and zero residual fixtures, security/privacy, browser/localization when applicable, final diff scope, and honest residual/post-merge status. The CLI refreshes the remote base, requires it to be an ancestor of the final head, and compares it with the PR base SHA. Four additional conditional safeguards cover official-checklist ownership, authoritative deployment-target mapping, live-test data safety, and saved-run/ checkpoint compatibility without forcing irrelevant work to execute them.
- Mark the result
pr_ready_for_reviewwhile GitHub checks or peer approval are pending. - Mark it
merge_readyonly after the observed PR head/base, CI checks, merge state, and human approval all pass.
The Web UI can refresh review and CI evidence with
POST /v1/agentic-loop/:run_id/recheck. Passing the local implementation gate
does not by itself make a PR merge-ready.
Start simy from the selected repository or a workspace containing it. Set
SIMY_REPO_ROOT when repositories live under a different root. The CLI
verifies the checkout's GitHub origin before starting an executor.
The real-PTY console E2E can be rerun with npm run test:e2e:console. It starts
from an empty CLI chat, types repository, branch, executor, attachment, and task
input through a pseudoterminal, creates the Web ledger, spawns real child
processes, and exercises HIL plus lifecycle controls. It regenerates the
screenshots, terminal frames, raw transcript, and SHA-256 manifest under
.artifacts/e2e/cli-agentic-loop-console/.
Delivery Loop design proposals
The minimum completion plan describes proposed checks for omitted requirements, inactive controls, unsaved results, and broken integrations at every delivery level, with workflow, sequence, and state diagrams. This is a design proposal, not a shipped completion guarantee.
Publishing
Publishing is handled by GitHub Actions in .github/workflows/publish.yml.
- Every push tag matching
v*.*.*runs checks and publishes to npm. - Manual workflow runs perform the same checks; set
publish=trueto publish. - The repository must define an
NPM_TOKENsecret with publish access to@awak-app/simy-cli. - The package check inspects the final npm tarball contents and fails if any source map file would be published.
Repository system scopes
Version 0.5.7 supports signed system_scope snapshots from Quality Monitor.
Each check pins the repository's connected component and graph revision. A new
check revalidates scoped bundles online; a stale or unavailable scope cannot fall
back to an older cached component. Minimum and full check proofs retain the
original snapshot. A scope revision change requires a fresh minimum check.
Deploy this CLI and the web capability proxy before enabling the backend scope
requirement. Older clients receive an explicit upgrade error, never an expanded
organization-wide bundle.
Minimum functional completion
Every CLI coding-loop promotion now requires a current two-layer observation matrix: the requested operation and result, plus independently checked final postcondition and integration. See the minimum completion contract for the executable format. Missing mappings, mismatching observations, changed artifacts or stale revisions block promotion even when the executor reports success. Independent review checks original-request coverage and the test oracle.
Existing v2 run history is retained; resumed/rechecked runs need the new evidence.
The same checker ships under src/orchestrator/delivery/scripts/ in the CLI
package, independent of Codex plugin installation. It does not establish deployment
completion or intercept another host application's final response.
PR closure and user-outcome transfer
The Web/CLI executor and independent auditor receive the same PR closure procedure. It requires a revision-bound before/after table, reachable normal-user paths (including no-workflow fallback), independent review, explicit authority, pre-close freshness checks and post-close remote reconciliation. A file match or successor link is not proof of UX inheritance; administrative transfer, approved retirement and complete inheritance have distinct decisions. Unresolved yellow/red comparison rows require remediation PRs (or exact existing open equivalents), mapped by criterion ID, owner and acceptance checks. Uncreated plans or unresolved authority/product decisions block transfer completion; creating a PR does not make a row green. Closure is not a new delivery completion state. This is runtime instruction and review guidance, not a GitHub hook that mechanically intercepts every external close command.
Submit repaired local delivery evidence to SQM
The integrated delivery state model connects delivery milestones, repair/wait/resume branches and actual controller states. Both executor and auditor receive its path. It documents existing control and evidence duties; it adds no scheduler or transition-enforcement mechanism.
The bundled Delivery Loop maintains an external acceptance checklist with separate user-journey and technical-invariant results. It refreshes the table after each evaluated action and marks evidence pending when the code or runtime context changes. A non-green table is a checkpoint, not completion: the executor must perform feasible authorized repairs, missing verification and bounded waits, then re-observe and update the same criteria until the requested endpoint passes. A genuine safety, authority, dependency or no-progress stop remains explicitly unfinished with affected IDs and a concrete resumption condition. No background monitor is implied. Both executor and independent auditor receive this continuation contract. The acceptance evidence guide documents the offline validator shipped inside the CLI runtime. Its result establishes record completeness and artifact integrity, not overall delivery or deployment. At the start of each new run, the loop recommends and confirms one of five delivery levels in chat. Level 5 is highest assurance; level 1 is lightest. Every selection or recommendation shows all five levels (5 to 1), their content and verification scope, task-specific hour estimates, and the recommended level with its rationale. Estimates distinguish execution from external waits and state assumptions and uncertainty; there are no fixed durations by level. Risk-based floors preserve mandatory assurance. Historical selections retain their numbering version and meaning. These are executor/auditor instructions, not a new deterministic CLI selection validator; selecting a level grants no additional authority.
simy sqm submit-evidence --file delivery.json --host https://app-dev.simy.one
submits an explicit bounded summary through your existing organization-bound device
session. --dry-run previews the exact document without uploading it. See the
local evidence contract.
The receipt identifies the source digest and analysis job; it is not a declaration
that an incident or blocking rule has been published.
Delivery Loop shared specification
UX descriptions throughout the Web/CLI lifecycle follow the behavioral-completeness standard. Observable interaction units explicitly cover intent/context, entry, perception, control, response, usable outcome, continuation, alternatives/recovery and evidence. Independent reconstruction, source coverage, causal continuity, branch closure, observability and traceability determine sufficient detail, not word counts or product-specific scenarios. Unknowns remain explicit; compact summaries link to the complete canonical specification. The same standard is supplied to the actual executor and independent auditor; it is not an additional automatic schema gate.
The Delivery Loop system specification explains its purpose, As-Is and proposed To-Be through twelve linked views. Keep these as shared team documents; each change updates affected views and checks cross-view consistency. The workflow requires independent design, implementation-to- To-Be and PR reviews, plus explicitly opted-in SQM before and after PR creation. The workflow reference explains baseline updates, historical regression coverage and evidence limits.
Desktop SQM history
SQM Tasks has separate Analysis and Checks lists on macOS and Windows.
Analysis retains local review summaries; Checks retains both simy sqm min-check
and simy sqm check outcomes, including checks run by the Delivery Loop.
Each desktop list displays the latest 100 runs. Durable summary files under
SIMY_HOME/sqm-history (default ~/.simy/sqm-history) are retained without
a count or age cutoff, isolated by web origin, account and organization. Adding
a run never deletes older history. Old job keys remain available through the
local task-detail lookup, including failure stage/code, HTTP status, provider
diagnostics and the matching web task link. The separate rotating diagnostic
log is a recent-event view; retained job summaries provide longer-term correlation.
Operators should include this directory in their disk usage and backup policy.
Deleting a scope directory explicitly removes its history; previously pruned
records cannot be recovered.
History remains visible with local analysis off and after restarting. A saved
in-flight analysis is shown as Interrupted until this client claims it again;
it does not contribute to the running count. The web workspace owns retry status.
Model/provider execution failures permit one recovery claim after a five-minute
cooldown, even when the app/account identity is unchanged. If that attempt fails
again, the pause survives reconnects and host restarts until a newer SIMY version
is installed. Sign-in, local installation, SIMY access, schema and persistence
errors do not open this automatic provider recovery path. The existing six-failure
circuit still protects other analysis failures.
Only summary metadata is retained, without analysis prompts, source inputs, model responses or credentials. Historical summaries are not a replacement for signed proof files or delivery acceptance. Records from versions that kept analysis only in memory cannot be recovered after that process exits.
The CLI Agentic Loop uses proportionate retry and review decisions: reasoning notes are advisory, and an independent reviewer may explicitly defer or dismiss minor/warning/info findings with a nonempty explanation while retaining them in the receipt. Required acceptance evidence and major safety findings remain protected.
Local coding attempt result collection
For executeProcessAttempt, native provider JSONL is read only from a successfully finished turn's final
assistant/result channel. Tool output and intermediate messages are never a
fallback result. Legacy line-anchored markers still work and may contain formatted
JSON; missing, ambiguous, truncated, invalid, oversized and failed-turn results
have distinct result_parse.code values. Invalid collection cannot become a
successful empty result. Plain-text markers are supported only for explicit
shell command overrides. This does not independently verify executor claims.
The parser can re-read a saved raw result without rerunning the provider.
Local Task and direct-task result parsers are separate boundaries and are not
covered by this coding-attempt guarantee.
Explicit desktop provider selection
Local coding requests may specify model / reasoning_effort and independently
audit_model / audit_reasoning_effort. Both selections survive recovery.
Codex receives the exact model and reasoning config; the CLI does not use a
catalog to reject unknown model names or silently replace a rejected model.
Claude receives the exact model; explicit reasoning is unsupported by this
adapter and fails before dispatch. A custom shell override cannot enforce an
explicit selection and is rejected instead of dropping it. Provider acceptance,
authentication and actual availability still require the provider's response.
This changes command construction only; it does not start any workflow.
Reusing a running service and launch recovery
simy and simy --daemon first inspect the private discovery record and verify
its live PID, instance, Web origin, CLI/API version and saved account,
organization and device. A compatible background service is reused: the entry
prints its existing address and opens SIMY Web (or its sign-in page). The local
interactive console is only created when this entry owns a new service.
SQM authorization also reuses that service and leaves it running when SQM exits.
Unreachable or incompatible live services produce diagnostics instead of a
second daemon. A stale record with a dead PID permits normal startup.
Web-led launches can omit their temporary challenge. The daemon requests a
credential for the existing queued run using its authenticated device. An
explicitly expired challenge can be renewed once; consumed, invalid or
cross-scope credentials cannot. Refresh the original run/device before retrying
a consumed credential because execution may have started before its ledger
status was uploaded. This recovery requires the companion Web
/api/local-cli/challenges/issue endpoint; existing valid challenges continue
to work against older Web versions.
Scope mismatch diagnostics identify the changed account, organization, device
or Web host and tell the caller to restore the original identity. Launch and
recovery operations do not switch identity automatically. The private
/v1/connection endpoint rejects browser requests and never returns tokens.
