@arronqzy/webllm-assistant
v0.1.16
Published
Offline WebLLM assistant runtime and action protocol for Abuilder
Readme
Offline WebLLM assistant for Abuilder (React)
Pure browser-side inference via @mlc-ai/web-llm. No cloud LLM API.
Models
- Default:
Qwen2.5-1.5B-Instruct-q4f16_1-MLC - Optional:
Qwen2.5-3B-Instruct-q4f16_1-MLC
Requires WebGPU (Chrome/Edge). First run downloads model weights into the browser cache; later runs work offline.
Install note
@mlc-ai/web-llm is vendored at vendor/bundled/web-llm (avoids root node_modules permission issues). Prefer a normal pnpm add @mlc-ai/web-llm when the monorepo store is writable.
Vite
Vite 5 can crash while transforming @mlc-ai/web-llm (Maximum call stack size exceeded in stripLiteral). Add the helper plugin so the huge bundle is excluded from commonjs / dep optimization:
import { webllmAssistant } from "@arronqzy/abuilder/vite";
export default defineConfig({
plugins: [webllmAssistant()],
});Use @arronqzy/abuilder/vite (not @arronqzy/webllm-assistant/vite) so pnpm can resolve the plugin from your existing @arronqzy/abuilder dependency.
Required for any Vite build that bundles @arronqzy/react-view / @arronqzy/webllm-assistant. Without this plugin, Vite 5 runs stripLiteral on @mlc-ai/web-llm/lib/index.js and fails with Maximum call stack size exceeded.
Do not put a Vite asset-query import of WebLLM in shared source. Webpack/Umi will bake the query into the async chunk filename, and the gzip size reporter then fails with ENOENT.
Webpack / Umi
Umi apps must not import @arronqzy/webllm-assistant/vite or use @mlc-ai/web-llm?url anywhere. Upgrade to @arronqzy/[email protected]+ (or @arronqzy/[email protected]+); the loader emits a normal mlc-web-llm async chunk instead of @mlc-ai-web-llm?url-lib.async.js.
