@adamcjm/dsh-vision
v0.1.0
Published
Vision subagent delegation for DeepSeek Harness: when the main session model cannot read images, delegate image understanding to a one-shot vision subagent; when the main model is natively multimodal, read directly and skip the subagent.
Maintainers
Readme
dsh-vision · 视觉子代理插件
English | 中文
A DeepSeek Harness (dsh) plugin that gives a text-only main model "eyes": when the session needs image understanding, the work is delegated to a one-shot vision subagent running on a vision-capable model; when the main model is natively multimodal, images are read directly and no subagent is ever spawned.
Installation
# npm (recommended)
dsh plugin --profile web add @adamcjm/dsh-vision
# or straight from GitHub
dsh plugin --profile web add github:adamcjm/dsh-vision
# or local development (live edits, no reinstall)
dsh plugin --profile web add link:~/.dsh/dev/dsh-visionThe CLI appends the bundle to dsh.profile.bundles automatically. Restart dsh web to activate.
Verify:
dsh --profile web --dump-config | grep -A6 "id: dsh-vision"How it works
- Automatic capability detection. The plugin checks the current session route's
inputModalities(the same gate the harness'sread_imageuses).- Main model declares
imageinput →vision_agentshort-circuits with "use read_image directly" — no subagent is spawned; the guidance section also tells the model to preferread_image. - Main model is text-only → the model calls
vision_agent(images, question); the plugin starts a one-shot foreground subagent on the configured vision route (the child inherits the session's preset, including its ownread_imagetool), and only the child's text answer returns to the main session.
- Main model declares
- Pasted images, zero message rewriting. In a text-only session the harness projects a pasted image into a stub like
[image omitted ...; attachment sha256:abcdef12].vision_agentaccepts thesha256:reference (full 64-hex id, truncated prefix, or the whole bracketed stub text — it is extracted automatically) and resolves it to the durable attachment's local object file, so the subagent reads the original bytes. No request messages are rewritten, which keeps the harness's agent-loop log-reconstruction invariant intact. - Hardened child. The subagent runs with
maxDepth: 0(no further delegation) andtoolFilterlimited toread_image; image bytes never enter the main session's context.
Configuration
The dsh-vision row in cordis.patch.yml (override by id in the profile's own cordis.patch.yml):
| Field | Default | Meaning |
| --- | --- | --- |
| provider | deepseek-official | Provider route for the vision subagent (must be registered in the LLM catalog) |
| model | deepseek-v4-flash-vision-exp | Vision model (must declare image input; shipped in the built-in catalog) |
| subagentProvider | spawn | Host subagent registry backend |
| maxDepth | 0 | The vision child may not delegate further |
Usage
- Paste an image and ask about it → the main model sees the
sha256:stub → callsvision_agent(["sha256:..." or the whole stub text], "your question")→ the subagent reads the image → the text answer comes back and the conversation continues. - Give image file paths (absolute or workspace-relative) → same flow.
- Main model supports images (e.g. switched to a vision model) → it reads with
read_imagedirectly, no subagent. - Image URLs → the model downloads them to the workspace first, then passes local paths.
Structure
dsh-vision/
├── package.json # bundle manifest (dsh.bundle.patch)
├── cordis.patch.yml # inserts the dsh-vision row (root layer, visible to every session)
└── lib/index.js # host plugin, zero external dependenciesLicense
中文
给纯文本主模型装上「眼睛」的 DeepSeek Harness(dsh)插件:会话需要识图时,自动把任务(识图意图 + 目标图片)交给运行在视觉模型上的一次性 vision sub-agent;主模型原生支持多模态时,直接读图,永不启动子代理。
安装
# npm(推荐)
dsh plugin --profile web add @adamcjm/dsh-vision
# 或直接从 GitHub 安装
dsh plugin --profile web add github:adamcjm/dsh-vision
# 或本地开发(改代码即时生效,无需重装)
dsh plugin --profile web add link:~/.dsh/dev/dsh-visionCLI 会自动把插件追加进 dsh.profile.bundles。重启 dsh web 生效。
验证:
dsh --profile web --dump-config | grep -A6 "id: dsh-vision"工作原理
- 能力自动判断:插件按 harness 的
read_image同款门禁检查当前会话路由的inputModalities。- 主模型声明
image输入 →vision_agent直接短路返回「请用 read_image」,不会 spawn 任何子代理;提示词区也引导模型优先read_image。 - 主模型不支持图片 → 模型调用
vision_agent(images, question),插件启动一次性前台子代理,路由钉在配置的视觉模型上(子代理继承会话预设,自带read_image),子代理读图后仅把文本答案返回主会话继续。
- 主模型声明
- 贴图闭环(零消息改写):纯文本会话中贴图会被官方适配器投影成
[image omitted ...; attachment sha256:abcdef12]占位符。vision_agent接受sha256:引用(完整 64 位、截断前缀、或整段方括号占位文本原样,自动提取),解析为持久化附件的本地对象路径后由子代理读原始字节。不改写任何请求消息,与 harness 的 agent-loop 日志一致性校验完全兼容。 - 安全边界:子代理
maxDepth: 0(禁止再委派)、toolFilter仅允许read_image,图片字节不进入主会话上下文。
配置
cordis.patch.yml 中的 dsh-vision 行(可在 profile 自己的 cordis.patch.yml 里按 id 覆盖):
| 字段 | 默认 | 说明 |
| --- | --- | --- |
| provider | deepseek-official | 视觉子代理的路由 provider(需已在 LLM 目录注册) |
| model | deepseek-v4-flash-vision-exp | 视觉模型(必须声明 image 输入;内置目录自带该模型) |
| subagentProvider | spawn | host subagent 注册表后端 |
| maxDepth | 0 | 子代理禁止再委派 |
使用
- 贴图 + 问题 → 主模型看到
sha256:占位符 → 调vision_agent(["sha256:..." 或整段占位文本], "问题")→ 子代理读图 → 文本答案返回主会话继续。 - 给图片文件路径(绝对或相对工作区)→ 同样流程。
- 主模型支持图片(如切到视觉模型)→ 直接
read_image,不启动子代理。 - 图片 URL → 模型先下载到工作区,再传本地路径。
结构
dsh-vision/
├── package.json # bundle 声明(dsh.bundle.patch)
├── cordis.patch.yml # 插入 dsh-vision 行(root 层,全 profile 会话可见)
└── lib/index.js # 零外部依赖的 host 插件