@phamkhachoabk/dsh-compaction-router
v0.1.0
Published
Reroute DeepSeek Harness compaction requests to a wider-context model, so /compact never truncates a summary just because the live model's context is too small to hold the replayed history.
Readme
Compaction Router for DeepSeek Harness
Reroute /compact (and session-title) requests to a second, wider-context
model, so a large session's summary never gets truncated just because the
live chat model's own context window is too small to hold the replayed
history plus the summary output.
Why this exists
Compaction replays the full history region into a request sent to the
live model, tagged purpose: 'compaction'. On a large session the
replayed history alone can consume nearly all of that model's context —
leaving only a few thousand tokens for the summary itself, which then gets
cut off (finish_reason: "length") no matter how high maxTokens is set
in the compaction config. Raising maxTokens does not help: the ceiling is
the model's total context, not the output budget. Pointing compaction at a
model with a genuinely larger context window is the actual fix.
Install
Published on npm as @phamkhachoabk/dsh-compaction-router. Or build a
tarball from a checkout:
pnpm install
pnpm run build
npm pack
dsh plugin --profile web add /abs/path/phamkhachoabk-dsh-compaction-router-0.1.0.tgzRestart the harness afterwards: bundle layers are composed at boot.
Configure
| Setting | Default | Meaning |
| --- | --- | --- |
| provider | gpt9router | Provider route for the compaction target model, exactly as configured on Settings → Models. Leave empty ("") to disable rerouting entirely. |
| model | cc/claude-sonnet-5 | Model id at that route. Needs enough context to hold the replayed session history — see "Choosing a target model" below. |
| maxTokens | 65536 | Output budget for the summary, replacing compaction's own default. |
Choosing a target model
The route and model need real, verified headroom over your largest sessions' replayed-history size — not just "bigger than the default context." Undersized targets fail exactly the same way the live model did. Verify a candidate against your largest session before trusting it in production.
How it intercepts
Same seam as @phamkhachoabk/dsh-vision-bridge: the harness's llm/stream
waterfall, the last boundary before a request leaves for the provider.
Only requests tagged purpose: 'compaction' are touched — ordinary
conversation requests and purpose: 'session-title' requests pass straight
through untouched, since a title generation call is small and does not
need the wider context. A matching request is re-dispatched once with
provider/model/maxTokens overridden; the rerouted copy is marked
(module-local WeakSet) so the interceptor recognizes it on the second
pass — the one its own re-dispatch causes — instead of rerouting it again.
Composes with vision-bridge regardless of registration order. Both
plugins rewrite by fully re-dispatching through ctx.llm.stream(...)
rather than mutating the request mid-chain via next(), so each rewrite
restarts the whole llm/stream waterfall from the top. A compaction
request carrying images gets its images stripped and its target rerouted
in either order — each plugin's own guard makes its second pass a no-op.
Cosmetic note: the harness's own compaction/summary log event
records the provider/model from the engine's configuration (the live
model), not the rerouted target, since the interceptor only changes what
actually gets called — not what the surrounding compaction code believes
it called. The event's token usage, however, is the real usage from the
rerouted call. Expect the provider/model fields in that log line to look
wrong; they are cosmetic only.
Development
pnpm install
pnpm run build
pnpm test