@shashimadushan/docx-editor-agent
v0.4.1
Published
AI agent for @shashimadushan/docx-editor-editor. Tool-calling agent that can read, write, insert, delete, format, and search document content via any LLM (OpenAI, Anthropic, or custom).
Maintainers
Readme
@shashimadushan/docx-editor-agent
AI agent for @shashimadushan/docx-editor-editor. Tool-calling agent that can read, write, insert, delete, format, and search document content via any LLM.
Install
npm install @shashimadushan/docx-editor-agent @shashimadushan/docx-editor-editor react react-domQuick start (React)
import * as React from 'react';
import { ReactDocxEditor, type DocxEditor } from '@shashimadushan/docx-editor-editor/react';
import '@shashimadushan/docx-editor-editor/style.css';
import { BackendAdapter } from '@shashimadushan/docx-editor-agent';
import { useDocxAgent } from '@shashimadushan/docx-editor-agent/react';
function App() {
const [editor, setEditor] = React.useState<DocxEditor | null>(null);
// Memoize the adapter — useDocxAgent rebuilds its DocxAgent (and resets
// the conversation) whenever the adapter *reference* changes, so passing
// `new BackendAdapter(...)` inline here would reset on every render.
// BackendAdapter talks to YOUR server (see the "BackendAdapter" section
// below) — no provider API key ever reaches the browser.
const adapter = React.useMemo(() => new BackendAdapter({ baseUrl: '/api/agent' }), []);
const { isRunning, log, run, reset, error } = useDocxAgent({ adapter }, editor);
return (
<>
<ReactDocxEditor onReady={setEditor} />
<button onClick={() => run('add heading called Summary')} disabled={isRunning || !editor}>
Add heading
</button>
<pre>{JSON.stringify(log, null, 2)}</pre>
</>
);
}Quick start (headless)
import { DocxEditor } from '@shashimadushan/docx-editor-editor';
import { DocxAgent, OpenAIAdapter } from '@shashimadushan/docx-editor-agent';
import '@shashimadushan/docx-editor-editor/style.css';
const editor = new DocxEditor({ element: document.getElementById('host')! });
const agent = new DocxAgent({
editor,
adapter: new OpenAIAdapter({
apiKey: process.env.OPENAI_API_KEY!,
model: 'gpt-4o-mini',
}),
});
// Run an autonomous task
const reply = await agent.run(
'Add a heading called "Summary" and a 3-bullet list of key points.',
);
console.log(reply);
// Inspect what happened
console.log(agent.history);LLM adapters
OpenAIAdapter
Works with OpenAI and any OpenAI-compatible provider (Together AI, Groq, Mistral, OpenRouter, vLLM, LM Studio, Ollama with OpenAI compatibility, etc.):
import { OpenAIAdapter } from '@shashimadushan/docx-editor-agent';
new OpenAIAdapter({
apiKey: process.env.OPENAI_API_KEY!,
model: 'gpt-4o-mini',
// Optional: use any OpenAI-compatible provider
baseURL: 'https://api.groq.com/openai/v1',
});AnthropicAdapter
import { AnthropicAdapter } from '@shashimadushan/docx-editor-agent';
new AnthropicAdapter({
apiKey: process.env.ANTHROPIC_API_KEY!,
model: 'claude-3-5-sonnet-20241022',
});BackendAdapter (recommended for anything shipped to users)
OpenAIAdapter/AnthropicAdapter need a provider API key in the browser — fine for a local prototype, not for production. BackendAdapter instead talks to two endpoints on your own server, which holds the real key and makes the provider call for you:
import { BackendAdapter } from '@shashimadushan/docx-editor-agent';
const adapter = new BackendAdapter({
baseUrl: '/api/agent', // relative or absolute; change anytime via adapter.setBaseUrl(url)
// headers: () => ({ Authorization: `Bearer ${myAppSessionToken}` }), // auth to YOUR api, not the LLM provider
});It POSTs the exact same { messages, tools, temperature, maxTokens } body DocxAgent always builds to:
POST {baseUrl}/chat— one-shot; your server responds with JSON{ message, finishReason, usage? }.POST {baseUrl}/chat/stream— streaming; your server respondstext/event-streamwithdata:lines shaped{ type: 'token'|'thinking', text }and a final{ type: 'done', message, finishReason, usage? }.
The simplest server implementation reuses this same package's OpenAIAdapter/AnthropicAdapter on the server, where holding a secret key is fine. Don't just forward whatever tools/system message the client sent, though — that puts the browser in charge of the agent's capabilities. Use @shashimadushan/docx-editor-agent/server's handleAgentChat/handleAgentChatStream instead: they build the tool schema + system prompt from this package's own registry (server-authoritative), call the adapter, and shape the response BackendAdapter expects — a route is a few lines:
// app/api/agent/chat/route.ts (Node or Edge runtime; any server framework works the same way)
import { handleAgentChat } from '@shashimadushan/docx-editor-agent/server';
import { OpenAIAdapter } from '@shashimadushan/docx-editor-agent';
const adapter = new OpenAIAdapter({ apiKey: process.env.OPENAI_API_KEY!, model: 'gpt-4o-mini' });
export async function POST(req: Request) {
return Response.json(await handleAgentChat(await req.json(), adapter));
}// app/api/agent/chat/stream/route.ts — same idea, streaming
import { handleAgentChatStream, AGENT_SSE_HEADERS } from '@shashimadushan/docx-editor-agent/server';
export async function POST(req: Request) {
return new Response(handleAgentChatStream(await req.json(), adapter), { headers: AGENT_SSE_HEADERS });
}See examples/demo/app/api/agent/chat/route.ts and .../chat/stream/route.ts for the complete reference implementation. handleAgentChat* take an optional third argument, { toolRegistry?, systemPrompt? }, if you need a custom tool set (e.g. legal tools merged in, or some builtins disabled) or a different prompt — pass systemPrompt: null to omit it entirely if you'd rather manage that yourself.
Custom adapter
Implement the LLMAdapter interface:
import type { LLMAdapter, LLMMessage, LLMResponse } from '@shashimadushan/docx-editor-agent';
class MyAdapter implements LLMAdapter {
async complete({ messages, tools }): Promise<LLMResponse> {
// Call your LLM...
return {
message: {
role: 'assistant',
content: 'Done.',
toolCalls: [
{ id: 'call_1', name: 'rewrite_range', arguments: { from: 3, to: 3, html: '<h2>Hello</h2>' } },
],
},
finishReason: 'tool_calls',
};
}
}Built-in tools (31)
31 tools across 13 categories, each carrying category + risk
(read / write / destructive) metadata used by the policy gate (see
"Security" below). ⚠ marks a destructive tool. For the full generated
listing, call buildToolCategoryGuide(tools) or read
skill/references/agent-api.md.
The tool count went 69 → 41 (three consolidation passes merging near-duplicate
single-property tools into dispatchers) → 31 with the HTML+diff rewrite
architecture: rewrite_range is now the primary content-authoring tool — the
LLM writes HTML for a block range instead of calling a dozen single-property
tools. The superseded ones (insert_paragraph, set_bold, set_heading,
append_html, batch_insert, batch_format, and others — see
tools/index.ts's doc comment for the full list) were REMOVED entirely, not
just unregistered — they're no longer importable, individually or otherwise.
This is a breaking change if you were importing one of them directly;
rewrite_range covers the same ground via HTML.
| Category | Tools |
|---|---|
| Inspection (read) | get_outline, get_word_count, list_blocks, get_block, get_block_range, find_block, get_selection, get_document, get_block_count |
| Batch | batch_delete ⚠ (non-contiguous indices), batch_reorder |
| Insert | rewrite_range ⚠ — the primary content-authoring tool |
| Insert media | insert_image, insert_link |
| Style | set_paragraph_style (semantic style-id, not HTML-representable) |
| Image layout | set_image_layout (alignment/wrap/size in one call) |
| Table | get_table_info, edit_table (set cell / add row / add column), delete_table_line ⚠ |
| Manipulation | insert_at_cursor, edit_selection (replace/mark on live selection) |
| Cursor-relative | get_cursor_block, delete_current_block ⚠ |
| Structure | move_block |
| Search | find_text, replace_text ⚠ (grounded single-match at or all-occurrences), clear_document ⚠, replace_placeholders |
| Page-level | set_watermark, set_page_html (header/footer) |
| Navigate | scroll_to_block |
Selected tool parameters:
Content authoring
| Tool | Description | Parameters |
|---|---|---|
| rewrite_range | Replace a contiguous block range [from, to] with HTML — parsed through the live editor's real schema, diffed against the prior content, and applied via the same splice primitive every other block tool uses. Pass to: from - 1 to INSERT before from without replacing anything (no need to re-type existing content); from equal to the block count appends at the end. No <img> (use insert_image/set_image_layout) or <ins>/<del> (reserved for track-changes). | from: number, to: number, html: string — returns { changedBlocks, insertOnly, citations: [{ blockIndex, before, after }] } |
Insert media (2)
| Tool | Description | Parameters |
|---|---|---|
| insert_image | Insert an image | src: string (URL or data URI), alt?: string, title?: string, index?: number |
| insert_link | Insert a hyperlink | href: string, text: string, index?: number (omit to append a new paragraph) |
Style / image layout
| Tool | Description | Parameters |
|---|---|---|
| set_paragraph_style | Convert a block to a different style | index: number, style: 'paragraph' \| 'h1'-'h6' \| 'blockquote' \| 'codeBlock' |
| set_image_layout | Set alignment/wrap/size on the most recently inserted image, in one call | alignment?: 'inline' \| 'left' \| 'center' \| 'right', wrap?: 'none' \| 'left' \| 'right', width?: string, height?: string |
Table (3)
| Tool | Description | Parameters |
|---|---|---|
| get_table_info | Inspect table structure | index: number |
| edit_table | Set a cell's text, or add a row/column | index: number, action: 'set_cell' \| 'add_row' \| 'add_column', ... |
| delete_table_line | Delete a row or column | index: number, axis: 'row' \| 'column', lineIndex: number |
Mutate structure
| Tool | Description | Parameters |
|---|---|---|
| move_block | Move an existing block to a new position (content unchanged) | from: number, to: number |
| delete_current_block | Delete the block the cursor is in (or anchorIndex) | anchorIndex?: number |
Search & clear (4)
| Tool | Description | Parameters |
|---|---|---|
| find_text | Find occurrences of a substring, or a regex pattern with regex: true (supports lookahead/lookbehind) | query: string, caseSensitive?: boolean (default: false), regex?: boolean, flags?: string |
| replace_text | Replace occurrences of a substring, or a regex pattern with regex: true (capture groups usable in replace via $1, $2, ...); pass at: { blockIndex, offset } from a find_text match to ground the edit to one exact occurrence | find: string, replace: string, caseSensitive?: boolean, firstOnly?: boolean, regex?: boolean, flags?: string, at?: { blockIndex, offset, length? } — returns { replaced, citation? } |
| clear_document | Empty the document | — |
| replace_placeholders | Fill ${name} placeholders everywhere from one values object | values: Record<string, string> |
Navigate (1)
| Tool | Description | Parameters |
|---|---|---|
| scroll_to_block | Scroll editor to a block | index: number |
DocxAgent options
interface DocxAgentOptions {
editor: DocxEditor; // the editor to operate on
adapter: LLMAdapter; // LLM adapter (OpenAI/Anthropic/Backend/custom)
systemPrompt?: string; // override the default system prompt
tools?: ToolRegistry; // custom tool registry (default: all builtins)
disableTools?: string[]; // disable specific builtin tools by name
customTools?: AgentTool[]; // additional tools on top of builtins
temperature?: number; // default: 0.2
maxTokens?: number; // default: 1024
maxRounds?: number; // max tool-call rounds per run() (default: 8)
// Security / tool safety (all optional):
disableToolsByRisk?: ('read' | 'write' | 'destructive')[]; // drop risk tiers from schema + execution
allowedCategories?: ToolCategory[]; // allow-list categories only
onConfirmDestructive?: (info: { toolName: string; args: any; preview: DestructivePreview }) => boolean | Promise<boolean>;
// Token/cost optimization (all optional, all opt-in):
pruneStaleDocumentDumps?: boolean; // replace oversized prior tool results with a placeholder each run()
smartToolSelection?: boolean; // advertise only tool categories relevant to each prompt (keyword heuristic, safe fallback to all)
}preview is a document-aware summary ({ summary, blocks: [{ index, type, textPreview }] })
built from the CURRENT document + the call's args before execution — e.g. for
batch_delete({ indices: [5, 10] }) it resolves each index to its block type
and a text snippet, so you can render "Delete 2 blocks: Heading (5), Paragraph (10)"
instead of showing raw indices.
Security
The agent applies defense-in-depth without touching the editor's mutation
flow. Full details in skill/references/agent-api.md ("Security model") and
ARCHITECTURE.md.
- Tool risk gating — every tool has a
risk(read/write/destructive). PassonConfirmDestructiveto confirm before destructive tools run, ordisableToolsByRisk/allowedCategoriesto drop them from both the schema and the executable set (e.g. a read-only reviewer agent).<AgentPanel>also accepts a top-levelonConfirmDestructiveprop (distinct from the one above) that renders a built-in confirmation modal from the samepreviewdata — use that instead of wiringagentOptions.onConfirmDestructiveby hand if you want a ready-made dialog rather than building your own. - Content sanitization —
rewrite_range/set_page_htmlsanitize HTML (strip scripts, event handlers, unsafe URLs;styleattributes are filtered to a layout/typography property allow-list rather than stripped outright);insert_image/insert_linkreject dangerous URL schemes. Swap in a stricter sanitizer viasetHtmlSanitizer. - Server request bounds —
validateAgentRequest(from/server) clampstemperature/maxTokensand caps payload size. Auth + rate limiting remain the consuming app's responsibility.
Disable specific tools
new DocxAgent({
editor,
adapter,
disableTools: ['clear_document'], // prevent the agent from clearing the doc
});Custom system prompt
new DocxAgent({
editor,
adapter,
systemPrompt: 'You are a strict legal-document editor. Only make changes the user explicitly requests.',
});Custom tools
import type { AgentTool } from '@shashimadushan/docx-editor-agent';
const translateTool: AgentTool = {
name: 'translate_document',
description: 'Translate all text in the document to the target language.',
parameters: {
type: 'object',
properties: {
language: { type: 'string', description: 'Target language, e.g. "Spanish".' },
},
required: ['language'],
},
execute: async (params, ctx) => {
const text = ctx.getText();
const translated = await myTranslateAPI(text, params.language);
// ... replace text in document via ctx.setDoc() ...
return { original: text, translated };
},
};
new DocxAgent({
editor,
adapter,
customTools: [translateTool],
});AgentTool interface
interface AgentTool<TParams = any, TResult = any> {
name: string; // unique tool name
description: string; // the LLM uses this to decide when to call
parameters: JSONSchema; // JSON Schema for params
execute: (params: TParams, ctx: AgentToolContext) => Promise<TResult> | TResult;
}
interface AgentToolContext {
editor: TipTapEditor; // the underlying TipTap editor
doc: any; // current document as JSON
setDoc(json: any): void; // replace the whole document
appendContent(content: string | any): void; // append HTML or JSON node
insertAtCursor(content: string | any): void; // insert at cursor
getText(): string; // plain text of the document
}useDocxAgent React hook
const { agent, isRunning, log, run, stop, reset, error } = useDocxAgent(options, editor);| Return | Type | Description |
|---|---|---|
| agent | DocxAgent \| null | The agent instance (null until editor is ready) |
| isRunning | boolean | True while the agent is processing |
| log | LogEntry[] | Conversation log |
| run(prompt) | Promise<void> | Send a user prompt to the agent |
| stop() | void | Cancel the in-flight run() call. A safe no-op when nothing is running. |
| reset() | void | Clear conversation history |
| error | string \| null | Last error message |
Cancelling a run (stop())
run() defaults to maxRounds: 8 (see DocxAgent options below), and every
round can call tools that mutate the document — so a user watching a
multi-round turn go somewhere unwanted needs a way out. stop() gives them
one: it aborts the AbortSignal threaded into DocxAgent.run() (see
AgentRunOptions.signal), which DocxAgent checks between tool-call rounds
and passes into the LLM adapter's request. Cancellation therefore takes
effect at the next opportunity, not instantly — a round already in progress
finishes its current tool calls (or its current LLM request rejects) before
isRunning flips back to false.
const { isRunning, run, stop } = useDocxAgent({ adapter }, editor);
<button onClick={isRunning ? stop : () => run(prompt)}>
{isRunning ? 'Stop' : 'Send'}
</button><AgentPanel> wires this up already — its Send button becomes a Stop
button while a run is in flight, no extra props needed.
A stopped run can have already edited the document. Earlier tool-call
rounds in the same turn may have run before the cancellation landed, and
DocxAgent does not automatically undo them just because the run was
cancelled (that's a deliberate choice — see AgentRunOptions.signal's doc in
src/agent/types.ts). When a run is cancelled, useDocxAgent appends a
'stopped' entry to log (see LogEntry below) instead of setting error
— the user asked for this, so it isn't treated as a failure — and its text
says plainly whether the document was already changed:
"Stopped. The document was already changed in N place(s) before this run was cancelled."— the cancellation landed between tool-call rounds andDocxAgentreturned grounded citations for what changed."Stopped. This turn already made changes to the document before it was cancelled."— same, but the mutating tool(s) that ran didn't report citations (not every tool does)."Stopped. Any changes from this turn were automatically rolled back."— the cancellation landed mid-request to the LLM adapter (the more common case in practice) andDocxAgent's defaultsnapshot.onError: truerestored the document, so nothing from this turn is actually standing."Stopped before any changes were made."— nothing had mutated the document yet when the stop took effect.
If you're building a fully custom UI on useDocxAgent (rather than using
<AgentPanel>), render 'stopped' log entries distinctly from both a normal
assistant reply and an error — it isn't either.
Memoize options.adapter. The hook rebuilds its internal DocxAgent
(and clears log/error) whenever editor or options.adapter changes
identity — that's what makes swapping providers at runtime (e.g. a dropdown
that switches between BackendAdapter/OpenAIAdapter/AnthropicAdapter, or
picks up a newly-entered API key or a changed baseUrl) actually take effect
instead of silently continuing to run whatever adapter was active on the
first render. The tradeoff: passing a fresh adapter instance on every render
(adapter: new BackendAdapter({ baseUrl }) inline in JSX) resets the
conversation on every render too. Wrap it in React.useMemo(() => new
BackendAdapter({ baseUrl }), [baseUrl]), keyed on whatever actually
determines the adapter (provider kind, API key/baseUrl, model), same as
examples/demo/app/page.tsx does.
LogEntry types
type LogEntry =
| { id: string; type: 'user'; text: string; ts: number }
| { id: string; type: 'assistant'; text: string; ts: number }
| { id: string; type: 'thinking'; text: string; ts: number; streaming: boolean }
| {
id: string; // the tool call's id — the matching result updates this same entry
type: 'tool';
name: string;
args: any;
status: 'running' | 'done' | 'error';
result?: any;
summary?: string;
ts: number;
}
| { id: string; type: 'stopped'; text: string; ts: number };A tool_call and its tool_result are merged into a single 'tool' entry
(matched by call id) rather than two separate log entries — the entry starts
as status: 'running' and updates in place once the result comes back. The
'thinking' entry streams in incrementally when the adapter supports it
(currently AnthropicAdapter with thinking: true); streaming is true
until that round's reasoning is complete. The 'stopped' entry appears when
stop() cancels a run — see "Cancelling a run (stop())" above for what
its text says and why.
How the agent loop works
When you call agent.run(prompt):
- Send
{ system prompt, conversation history, tool definitions }to the LLM - If the LLM returns a tool call, execute it and append the result to history
- Repeat (up to
maxRounds) until the LLM returns a plain text response with no tool calls - Return the final text response
User: "Add a heading called Summary"
↓
LLM: tool_call(rewrite_range, { from: 3, to: 3, html: "<h2>Summary</h2>" })
↓
Agent executes rewrite_range → returns { changedBlocks: 1, citations: [...] }
↓
LLM: "Done — I added a 'Summary' heading to the document."
↓
return "Done — I added a 'Summary' heading to the document."Example: full agent UI
function AgentPanel({ editor }: { editor: DocxEditor | null }) {
const adapter = React.useMemo(() => new BackendAdapter({ baseUrl: '/api/agent' }), []); // see note above — memoize this
const { isRunning, log, run, reset, error } = useDocxAgent({ adapter }, editor);
const [prompt, setPrompt] = React.useState('');
const send = async () => {
if (!prompt.trim()) return;
await run(prompt);
setPrompt('');
};
return (
<div style={{ width: 340, display: 'flex', flexDirection: 'column' }}>
<div style={{ flex: 1, overflow: 'auto' }}>
{log.map((entry, i) => (
<div key={i}>
{entry.type === 'user' && <p><b>You:</b> {entry.text}</p>}
{entry.type === 'thinking' && <p><i>Thinking: {entry.text}</i></p>}
{entry.type === 'assistant' && <p><b>Agent:</b> {entry.text}</p>}
{entry.type === 'tool' && (
<pre>→ {entry.name} ({entry.status}): {entry.summary ?? JSON.stringify(entry.args)}</pre>
)}
</div>
))}
</div>
<input
value={prompt}
onChange={(e) => setPrompt(e.target.value)}
onKeyDown={(e) => e.key === 'Enter' && send()}
placeholder="Ask the agent…"
disabled={isRunning}
/>
<button onClick={send} disabled={isRunning}>Send</button>
<button onClick={reset}>Reset</button>
{error && <p style={{ color: 'red' }}>{error}</p>}
</div>
);
}Package structure
src/
├── agent/ # DocxAgent orchestrator (LLM ↔ tool loop)
│ ├── index.ts # public barrel — import from here (or '@shashimadushan/docx-editor-agent')
│ ├── core.ts # the DocxAgent class: run() loop, snapshot/rollback, target memory
│ ├── types.ts # DocxAgentOptions, AgentRunResult, AgentRunCallbacks, etc.
│ ├── system-prompt.ts # DEFAULT_SYSTEM_PROMPT
│ └── target-memory.ts # cross-round "last resolved target" memory (internal)
├── tools/ # 31 tools + ToolRegistry + risk/category policy, one file per category
├── adapters/ # OpenAIAdapter, AnthropicAdapter, BackendAdapter
├── security/ # HTML/URL sanitizer + validateAgentRequest
├── react/ # React bindings
│ ├── AgentPanel.tsx # the drop-in chat panel component
│ ├── DestructiveConfirmModal.tsx # built-in destructive-action confirm dialog
│ ├── ConversationTurn.tsx # per-turn log grouping/rendering
│ ├── ThinkingBlock.tsx # collapsible "thinking" trace
│ ├── ToolTimeline.tsx # collapsible tool-call timeline
│ ├── tool-labels.ts # tool-name → plain-language label lookup
│ ├── Dot.tsx # pulsing loading-dot indicator
│ ├── useDocxAgent.ts # headless hook (build your own UI)
│ ├── abortable-run.ts # React-free AbortController bookkeeping behind stop()
│ ├── stopped-message.ts # builds the 'stopped' log entry's text honestly (edits may have landed)
│ └── log-types.ts # LogEntry union used by both AgentPanel and useDocxAgent
├── index.ts / react.ts / legal.ts / server.ts # the 4 published entry points (see `package.json` "exports")Tests are colocated with the source they cover (*.test.ts next to its
.ts, e.g. agent/core.test.ts, tools/search.test.ts), same convention
throughout the package — there's no separate test/ tree.
License
MIT
