docx-edit
v0.3.2
Published
A JS library that parses DOCX into a virtual component tree and writes paragraph-level text changes back to split OOXML runs.
Maintainers
Readme
docx-edit
一个基于 JavaScript 的 .docx 解析与修改库。核心思路是把 Word 文档解析成一棵虚拟树,所有修改都收敛到虚拟树 diff / patch,再同步回底层 OOXML。相比传统解析库,本库对"动态修改"更友好,支持全文高精度匹配替换、样式建模、SDT 与域代码识别等。
特性
- 解析正文、页眉、页脚、批注、脚注、尾注
- 真正的虚拟树
diff / patch(文本更新、结构增删替换重排) - 段落 / run 样式建模、修改、迁移;样式档案 JSON 导入导出
- 段落跨多个
w:t的整段文本读写(保留脚注引用、数学公式占位符) - 表格、图片、文本框、数学公式读写
- 上标 / 下标、脚注引用、尾注引用、新建脚注
- 结构化文档标签(SDT)精确识别 —— 一行判定目录 / 图表目录
- 域代码(field)保留与读取 ——
TOC/PAGEREF/SEQ等域指令可读 - 提取全文为 HTML;解析标题级别(中英文样式 ID)
安装
npm install docx-edit本库使用 CommonJS 导出,建议 Node.js >=18。
快速开始
const { loadDocx } = require("docx-edit");
async function main() {
const doc = await loadDocx("./sample.docx");
// 全文替换
doc.replaceAll("旧词", "新词");
// 精确判定目录(靠 SDT 元数据,零误判)
const tocTags = doc
.getStructuredDocumentTags()
.filter((sdt) => sdt.getGalleryType() === "Table of Contents");
await doc.saveAs("./sample.modified.docx");
}
main();虚拟树模型
文档会被解析成一棵虚拟树,典型结构:
document
body
paragraph
run
text
sdt ← 结构化文档标签(目录等)
sdtContent
paragraph
table
table-row
table-cell
paragraph
header / footer / comments / footnotes / endnotes支持的节点类型:
document、body、header、footer、footnotes、endnotes、comments、paragraph、run、text、table、table-row、table-cell、hyperlink、tab、break、text-box、comment、footnote、endnote、footnoteReference、endnoteReference、footnoteRef、endnoteRef、image、math、sdt、sdtContent、fldChar、instrText
内部写入流程
无论调用控制器接口还是 doc.patch(nextTree),内部流程一致:
- 从当前文档生成虚拟树副本
- 在副本上修改目标节点
- 调用
doc.patch(nextTree) - patch 引擎执行
INSERT / REMOVE / REPLACE / MOVE / PROPS/TEXT_UPDATE - 将结果同步回底层 OOXML
- 从 XML 重新建树并重建索引
段落文本修改保留 ParagraphTextModel 策略:尽量保留原有 w:r / w:t 和 tab / break,只把新文本重新分配回原有文本节点。
const tree = doc.toComponentTree();
const body = tree.children.find((node) => node.type === "body");
body.children[0].props.text = "新的第一段";
const result = doc.patch(tree);
console.log(result.operations);📚 详细文档
README 只覆盖核心概念,完整 API 与示例见 docs/:
| 文档 | 内容 |
|---|---|
| docs/API.md | 完整 API 参考:加载保存、虚拟树 patch、文档级方法、样式模型、样式档案、所有控制器、createVNode |
| docs/sdt-and-fields.md | 结构化文档标签(SDT)与域代码(Field)的精确识别、读取与创建 |
导出 API
const {
loadDocx,
VirtualWordDocument,
VNode,
createVNode,
cloneVNode,
DocumentPartController,
ParagraphController,
RunController,
TableController,
TableRowController,
TableCellController,
TextBoxController,
StructuredEntryController,
StructuredDocumentTagController,
} = require("docx-edit");常用入口:
doc.getParts();
doc.getBody();
doc.getHeaders();
doc.getFooters();
doc.getParagraphs();
doc.getTables();
doc.getTextBoxes();
doc.getStructuredDocumentTags(); // 所有 SDT
doc.getImages();
doc.getMaths();
doc.getFootnotes();
doc.getEndnotes();
doc.getComments();
doc.replaceAll("旧词", "新词");
doc.addFootnote("脚注内容");
doc.extractHtml();控制器一览(方法签名详见 docs/API.md):
DocumentPartController— body / header / footerParagraphController—getText/setText/replace/getFields/getStyle/getRuns...RunController—getText/getStyle/isFieldBegin/getFieldCode...TableController/TableRowController/TableCellControllerTextBoxControllerStructuredEntryController— comment / footnote / endnoteStructuredDocumentTagController— SDT,详见 docs/sdt-and-fields.md
段落文本中的占位符
读取段落文本时,脚注引用和数学公式以占位符形式出现,修改文本时自动保留:
[[FOOTNOTE_REF:id]]— 脚注引用[[ENDNOTE_REF:id]]— 尾注引用[[MATH:text]]— 数学公式
域标记(fldChar / instrText)不出现在文本里,但域的缓存结果作为普通文本可见。
测试
npm test # 运行全部测试
npm run example # 运行示例脚本当前测试覆盖:段落读写、tab/break 保留、虚拟树 patch(文本/结构/重排)、表格、header/footer/comment/text-box 持久化、段落与 run 样式、样式迁移、styles.xml 解析与继承链、样式档案导入导出、真实样本文档回归、HTML 全文提取、标题级别解析、上下标、脚注/尾注引用、数学公式、新建脚注、SDT 元数据与 round-trip、域代码读取与 setText 不破坏域。
已知边界
- 覆盖常见文本相关 OOXML 节点,不是完整的 Word OOXML 实现
- 样式建模主要覆盖段落和 run 的常用属性
- 对未知节点的策略是尽量保留,而不是细粒度理解和编辑
doc.patch(nextTree)期望目标树由当前树演化而来,不保证支持任意非法结构- 嵌套域(域中域)按线性配对处理,不展开为独立
getFields()项
License
ISC
