@lumifai/harness-tool-pack-file-processing
v0.3.0
Published
Extract workspace PDF files into Markdown or text files, with optional document metadata.
Readme
@lumifai/harness-tool-pack-file-processing
Extract workspace PDF files into Markdown or text files, with optional document metadata.
import { createFileProcessingToolPack } from '@lumifai/harness-tool-pack-file-processing';
createFileProcessingToolPack();process_pdf reads a workspace PDF, runs OCR-aware extraction via LiteParse, and returns an output_path reference instead of inline content. Markdown is the default format and preserves more document structure, including tables and links. Bundled agent skills ship in this package and are merged into workspace.skills when the pack is enabled.
The tool applies consistent OCR-aware parsing and returns metadata, annotations, form fields, and page complexity by default. The small set of meaningful optional inputs is:
format:markdown(default) ortext.include_images: extracted image metadata and workspace paths to image files.include_complexity: per-page OCR and layout complexity signals (defaulttrue).
Tool pack id: file-processing. Disabled by default (enabledByDefault: false). Default policy: process_pdf → ask.
