annotaitr
v1.4.0
Published
Browser-based image and Markdown annotator for AI-assisted review, with built-in web page capture
Downloads
666
Maintainers
Readme
An AI coding agent plugin that opens images, captured web pages, videos, GIFs, PDFs, or Markdown files in a browser-based annotator.

An AI coding agent can read source code but has no way to point at a rendered page or a prose document and say "this, right here." annotaitr closes that gap: it captures a web page or opens an image or Markdown file in the browser, lets a person mark it up with boxes, arrows and comments or with text selections, and hands the agent back structured feedback it can act on directly, instead of the person writing out pixel coordinates or line numbers by hand.
✨ Features
Which mode runs is auto-detected from the target: see Usage below and How it works for the mechanism behind both.
Image and web page review:
- Web page capture: full-page screenshot of any
http(s)URL via Playwright, at a chosen viewport and optional delay; switch viewport, section or delay from the open tab to capture again in place - Clipboard support: run with no target to annotate whatever screenshot is on the (macOS) clipboard, or pass a screenshot pasted into the Claude Code chat
- Drawing tools: boxes, arrows (with an optional dimension-line style for marking distance/spacing), freehand marks, highlighter marks, and numbered comment pins, each with an optional comment and color
- Coarse position descriptions: feedback names each annotation's plain-language position, and flags annotations positioned close together
- Page element names: on a captured URL, the Element tool picks a page element straight from the screenshot, and feedback names the element under each mark (tag, alt or text, media file, short selector), so the agent can find it in the source
- Annotated screenshot export: submitting bakes the markup into a copy of the image and passes its path to the agent; the annotator can also copy that image or the feedback as Markdown, or save the image, for a ticket or a colleague
- Voice notes: speak a comment instead of typing it, transcribed locally with whisper.cpp when it is installed
- Videos and GIFs: annotate screen recordings on a timeline, as single moments or spans; the agent gets each annotated frame as a PNG plus a strip per span and an overview
- PDFs: review a generated slide deck or document page by page; the agent gets one annotated image per page, feedback grouped by page and the source file to edit
Markdown and plain-text review:
- Multi-file support: review multiple files in one session, with a Files overview of notes and reviewed files
- Config and data files: annotate YAML, JSON, TOML, CSV, XML and more as raw source with line numbers
- Linked navigation: click relative
.mdlinks to add them to the review - LaTeX math, Mermaid, PlantUML, and Kroki diagrams: rendered inline, annotatable as a whole
- File references: type
@in a comment to autocomplete other project files - Quick labels: categorize a selection instantly with a predefined label
- Annotation persistence: annotations auto-save to the server and survive page reloads
Both modes:
- Export and import: annotations as Markdown or JSON, to continue a review later
- Dark mode, undo/redo, and an auto-close timer after submitting
- Iterative review: the agent applies your feedback and re-opens the annotator for another round until you approve
🔥 Installation
[!IMPORTANT] Requires Node.js 22.13+ and npm. Image mode additionally needs
playwrightand@napi-rs/canvas, PDFs needpdfjs-dist, alloptionalDependenciesinstalled by default. A markdown-only install can skip them and gets an actionable error if image mode is ever invoked without them. Web page capture also needs a Chromium build for playwright, which npm does not download:npx playwright install chromium.
Upgrading from md-annotator? See docs/migration.md.
Claude Code plugin
claude plugin marketplace add konradmichalik/annotaitr
claude plugin install annotaitr@annotaitrOr via the installer script, which also installs the standalone CLI:
curl -fsSL https://konradmichalik.github.io/annotaitr/install.sh | bashOpenCode plugin
Markdown-only for now. Add to opencode.json:
{
"plugin": ["annotaitr-opencode@latest"]
}Mistral Vibe skill
Markdown-only for now, and drives the standalone CLI, so install that too:
curl -fsSL https://konradmichalik.github.io/annotaitr/install.sh | bash
cp -r apps/vibe/skills/annotate ~/.vibe/skills/annotateStandalone CLI
npm install -g annotaitr🚀 Quick start
annotaitr README.mdOpens README.md in the browser; approve it or leave annotations, and
annotaitr prints the result to stdout once you're done.
⚡ Usage
Claude Code:
/annotaitr:review ./anything # auto-detects image vs. markdownOr force a mode directly: /annotaitr:md README.md, /annotaitr:image ./mockup.png.
A screenshot pasted into the chat works as a target too: /annotaitr:image [Image #1]. Claude looks up the saved image and opens the annotator on it in the background, which takes a few seconds longer than passing a path.
OpenCode:
/annotate:md README.mdor the tool directly: annotate_markdown({ filePath: "/path/to/file.md" }).
Mistral Vibe:
/annotate README.mdStandalone CLI, mode auto-detected from the target:
annotaitr README.md # markdown
annotaitr ./mockup.png # image, local file
annotaitr http://localhost:3000 # image, capture
annotaitr ./bug-recording.mov # image, video on a timeline
annotaitr ./deck.pdf --source ./deck.pptx # image, PDF page by pageFull flag and environment variable reference: docs/usage.md.
📚 Documentation
| Topic | What's inside | |-------|----------------| | Usage | Every flag, environment variable, exit code, and the mode-detection rules | | How it works | The annotation and review-loop mechanism behind each mode | | Development | Local setup, build commands, plugin testing | | Design | Design rules for the UI and the reference screens |
🧑💻 Contributing
Please have a look at CONTRIBUTING.md.
💎 Credits
Heavily inspired by plannotator, which pioneered this general browser-markup-and-feedback approach for reviewing planning documents; annotaitr applies it to images, captured web pages, and Markdown/plain-text files.
⭐ License
This project is licensed under MIT.
