@worthy-ventures/metaglotta-incontext
v1.0.1
Published
Screen capture for in-context translation editing: one permission per session, region-cropped to the tab.
Readme
@worthy-ventures/metaglotta-incontext
In-context translation editing: the browser-extension contract, session bootstrapping, and
screenshot capture. Consumed by @worthy-ventures/metaglotta-ngx, which is what an application
normally installs — you need this package directly only to start a session yourself.
import { provideInContext } from '@worthy-ventures/metaglotta-incontext';The extension contract is frozen
The browser extension talks to the page through two sessionStorage keys and a small
postMessage vocabulary. Those literal values are the extension's own — a protocol this project
does not own — and are spelled out in exactly one place each, EXTENSION_STORAGE_PREFIX and
EXTENSION_MESSAGES in src/, with tests in src/extension.test.ts pinning them.
| What | Where | Constant |
| -------------------------------- | ---------------- | --------------------- |
| API key | sessionStorage | API_KEY_STORAGE_KEY |
| API url | sessionStorage | API_URL_STORAGE_KEY |
| Detection, handshake, screenshot | postMessage | EXTENSION_MESSAGES |
They still carry the upstream project's name, and renaming them to match this project's
vocabulary would simply mean the extension never answers. They are the two files the repo's
assert:debranded check deliberately exempts.
A translator installs the extension once, generates their own key under their own permissions,
and activates it on a tab. The handshake reports uiPresent: true, which tells the extension
not to inject its own copy of the editor UI.
Starting a session without the extension
startInContextSession() and bootstrapFromUrl() write the same two sessionStorage keys, so
a session started from a bookmarklet or a ?__i18n=1 link is indistinguishable from one the
extension started.
If no extension answers the detection ping, this package installs its own responder and captures
screenshots with getDisplayMedia:
import { enableScreenshots } from '@worthy-ventures/metaglotta-incontext';
// Must be called from a user gesture. One permission prompt per session.
await enableScreenshots();It never competes with the extension: the probe runs first, and if a real extension replies this stays silent — two responders on one message would race, and the extension's native tab capture is better anyway.
Chromium only. Safari has no tab capture or ImageCapture; Firefox has no
preferCurrentTab.
Why capture is tab-only, and armed once per session
The track is armed once per session rather than prompting per screenshot, because a screenshot request is abandoned after 3000 ms and no one answers a picker that fast. Each capture then grabs a frame in tens of milliseconds.
Capture is constrained to the browser tab and validated afterwards
(track.getSettings().displaySurface === 'browser'), with whole monitors removed from the
picker. This is load-bearing, not caution: key positions come from getBoundingClientRect() —
viewport-relative CSS pixels — and are scaled by imageSize / windowSize per axis with no
translation term. A window capture adds browser chrome at the top and shifts every position
box, and nothing downstream can correct for it, so a non-tab capture is rejected rather than
uploaded misaligned. Where available, Region Capture crops to the layout viewport, which makes
both scale factors exactly devicePixelRatio.
A known upstream issue
useGallery.ts in the upstream web SDK computes key positions after reverting the previewed
text:
try {
screenshot = await takeScreenshot();
} finally {
revert();
setTakingScreenshot(false);
}
const positions = uiProps.findPositions(key, ns); // runs AFTER revert()finally runs before the statement following the try, so the boxes are measured against the
reverted text while the image shows the new text. It only matters when a screenshot is taken in
the same action as a length-changing edit. If it becomes a problem, one patch-package patch on
the unminified bundle fixes it — still far cheaper than forking.
