@molecule/api-text-provenance
v1.0.1
Published
Attributes a document's paragraphs to the agent sessions that wrote them: human or AI, the prompt and model behind each, and the AI word share
Maintainers
Readme
@molecule/api-text-provenance
Auto-generated, AI-first package reference for the molecule.dev ecosystem. It is written to be read by coding agents as much as by people, and is generated from this package's source — edit
src/index.tsJSDoc, not this file.
Which paragraphs of a document an AI wrote — and the prompt and model behind each.
Give it the document's blocks and the coding-agent sessions that wrote them
(read with @molecule/api-agent-transcript); it returns each block as
human or ai, the person's message each AI block answered (as typed),
the model that wrote it, and the document's AI word share. The example is
the whole build step for a blog or docs site: a post's markdown file and its
folder of transcript exports in, provenance.json out — per-block
{ origin, prompt, model } plus aiShare. @molecule/api-text-provenance-overlap
is the provider.
Quick Start
// provenance.ts — runs in Node at build time (a build script or a Vite plugin), never in the page.
import { existsSync, mkdirSync, readdirSync, readFileSync, writeFileSync } from 'node:fs'
import { dirname, join } from 'node:path'
import {
canReadTranscript,
readTranscript,
setProvider as setTranscriptReader,
} from '@molecule/api-agent-transcript'
import { provider as anyTranscript } from '@molecule/api-agent-transcript-autodetect'
import { attributeText, setProvider as setAttribution } from '@molecule/api-text-provenance'
import { provider as wordOverlap } from '@molecule/api-text-provenance-overlap'
setTranscriptReader(anyTranscript) // reads Claude Code, Codex and Molecule IDE exports
setAttribution(wordOverlap)
// One block of the post, in page order. `prompt` and `model` are on every ai span.
export interface ProvenanceSpan {
text: string // the block's markdown: a paragraph, heading, list or code block
origin: 'human' | 'ai'
prompt?: string // the person's message the AI was answering, as typed
model?: string // the model that wrote it, as the transcript names it
}
// What /<slug>/provenance.json holds.
export interface Provenance {
aiShare: number // 0..1, the share of the post's words the AI wrote
words: number
aiWords: number
prompts: string[] // the distinct prompts behind the ai spans, in page order
spans: ProvenanceSpan[]
}
// The post's top-level blocks: front matter dropped, split on blank lines, never inside a code fence.
export function markdownBlocks(markdown: string): string[] {
const body = markdown.replace(/^---\r?\n[\s\S]*?\r?\n---\r?\n/, '')
const blocks: string[] = []
let lines: string[] = []
let inFence = false
for (const line of body.split(/\r?\n/)) {
if (/^\s*(`{3}|~{3})/.test(line)) inFence = !inFence
if (!inFence && line.trim() === '') {
if (lines.length > 0) blocks.push(lines.join('\n'))
lines = []
} else {
lines.push(line)
}
}
if (lines.length > 0) blocks.push(lines.join('\n'))
return blocks
}
// Attribute one post from its markdown file and the folder holding its transcript exports.
// A missing or empty folder is a 100% human post; files that are not transcripts are skipped.
export function postProvenance(markdownFile: string, transcriptDir: string): Provenance {
const blocks = markdownBlocks(readFileSync(markdownFile, 'utf8'))
const files = existsSync(transcriptDir)
? readdirSync(transcriptDir, { withFileTypes: true })
.filter((entry) => entry.isFile())
.map((entry) => entry.name)
.sort()
: []
const sessions = files
.map((name) => ({ text: readFileSync(join(transcriptDir, name), 'utf8'), fileName: name }))
.filter((input) => canReadTranscript(input))
.map((input) => readTranscript(input))
const result = attributeText({ paragraphs: blocks, sessions })
return {
aiShare: result.aiShare,
words: result.words,
aiWords: result.aiWords,
prompts: result.prompts,
spans: result.paragraphs.map((p): ProvenanceSpan => {
const text = blocks[p.index]
return p.origin === 'ai'
? { text, origin: 'ai', prompt: p.prompt, model: p.model }
: { text, origin: 'human' }
}),
}
}
// Write provenance.json, creating its folder.
export function writeProvenance(outFile: string, provenance: Provenance): void {
mkdirSync(dirname(outFile), { recursive: true })
writeFileSync(outFile, `${JSON.stringify(provenance, null, 2)}\n`)
}
// In the build, for each PUBLISHED post (skip drafts), after the site's own build has written dist/:
// writeProvenance('dist/my-post/provenance.json', postProvenance('posts/my-post.md', 'transcripts/my-post'))Type
core
Installation
npm install @molecule/api-text-provenance @molecule/api-agent-transcript @molecule/api-bond @molecule/api-i18nAPI
Interfaces
Attribution
A document's attribution.
interface Attribution {
/** One entry per input paragraph, in order. */
paragraphs: ParagraphAttribution[]
/** Words in the whole document. */
words: number
/** Words in the paragraphs attributed to the AI. */
aiWords: number
/** `aiWords / words` (0 for an empty document). */
aiShare: number
/** The distinct prompts behind the AI paragraphs, in the order they first appear in the document. */
prompts: string[]
}AttributionInput
What to attribute: the document's paragraphs and the sessions that may have written them.
interface AttributionInput {
/** The document's paragraphs as plain text, in order (headings and list items may be passed as their own blocks). */
paragraphs: readonly string[]
/** The agent sessions behind the document (`@molecule/api-agent-transcript`). None = all human. */
sessions: readonly AgentSession[]
/** Tuning. */
options?: AttributionOptions
}AttributionOptions
Tuning for an attribution.
interface AttributionOptions {
/**
* The share of a paragraph's words that must come from the AI for the
* paragraph to count as AI-written. Default 0.5.
*/
minAiShare?: number
}ParagraphAttribution
One paragraph's attribution.
interface ParagraphAttribution {
/** The paragraph's position in the input. */
index: number
/** `ai` when at least `minAiShare` of its words came from the AI and not from the user. */
origin: ParagraphOrigin
/** How many words it has. */
words: number
/** How many of them came from the AI (and were not typed by the user). */
aiWords: number
/** For an `ai` paragraph: the user message, as typed, that the writing answered. */
prompt?: string
/** For an `ai` paragraph: the model that wrote it, when the session records one. */
model?: string
/** For an `ai` paragraph: which session (index into `sessions`) and turn (index into its `turns`) wrote it. */
source?: { session: number; turn: number }
}TextProvenanceProvider
The contract every attribution bond implements.
interface TextProvenanceProvider {
/**
* Attribute each paragraph to the human or to the AI.
*
* @param input - The paragraphs, the sessions, and tuning.
* @returns The attribution.
*/
attribute(input: AttributionInput): Attribution
}Types
ParagraphOrigin
Who wrote a paragraph.
type ParagraphOrigin = 'human' | 'ai'Functions
attributeText(input)
Attribute a document's paragraphs to the human or to the AI with the bonded provider.
function attributeText(input: AttributionInput): Attributioninput— The paragraphs, the agent sessions behind them, and tuning.
Returns: Per paragraph: origin, prompt and model; for the document: word counts, AI share and prompts.
getProvider()
Retrieves the bonded attribution provider, throwing if none is configured.
function getProvider(): TextProvenanceProviderReturns: The bonded provider.
hasProvider()
Checks whether an attribution provider is bonded.
function hasProvider(): booleanReturns: true if a provider is bonded.
setProvider(provider)
Registers an attribution provider as the active one. Called during application startup.
function setProvider(provider: TextProvenanceProvider): voidprovider— The provider to bond.
Available Providers
| Provider | Package |
| ------------------------------- | --------------------------------------- |
| Text provenance by word overlap | @molecule/api-text-provenance-overlap |
Injection Notes
Requirements
Peer dependencies:
@molecule/api-agent-transcript^1.0.0@molecule/api-bond^1.0.1@molecule/api-i18n^1.0.1
Runtime Dependencies
@molecule/api-agent-transcript@molecule/api-bond@molecule/api-i18nDo NOT write your own transcript parser, paragraph matcher or prompt pairing. Bond the packages exactly as the example does, and install all four:
npm install @molecule/api-agent-transcript @molecule/api-agent-transcript-autodetect @molecule/api-text-provenance @molecule/api-text-provenance-overlap.Do NOT import these from page or client code. They are server-only and throw in a browser bundle. Run the example at build time; the page reads its output (the file, or the spans handed to the prerender).
Do NOT write into
dist/before the site's own build — Vite empties it. Write after the build, or emit the JSON from a Vite plugin:this.emitFile({ type: 'asset', fileName: 'my-post/provenance.json', source: JSON.stringify(map) }).Do NOT look a span's prompt up in
promptsby position. Everyaispan carries its ownpromptandmodel;promptsis the distinct list, for the count ("from 3 prompts").spans[i]is blockiofmarkdownBlocks(file). Render the page from these same spans (see@molecule/app-margin-notes-react, whose example takes this file as its input) so the page and the map always agree.A block made of words the person typed — a heading copied from the prompt, a paragraph pasted into it — is
human. That is correct, not a miss.A light human edit keeps a block
ai; a rewrite that keeps less thanminAiShare(default 0.5) of the AI's words makes ithuman. The share counts words, and a block counts wholly one way.No sessions (a missing or empty transcript folder) means every span is
humanandaiShareis 0.
