hyphal
v0.3.0
Published
Extract documents into a queryable RDF knowledge graph, with persistent CLI commands and automatic Claude Code integration
Maintainers
Readme
hyphal
Turn your documents into something your agents can actually reason over.
Hand hyphal your documents. It builds an RDF ontology and a validated fact graph as it reads, saves both to disk, and lets you — or your Claude Code sessions — ask precise questions against everything it's read, without re-extracting anything.
export ANTHROPIC_API_KEY=sk-...
npx hyphal@latest extract report.pdf contract.docx notes.txt
npx hyphal@latest ask "what's the payment term in the contract?"Always use
@latest.npx hyphalwithout a version tag can silently run a cached older version on your machine —@latestforces npx to fetch the current published release every time.
Runs entirely on your machine, on your own Anthropic key. No server, no hosted service.
What happens
hyphal extract reads each document (.txt, .md, .pdf, .docx), builds or extends an ontology and fact graph, and saves both as ontology.ttl and facts.ttl in your working directory. Run it again later on a new document — it loads what's already there and only adds what's genuinely new, reusing schema and entities it already knows instead of duplicating them.
hyphal ask "<question>" loads those two files, gives an agent the schema as context, and lets it write and run real SPARQL queries against your facts. No LLM cost on the query itself — only the reasoning turn that decides what to ask. Ask it a hundred questions; you pay for extraction once.
Wiring it into Claude Code
Drop this repo's CLAUDE.md into your own project root. Claude Code reads it automatically at the start of every session and knows, without you re-explaining, to run hyphal ask for questions and hyphal extract when a new document needs to be added to the graph. Nothing to configure beyond having the file present and the two .ttl files it points to.
Architecture
flowchart TD
A[Documents: txt, md, pdf, docx] --> B[hyphal extract]
B --> C[ontology.ttl + facts.ttl]
C --> D[hyphal ask]
D --> E[Claude Code session, via CLAUDE.md]The mechanism
Every extraction call sends the LLM the ontology's current serialized state as context. Whatever it proposes is diffed against exactly that snapshot — an exact structural comparison, not fuzzy or embedding-based. Anything it restates that already exists is dropped before it's ever written. The same discipline applies to entities: before extracting facts, the model is shown a summary of what's already known, so the same real-world entity mentioned across different documents resolves to one node, not duplicates.
Run two documents through the same graph and this is directly observable: the second contributes a noticeably smaller delta than the first, because the concepts it needs mostly already exist.
src/diff.ts (complementInserts) and test/diff.test.ts prove this deterministically, no API key required.
CLI reference
npx hyphal@latest extract <file1> [file2 ...] [--api-key sk-...]Reads ANTHROPIC_API_KEY from the environment by default. Loads any existing ontology.ttl/facts.ttl first, so calling this repeatedly over time builds one continuous graph.
npx hyphal@latest ask "<question>" [--api-key sk-...]Loads the current graph fresh on every call — always reflects the latest extraction, across separate process runs, no manual reload step.
Advanced: using it as a library
For custom integrations — your own agent loop, a different tool-calling setup, embedding it in a larger app:
npm install hyphalimport { OntologyStore, FactsStore, evolveOntology, extractFacts, queryGraph, searchEntities } from "hyphal";
const prefixes = {
ex: "https://example.org/ontology#",
cd: "https://example.org/facts#",
rdf: "http://www.w3.org/1999/02/22-rdf-syntax-ns#",
rdfs: "http://www.w3.org/2000/01/rdf-schema#",
xsd: "http://www.w3.org/2001/XMLSchema#",
};
const ontologyStore = new OntologyStore(prefixes);
const factsStore = new FactsStore(prefixes);
const apiKey = process.env.ANTHROPIC_API_KEY!;
// Persist/reload across process runs:
ontologyStore.loadTurtle(await readFile("ontology.ttl", "utf-8"));
factsStore.loadTurtle(await readFile("facts.ttl", "utf-8"));
await evolveOntology(ontologyStore, documentText, apiKey);
await extractFacts(ontologyStore, factsStore, documentText, apiKey);
const rows = await queryGraph(
`SELECT ?bond ?faceValue WHERE { ?bond a ex:CorporateBond ; ex:faceValue ?faceValue }`,
[ontologyStore.rdfSource, factsStore.rdfSource],
prefixes
);
const matches = searchEntities("acme", factsStore);queryGraph runs on Comunica against N3-backed stores, entirely locally — real SPARQL, real joins, no LLM in the query path, and it auto-injects PREFIX declarations if your query text doesn't already include them.
Honest scope
Single process, sequential extraction — no parallel chunking for very large documents yet. No SHACL or shape validation, so a malformed triple isn't caught beyond the LLM's own structured-output compliance. Conflicting facts across documents (e.g. two different values for the same property) currently both get stored, with no detection or resolution — treat facts.ttl as cumulative evidence, not automatically reconciled truth. Entity resolution (searchEntities) is a plain label substring match, not semantic search. Each of these is a scoped next step, not an oversight.
License
MIT
