npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@x12i/candidate-retrival-service

v0.1.0

Published

Generic org-isolated lexical, semantic, and hybrid retrieval companion for Memorix

Readme

@x12i/candidate-retrival-service

Generic retrieval companion for the Memorix data plane. It builds a derived, org-isolated search index from canonical Memorix V2 records and exposes lexical, vector, hybrid, lineage, and batched candidate retrieval without extending SafeFilter or becoming a second source of truth.

KnowX is supported through two declarative profiles; it is not hard-coded into search or storage.

Runtime shape

  • API: http://127.0.0.1:6560 (memorix-retrieval ports-manager zone)
  • PostgreSQL + pgvector: host port 6562 in the supplied Compose stack
  • Memorix proxy env: MEMORIX_RETRIEVAL_URL=http://127.0.0.1:6560
  • Health: /health; status UI (@x12i/api-live-view, matches other Memorix services): /_live; OpenAPI: /openapi.json; metrics: /metrics
  • Required scope on every /v1/* call: x-memorix-org-id and plural x-memorix-agent-ids

Start locally

npm run up

Builds the image, starts Postgres + the API via Compose, waits for /health to report ready, and prints the status UI URL. Equivalent to cp .env.example .env && docker compose up --build -d plus a health wait. Other lifecycle scripts: npm run down (stop the stack), npm run logs (tail the API container).

Open http://127.0.0.1:6560/_live for the live status dashboard (service info, health, and a rolling view of inbound/outbound traffic).

For a non-durable, dependency-free development process:

MEMORIX_RETRIEVAL_STORE=memory npm run dev

memory is never a production setting. Production uses PostgreSQL 17 with the pgvector extension. Schema migration is idempotent and runs at startup.

Index a Memorix record

Direct V2 envelopes use the best matching declarative profile:

curl -sS http://127.0.0.1:6560/v1/index/upsert \
  -H 'content-type: application/json' \
  -H 'x-memorix-org-id: sandbox-demo' \
  -H 'x-memorix-agent-ids: ops' \
  -d '{
    "records": [{
      "format": "memorix-record/2",
      "recordId": "PROC-2",
      "revision": 1,
      "objectType": "procedure",
      "contentType": "snapshot",
      "dataCategory": "knowledge",
      "concept": {
        "title": "Queue recovery",
        "identifiers": { "primary": { "kind": "procedureId", "value": "PROC-2" } }
      },
      "data": { "title": "Queue recovery", "text": "Stop the worker, drain the queue, then restart it." },
      "_system": {
        "state": "active",
        "createdAt": "2026-08-03T00:00:00.000Z",
        "modifiedAt": "2026-08-03T00:00:00.000Z"
      }
    }]
  }'

An item may instead be { "record": <V2 envelope>, "profileId": "pack.profile", "agentIds": [...], "embedding": [...], "embeddingModel": "..." }. An agentIds override must be a subset of the effective scope header, preventing visibility escalation.

Every envelope is validated with validateMemorixRecordV2 from @x12i/memorix-format; malformed, legacy, and unknown-root records are rejected before mapping or vectorization.

Upserts are revision-safe. A lower revision is skipped; a newer record with no searchable text removes the older derived documents. Deletes and source tombstones apply across all model generations.

Search

curl -sS http://127.0.0.1:6560/v1/search \
  -H 'content-type: application/json' \
  -H 'x-memorix-org-id: sandbox-demo' \
  -H 'x-memorix-agent-ids: ops' \
  -d '{
    "q": "how do I recover a stuck queue?",
    "mode": "auto",
    "filters": { "objectTypes": ["procedure"] },
    "limit": 10
  }'

Modes:

  • lexical: PostgreSQL full-text search.
  • vector: cosine similarity over pgvector HNSW; accepts embedding or vectorizes q.
  • hybrid: weighted reciprocal rank fusion of independent lexical and vector pools.
  • auto: hybrid when a query vector is possible, otherwise explicit lexical fallback with a warning.

Exact filters cover profiles, document kinds, OT/CT/data category, record/source/item ids, time range, and profile-defined facets. Search hits include source ids, hashes, chunk character spans, score components, and a compact or full record payload for deep links and citations.

POST /v1/candidates accepts up to 100 {id,text,embedding?,filters?} items and returns top-k candidates per item. This is the generic integration point for classify, dedupe, association, recommendation, or grounding pipelines.

FuncX and AI Profiles

Deterministic retrieval never needs an LLM. Optional semantic reranking is requested with:

{
  "q": "recovery after queue saturation",
  "mode": "hybrid",
  "rerank": { "enabled": true, "model": "cheap/default", "candidateCount": 15 }
}

The reranker uses rank from @x12i/funcx/functions. Its model must be an @x12i/ai-profiles profile/choice; AI Profiles validates the intent and FuncX resolves it again at execution. Set OPENROUTER_API_KEY to enable it. Reranking is opt-in because it adds remote-data exposure, cost, and latency.

Current FuncX has no embedding transport, and current AI Profiles declares but does not populate an embeddings lane. Vectorization therefore remains behind a generic adapter. Callers can supply 1536-dimensional vectors, or configure the included OpenAI-compatible adapter with EMBEDDING_BASE_URL, EMBEDDING_API_KEY, and EMBEDDING_MODEL. No LLM is asked to fabricate embeddings.

Retrieval profiles

Profiles map arbitrary pack records onto the generic index. Built-ins are:

  • memorix.default: common V2 title/label/summary/text fields.
  • knowx.knowledge: KnowX labels/summaries, epistemic facets, and lineage.
  • knowx.snapshot: Markdown chunking with source and Drive-item citations.

Add pack profiles through RETRIEVAL_PROFILES_PATH pointing at a JSON array:

[
  {
    "id": "ops.runbook",
    "description": "Operations runbook passages",
    "priority": 200,
    "selector": { "objectTypes": ["procedure"], "contentTypes": ["runbook"] },
    "documentKind": "passage",
    "titlePaths": ["data.name"],
    "summaryPaths": ["data.summary"],
    "textPaths": ["data.markdown"],
    "chunk": { "path": "data.markdown", "targetTokens": 350, "overlapTokens": 60 },
    "facets": { "severity": "data.severity", "state": "data.state" },
    "source": { "recordIdPath": "data.sourceId", "hashPath": "data.sourceHash" },
    "includeRecord": false
  }
]

Paths are dot-separated. Arrays are flattened automatically; [] may be written explicitly. Custom profiles take precedence by priority and may override a built-in by using the same id.

Backfill

Backfill uses only the public Memorix list API and the retrieval API. OT/CT pairs are explicit so a pack controls scope:

npm run build
node dist/cli.js backfill \
  --org sandbox-demo \
  --agents knowx \
  --pair content/snapshots=knowx.snapshot \
  --pair content/knowx=knowx.knowledge \
  --pair people/knowx=knowx.knowledge \
  --pair work-items/knowx=knowx.knowledge

The command follows pageInfo.nextCursor, reads at most 200 records per page, and reports progress. It is idempotent because record revision is part of the write lifecycle.

Memorix integration

Browsers call Memorix only. memorix-service should own thin scoped proxy routes and instantiate the exported RetrievalClient with MEMORIX_RETRIEVAL_URL and the optional companion token. Preserve x-correlation-id and canonical scope headers. Add a retrieval block to /api/diagnostics/scope with configured URL and live health.

After a successful canonical data write, enqueue an idempotent index-upsert job containing the persisted record revision. On remove/supersede, enqueue a record or source delete. Pipeline consumers call /v1/candidates; they do not send full collection dumps to FuncX/skills. SafeFilter remains unchanged.

See the functional and technical specification and OpenAPI. The internally controlled dependency-chain audit is recorded as an informational operational dependency note.

Verification

npm run check
docker compose up -d postgres
DATABASE_URL=postgres://memorix:[email protected]:6562/memorix_retrieval npm run test:postgres

The default suite uses the in-memory store to test contracts without external state. PostgreSQL/pgvector is exercised separately as an integration target.