@volter/twin-ahrefs
v0.1.35
Published
Local Ahrefs API v3 twin (Site Explorer read surface, Public, Subscription Info) built on @volter/world-core.
Readme
@volter/twin-ahrefs
A local, stateful, vendor-faithful twin of the Ahrefs API v3 (https://api.ahrefs.com/v3),
built on the shared @volter/world-core kernel.
Point an unmodified HTTP client at it and get vendor-correct answers offline: the Site Explorer read
surface (Domain Rating, backlinks stats, backlinks, referring domains, anchors, organic keywords and
the history series), the Public endpoints, and Subscription Info — with the real Authorization:
Bearer gate, both of the vendor's error envelopes, and the documented API-units accounting
including the five x-api-* response headers.
bunx world-ahrefs serve --port 8787 # start the twin
bunx world-ahrefs serve --read-only # a pure mirror: reads served, writes 405
bunx world-ahrefs conformance # certify every routed endpointcurl -H 'Authorization: Bearer any-key' \
'http://127.0.0.1:8787/v3/site-explorer/domain-rating?target=ahrefs.com&date=2026-01-01'
# {"domain_rating":{"domain_rating":56.1,"ahrefs_rank":8714}}Coverage
Coverage is API + connector. The capability manifest
(src/ahrefs-capabilities.ts) is the real vendor surface as the denominator, enumerated
top-down from Ahrefs' own first-party OpenAPI document — https://docs.ahrefs.com/openapi.json
(info.title "Ahrefs API", info.version 3.0.0), 129 path items / 148 operations across its 13
tags: Site Explorer (28), Web Analytics (34), Management (22), Brand Radar (20), GSC Insights (12),
Social Media (9), Keywords Explorer (6), Rank Tracker (6), Site Audit (4), Public (4), SERP Overview
(1), Batch Analysis (1), Subscription Information (1).
Every operation the twin does not serve is filed as an explicit todo naming its method and path,
so coverage reads honestly partial and grows as the twin does. Seventeen endpoints are modeled
today.
Modeled endpoints
| Area | Endpoints |
|---|---|
| Site Explorer — overview | domain-rating, backlinks-stats, outlinks-stats, metrics |
| Site Explorer — history | metrics-history, domain-rating-history, refdomains-history |
| Site Explorer — backlink profile | all-backlinks, broken-backlinks, refdomains, anchors |
| Site Explorer — organic search | organic-keywords |
| Public | domain-rating-free, domain-rating-top-domains, crawler-ips, crawler-ip-ranges |
| Subscription Information | limits-and-usage |
No UI mirror
The API is the product here. Ahrefs has a real web app that SEO practitioners work in, but this
pack twins the API v3, which is a separate surface an integrator calls from code — it exposes
none of the app's workflow, and every consumer grounding for it is server-side HTTP. The concrete
grounding is Dub's apps/web/lib/ahrefs/client.ts: a server-side client that calls
GET /v3/public/domain-rating-free from a cron job and never renders anything. There are therefore
no ui capabilities in the manifest; fabricating a dashboard the API cannot data-couple to would be
a false green.
Planned (todo)
csv/xml/phpserializations — render the same rows in those formats; today the twin renders JSON only and refuses the others in the vendor's error envelope rather than serving JSON under acsvrequest. The row model is flat enough to render.- Units cost past the generation window — price
x-api-units-cost-totalon the requestedlimit, not only on the rows the 500-row generation window materializes. - Credential gate on the twin-only routes —
/twin/target,/twin/auditand/twin/usagehave no vendor credential to check, so they are currently ungated (see below).
What this twin gets right that a mock would not
Two error envelopes, not one. The vendor's openapi.json declares a single shape for every
non-200 on all 148 operations — {"error": "<message>"}. The live API has two, split by which
layer refused, and the split was confirmed against api.ahrefs.com directly:
| Situation | Status | Body |
|---|---|---|
| No Authorization header | 403 | ["Error","Forbidden"] |
| Authorization: Bearer <invalid> | 401 | ["Error","Unauthorized"] |
| Bad parameter | 400 | {"error":"bad output: 'NOPE'"} |
| Unknown path | 404 | Not found (bare text, not JSON) |
| Wrong method | 405 | {"error":"POST"} |
Note the inversion against the usual HTTP convention: a missing credential is 403 and a
present but invalid one is 401. A client that branches on 401-vs-403 to decide "refresh the
token" versus "upgrade the plan" behaves differently against a twin that guesses. None of 404,
405, or the auth tuple appears anywhere in the vendor's published spec.
The API-units accounting is real, not decorative. Ahrefs prices a request
max(base_cost, per_row_cost * num_rows) with a 50-unit base and per-field tiers of 1, 5 or 10
units — and "requests served from cache do not consume units". The twin implements that formula
against the per-field prices read out of the vendor's own spec annotations, and emits all five
documented response headers (x-api-rows, x-api-units-cost-row, x-api-units-cost-total,
x-api-units-cost-total-actual, x-api-cache), which are documented in prose only and appear in no
generated client. /v3/subscription-info/limits-and-usage reports a balance that the requests you
actually made moved, and a repeated request is served from cache for free — so -total and
-total-actual diverge exactly where the vendor says they do.
One documented contradiction, resolved deliberately: the prose page's worked Example 2 prices
refdomains_sourceat 1 unit while the machine-readable spec annotates it(5 units). The twin follows the spec, because it is the vendor's generated artifact; the divergence is recorded insrc/ahrefs-units.tsrather than silently picked. Example 1 (domain-rating → 50 units) is reproduced exactly.
No pagination, because Ahrefs has none. Site Explorer's list endpoints take limit (vendor
default 1000), order_by, where and a mandatory select — and no offset, no cursor, no page
parameter. Across all 148 operations offset appears on exactly two, neither of them Site
Explorer (/site-audit/page-explorer and /social-media/posts). A twin that invented an offset
would teach a consumer a paging strategy that fails against the real API.
where and order_by run over the full row, before projection — which is why the vendor's docs
warn that the where field list differs from the select one. An unknown column, an unknown
operator or malformed JSON is a 400, never a silently-true predicate.
Wrapper keys are overloaded, and stay overloaded. metrics is an object on backlinks-stats
and an array on metrics-history; backlinks covers two different schemas (all-backlinks and
broken-backlinks, which alone carries http_code_target and last_visited_target). Reproducing
that is the point — a client sharing a deserializer by wrapper key breaks on exactly this.
What adversarial review changed
This pack went through the two-round §9 skeptic pass. Round one refuted five done claims and they
were fixed rather than demoted — recorded here because the bugs are the interesting part:
| Refuted | The bug | The fix |
|---|---|---|
| mode_scoping, protocol_scoping | The per-target jitter was keyed on targetKey, which folds in mode/protocol. A +/-25% per-scope jitter is wider than the gap between neighbouring scale factors, so mode=domain reported more referring domains than mode=subdomains for 28% of targets — an impossible answer for nested scopes. Both verifies probed one lucky target. | Jitter keyed on the target only; scope enters solely through the scale. Both verifies now sweep 30 targets. |
| anchors | refdomains, refpages and dofollow_links were minted from independent hashes, so 716 of 8,000 rows claimed more referring domains than the anchor had links. | Every count derived from links_to_target; the verify sweeps 5 targets at limit=500. |
| outlinks_stats | Its three inequalities were guaranteed by construction — strip them and nothing failable remained. | A pinned value plus the mode/protocol assertions it never had. |
| refdomains | Claimed agreement with backlinks-stats unconditionally; false above the 500-row generation cap. | Title and verify narrowed; the gap is filed as reports_beyond_generation_cap. |
| (new) unmodeled options | history=live returned the byte-identical all-time rows under a 200 — the caller believed they had filtered and had not. Same for traffic_mode, volume_mode, date_compared. | Refused with the vendor's {error} envelope, like output=csv already was. |
Also fixed: a where type mismatch quietly matched nothing instead of erroring; country accepted
any two lowercase letters instead of the vendor's 210-value enum (zz returned 200);
broken-backlinks accepted the eleven columns only all-backlinks declares; and an empty result set
made every select a 400 because the column schema was read from rows[0].
Deterministic data
The twin cannot know the real web, so it generates a stable, deterministic profile per target.
Every value is a pure function of a sha256 of the normalized target (plus mode and protocol,
which really do change what Ahrefs reports) — no PRNG, no clock, no module-level state. The same
domain therefore yields the same Domain Rating in this process, the next one, and on another machine.
The generated numbers hang together rather than being independently random:
- Domain Rating anchors everything: referring-domain counts grow with it, and Ahrefs Rank falls as it rises ("#1 being the strongest"). Measured over 8,000 targets there are zero cases of a higher rating earning a worse rank; adjacent tenths of a DR point above ~88 can tie, because the rank base collapses toward 10 at the top of the scale. The vendor's rank is a globally unique ordinal over its real index — this reproduces the scale, not the uniqueness, because each target's profile is a deterministic per-target function rather than a global index.
- Backlinks are drawn from the target's referring-domain set, so
liveexceedslive_refdomainsandaggregation=1_per_domainhas something real to collapse. - A full
refdomainslisting has exactly as many rows asbacklinks-statsreports referring domains, and each row's domain is distinct. modeandprotocolscope monotonically for every target:exact<=prefix<=domain<=subdomains, andhttp<=https<=both. The scopes are nested, so a narrower one can never report more than a wider one; ties happen only where rounding collapses a small profile.- Keyword traffic follows a click-through curve off
best_position, and never exceeds the keyword's own search volume;best_position_setand its boolean flags are derived frombest_position. all_timecounts never fall belowliveones, and referring-domain history never goes backwards.- Every count is a subset of the one containing it: an anchor's
refdomains<=refpages<=links_to_target, anddofollow_links/new_links/lost_linksnever exceed the links they are drawn from.
Row generation is capped at 500 rows per report (GENERATION_CAP); limit truncates within that,
and x-api-rows reports what was really returned.
Twin-only seed routes
Ahrefs is a read-only vendor — there is no write endpoint to seed through — so the twin adds a small set of routes outside the vendor's surface, which a real client never calls and which are deliberately absent from the capability manifest (counting them would pad the denominator):
| Route | Purpose |
|---|---|
| POST /twin/target | Override a target's Domain Rating / metrics / backlinks stats |
| GET /twin/audit | The request audit trail |
| GET /twin/usage | The local unit-usage meter |
Connector
pull fetches Domain Rating and backlinks stats per target over an injected client (the real API
in prod, a fake in tests) and folds them into local state as target overrides that every endpoint
then serves. It is idempotent: a re-pull of identical state appends no deltas. Two scalar calls per
target, never a row-returning list — establish identity cheaply and let the twin generate the
expensive detail.
Rate budget
Every live call goes through one guarded client (liveAhrefsExecute) with a persistent,
fail-closed spend ledger. Ahrefs documents 60 requests/minute for API v3, so the ceiling is 40
weighted units per 60s — two thirds of it — with row-returning list endpoints priced at 2 because
their unit cost scales with the response. The vendor's other meter, the subscription-period unit
quota, is not something a rolling window can enforce; src/ahrefs-budget.ts says so rather than
pretending otherwise. Retries are off. The ledger is keyed by API key and survives a restart.
No official SDK
Ahrefs publishes no official JavaScript SDK for API v3. The npm package ahrefs is a
third-party client last released in 2015 against the long-dead v1/v2 API, and ahrefs-v3 is another
individual's package; @ahrefs/mcp is the vendor's own MCP server, not a library to import. So no
vendor SDK is declared as a devDependency and none is invented — the fidelity test
(src/ahrefs-sdk.integration.test.ts) drives an unmodified real-transport fetch client, written
from the published reference, at the running twin. It includes Dub's exact production
getDomainRating() call path, verbatim.
Authentication
Authorization: Bearer <API key>, as the vendor's spec declares globally. Any non-sentinel key is
accepted; a few sentinels drive the deterministic failure paths so verifies can exercise them
offline:
| Key | Effect |
|---|---|
| (no header) | 403 ["Error","Forbidden"] |
| invalid / expired / revoked | 401 ["Error","Unauthorized"] |
| rate-limited / over-quota | 429 {"error":"Too Many Requests"} |
/v3/public/crawler-ips and /v3/public/crawler-ip-ranges need no key at all — the vendor
documents both as "free and do not require an API key", and a live request with no Authorization
header really does return 200. /v3/public/domain-rating-free is free of units but still requires
a key, and the twin keeps that distinction.
