@veltrosecurity/ai
v0.1.0
Published
Veltro bring-your-own AI provider contract: config validation, redaction, structured output, audit digests.
Readme
@veltrosecurity/ai
Apache-2.0 shared client for the Veltro AI provider contract (veltro-ai-provider/v1).
- Validate BYO provider config (
openai_compatible|anthropic|ollama) - Redact secrets/tokens before prompts leave the product
- Enforce JSON-schema structured outputs
- Build
ai.invocationaudit records (digests only — never raw prompts) - Call providers through an injectable transport (tests use recorded fixtures)
AI is off until configured. Products store API keys in their own secret stores and pass only api_key_ref handles through settings. Suite identity never fans out secrets.
parseStructuredOutput requires a Draft 2020-12 JSON Schema and returns only
after the model payload satisfies it:
const verdict = parseStructuredOutput(modelText, {
type: "object",
properties: {
verdict: { type: "string", enum: ["true_positive", "false_positive"] },
confidence: { type: "number", minimum: 0, maximum: 1 },
},
required: ["verdict", "confidence"],
additionalProperties: false,
});Malformed schemas raise ai_invalid_config. Invalid model payloads raise
ai_structured_output_invalid without including model values in the error.
Cross-language guarantee. For a provider config, a base_url, a text to
redact, or a structured-output schema and payload, this package and
veltro-ai return the same value or the same refusal code — never opposite
answers. Where the two runtimes cannot be made to agree on a construct, both
refuse it rather than answering differently, and every such construct is named
below. One difference remains, and it is not in the answer: a payload or
schema nested deeper than the runtime's own stack is refused
(ai_structured_output_invalid for a payload, ai_invalid_config for a
schema), and the depth at which each runtime gives up is a property of the
runtime, not of this contract.
The schema is always enforced as Draft 2020-12, whatever dialect a $schema
announces, and format is an annotation only — veltro-ai behaves the same
way, and conformance/veltro-ai-provider-v1/vectors.json pins both. A
validator is compiled per schema object and cached weakly, so passing an inline
object literal per call retains nothing; two schemas that declare the same
$id do not collide.
Draft 2020-12 is enforced exactly: draft-07 dependencies is an inert unknown
keyword (use dependentRequired/dependentSchemas), a draft-04 id is
refused as ai_invalid_config, and an if with no then/else asserts
nothing. Every $ref a validator could reach is resolved before the first
payload is seen, so a reference that does not resolve inside the schema —
remote, or dangling into $defs — is refused as ai_invalid_config whatever
the payload looks like, and is never retrieved over the network. Neither
client has ever retrieved one: Ajv resolves only what the document carries,
and veltro-ai passes jsonschema the bundled specification registry, whose
role is resolving the dialect metaschemas — its default registry is empty and
refuses offline too. A $ref chain that returns to itself through
allOf/anyOf/oneOf/not alone can never validate anything and is
refused; a genuinely recursive schema (through properties, items, …) is
accepted. A broken reference in an unreferenced $defs entry, or a
$ref-shaped const value, is data and stays accepted. There is exactly one
position none of that reaches: a subschema under an if with no then/else.
That keyword asserts nothing, and both clients drop it before validating — Ajv
never compiles it and veltro-ai deletes the key — so a remote $ref, a
$ref cycle, a $dynamicRef or a pattern outside the accepted subset is
accepted and inert there in both clients, and nothing under it is ever
resolved or compiled. Two vectors pin that, and every refusal below is about
the schema both clients validate, which never carries a lone if.
$dynamicRef, $dynamicAnchor, $recursiveRef and $recursiveAnchor are
refused as ai_invalid_config — everywhere except under the lone if above.
Their target depends on the dynamic scope of the instance, and Ajv and
jsonschema disagree about a dangling one, so neither client accepts either
spelling: use $ref into $defs.
pattern, patternProperties keys and propertyNames.pattern are ECMA-262
expressions, and only the subset both engines read identically is accepted —
anything else is ai_invalid_config, in both clients, wherever it sits in the
document (including an unreferenced $defs), the lone if above excepted.
Accepted: literals, ., ^, $, |, groups (…)/(?:…), lookahead
(?=…)/(?!…), character classes
with ranges and negation, \d/\D/\w/\W (\D/\W outside a class),
\b/\B, \n/\r/\t/\f/\v, \xHH, \uHHHH, escaped
metacharacters, and quantifiers */+/?/{m}/{m,}/{m,n} with an
optional lazy ? and bounds ≤65535. Refused: \s/\S (the two whitespace
sets differ), \p{…}, \A/\Z, named groups in either spelling, lookbehind,
inline flags, (?#…), backreferences, possessive quantifiers, a{,n}, [],
[^], an unescaped [/]/{/}, a lone surrogate escape, and a range
whose ends are not single characters. Inside the accepted subset the two
clients agree on the answer, not only on acceptance: $ is end of input (so
^ok$ refuses "ok\n"), . excludes every line terminator, and \d, \w
and \b are ASCII. veltro-ai translates the same subset into the Python
spelling of the same expression, and
conformance/veltro-ai-provider-v1/vectors.json pins every one of those
decisions in both suites.
base_url is http/https with an ASCII host (a hostname, a dotted IPv4, or
a bracketed plain-hex IPv6), an optional port in 1..65535, and an optional
path; userinfo, a query, a fragment, whitespace, a control character, a
percent-escape in the host and a non-ASCII host are refused as
ai_invalid_config. That one shape — not new URL(), and not Python's
urlparse — is what both clients apply, so a stored value is accepted or
refused identically, and isAiConfigured/isAiEnabled/buildAiStatus answer
false for a bad one instead of throwing.
redactSecrets masks provider and cloud credential shapes as issued today —
sk-/sk-proj-/sk-svcacct-/sk-ant-api03-, Google AIza, GitHub ghp_
and github_pat_, Slack xox*-, AWS AKIA/ASIA key ids, Bearer values,
base64-shaped Basic credentials, the value of an
Authorization/Proxy-Authorization header, an AWS4-HMAC-SHA256 credential
string wherever it sits, a signature parameter of sixteen characters or
more, JWTs, and …secret…/…token…/…key… assignments — and runs both on
every provider response body before it can reach an ai_provider_error
message and on every message and system prompt before the transport sees it.
Ordinary prose, digests and commit SHAs pass through unchanged.
What "the value" means. The name is one that ends in Authorization or
Proxy-Authorization after a character that is not a letter, digit or _, so
a dashed vendor prefix comes along (X-Authorization:,
x-amz-authorization:) and a name that merely contains the word does not
(deauthorization:, my_authorization:). It is read as name: value,
name=value, headers["name"] = "value", or the JSON, escaped-JSON and
curl -H quoting of any of those, with the value on the name's own line. The
mask covers every ,/;-delimited parameter of the value — so an RFC 7616
Digest response=, a SigV4 Signature= and a Hawk
mac= go with the parameter that precedes them instead of surviving in the
outbound prompt — and a space extends it only into another credential-shaped
run, so Authorization: NTLM TlRMTVNTUAABAAAA rejected by upstream keeps its
last three words.
What has to look like a credential. Because redactMessages runs on every
message, exactly one thing is shape-tested: the value itself, or the token
after a recognised scheme. The recognised set is exported as
AI_AUTH_SCHEMES — AI_REGISTERED_AUTH_SCHEMES and
AI_VENDOR_AUTH_SCHEMES are its two halves — so a consumer can check its own
header instead of guessing. It is all fourteen names in the IANA HTTP
Authentication Scheme registry (Basic, Bearer, Concealed, DPoP,
Digest, GNAP, HOBA, Mutual, Negotiate, OAuth, PrivateToken,
SCRAM-SHA-1, SCRAM-SHA-256, vapid) plus twenty-one schemes real services
issue that no registry governs, each taken from its vendor's own
documentation: Okta's SSWS, Discord's Bot, Splunk, Zoho's
Zoho-oauthtoken and Zoho-authtoken, Opsgenie's GenieKey,
acquia-http-hmac, the OAuth MAC draft, Klaviyo-API-Key, GitHub's
Token, Elasticsearch's ApiKey, Azure's SharedKey, SharedKeyLite and
SharedAccessSignature, AWS4-HMAC-SHA256 and its predecessor AWS,
Kerberos, NTLM, Hawk, WSSE and Signature. Credential shape is eight
characters or more, or a key=value parameter, and not ordinary text.
Ordinary text is one word — ALL-CAPS,
Capitalised or lower-case, at most fifteen characters a part, up to five parts
joined by -, _ or an apostrophe, an optional bracketed word the way
Optional[str] spells one, and optional closing sentence punctuation — or one
key=value whose key is such a word and whose value is such a word or a
number of up to four digits, with no further parameter after it. Case and
length are ASCII notions, so a value carrying any character outside ASCII is
ordinary text whatever its case: approuvé, одобрено, 承認済み and
อนุมัติ all survive, and none of them survives on a character count.
So authorization: approved tenant=acme, authorization: denied
reason=policy_violation, authorization: approved 2026-09-06T19:00:00Z,
Authorization: Token expired, authorization: role_admin,
authorization: scope=read, Authorization: n/a,
authorization: 3 of 5 approvers signed and {"authorization": "approved"}
survive byte for byte — in an outbound prompt and in the ai_provider_error
built from a failing response — while Authorization: NTLM TlRMTVNTUAABAAAA, a
reordered authorization: qop=auth, response="…" and a schemeless
authorization: 4a5b6c7d8e9f0a1b2c3d4e5f60718293 do not.
One accepted over-masking edge is contractual: when an otherwise ordinary
value starts with Bearer, PrivateToken, GenieKey, SharedKey,
SharedKeyLite, SharedAccessSignature, SCRAM-SHA-1, SCRAM-SHA-256 or
AWS4-HMAC-SHA256, that recognised leading scheme name is itself masked (for
example, authorization: PrivateToken expired becomes
authorization: [REDACTED_AUTH] expired). The shared vectors pin this outcome;
it is not a promise that every recognised name survives when followed by prose.
Where it stops. This is defence in depth behind a product's own secret
handling, not a sanitiser, and it does not promise to find a credential that is
not spelled like one. It does not read a name that merely starts with the word
(Authorization-Header:), a name held away from its : or = by anything but
spaces and tabs, or a value that begins on the next line. It does not read the
WWW-Authenticate: and Proxy-Authenticate: challenges, which carry a
server's realm rather than a client's proof — though an AWS4-HMAC-SHA256
string inside one is still masked. Its largest gap is deliberate: a scheme
that is not in AI_AUTH_SCHEMES and whose name is itself ordinary text
leaves its credential alone (Authorization: Vendor abc.def-ghi, or whatever
the next vendor invents — Authorization: Acme 4a5b6c7d8e9f0a1b), while one
that is not ordinary text is masked whole (Authorization: X-Vendor2
abc.def-ghi). Re-testing the second token whenever the first is merely
word-shaped is what closes that gap, and it was measured: it changes 501 of
2,570 realistic prose strings, deleting reason=, actor=, request ids,
URLs and timestamps out of log lines. This client is defence in depth behind a
product that is not supposed to put credentials in prompts at all, and it
reads log lines for a living, so the gap stays and the set grows by name
instead — check yours against AI_AUTH_SCHEMES. It also leaves a schemeless
value that is ordinary text by the rule above
(Authorization: SECRETPASSWORD, Authorization: deadbeefcafe,
Authorization: SECRET_PASSWORD_VALUE) — the price of letting a verdict
through.
conformance/veltro-ai-provider-v1/vectors.json
pins all of it in both directions, and veltro_ai.redact_secrets applies the
same table.
normalizeLegacyAiSettings denies every data class unless the stored blob
grants one: a stored allow_data_classes array is authoritative (an empty one
is an exact deny-all), otherwise ai_allow_log_samples: true — and only
exactly true — grants events. Product-specific widening belongs in the
product adapter, not here.
TeamAiSettings, mapStoredProviderToKind and normalizeTeamAiSettings have
Python counterparts with the same value-shape contract. Both adapters map the
same stored provider spellings, defaults, key-presence marker and pipeline-data
grant, and both execute the shared team vectors.
A transport that resolves to anything other than an object raises
ai_provider_error, so no adapter can leak a raw TypeError to a caller.
