@mosadd/voice-analyzer-core
v0.1.0-alpha.2
Published
Deterministic voice-deepfake feature-extractor + scorer. Pure TS, no DOM / no Web Audio. Used by Demo112 voice analyzer and the analyze-call-audio edge function — both run the same code on the same inputs.
Downloads
411
Readme
@mosadd/voice-analyzer-core
Pure-TS deterministic voice-deepfake feature extractor + scorer. No DOM, no Web Audio, no FFT library — runs identically in browser, React Native, Deno edge functions and Node tests.
License: MIT.
Why this exists
mosadd needs the same scoring function in three different runtime environments without divergence:
apps/webDemo112 voice analyzer — minister demo, callsanalyzeAudio(samples, sampleRate)on Web Audio captures.supabase/functions/analyze-call-audio— LiveKit webhook handler, decodes egress recordings and emits session_events with the same shape.apps/gov-edgeRust appliance (future) — implements the same weighting in Rust; CI runs the parity fixtures from this package.
If the scoring drifts between consumer banner and operator console, an operator could see "0.91 synthetic" while the user sees "0.06 fine". That is the bug we are preventing.
API
import { analyzeAudio } from "@mosadd/voice-analyzer-core";
const result = analyzeAudio(pcmSamples, 16000);
// {
// confidence: 0.87,
// verdict: "VOICE_DEEPFAKE_HIGH_CONFIDENCE",
// features: { spectralTiltStdDev: 0.05, zcrJitter: 0.002, ... },
// topReasons: ["heuristic_score=0.87", "prosody_flatness=0.92", ...],
// analyzerVersion: "heuristic-1.0"
// }Scoring heuristic
The MVP scorer (version heuristic-1.0) weights five cues:
| Cue | Weight | What it measures |
|-----|--------|------------------|
| prosodyFlatness | 0.28 | Flat energy envelope = unnatural |
| zcrLowJitter | 0.22 | Synthetic voices have unnaturally smooth ZCR |
| spectralTiltLowVariance | 0.18 | Real speech tilt wanders, synth holds steady |
| noBreathPauses | 0.18 | Real voice has breaths between phrases; synth doesn't |
| f0LowVariance | 0.14 | Monotone pitch curve |
All five cues are deterministic time-domain or narrow-band-Goertzel measures — no FFT library, no model file.
Faza B — replacing with AASIST ONNX
When LINEAR-2009 Faza B lands a real AASIST-PL model, this package
keeps the same exports but scoreFeatures becomes a thin wrapper around
ONNX inference. The wire shape (VoiceAnalysisResult) does not change,
so neither the operator console nor the consumer banner needs changing.
Synthesizers
For offline reproducible demos and tests we ship three signal synthesizers (NOT real recordings):
synthesizeHumanPolish(seed)— varied F0, natural ZCR jitter, embedded breath pauses, dynamic envelope.synthesizeElevenLabsLikeSynth(seed)— uniform F0, near-zero ZCR jitter, zero breaths, flat envelope.synthesizeVoiceCloneChild(seed)— child-pitch, synthetically smooth.
Real recordings can later land as apps/web/public/audio/*.wav and be
decoded by the AudioContext to feed analyzeAudio() — the function
itself does not care about the audio source.
