npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

react-native-nitro-audio-anvil

v2.0.2

Published

Advanced Audio Recording And PCM Streaming of the same audio bytes which can be used for Cloud Transcription

Readme

react-native-nitro-audio-anvil

React Native Nitro Module for corruption-proof, long-form microphone recording — PCM straight to disk, fsynced and segmented, with live PCM and speaker-window streams. Built for the 60–90 minute recordings that must survive incoming calls, backgrounding, force-quits and dead batteries.


[!NOTE]

  • This library was originally created for my production app, where we record long conversations — 60 to 90 minutes — on the phone that is also, well, a phone.
  • We started on expo-audio, and it served its purpose: it got us recording in an afternoon, and for short clips it is exactly the right tool. Then a customer took an incoming call 40 minutes into a recording. The encoder was torn down mid-write, the .m4a never got its index written, and the file was unrecoverable. Nobody did anything wrong — a container that needs a finalize step is simply the wrong shape for a recording that can be interrupted at any second.
  • Losing an audio file is worse than most other failures, because there is no retry. The customer already spoke. The moment is gone. If we lose the bytes, we lose the meeting — and with it any transcript, summary, action item or downstream analysis that depended on it. Everything else in a recording pipeline (transcription, upload, storage) can be retried. The recording itself cannot.
  • Anvil is built with fault tolerance as the first design constraint, not a nice-to-have. There is no encoder and no finalize step: raw PCM goes to a WAV file whose header is patched and fsynced twice a second, in 30-second segments. A call, a crash or a power cut costs you at most half a second of audio — never a file. On next launch, discoverOrphanedRecordings() repairs any headers that never got a final patch and hands you back everything on disk. RecordingService groups sessions across the crash boundary so a 90-minute meeting interrupted mid-way still comes back as one logical recording.

What you get out of the box:

  • Mono 16-bit WAV capture that is a valid file at every instant, not just at stop()
  • Segmentation every segmentDurationMs (default 30 s), rotated on pause, interruption and input-device change
  • Native handling of calls, Siri, alarms, media-server reset (iOS), audio-focus loss and capture-silenced (Android)
  • Auto-resume with exponential backoff when another app releases the mic — WhatsApp, Voice Memos, Siri, phone calls
  • Deferred-resume-on-foreground for the case where the interruption ends while your app is backgrounded (both platforms have known bugs here — Anvil works around them)
  • PCMChunk stream (default 100 ms) for streaming speech-to-text, with sequence numbers so gaps are detectable
  • Overlapping SpeakerWindow stream (default 1.5 s / 750 ms hop) for speaker labelling or diarization
  • extractRange(startMs, endMs) to re-read any span from disk, even while recording — useful when a streaming socket drops
  • concatenate(paths, output) to stitch segments into one WAV without re-encoding
  • discoverOrphanedRecordings(dir) with WAV header repair on relaunch
  • createRecordingService() to group sessions across crashes under your own id (meetingId, callId, …)
  • SHA-256 per finalized segment for integrity checks
  • Foreground-service notification on Android, audio background mode on iOS

What this library does NOT do (by design):

  • No encoding. Nothing to opus, aac or mp3 on device. WAV out. Encode server-side if you want smaller files — do it after the bytes are safely off the device, never before.
  • No transcription, no VAD, no speaker embedding, no summarization. The PCMChunk and SpeakerWindow streams hand you the bytes; you pick the model and where it runs (cloud, on-device with ExecuTorch, whatever).
  • No upload. Pair with react-native-nitro-cloud-uploader for S3-compatible multipart uploads, or roll your own. The example app wires both.
  • No playback. Pair with react-native-nitro-player — every WAV Anvil writes plays as-is.

If your app needs to record something long, on the same device that can be interrupted at any second, and you cannot afford to lose it — this is the recorder.


📦 Installation

yarn add react-native-nitro-audio-anvil react-native-nitro-modules
cd ios && pod install

[!IMPORTANT]

  • iOS: Fully tested and production-ready ✅
    • AVAudioEngine capture, AVAudioSession interruption / route / media-server-reset handling
    • CallKit call detection, audio background mode
    • Auto-resume with exponential backoff (200 ms → 400 ms → 800 ms → 1.6 s → 3.2 s) when foreground
    • See iOS quirks for platform-specific gotchas (background reactivation bug, CarPlay, Bluetooth chaos)
  • Android: Fully tested and production-ready ✅
    • AudioRecord on a dedicated audio thread
    • Microphone foreground service
    • Audio focus + isClientSilenced interruption detection
    • Auto-resume with exponential backoff (300 ms → 600 ms → 1.2 s → 2.4 s → 4.8 s) when foreground
    • Requires Android 7.0+ (API 24+)
    • See OEM quirks for per-brand walkthroughs (Xiaomi, Huawei, Oppo, Vivo, OnePlus, Samsung, and more)
  • Tested on React Native 0.85+ with the New Architecture (required by Nitro Modules). PRs welcome for lower RN versions.

🎥 Demo

The example app records with Anvil, plays the result with react-native-nitro-player and uploads it with react-native-nitro-cloud-uploader — the whole capture → play → upload flow on Nitro Modules.

[!NOTE]

The example uploads to my Cloudflare R2 bucket test-bucket via a public Worker at https://api.gauthamvijay.com, so you can run it end-to-end without setting up any backend. Uploaded files are automatically deleted after 3 days.

const BASE_URL = 'https://api.gauthamvijay.com';
const CREATE_UPLOAD_URL = `${BASE_URL}/create-and-start-upload`;
const COMPLETE_UPLOAD_URL = `${BASE_URL}/complete-upload`;
const ABORT_UPLOAD_URL = `${BASE_URL}/abort-upload`;
const SINGLE_UPLOAD_URL = `${BASE_URL}/single-upload`;

📚 Documentation

For long-form microphone recording, the library itself is only half the story. The other half is knowing the platform quirks that affect background audio.

| Doc | When to read | | -------------------------------------------- | ------------------------------------------------------------------------------------- | | iOS quirks | Before shipping on iOS — covers the background reactivation bug, CarPlay, Bluetooth | | OEM quirks | Before shipping on Android — per-brand setup for Xiaomi, Huawei, Oppo, Vivo, and more | | Troubleshooting | When something breaks — symptom-first debugging with hypothesis and fix per symptom | | Recovery | When integrating RecordingService for cross-crash session grouping |

Every real-world quirk we have hit — background reactivation permanent-fail on iOS, HyperOS killing foreground services, WhatsApp holding the mic HAL, sample rate changes on Bluetooth route — is documented in one of these files. If you hit something not covered, file an issue and it will land here.


🧠 Overview

| Feature | Implementation | | --------------------------- | -------------------------------------------------------------------------------- | | Format | Mono 16-bit PCM WAV, no encoder, no finalize step | | Durability | Header patched + fsync every fsyncIntervalMs (default 500 ms) | | Segmentation | New file every segmentDurationMs, on pause, interruption and route change | | Phone calls / Siri / alarms | Segment finalized before the OS takes the mic; event emitted | | Auto-resume | Exponential backoff retry when the OS releases the mic — foreground-gated | | Bluetooth / headset changes | Route event + segment rotation so no file mixes two input devices | | Background recording | iOS audio background mode / Android microphone foreground service | | Crash & force-quit recovery | discoverOrphanedRecordings() repairs headers; RecordingService re-groups | | Live PCM stream | PCMChunks (default 100 ms) for streaming speech-to-text, with sequence numbers | | Speaker windows | Overlapping SpeakerWindows (default 1.5 s / 750 ms hop) for speaker labelling | | Range extraction | extractRange(startMs, endMs) re-reads any span from disk, even while recording | | Stitching | concatenate(paths, output) joins segments into one WAV without re-encoding | | Integrity | SHA-256 per finalized segment | | Storage guard | Warning event below a configurable free-space threshold | | Threading | One owner thread per recorder, no locks, no JS-thread blocking |


🛡️ Fault tolerance, in detail

Every design decision in Anvil starts from the question "what happens if the process disappears right now?" Here is the answer for each failure mode:

| Failure | What Anvil does | What you get back | | ------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------ | | Incoming phone call | iOS AVAudioSession.interruptionNotification / Android audio focus loss → current segment is finalized (patched, fsynced, hashed) before the OS takes the mic | An interruption event with a valid WAV path, then optional auto-resume | | Another app takes the mic (WhatsApp, Voice Memos, Siri) | Segment finalized before the OS reassigns the mic; when the other app releases it, native retries with exponential backoff until it succeeds or gives up | Recording resumes seamlessly when the other app is done | | Bluetooth headset connect / disconnect | Route change → current segment finalized so no file mixes two input devices | A routeChange event and a fresh segment for the new device | | App backgrounded / screen locked | iOS audio background mode / Android microphone foreground service keeps the capture running | Recording continues; timer keeps advancing | | Interruption ends while app is backgrounded | Both platforms have OS bugs blocking background auto-resume (documented). Native emits a "deferred" error; JS wires a foreground listener to retry on return | Recording resumes the moment the user opens the app | | App force-quit | Whatever was fsynced is on disk. On next launch, discoverOrphanedRecordings repairs any headers that never got patched | Every segment written, up to the last 500 ms | | Process crash / OOM kill | Same as force-quit — nothing to finalize, nothing to lose except the last 500 ms | Same as above | | Device reboot / battery dies | Same as force-quit | Same as above | | Streaming STT socket drops | extractRange(startMs, endMs) re-reads exactly the missing span from disk | A WAV you can upload to a batch transcription endpoint | | Free space low | Warning event on start() and every rotation, before it becomes an error | Time to prompt the user or rotate off the device |

There is no moov atom, no encoder state, no finalize step to skip. The file on disk is always a valid WAV, at every instant.


🔁 Interruption handling, in detail

Real-world microphone interruptions are messier than the OS docs suggest. Anvil handles the full matrix:

Native side (both platforms):

  • On interruption begin: stop capture, finalize the current segment, emit interruption event with phase: 'began' and a valid WAV path for what was recorded up to that moment
  • On interruption end with the OS-provided shouldResume flag: check foreground state, then retry resume() with exponential backoff (5 attempts, ~6-9 seconds total) until it succeeds
  • If foreground check fails: emit an error with the message "Auto-resume deferred — bring app to foreground to continue" and stop trying. Retrying while backgrounded wastes CPU on iOS (Apple platform bug 560557684) and battery on Android (aggressive OEMs like Xiaomi kill background retries anyway).

JS side (your app):

  • Wire an AppState listener that watches for foreground transitions
  • On any transition to 'active', if the recorder is paused and a deferred-resume flag is set, call recorder.resume() explicitly with a 300 ms settle delay
  • The deferred flag is set by the addErrorListener when it sees the deferred / abandoned messages

Here is the pattern:

import { AppState } from 'react-native';

const resumeDeferredRef = useRef(false);

// Watch for the deferred signal from native.
recorder.addErrorListener((error) => {
  if (
    error.message?.includes('Auto-resume deferred') ||
    error.message?.includes('Auto-resume abandoned')
  ) {
    resumeDeferredRef.current = true;
  }
});

// When the user comes back to the app, retry.
useEffect(() => {
  const sub = AppState.addEventListener('change', async (state) => {
    if (state !== 'active' || !resumeDeferredRef.current) return;
    resumeDeferredRef.current = false;

    // Let the OS finish handing focus back before hitting the mic.
    await new Promise((r) => setTimeout(r, 300));

    try {
      await recorder.resume();
    } catch (err: any) {
      // Rare — surface a toast so the user can tap Resume manually.
      console.log('resume failed:', err?.message);
    }
  });

  return () => sub.remove();
}, []);

The end result is that every real-world interruption scenario resolves cleanly:

  • Short interruptions (Siri, quick calls): instant auto-resume
  • Medium interruptions (WhatsApp voice notes): retry with backoff, resumes within a few seconds
  • Long interruptions with your app backgrounded: deferred, resumes the moment the user returns to your app
  • Uncooperative other apps holding the mic too long: 5 tries with backoff, then user taps Resume manually

The example app wires all of this. See example/App.tsx for the reference implementation.

[!TIP]

  • iOS-specific quirks (background reactivation permanent-fail bug, CarPlay routing chaos, media services reset, etc.) are documented in docs/ios-quirks.md. Read this before shipping on iOS.
  • Android OEM quirks (Xiaomi/HyperOS, Huawei, Oppo, Vivo, Realme, OnePlus, Samsung) may still kill your foreground service on screen-off despite everything the library does. This is a device-level setting the user has to change — see docs/oem-quirks.md for a per-brand walkthrough, or link users to dontkillmyapp.com which stays up to date with each OEM's UI changes.
  • Something not working? See docs/troubleshooting.md for symptom-first debugging.

⚙️ Basic Usage

import { Anvil, type AnvilRecorder } from 'react-native-nitro-audio-anvil';

if ((await Anvil.requestPermission()) !== 'granted') return;

const recorder: AnvilRecorder = await Anvil.createRecorder({
  outputDirectory: `${documentDirectory}/recordings`, // plain path or file:// URL
  segmentDurationMs: 30_000,
  fsyncIntervalMs: 500,
  sampleRate: 16000,
  streamChunkMs: 100,
  speakerWindowMs: 1500,
  speakerWindowHopMs: 750,
  onInterruption: 'resume',
  keepAwakeInBackground: true,
  storageWarningBytes: 200 * 1024 * 1024,
  notification: { title: 'Recording', text: 'Tap to return' },
});

// Stream 1 → your streaming speech-to-text socket (pcm16, 16 kHz, mono — send the buffer as-is)
const pcm = recorder.addPCMListener((chunk) => socket.send(chunk.buffer));

// Stream 2 → your speaker-embedding model → label who is talking
const speaker = recorder.addSpeakerWindowListener(async (window) => {
  if (window.rms < 0.01) return; // silence
  const label = await labelSpeaker(window.buffer, window.startMs, window.endMs);
});

recorder.addInterruptionListener((e) => {
  // e.phase === 'began': e.segmentPath is already a valid file on disk
  // e.phase === 'ended' && !e.shouldResume: call recorder.resume() when you want
});
recorder.addSegmentCompletedListener((segment) => uploader.enqueue(segment));
recorder.addErrorListener((error) => log.error(error.code, error.message));

await recorder.start();
// ...
const segments = await recorder.stop();

pcm.remove();
speaker.remove();

One file instead of segments

const full = await Anvil.concatenate(
  segments.map((s) => s.filePath),
  `${documentDirectory}/recordings/meeting-full.wav`
);
// full.filePath, full.durationMs, full.fileSize, full.sha256

Recovering a gap in the stream

// Streaming socket dropped from media time 120000 to 135000 ms
const path = await recorder.extractRange(120_000, 135_000);
await transcribeFile(path); // your batch transcription endpoint

After a crash

const orphaned = await Anvil.discoverOrphanedRecordings(recordingsDir);
for (const session of orphaned) uploader.enqueueAll(session.segments);

Headers are repaired and markers cleared before the sessions are returned; the WAV files stay on disk for you. To group sessions back into your own ids (meetingId, callId…) across a crash — so a 90-minute meeting interrupted mid-way comes back as one logical recording — use createRecordingService. See docs/recovery.md.


🔁 Record → Play → Upload, all Nitro

Anvil produces plain WAV files, so the rest of the pipeline is whatever you already use. The example app wires it like this:

// Record
const segments = await recorder.stop();
const full = await Anvil.concatenate(
  segments.map((s) => s.filePath),
  outputPath
);

// Play — react-native-nitro-player
await PlayerQueue.addTrackToPlaylist(playlistId, {
  id: full.filePath,
  title: 'Recording',
  artist: 'Anvil',
  album: 'Recordings',
  duration: full.durationMs / 1000,
  url: `file://${full.filePath}`,
});
await TrackPlayer.playSong(full.filePath, playlistId);

// Upload — react-native-nitro-cloud-uploader (multipart presigned URLs, background, resumable)
await CloudUploader.startUpload(uploadId, full.filePath, uploadUrls, 3, true);

Every step is a Nitro Module and nothing crosses the old bridge. See example/ for the full app with a player card, an upload progress bar and "play the uploaded URL".


🔐 Permissions

iOS — Info.plist

<key>NSMicrophoneUsageDescription</key>
<string>Records your conversations</string>
<key>UIBackgroundModes</key>
<array>
  <string>audio</string>
</array>

Without UIBackgroundModes = audio, iOS suspends your app within 30 seconds of backgrounding and your recording stops. See iOS quirks for the full explanation.

Android

Declared by the library and merged automatically:

  • RECORD_AUDIO — microphone capture
  • FOREGROUND_SERVICE + FOREGROUND_SERVICE_MICROPHONE — background recording (keepAwakeInBackground: true)
  • WAKE_LOCK — keep the CPU awake while recording in the background
  • POST_NOTIFICATIONS — the foreground-service notification (Android 13+)

A typical app manifest that works out of the box with Anvil, a player and an uploader:

<uses-permission android:name="android.permission.INTERNET" />
<uses-permission android:name="android.permission.MODIFY_AUDIO_SETTINGS" />
<uses-permission android:name="android.permission.POST_NOTIFICATIONS" />
<uses-permission android:name="android.permission.RECORD_AUDIO" />
<uses-permission android:name="android.permission.WAKE_LOCK" />
<uses-permission android:name="android.permission.READ_EXTERNAL_STORAGE" android:maxSdkVersion="32" tools:replace="android:maxSdkVersion" />
<uses-permission android:name="android.permission.WRITE_EXTERNAL_STORAGE" android:maxSdkVersion="32" tools:replace="android:maxSdkVersion" />

MODIFY_AUDIO_SETTINGS is recommended: some OEMs need it for audio-focus and routing calls to behave. The storage permissions are only needed for apps that write outside their sandbox — Anvil writes to whatever directory you give it and needs none of them for the app's own document directory.

Runtime permission for Android 13+: request POST_NOTIFICATIONS before starting a background recording so the foreground-service notification is visible. Recording works either way; only the notification is hidden if denied.

import { PermissionsAndroid, Platform } from 'react-native';

if (Platform.OS === 'android' && Platform.Version >= 33) {
  await PermissionsAndroid.request(
    PermissionsAndroid.PERMISSIONS.POST_NOTIFICATIONS
  );
}

keepAwakeInBackground: true requires notification in the config and must be started while the app is in the foreground (Android 14+ rule).

Aggressive Android OEMs (Xiaomi/HyperOS, Huawei, Oppo, Vivo, Realme, older OnePlus) may still kill your foreground service on screen-off despite these permissions. See OEM quirks for per-brand user setup steps.


📡 Events

| Listener | When | | ----------------------------- | ------------------------------------------------------------------------------------------------- | | addPCMListener | every streamChunkMs while recording | | addSpeakerWindowListener | every speakerWindowHopMs once a full window exists | | addInterruptionListener | OS took / returned the mic (call, muted, route, reset, focus, other) | | addRouteChangeListener | input device changed; segment rotated when the active input changed | | addPermissionChangeListener | mic permission differs from last check (checked on every start/resume) | | addStorageWarningListener | free space below storageWarningBytes (checked at start and every rotation) | | addSegmentCompletedListener | a WAV file was finalized, with sha256 | | addErrorListener | pipeline failure OR deferred-resume signal; recorder moves to interrupted, data on disk is safe |

RecorderState: idle → recording ⇄ paused / interrupted → stopped. stop() always resolves with every segment.

Timeline

All timestamps (PCMChunk.timestampMs, SpeakerWindow.startMs, RecordingSegment.mediaStartMs, extractRange) are media time: milliseconds of captured audio, which only advance while capturing. That is the timeline a streaming transcription service sees, so joining transcript segments with speaker labels is a plain interval overlap.

Note on resumed recordings: after an interruption + auto-resume, media time resumes from where it left off (the samples pause too). If you're rebuilding a wall-clock timeline for the UI, use Date.now() at each turn rather than media time — media time is a captured-audio counter, not a real-world one.


🧩 Supported Platforms

| Platform | Status | | -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | | iOS | ✅ Fully Supported | | Android | ✅ Fully Supported | | iOS Simulator | ⚠️ Partial — records fine but interruption / route / background behaviors don't fire realistically. See iOS quirks. | | Android Emulator | ⚠️ Partial — audio focus events unreliable when other apps take the mic. Test on real device for interruption flows. |

Interruption, route-change and background behaviour cannot be verified on simulators — test on a real device before shipping. On both platforms, the simulator/emulator's virtual audio HAL does not behave like real hardware, so AUDIOFOCUS_GAIN (Android) or interruption end notifications (iOS) may not fire when another emulator app releases the mic.


🔧 Building the library

yarn install
yarn nitrogen        # generates nitrogen/generated from src/specs/*.nitro.ts
yarn typecheck

nitrogen/generated/ must be committed and shipped in the npm package.


🤝 Contributing

Contributions are welcome!


🪪 License

MIT © Gautham Vijayan


Made with ❤️ and Nitro Modules