npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

sine-wave-tts

v0.1.0

Published

Makes text speak in beeping sine-wave tones. A voice for robots, AI characters, and mascots that aren't human.

Readme

Sine Wave TTS

Read this in other languages: 日本語

A TypeScript library that makes text "speak" in beeping sine-wave tones.

The result sounds like a robot or a video-game character chattering in electronic beeps. Instead of imitating a human voice, the library moves pitch and rhythm to follow the readings and accents of the input text. No actual words come through, but it still sounds like someone talking. It was built to give a voice to AI characters and mascots that aren't human.

  • The same text always produces the same sound, so each character keeps a consistent, recognizable voice
  • Understands Japanese readings and accents and reflects them in the sound (kuromoji.js, with a built-in lightweight fallback)
  • 5 voice characters and 7 emotions that can be combined in any pairing
  • Outputs 44.1 kHz / mono / 16-bit WAV plus per-unit timings for subtitles and lip sync
  • Works in both Node.js and the browser

Requirements

  • Node.js 20 or later
  • npm (bundled with Node.js)

Corepack, pnpm, and Yarn are not used.

Installation

To try it from the repository:

npm install
npm test

To use it as a library once published:

npm install sine-wave-tts

Basic API

import {
  loadKuromojiAnalyzer,
  synthesizeAsync,
} from "sine-wave-tts";

await loadKuromojiAnalyzer({ throwOnError: true });

const result = await synthesizeAsync("こんにちは、サイン波の声です。", {
  speaker: "songful",
  emotion: "joy",
  speed: 1,
  pitch: 1,
  volume: 0.8,
});

console.log(result.durationMs, result.timings);
const wav = result.toWav();

synthesize(text, options?)

Generates PCM synchronously. When the kuromoji analyzer is not loaded yet, it still synthesizes immediately using the built-in approximate analysis.

synthesizeAsync(text, options?)

A Promise-based wrapper that makes it easy to swap in for a real TTS service. Loading the analyzer itself is done explicitly with loadKuromojiAnalyzer().

SynthesisOptions

| Field | Type | Default | Description | |---|---|---|---| | speaker | string \| SpeakerPreset | "default" | Register, tempo, harmonics, portamento | | emotion | string \| EmotionPreset | "neutral" | Pitch, range, tempo, phrase endings, envelope | | speed | number | 1 | Overall speed. Positive number | | pitch | number | 1 | Overall pitch multiplier. Positive number | | volume | number | 0.8 | Output volume, 0..1 |

SynthesisResult

| Field | Type | Description | |---|---|---| | pcm | Float32Array | 44.1 kHz mono PCM | | sampleRate | number | Currently 44100 | | durationMs | number | Length of the synthesized audio | | timings | UnitTiming[] | Start / end time of each utterance unit | | toWav() | () => ArrayBuffer | Encode as 16-bit PCM WAV |

Presets

Speakers define voice character and base rhythm; emotions define expressive modulation. The two axes are independent, so any speaker works with any emotion.

| Speaker | Character | |---|---| | default | Standard signal voice, 32 tones over ~2.6 octaves | | chirpy | High register, fast, bright harmonics, small-creature-like | | deep | Low register, slow, near-pure tone, long glides | | robotic | Mid register, odd harmonics, no portamento | | songful | 3 octaves, standard vibrato, singing-like glides |

Available emotions are neutral, joy, sad, angry, surprise, calm, and fear. The registered lists can be read at runtime from supportedSpeakers and supportedEmotions.

A custom SpeakerPreset specifies an ascending frequency scale, baseTempo in morae per second, harmonics, vibrato, ADSR, and portamento time. Comments in the type definitions describe which direction to adjust each parameter.

Web demo (sample playback)

npm run webdemo

Open the printed local URL in your browser. Before startup, the kuromoji dictionary is automatically synced into demo/public/dict/. While the dictionary is loading, playback works via the fallback analyzer; once loading finishes, the status indicator switches to "Kuromoji analyzer ready".

To build the static production files:

npm run webdemo:build

HTTP API server

Start the dependency-free Node.js server (default: 127.0.0.1:50021):

npm run serve

Select another port with PORT=51000 npm run serve or npm run serve -- --port 51000. The command-line flag takes precedence. The server loads kuromoji before listening and reports whether the analyzer is ready or using the fallback. All routes allow CORS.

Native API

| Method and path | Response | |---|---| | POST /v1/synthesize | WAV; send Accept: application/json for Base64 WAV and timings | | GET /v1/speakers | Speaker presets and parameter summaries | | GET /v1/emotions | Emotion presets | | GET /v1/health | Server and analyzer status |

curl -sS -X POST http://127.0.0.1:50021/v1/synthesize \
  -H 'Content-Type: application/json' \
  --data '{"text":"こんにちは、APIです。","speaker":"chirpy","emotion":"joy"}' \
  --output native.wav

OpenAI-compatible API

OpenAI TTS clients can use POST /v1/audio/speech. The model string is accepted for compatibility and does not change synthesis. voice accepts a speaker name such as chirpy or a speaker:emotion pair such as chirpy:joy. Read the available names from GET /v1/speakers and GET /v1/emotions.

curl -sS -X POST http://127.0.0.1:50021/v1/audio/speech \
  -H 'Authorization: Bearer local' \
  -H 'Content-Type: application/json' \
  --data '{"model":"tts-1","input":"こんにちは、OpenAI互換APIです。","voice":"chirpy:joy","speed":1,"response_format":"wav"}' \
  --output openai.wav

response_format defaults to wav, intentionally differing from OpenAI's MP3 default. The other supported value is pcm (raw 44.1 kHz, mono, 16-bit little-endian PCM); MP3, Opus, AAC, and FLAC return an OpenAI-shaped 400 error. speed must be between 0.25 and 4.0. GET /v1/models lists the local compatibility model.

VOICEVOX-compatible API

The compatibility layer exposes /audio_query, /synthesis, /speakers, and /version. Each speaker × emotion pair is assigned a stable numeric style ID; read /speakers instead of hard-coding IDs.

curl -sS -X POST \
  'http://127.0.0.1:50021/audio_query?text=こんにちは&speaker=0' \
  --output query.json

curl -sS -X POST \
  'http://127.0.0.1:50021/synthesis?speaker=0' \
  -H 'Content-Type: application/json' \
  --data-binary @query.json \
  --output voicevox.wav

speedScale, pitchScale, and volumeScale in the AudioQuery are mapped to the native synthesis controls. This is a practical playback-compatible subset, not a complete implementation of every VOICEVOX feature.

Architecture

text
  → analyzer      readings, POS, accent phrases (falls back on failure)
  → contour       mora timing, accent, declination, scale quantization
  → prosody       speaker × emotion × phrase-final punctuation
  → synthesizer   harmonics, ADSR, portamento, vibrato → PCM
  → wav/timings   WAV encoding and sync information

The synthesis core does not depend on the Web Audio API; browser playback is handled by the demo. Only the Japanese analysis layer uses kuromoji.js and its dictionary data.

Development

npm run typecheck
npm test
npm run build
npm run webdemo:build
npm run serve