@cococopi/trackforge
v0.1.0
Published
Search an artist, pick the tracks you want, and forge them into a normalised, segmented, manifest-backed audio dataset for training AI models.
Maintainers
Readme
trackforge
Search an artist, pick the tracks you want, and forge them into a training-ready audio dataset: normalised, resampled, segmented, checksummed and described by a manifest your loader can read.
Built for the boring middle of the workflow — the part between "I have some songs" and "I can actually train on this".
trackforge "Sai Abhyankar"That searches, shows you the results, asks which ones you want, downloads them, and writes a dataset.
Requirements
Two external tools do the heavy lifting. trackforge orchestrates them; it does not reimplement them.
| tool | why | install |
| --- | --- | --- |
| yt-dlp | search + download | pip install -U yt-dlp |
| ffmpeg (with ffprobe) | probe, resample, normalise, cut | pkg install ffmpeg (Termux) · apt install ffmpeg · brew install ffmpeg |
| demucs | optional, --stems only | pip install demucs |
Run the built-in check at any time:
trackforge doctorIt reports each tool, its version, and the exact install command for your platform.
Install
npm install -g @cococopi/trackforgeOr run it without installing anything global:
npx @cococopi/trackforge "Sai Abhyankar"Quick start
# Search, pick interactively, build ./trackforge-dataset
trackforge "Sai Abhyankar"
# Non-interactive: take every result
trackforge "Sai Abhyankar" --all
# Exactly these results, 20s clips, 16 kHz mono for a speech model
trackforge "Sai Abhyankar" --pick 1,3,5-8 --segment 20 --sample-rate 16000
# Look before you download
trackforge "Sai Abhyankar" --dry-runSelection accepts 1,3,5-8, all, or nothing to cancel. With --pick, --all or --yes it never prompts, so it drops straight into a script.
What you get
trackforge-dataset/
audio/ full tracks: resampled, channel-mixed, loudness-normalised
segments/ fixed-length clips cut from those tracks
raw/ untouched downloads (deleted unless --keep-raw)
manifest.jsonl one JSON record per training clip
dataset.json summary statistics + the config that produced them
dataset-card.md human-readable card, with a loader snippetOne manifest line:
{
"id": "kJ8xQ_0003",
"split": "train",
"audio": {
"path": "segments/kJ8xQ/kJ8xQ_0003.wav",
"format": "wav",
"sample_rate": 44100,
"channels": 1,
"duration": 30,
"bytes": 2646044,
"sha256": "9f2c…"
},
"segment": { "index": 3, "start": 90, "end": 120 },
"source": {
"platform": "youtube",
"video_id": "kJ8xQ",
"url": "https://www.youtube.com/watch?v=kJ8xQ",
"title": "Tuzi Yaad",
"artist": "Sai Abhyankar",
"upload_date": "20240118",
"duration": 245.6
},
"text": null,
"loudness_target_lufs": -16
}Why these choices
Splits are assigned per track, not per segment. Two clips from the same song are near-duplicates, so letting them straddle train and validation leaks and quietly inflates every number you measure afterwards. splitForTrack hashes the track id, so the split is stable across runs.
Loudness normalisation is single-pass EBU R128, targeting -16 LUFS with a -1.5 dBTP ceiling. Two-pass is marginally more accurate, but the difference is well under the noise floor of most training setups, and the true-peak ceiling is what actually keeps samples from clipping.
text is reserved, not filled. It is null here, so an ASR or forced-alignment pass can populate it later without a schema change.
Resampling happens once. The source stream is downloaded untouched and decoded a single time by ffmpeg straight into the target rate, channel count and normalisation.
Loading it
import json
from pathlib import Path
import torchaudio
root = Path("trackforge-dataset")
records = [json.loads(line) for line in (root / "manifest.jsonl").read_text().splitlines() if line.strip()]
train = [r for r in records if r["split"] == "train"]
val = [r for r in records if r["split"] == "validation"]
waveform, sample_rate = torchaudio.load(root / train[0]["audio"]["path"])
print(waveform.shape, sample_rate)Anything that reads JSONL works — this is not tied to a framework.
Options
| flag | default | meaning |
| --- | --- | --- |
| --out <dir> | ./trackforge-dataset | output directory |
| --limit <n> | 20 | search results to consider |
| --pick <expr> | — | 1,3,5-8 |
| --all / --yes | — | take every result, never prompt |
| --dry-run | — | search and select only |
| --format wav\|flac | wav | container |
| --sample-rate <hz> | 44100 | use 16000 for speech models |
| --channels 1\|2 | 1 | mono is what most audio models want |
| --no-normalize | — | skip loudness normalisation |
| --lufs <n> | -16 | loudness target |
| --segment <sec> | 30 | clip length, 0 disables segmentation |
| --overlap <sec> | 0 | overlap between clips |
| --min-segment <sec> | 5 | drop runt tails shorter than this |
| --val-split <0..1> | 0.1 | validation fraction |
| --stems | — | also run Demucs vocal separation |
| --keep-raw | — | keep untouched downloads |
| --cookies-from-browser <b> | — | passed to yt-dlp for restricted content |
| --retries <n> | 3 | network retries per track |
| --quiet | — | hide yt-dlp progress |
trackforge --help prints the same list.
As a library
import { searchTracks, forgeTracks } from "@cococopi/trackforge";
const found = await searchTracks("Sai Abhyankar", { limit: 20 });
const result = await forgeTracks(found.slice(0, 3), "Sai Abhyankar", {
out: "./dataset",
sampleRate: 44100,
channels: 1,
format: "wav",
normalize: true,
targetLufs: -16,
segmentSeconds: 30,
overlapSeconds: 0,
minSeconds: 5,
valSplit: 0.1,
stems: false,
keepRaw: false,
showProgress: false,
});result.failures lists per-track problems, so a batch run does not die on one bad video.
Zero dependencies
The package ships no runtime dependencies. yt-dlp and ffmpeg are invoked as subprocesses, which means you get their full capability — cookies, proxies, formats, filters — instead of a wrapper's subset, and this package never needs to be rebuilt when they update.
Rights
Downloading audio does not grant you rights to it. The people who wrote and recorded these songs still own them. Use this on material you are permitted to use, respect the source platform's terms of service, and do not redistribute a dataset or ship a model trained on one unless the rights holders' terms allow it. dataset-card.md records this too, so the warning travels with the data.
License
MIT
