npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@simo777/youtube-smart

v1.0.13

Published

Production-ready YouTube Smart Recommendation SDK for LearnPilot AI

Readme

YouTube Bridge API (@simo777/youtube-smart)

A direct, metadata-focused YouTube search and ranking SDK. Designed to separate data collection from scoring, making it perfect for systems that save video metadata to a database and rank them later.

Installation

npm install @simo777/youtube-smart
# or
pnpm add @simo777/youtube-smart
# or
yarn add @simo777/youtube-smart

Core Principles

  1. Simple & Direct: No complex "Learning Context" objects. Just search with queries.
  2. Decoupled Ranking: Fetch data now, save to DB, and rank whenever you need.
  3. Metadata Driven: Ranking uses real engagement signals (views, likes, recency) instead of unreliable AI evaluations.
  4. AI Transcript Support: Keep the power of AI where it matters—extracting text from videos.

Quick Start

1. Initialize the Client

import { LearnPilotYouTube } from "@simo777/youtube-smart";

const youtube = new LearnPilotYouTube({
  youtubeApiKey: process.env.YOUTUBE_API_KEY!,
  ai: {
    // Optional: URL for your AI transcript service
    freellmUrl: "http://localhost:3015/v1",
    freellmApiKey: "your-key"
  }
});

2. Search for Videos (Data Collection)

This method returns raw metadata suitable for saving to your database. It does not perform ranking.

const videos = await youtube.findVideos({
  query: "Quantum Physics for beginners",
  options: {
    candidates: 20,
    regionCode: "US"
  }
});

// Save 'videos' to your Database...

3. Rank Stored Videos

Use the RankingEngine to calculate scores and sort videos retrieved from your database.

import { RankingEngine } from "@simo777/youtube-smart";

const rankingEngine = new RankingEngine();

// Data from your DB mapped to the VideoRankingData interface
const videosFromDb = [
  {
    videoId: "abc123",
    viewCount: 150000,
    likeCount: 5000,
    publishedAt: "2023-10-01T10:00:00Z",
    caption: true,
    matchedQueries: ["Quantum Physics"]
  }
  // ... more videos
];

const scores = rankingEngine.calculateScores(videosFromDb);
// Returns: Array of { videoId, final, engagement, likeRatio, freshness, metadata }

API Reference

LearnPilotYouTube

findVideos({ query, options })

Returns a list of RecommendedVideo metadata objects.

  • options.candidates: Number of results to fetch (default 30).
  • options.regionCode: ISO 3166-1 alpha-2 country code.

findPlaylists({ query, options })

Returns a list of RecommendedPlaylist metadata.

scoreVideos(videos: VideoRankingData[])

A convenience method on the client that uses the internal RankingEngine to return scores for a list of video data.

getVideoTranscript(videoId, lang, fallbackToAI)

Fetches video transcripts.

  • fallbackToAI: If true, calls your configured AI service if YouTube has no official transcript. It transcribes the full audio directly. Throws an error if the audio duration exceeds maxAudioDurationMinutes (default 40 min).

getVideoTranscriptAI(videoId, lang)

Fetches video transcripts directly using the AI pipeline (full audio), bypassing official platform transcription entirely.

chunkAudio(videoId, maxChunkMinutes)

Manually segments a downloaded audio file into chunks. Useful if you want to process long videos yourself.

  • Returns: Promise<string[]> (list of chunk filenames like videoId_001.mp3).

transcribeChunks(chunkFiles, lang)

Transcribes a specific list of audio chunks.

deleteVideoAudio(videoId)

Cleans up the disk space by deleting parent audio and associated chunks.


AI Transcription Strategy

  1. Direct Processing: By default, the AI pipeline transcribes the full audio file directly.
  2. Resilient Loop Execution: Rotates through multiple compatible models (Whisper v3, etc.) with a 30-second timeout per attempt.
  3. Safety Limit: If a video is longer than the configured limit (default 40 mins), the system throws an error to prevent massive API costs or timeouts. You can then choose to use chunkAudio manually.

Error Handling & Tracking

When using the AI pipeline, the response object returns specific statuses to allow storage and recovery inside your application's DB layer:

  1. Critical Audio Download Failures: If the full video audio extraction fails via yt-dlp, the SDK returns:

    {
      "videoId": "KhFsKup7tUk",
      "items": [],
      "fullText": "",
      "method": "ai",
      "model": "whisper-large-v3-turbo",
      "error_type": "YT_DOWNLOAD_AUDIO_FAILED",
      "error": "Command failed: yt-dlp ...",
      "chunks": []
    }
  2. Chunk-Level Provider Failures: If rate-limits or server failures occur midway, the audio chunk is kept intact on disk. The response lists the status of all segments so you can save failed chunks to your DB:

    {
      "videoId": "KhFsKup7tUk",
      "method": "ai",
      "model": "whisper-large-v3",
      "chunks": [
        { "chunk_id": "KhFsKup7tUk_000", "status": "success" },
        { "chunk_id": "KhFsKup7tUk_001", "status": "fail" }
      ]
    }

    You can easily recover by feeding the failed array later into retryFailedChunks(["KhFsKup7tUk_001.mp3"]).


AI Transcription & Chunking Strategy

When official YouTube transcripts fail or getVideoTranscriptAI is invoked, the SDK applies the following robust architecture:

  1. Audio Extraction: Extracts full high-efficiency audio via yt-dlp.
  2. Local Splitting: If the audio is long, it segments it into 40-minute chunks (videoId_00*.mp3) using ffmpeg locally inside the audio/ folder.
  3. Resilient Loop Execution: Processes chunks sequentially with a configurable delay (chunkDelayMs) to minimize API rate-limits.
  4. Robust Model Fallback: For each chunk, it rotates through multiple compatible OpenAI/Whisper models. Each request enforces a strict 30-second timeout to dynamically switch models if a provider hangs or throws a 502.
  5. No Auto-Deletion: Chunks are preserved on disk until explicitly deleted, allowing fail/success tracking status to be written directly into your database.

Docker Deployment Guide (e.g., Next.js inside Docker)

Because Docker containers are isolated, headless Linux environments, they do not have access to your local computer's Chrome installation or profiles. If you invoke the SDK inside a Docker container without a configuration, yt-dlp will try to scrape cookies from Chrome and crash.

To successfully run this inside Docker (such as deploying a Next.js production app), follow these steps:

1. Extract YouTube Cookies Locally

  1. Install a browser extension like Get cookies.txt LOCALLY (for Chrome/Firefox) on your developer machine.
  2. Visit youtube.com and ensure you are logged in.
  3. Open the extension, click export, and save the file as cookies.txt inside your project root. Note: YouTube cookies expire after a few months, or instantly if you sign out from that browser session. If you get 429 Too Many Requests or download failure errors later, simply replace it with a fresh cookies.txt.

2. Configure the SDK with cookiesPath

Pass the path where the container can find your cookie file:

const youtube = new LearnPilotYouTube({
  youtubeApiKey: process.env.YOUTUBE_API_KEY!,
  cookiesPath: "/app/cookies.txt", // Absolute path inside the Docker container
  ai: {
    freellmUrl: "http://your-whisper-backend/v1",
    freellmApiKey: "your-api-key",
    chunkDelayMs: 2000 // 2 seconds between chunks to prevent server strain
  }
});

3. Add to Docker Environment

Option A: Copy inside Dockerfile (Best for cloud deployments)

Add a layer in your production multi-stage Dockerfile to pull the file inside your active application context:

# Inside your Dockerfile
WORKDIR /app
COPY cookies.txt /app/cookies.txt

⚠️ Security Reminder: Make sure your Docker image repository is private, as your active YouTube session cookies will be baked directly into the image structure.

Option B: Mount at Runtime via Volumes (Best for local dev/testing)

Avoid copying the file into the build layers. Mount the text file from the host filesystem directly when initiating the container:

docker run -d \
  -p 3000:3000 \
  -v $(pwd)/cookies.txt:/app/cookies.txt \
  -e YOUTUBE_API_KEY="your_api_key" \
  your-nextjs-image

4. Ensure Dependencies Exist in your Base Image

Your Docker image container must have ffmpeg and ffprobe installed on the underlying OS. Ensure your Dockerfile installs them:

# For Debian/Ubuntu-based Node images
RUN apt-get update && apt-get install -y ffmpeg yt-dlp && rm -rf /var/lib/apt/lists/*

Ranking Algorithm

The RankingEngine uses a weighted formula to calculate the final score (0.0 to 1.0):

| Signal | Weight | Description | | :--- | :--- | :--- | | Engagement | 40% | Log-normalized view count relative to the batch. | | Like Ratio | 30% | Quality signal (Likes/Views). Normalized against a 5% "excellent" threshold. | | Freshness | 20% | Rewards recent videos; decays for content older than 1-5 years. | | Metadata | 10% | Rewards videos with Captions (CC) and multiple query matches. |


Configuration

const youtube = new LearnPilotYouTube({
  youtubeApiKey: "...",
  ai: {
    freellmUrl: "...",
    freellmApiKey: "...",
    chunkDelayMs: 2000
  },
  ranking: {
    weights: {
      engagement: 0.5, // Increase importance of views
      freshness: 0.1   // Decrease importance of date
    }
  },
  timeoutMs: 30000
});

Development

npm run build
npm run lint

License

ISC