npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

media-classifier-loop

v1.1.2

Published

Autonomous media classification loop via remote Ollama and HTTPS SQLite wrapper

Readme

classy

Classify a private media library with Ollama and computer vision models.

The indexer scans the data/ directory recursively, fingerprints each file with SHA-256, extracts local metadata, and stores classification results in a remote SQLite-compatible database module.

How It Works

  1. The indexer initializes the media_signatures and file_locations tables.
  2. It verifies that the configured Ollama models are available.
  3. It scans files below data/.
  4. It skips paths already recorded in file_locations.
  5. It reuses metadata for duplicate file contents based on SHA-256.
  6. It classifies photos, videos, and document text with Ollama.
  7. It extracts audio and document metadata with the Python helpers.

Supported categories are determined from MIME types:

  • Photos: image/*
  • Videos: video/*
  • Audio: audio/*
  • PDF files: application/pdf
  • EPUB and Mobipocket files

Requirements

  • Node.js with native fetch support
  • Python 3
  • FFmpeg and FFprobe for video files
  • An Ollama server reachable from the indexer
  • A database module importable by Node.js through DATABASE_URL

Install the JavaScript and Python dependencies with:

pnpm install
pip3 install -r requirements.txt

Configuration

Set these environment variables before starting the indexer:

| Variable | Required | Default | Description | | --- | --- | --- | --- | | DATABASE_URL | Yes | None | Importable database module exposing run() and get() methods | | OLLAMA_URL | Yes | None | Ollama chat API base URL | | VISION_MODEL | No | qwen2.5-vl:7b | Model used for photos and video frames | | TEXT_MODEL | No | gemma2:27b | Model used for PDFs and EPUBs | | PORT | For web process | None | Port used by the health-check server |

The input directory is fixed to data/ relative to the process working directory. Mount or copy the media library there.

Running Locally

Start the classification worker:

DATABASE_URL="your-database-module" \
OLLAMA_URL="http://localhost:11434/v1" \
pnpm start

The application starts the HTTP health process and performs an initial scan. Trigger another scan with:

curl -X POST http://localhost:3000/scan

Monitor the current scan with:

curl http://localhost:3000/status

The root endpoint responds with OK.

Docker

The included Dockerfile installs Node.js dependencies, Python dependencies, and FFmpeg. Build and run it with a mounted media directory and the required environment variables:

docker build -t classy .
docker run --rm \
  -e DATABASE_URL="your-database-module" \
  -e OLLAMA_URL="http://ollama:11434/v1" \
  -v "$PWD/data:/app/data" \
  classy

The exact mount path must match the container's process working directory, because the application resolves data/ from process.cwd().

Deployment

The deployed health endpoint is available at:

https://classy.api.apphor.de

Check the deployment with:

curl https://classy.api.apphor.de

The expected response is OK.

Start a scan and inspect its progress with:

curl -X POST https://classy.api.apphor.de/scan
curl https://classy.api.apphor.de/status

The archive UI is available at the same URL. It is responsive, supports image, audio, and video previews, and can be installed as a PWA from a compatible browser.

Media API

  • GET /api/media: paginated indexed media. Supports search, type, category, limit, and offset.
  • GET /api/media/:id: full metadata for one indexed file.
  • GET /api/media/:id/content: stream the original file, including byte ranges for video and audio playback.
  • GET /api/status: scan counters and the current index location.
  • POST /scan: trigger a new scan.

Database Tables

media_signatures

Stores one classification record per unique SHA-256 file content:

  • sha256: primary key
  • file_type: detected media type
  • extracted_date: file or embedded metadata date
  • ai_category: broad classification
  • ai_summary: generated description
  • raw_metadata: JSON-encoded extractor metadata

file_locations

Stores every indexed path:

  • id: auto-incrementing identifier
  • sha256: content hash
  • file_path: unique indexed path
  • file_size: size at indexing time

Python Extractors

  • extract_audio.py reads ID3 metadata with eyed3 and falls back to artist and album names inferred from parent directories.
  • extract_doc.py reads the first two PDF pages or the first EPUB document contents and returns a text excerpt for classification.

Extractor failures are returned as empty metadata rather than stopping the entire scan.

Operational Notes

  • A path is considered processed once it is inserted into file_locations. Replacing a file at an already indexed path does not currently trigger a re-scan.
  • Duplicate content is classified once; additional paths only receive a location record.
  • Video processing creates temporary frame files under /tmp/frames.
  • The Ollama model pull step can take significant time, especially for the default 27B text model.
  • There are currently no automated tests in the repository.