media-classifier-loop
v1.1.2
Published
Autonomous media classification loop via remote Ollama and HTTPS SQLite wrapper
Readme
classy
Classify a private media library with Ollama and computer vision models.
The indexer scans the data/ directory recursively, fingerprints each file with
SHA-256, extracts local metadata, and stores classification results in a remote
SQLite-compatible database module.
How It Works
- The indexer initializes the
media_signaturesandfile_locationstables. - It verifies that the configured Ollama models are available.
- It scans files below
data/. - It skips paths already recorded in
file_locations. - It reuses metadata for duplicate file contents based on SHA-256.
- It classifies photos, videos, and document text with Ollama.
- It extracts audio and document metadata with the Python helpers.
Supported categories are determined from MIME types:
- Photos:
image/* - Videos:
video/* - Audio:
audio/* - PDF files:
application/pdf - EPUB and Mobipocket files
Requirements
- Node.js with native
fetchsupport - Python 3
- FFmpeg and FFprobe for video files
- An Ollama server reachable from the indexer
- A database module importable by Node.js through
DATABASE_URL
Install the JavaScript and Python dependencies with:
pnpm install
pip3 install -r requirements.txtConfiguration
Set these environment variables before starting the indexer:
| Variable | Required | Default | Description |
| --- | --- | --- | --- |
| DATABASE_URL | Yes | None | Importable database module exposing run() and get() methods |
| OLLAMA_URL | Yes | None | Ollama chat API base URL |
| VISION_MODEL | No | qwen2.5-vl:7b | Model used for photos and video frames |
| TEXT_MODEL | No | gemma2:27b | Model used for PDFs and EPUBs |
| PORT | For web process | None | Port used by the health-check server |
The input directory is fixed to data/ relative to the process working
directory. Mount or copy the media library there.
Running Locally
Start the classification worker:
DATABASE_URL="your-database-module" \
OLLAMA_URL="http://localhost:11434/v1" \
pnpm startThe application starts the HTTP health process and performs an initial scan. Trigger another scan with:
curl -X POST http://localhost:3000/scanMonitor the current scan with:
curl http://localhost:3000/statusThe root endpoint responds with OK.
Docker
The included Dockerfile installs Node.js dependencies, Python dependencies, and FFmpeg. Build and run it with a mounted media directory and the required environment variables:
docker build -t classy .
docker run --rm \
-e DATABASE_URL="your-database-module" \
-e OLLAMA_URL="http://ollama:11434/v1" \
-v "$PWD/data:/app/data" \
classyThe exact mount path must match the container's process working directory,
because the application resolves data/ from process.cwd().
Deployment
The deployed health endpoint is available at:
Check the deployment with:
curl https://classy.api.apphor.deThe expected response is OK.
Start a scan and inspect its progress with:
curl -X POST https://classy.api.apphor.de/scan
curl https://classy.api.apphor.de/statusThe archive UI is available at the same URL. It is responsive, supports image, audio, and video previews, and can be installed as a PWA from a compatible browser.
Media API
GET /api/media: paginated indexed media. Supportssearch,type,category,limit, andoffset.GET /api/media/:id: full metadata for one indexed file.GET /api/media/:id/content: stream the original file, including byte ranges for video and audio playback.GET /api/status: scan counters and the current index location.POST /scan: trigger a new scan.
Database Tables
media_signatures
Stores one classification record per unique SHA-256 file content:
sha256: primary keyfile_type: detected media typeextracted_date: file or embedded metadata dateai_category: broad classificationai_summary: generated descriptionraw_metadata: JSON-encoded extractor metadata
file_locations
Stores every indexed path:
id: auto-incrementing identifiersha256: content hashfile_path: unique indexed pathfile_size: size at indexing time
Python Extractors
extract_audio.pyreads ID3 metadata witheyed3and falls back to artist and album names inferred from parent directories.extract_doc.pyreads the first two PDF pages or the first EPUB document contents and returns a text excerpt for classification.
Extractor failures are returned as empty metadata rather than stopping the entire scan.
Operational Notes
- A path is considered processed once it is inserted into
file_locations. Replacing a file at an already indexed path does not currently trigger a re-scan. - Duplicate content is classified once; additional paths only receive a location record.
- Video processing creates temporary frame files under
/tmp/frames. - The Ollama model pull step can take significant time, especially for the default 27B text model.
- There are currently no automated tests in the repository.
