zetago-reddit-scraper
v1.2.2
Published
High-performance Reddit Scraper & OSINT Engine for Node.js and CLI. Scrapes Subreddits, Posts, Nested Comments, User Profiles, and Global Search with Anti-Bot & OAuth2 Support.
Maintainers
Readme
🔥 Reddit Scraper & OSINT Engine
High-performance, production-grade Reddit Scraper, REST API Microservice & Chatbot Toolkit for Node.js, TypeScript, and CLI.
Quick Start • Chatbot Integration • REST API Server • CLI Usage • API Reference • Instant Mock Testing • Output Examples • Authentication Guide
🌟 Key Features
- 🔍 Global & Subreddit Search: Search posts across all of Reddit or restrict to specific subreddits with filters (relevance, hot, top, new, comments).
- 📜 Full Subreddit Scraping: Extract
hot,new,top, andrisingposts with pagination tokens (after). - 📥 Raw Media Scraping & Auto-Downloader: Scrape direct image links, full multi-image gallery albums, and native Reddit videos with automatic DASH audio muxing via FFmpeg (
scrapeMedia,downloadMedia,reddit-scraper download). - 🤖 Chatbot Ready: Built-in helpers to format posts for WhatsApp (Baileys), Discord, and Telegram with direct media URL extraction (
extractMedia,formatForBot). - 📡 Subreddit Event Stream Watcher: Event-driven watcher that continuously polls subreddits for incoming new submissions (
watchSubreddit). - 🚀 Zero-Dependency REST API Microservice: Built-in CORS-enabled micro HTTP server (
createApiServer,reddit-scraper serve). - 💬 Deep Comment Tree Parsing: Recursively extracts nested comment threads with user scores, authors, and timestamps up to arbitrary depth.
- 👤 User Profile & OSINT: Fetch user metadata, link/comment karma breakdown, moderator status, submitted posts, and user comments.
- 🧪 Offline Mock Testing: Pass
{ mock: true }or--mockfor instant testing in CI/CD without hitting Reddit rate limits or needing API keys. - 🛡️ Anti-Bot & OAuth2 Support: Seamless fallback between public
.jsonendpoints, browser header impersonation, and official OAuth2 application tokens. - ⚡ Zero-Config CLI & NPX: Run instantly without installing via
npx zetago-reddit-scraper search "AI". - 📦 Dual Module & TypeScript Native: Full CommonJS (
require), ES Modules (import), and comprehensive TypeScript typings (.d.ts).
📦 Installation
Via NPM:
npm install zetago-reddit-scraperGlobal CLI Installation:
npm install -g zetago-reddit-scraperInstant Execution with NPX:
npx zetago-reddit-scraper --help🚀 Quick Start
Basic Scraping (ESM or CommonJS)
import { RedditScraper } from 'zetago-reddit-scraper';
// Or CommonJS: const { RedditScraper } = require('zetago-reddit-scraper');
const scraper = new RedditScraper();
// 1. Search posts
const posts = await scraper.search({
query: 'Artificial Intelligence',
subreddit: 'technology',
sort: 'top',
timeFilter: 'week',
limit: 10,
});
posts.forEach((post) => {
console.log(`[${post.score} pts] ${post.title} (by u/${post.author})`);
});🤖 Chatbot Integration (WhatsApp, Discord, Telegram)
zetago-reddit-scraper has first-class helpers specifically designed for bots (e.g. WhatsApp with Aurum-Baileys, Discord.js, Telegraf):
import { RedditScraper, formatForBot, extractMedia } from 'zetago-reddit-scraper';
const scraper = new RedditScraper();
// Example: Fetch random meme for "!meme" bot command
const meme = await scraper.getRandomPost('memes');
const media = extractMedia(meme);
// 1. Send to WhatsApp (Baileys)
await sock.sendMessage(jid, {
image: { url: media.url },
caption: formatForBot(meme, { platform: 'whatsapp' }),
});
// 2. Send to Discord
channel.send({
content: formatForBot(meme, { platform: 'discord' }),
files: media.url ? [media.url] : [],
});
// 3. Send to Telegram
bot.telegram.sendPhoto(chatId, media.url, {
caption: formatForBot(meme, { platform: 'telegram' }),
parse_mode: 'HTML',
});📡 Live Subreddit Event Watcher
Listen for new posts in real time and automatically forward them to your bot channels:
const watcher = scraper.watchSubreddit({
subreddit: 'technology',
intervalMs: 15000, // poll every 15s
});
watcher.on('post', (post) => {
console.log('New post submitted:', post.title);
// Broadcast to WhatsApp / Telegram group
});
watcher.on('error', (err) => console.error(err.message));
// To stop listening:
// watcher.stop();📡 Built-In REST API Server
Need an API endpoint for your frontend, mobile app, or webhook? Start a local microservice with zero external server dependencies and CORS enabled:
Via Code:
import { createApiServer } from 'zetago-reddit-scraper';
const { start } = createApiServer({ port: 3000 });
start(({ port, host }) => console.log(`API running at http://${host}:${port}`));Via CLI:
# Live API Server
reddit-scraper serve --port 3000
# Mock API Server (for local frontend/bot offline development)
reddit-scraper serve --port 3000 --mockReady-To-Use REST Endpoints:
| Endpoint | Method | Description |
| :--- | :--- | :--- |
| /health | GET | Server status and documentation index |
| /api/search?q=<query>&subreddit=<name>&limit=<num> | GET | Search posts with filtering |
| /api/r/:subreddit?sort=<hot\|top\|new>&limit=<num> | GET | Fetch subreddit submissions |
| /api/r/:subreddit/random | GET | Get a random post with parsed media |
| /api/r/:subreddit/about | GET | Get subreddit rules & subscriber metrics |
| /api/post/:id?depth=3&limit=50 | GET | Fetch post with full comment tree |
| /api/user/:username?posts=true&comments=true | GET | Fetch user profile, karma, and activity |
| /api/bot/random?subreddit=memes&platform=whatsapp | GET | Pre-formatted payload ready for bot sending |
🧪 Zero-Network Mock Testing
Test your bots, CLI tools, and unit tests 100% offline without hitting Reddit rate limits or needing API keys:
import { RedditScraper } from 'zetago-reddit-scraper';
// Simply pass { mock: true }
const testScraper = new RedditScraper({ mock: true });
const posts = await testScraper.search({ query: 'offline test' });
console.log(`Received ${posts.length} mock posts instantly!`);Or test in the CLI anytime by adding --mock:
reddit-scraper search "cybersecurity" --mock
reddit-scraper random memes --mock --platform whatsapp
reddit-scraper user ZetaGo-Aurum --mock📥 Raw Media Scraping & Direct Downloader
Reddit hides raw media URLs behind complex metadata structures and separate audio/video DASH streams. zetago-reddit-scraper provides native raw media extraction and automated downloads:
1. Extract Raw Direct URLs (Images, Full Galleries, Video & Audio Streams)
import { RedditScraper } from 'zetago-reddit-scraper';
const scraper = new RedditScraper();
const media = await scraper.scrapeMedia('https://reddit.com/r/pics/comments/1cv9a01');
console.log('Media Type:', media.mediaType); // 'video' | 'gallery' | 'image' | 'gif'
console.log('Direct Files:', media.files);
// [
// { type: 'image', url: 'https://i.redd.it/photo1.jpg', width: 1920, height: 1080 },
// { type: 'image', url: 'https://i.redd.it/photo2.jpg', width: 1920, height: 1080 }
// ]2. Download Raw Media Directly to Disk (with Automatic FFmpeg Audio Muxing)
When downloading native Reddit videos (v.redd.it), the engine automatically pulls the separate video and audio streams and merges them into a complete, synchronized MP4 using FFmpeg:
const result = await scraper.downloadMedia('1cv9a01', {
outputDir: './downloads',
mergeAudio: true, // Auto-merges video + audio stream into single playable MP4
});
console.log('Saved files:', result.savedFiles);
// ['/home/.../downloads/Awesome_Clip_1cv9a01.mp4']3. Via CLI
# Inspect direct raw media links and video/audio stream URLs
reddit-scraper media 1cv9a01
# Download full gallery or video (with audio merged) to a folder
reddit-scraper download 1cv9a01 -o ./my_downloads💻 CLI (Command Line Interface)
The package ships with the reddit-scraper (or zetago-reddit) binary:
reddit-scraper [command] [options]1. Search Posts
# Global search
reddit-scraper search "machine learning" --limit 15 --format table
# Restricted search in a subreddit
reddit-scraper search "quantum computing" -s science --sort top -t month
# Export search results to JSON or CSV
reddit-scraper search "cybersecurity" -f json -o results.json
reddit-scraper search "linux kernel" -f csv -o linux_posts.csv2. Random Post & Media (for Chatbots)
# Get random post formatted for WhatsApp
reddit-scraper random memes --platform whatsapp
# Get direct media URL only
reddit-scraper random anime --media3. Live Subreddit Watcher
reddit-scraper watch technology --interval 15 --platform whatsapp4. Start REST API Server
reddit-scraper serve --port 30005. Subreddit Feeds
reddit-scraper subreddit programming --limit 20
reddit-scraper about webdev6. Post and Comment Tree
reddit-scraper post 1cv9a01 --depth 3 --limit 50 -f json -o post_tree.json7. User Profile OSINT
reddit-scraper user ZetaGo-Aurum
reddit-scraper user spez --posts --limit 10📚 Programmatic API Reference
new RedditScraper(options?)
| Method | Parameters | Return Type | Description |
| :--- | :--- | :--- | :--- |
| search(options) | { query, subreddit?, sort?, timeFilter?, limit?, after? } | Promise<Post[]> | Search posts on Reddit |
| getSubredditPosts(options) | { subreddit, sort?, timeFilter?, limit?, after? } | Promise<Post[]> | Fetch posts from subreddit feed |
| getSubredditAbout(subreddit) | subreddit: string | Promise<SubredditInfo> | Fetch subreddit rules & subscriber metrics |
| getRandomPost(subreddit?, sort?) | subreddit = 'memes', sort = 'hot' | Promise<Post> | Fetch 1 random post (ideal for bots) |
| getPost(postIdOrUrl, options?) | postIdOrUrl, { sort?, limit?, depth? } | Promise<Post> | Fetch post with complete comment tree |
| getUserProfile(username) | username: string | Promise<UserProfile> | Fetch user profile & karma stats |
| getUserPosts(username, options?) | username, { sort?, limit? } | Promise<Post[]> | Fetch posts submitted by user |
| getUserComments(username, options?) | username, { sort?, limit? } | Promise<Comment[]> | Fetch comments written by user |
| watchSubreddit(options) | { subreddit, intervalMs?, limit? } | SubredditWatcher | Listen for new posts as EventEmitter |
| formatForBot(post, options?) | post, { platform?, maxTextLength? } | string | Format markdown for WhatsApp/Discord/TG |
| extractMedia(post) | post | MediaAttachment | Get direct image/video/gallery links |
| export(data, options) | data, { format, filePath, title? } | string | Export dataset to JSON, CSV, or MD |
🛡️ Reddit Anti-Bot & OAuth2 Guide
Reddit frequently responds with HTTP 403 Forbidden to datacenter IPs, VPNs, or unauthenticated scrapers.
Option 1: Free Reddit Script App (Recommended & 100% Reliable)
- Go to https://www.reddit.com/prefs/apps
- Click "are you a developer? create an app..." at the bottom.
- Select "script".
- Set Name:
my_reddit_scraperand Redirect URI:http://localhost:8080. - Copy your Client ID (string under the app name) and Client Secret.
- Set in your
.envor system environment:
REDDIT_CLIENT_ID=your_client_id_here
REDDIT_CLIENT_SECRET=your_client_secret_here
REDDIT_USER_AGENT=nodejs:com.zetagoaurum.reddit-scraper:v1.1.0 (by /u/your_user)The scraper will automatically acquire OAuth2 Bearer tokens and rotate them transparently!
Option 2: Session Cookie
If you need to scrape private/quarantined subreddits or NSFW feeds without OAuth:
REDDIT_SESSION_COOKIE=your_reddit_session_cookie📊 Example Outputs
- Search Results (JSON)
- Subreddit Feed (CSV)
- Subreddit Markdown Table
- Post with Nested Comments (JSON)
- User Profile (JSON)
🐍 Python Implementation
A native Python SDK version is also included in the python/ directory:
cd python
pip install -r requirements.txt
pip install -e .
reddit-scraper-py search "Artificial Intelligence"⚖️ Watermark, Credits & License
====================================================================
REDDIT SCRAPER CORE & CLI ENGINE
====================================================================
Author : ZetaGo-Aurum
GitHub : https://github.com/ZetaGo-Aurum
Repository : https://github.com/ZetaGo-Aurum/reddit-scraper
License : MIT
[NOTICE & WATERMARK]
DO NOT REMOVE THIS WATERMARK OR AUTHOR CREDITS!
This software is created and maintained by ZetaGo-Aurum.
All rights reserved. Unauthorized removal of this header is prohibited.
====================================================================Released under the MIT License.
Copyright (c) 2024-2026 ZetaGo-Aurum.
