MediaSearch

Archive · Jul 2026

A film-news search engine built from scratch.

Seeded crawler, polite delays, SimHash dedupe, SQLite FTS5 index, and a tiny Flask front end — then the reels stopped.

Final tally

What the crawler pulled

Live from June 7 to July 21, 2026 across breaking, daily, and weekly seed tiers.

Architecture

Seeds → pages → index → search

No SaaS search API. Just a BFS crawler that respected robots.txt, fingerprinted near-duplicates, and rebuilt an inverted index after every run.

Seed lists

Breaking, daily, and weekly URL lists for trade press, genre sites, and niche outlets.

Crawl

Per-domain delay, depth caps, URL normalization, and SimHash near-dupe skips.

Index

HTML stripped to text, upserted into SQLite with FTS5 Porter stemming.

Serve

Flask + Gunicorn at mediasearch.online with BM25 ranking and snippets.

Schedule

Three crawl tiers

High-churn trades every two hours; deeper genre and niche runs on slower cadences. All cron jobs disabled on shutdown day.

Coverage

Where the index came from

Document counts by domain in the final search.db snapshot.

Sample index

Headlines from the last crawl

A slice of titles sitting in the FTS index when the crawler was killed — links go to the original publishers.

Build

Stack & leftovers

Backups of search.db, crawl logs, seeds, and gzipped page HTML live under backups/. This page is the public face of that archive.