No crawler of our own. A topic slice of Common Crawl, plus the feeds and APIs of sites that block crawlers entirely.
The per-crawl CDX index is binary-searched, then only the needed WARC ranges are fetched — politely, five requests a second, days not minutes.
A deterministic BM25 build: same pages in, byte-identical index out. Ranking adds source weight, page shape and, for jobs, freshness.
Title, a short attributed snippet, and a link out. Never a cached page. Removal requests are honoured on the next build.