Design a distributed web crawler that fetches billions of web pages daily, extracts hyperlinks, respects web server politeness rules, and avoids duplicate content.
Bootstraps crawler with top domains and caches DNS lookups in RAM to eliminate DNS blocking.
Manages URL prioritization and host politeness queues using FIFO per domain.
Downloads web pages via async HTTP, enforces timeout limits and SSL checks.
Caches downloaded robots.txt files per domain to verify crawling permissions.
Parses DOM tree, extracts absolute hyperlinks, and normalizes URLs.
Checks if candidate URL or page content has already been processed.
Stores compressed raw HTML page content and document metadata.
Frontier selects host queue respecting 1-second politeness delay.
Verify user-agent is permitted to crawl path.
Download raw HTML payload.
If near-duplicate content exists, discard page.
Extract 50 new outgoing links on page.
Push unvisited URLs into Frontier queue for next crawl cycle.