Domains Indexed
URLs Crawled
Data Archived
Datasets Curated
What's In Every Record
Each of the 50B+ URLs in the index carries the full story of its crawl.
Requested URL, final URL after redirects, and canonical link — deduplicated across the index.
Crawl timestamp, HTTP status, response headers, and content type for every fetch.
The page as our crawler received it, so you can run your own parsers and extractors.
Boilerplate-stripped body text with detected language — ready for search, analytics, and model training.
Title, meta description, headings, and outbound links for graph and SEO analysis.
Delivered Your Way
Take the index in bulk, query it on demand, or both — the data fits your architecture, not the other way around.
S3 Bucket Sync
Partitioned Parquet or WARC files synced to your bucket on every refresh. Pull only the domains, languages, or date ranges you need.
REST API
Targeted lookups by URL, domain, or query — fetch the latest crawl of any page in the index without moving bulk data.
Direct Transfer
For petabyte-scale needs, we arrange direct transfer into your storage — one-time snapshots or standing replication.
Full-index access is currently onboarding a limited number of customers. Tell us about your use case and we'll prepare a sample slice of the index for evaluation. Request Access
Need a Slice of the Web?
Every data project is different. Reserve 15 minutes to describe yours, and we'll recommend the right datasets, tell you what the full index would add, and prepare a custom proposal — free, with no obligation.