Architecture of Web Search
Diffbot Web Search is three stages: a web index built from Diffbot's global index, a multi-strategy retrieval pass, and a two-stage reranker.
┌─────────────────────┐ ┌───────┐ │ Global Index ~150TB │ │ Query ├────────────┐ └──────────┬──────────┘ └───────┘ │ └────────┐ │ ▼ │ ┌────────────────┐ │ │ Web Index ~8TB │ │ └────────┬───────┘ │ │ 1,000,000,000 candidates │ ▼ │ ┌──────────────────────┐ │ │ CatBoost fusion of 5 │ │ │ retrieval strategies │◄──────────────┘ └───────────┬──────────┘ │ 400 candidates ▼ ┌─────────────────────────┐ │ Crossencoder + CatBoost │ │ rerank │ └────────────┬────────────┘ │ 10 candidates ▼ ┌─────────┐ │ Results │ └─────────┘
Index
Diffbot maintains a 150TB Global Index of the public web that feeds the Knowledge Graph and a filtered 8TB snapshot index containing about 1 billion documents across all domains.
- Ingestion. Documents sourced from the Diffbot Global Index (~150TB).
- Filtering. Documents are filtered and sorted into buckets by freshness, content type (e.g. Wikipedia, documentation, articles), and quality.
- Sharding. Buckets are built into temporal and evergreen shards, so the index can be rebuilt in parallel and incrementally.
- Enrichment. Each document gets PageRank, deduplication, text embeddings, inlinks, domain age, site metadata, and keyword embeddings.
Retrieval
Five retrieval strategies run in parallel. Their results are unioned and deduplicated into ~1700 candidates.
| Strategy | Matches on |
|---|---|
| BM25 | Query terms, with top-k term selection, rare term protection, and positional phrase/proximity boosts |
| Document embedding | Dense vector similarity to the document's text embedding |
| Keyword embedding | Dense vector similarity to the document's keyword embedding |
| Inlinks | Anchor text of links pointing to the candidate page |
| URL words | Query words that appear in the candidate's URL |
A CatBoost fusion model then ranks the candidates using the five strategy scores, PageRank, and 41 term match features. The list is truncated to the top 400.
Reranking
Reranking runs in two stages.
- Document rerank. A crossencoder model scores each query–document pair. A CatBoost ensemble then rescores the results with 37 additional features, including term proximity, title and meta match, PageRank, site inlinks, URL structure, and headings and length.
- Highlights. Each returned document is sliced into meaningful phrases, and a chunk-level crossencoder ranks them to form the result's highlights.
Interface
Diffbot Web Search is accessed natively via REST API. An MCP server and Agent Skills are available for agentic applications.
curl --request GET \
--url 'https://llm.diffbot.com/api/v1/web_search?text=diffbot' \
--header 'Accept: application/json' \
--header 'Authorization: Bearer <DIFFBOT_TOKEN>'