Skip to content

Architecture of Web Search ​

Diffbot Web Search is three stages: a web index built from Diffbot's global index, a multi-strategy retrieval pass, and a two-stage reranker.

 ┌─────────────────────┐   ┌───────┐
 │ Global Index ~150TB │   │ Query ├────────────┐
 └──────────┬──────────┘   └───────┘            │
            └────────┐                          │
                     ▼                          │
            ┌────────────────┐                  │
            │ Web Index ~8TB │                  │
            └────────┬───────┘                  │
                     │ 1,000,000,000 candidates │
                     ▼                          │
         ┌──────────────────────┐               │
         │ CatBoost fusion of 5 │               │
         │ retrieval strategies │◄──────────────┘
         └───────────┬──────────┘
                     │ 400 candidates
                     ▼              
        ┌─────────────────────────┐
        │ Crossencoder + CatBoost │
        │         rerank          │
        └────────────┬────────────┘
                     │ 10 candidates
                     ▼             
                ┌─────────┐
                │ Results │
                └─────────┘

Index ​

Diffbot maintains a 150TB Global Index of the public web that feeds the Knowledge Graph and a filtered 8TB snapshot index containing about 1 billion documents across all domains.

  1. Ingestion. Documents sourced from the Diffbot Global Index (~150TB).
  2. Filtering. Documents are filtered and sorted into buckets by freshness, content type (e.g. Wikipedia, documentation, articles), and quality.
  3. Sharding. Buckets are built into temporal and evergreen shards, so the index can be rebuilt in parallel and incrementally.
  4. Enrichment. Each document gets PageRank, deduplication, text embeddings, inlinks, domain age, site metadata, and keyword embeddings.

Retrieval ​

Five retrieval strategies run in parallel. Their results are unioned and deduplicated into ~1700 candidates.

StrategyMatches on
BM25Query terms, with top-k term selection, rare term protection, and positional phrase/proximity boosts
Document embeddingDense vector similarity to the document's text embedding
Keyword embeddingDense vector similarity to the document's keyword embedding
InlinksAnchor text of links pointing to the candidate page
URL wordsQuery words that appear in the candidate's URL

A CatBoost fusion model then ranks the candidates using the five strategy scores, PageRank, and 41 term match features. The list is truncated to the top 400.

Reranking ​

Reranking runs in two stages.

  1. Document rerank. A crossencoder model scores each query–document pair. A CatBoost ensemble then rescores the results with 37 additional features, including term proximity, title and meta match, PageRank, site inlinks, URL structure, and headings and length.
  2. Highlights. Each returned document is sliced into meaningful phrases, and a chunk-level crossencoder ranks them to form the result's highlights.

Interface ​

Diffbot Web Search is accessed natively via REST API. An MCP server and Agent Skills are available for agentic applications.

curl --request GET \
  --url 'https://llm.diffbot.com/api/v1/web_search?text=diffbot' \
  --header 'Accept: application/json' \
  --header 'Authorization: Bearer <DIFFBOT_TOKEN>'