A classical lexical (keyword-based) ranking function for full-text search — counts term occurrences with diminishing returns and document-length normalization. Still the strongest non-neural baseline for search.
BM25 (Best Match 25) scores documents based on query-term frequency, inverse document frequency (rare terms score higher), and document length (long documents are penalized). Unlike pure TF-IDF, BM25 has saturation — additional occurrences of a term contribute diminishing weight. Despite being from 1994, BM25 is still the default in Elasticsearch, Lucene, and most production search systems, often combined with semantic search in hybrid retrieval pipelines.
Using BM25 as the lexical half of a hybrid search system, combined with vector search via reciprocal rank fusion for both keyword precision and semantic recall.
BM25 remains the default text-search ranking algorithm because it's simple, fast, and surprisingly hard to beat — particularly for queries with rare, specific terms.
Need help implementing this in your business?
Get Started