⚡ CXM25: better than BM25, still blazing fast on CPU

CXM25 is a BM25-inspired lexical retriever with no embeddings, no GPU — it adds character-gram, phrase and title signals on top of BM25 and measurably outranks it. CXM25-LTR is the learned refinement (same features, a gradient-boosted ranker trained on 12M triples).

Links: GitHub · PyPI

Enter a query and one document per line, then compare how the three models rank them and how long each takes. Pairwise accuracy on the reference benchmark: BM25 67.3% · CXM25 70.5% · CXM25-LTR 77.8%.

Note on CXM25-LTR: as a learned reranker it optimises pairwise accuracy, not literal term matching. On very short queries (a single word, e.g. verão) it can surface thematically-related documents that don't contain the word itself — that is the model working as trained, and it is still the most accurate model overall.

1 10

BM25 · 67.3%

BM25 · 67.3%
Rank
Score
Document

CXM25 · 70.5%

CXM25 · 70.5%
Rank
Score
Document

CXM25-LTR · 77.8%

CXM25-LTR · 77.8%
Rank
Score
Document
Examples
Query Documents (one per line)