⚡ CXM25: better than BM25, still blazing fast on CPU
CXM25 is a BM25-inspired lexical retriever with no embeddings, no GPU — it adds character-gram, phrase and title signals on top of BM25 and measurably outranks it. CXM25-LTR is the learned refinement (same features, a gradient-boosted ranker trained on 12M triples).
Enter a query and one document per line, then compare how the three models rank them and how long each takes. Pairwise accuracy on the reference benchmark: BM25 67.3% · CXM25 70.5% · CXM25-LTR 77.8%.
Note on CXM25-LTR: as a learned reranker it optimises pairwise accuracy, not literal term matching. On very short queries (a single word, e.g.
verão) it can surface thematically-related documents that don't contain the word itself — that is the model working as trained, and it is still the most accurate model overall.
BM25 · 67.3%
Rank | Score | Document |
|---|---|---|
CXM25 · 70.5%
Rank | Score | Document |
|---|---|---|
Rank | Score | Document |
|---|---|---|
CXM25-LTR · 77.8%
Rank | Score | Document |
|---|---|---|
Rank | Score | Document |
|---|---|---|
| Query | Documents (one per line) |
|---|