Your Vector Database Cannot Find An ISIN: Hybrid Retrieval, Reciprocal Rank Fusion And Why BM25 Beats Embeddings On Filings
Here is a result that surprises teams who have spent a year building vector search: on financial documents, BM25 has been found to outperform state-of-the-art dense retrieval. Not complement it - beat it. The reason is structural rather than incidental. Financial text is dense with identifiers, tickers, ISINs, CUSIPs, section references and exact figures, and embeddings are trained to smooth precisely those differences away. Vector search alone reaches around 78% recall@10 and misses exact identifiers; BM25 alone hits about 65% and cannot match meaning. A two-stage hybrid pipeline with neural reranking reaches Recall@5 of 0.816 on text-and-table financial documents. This is how to build one.
AlchmAI Engineering14 min read
0.816
Recall@5 achieved by a two-stage hybrid plus neural reranking pipeline on financial text-and-table documents
78%
Recall@10 from vector search alone - which still misses exact identifiers like product codes and error strings
65%
Recall@10 from BM25 alone, which cannot match semantic meaning but nails the identifiers embeddings lose
Ranks
What reciprocal rank fusion operates on - solving the score-incompatibility problem that breaks naive weighted averaging
A support ticket we have received, in some form, from four separate clients: the research assistant cannot find the document about GB00B03MLX29. The document exists. It is indexed. A colleague found it in three seconds using the search box on the old intranet. The expensive new retrieval-augmented system, with its embeddings and its vector database, returns four documents about unrelated gilt issuance and one about a different company entirely.
This is not a bug in anyone's implementation. It is the predictable consequence of using a tool built for semantic similarity on text whose most important tokens carry no semantics at all. The research now says so plainly: BM25 has been found to outperform state-of-the-art dense retrieval on financial documents, challenging the common assumption that semantic search universally dominates, because financial documents are full of identifiers, numbers and exact terminology that dense embeddings smooth over. Measured in isolation, vector search reaches around 78% recall@10 and misses exact identifiers; BM25 reaches about 65% and cannot match meaning. Neither is sufficient, and the combination with reranking reaches Recall@5 of 0.816 on mixed text-and-table financial documents.
The Pipeline: Two Retrievers, One Fusion, One Reranker
The architecture that works is not complicated, and its value is entirely in a few details that are easy to get wrong. Retrieve independently with a sparse and a dense retriever, fuse the two ranked lists, then rerank the fused candidates with a cross-encoder that reads the query and each candidate together.
from dataclasses import dataclass
@dataclass(frozen=True)
class Hit:
chunk_id: str
text: str
doc_id: str
known_from: str # point-in-time: when this became available to us
def reciprocal_rank_fusion(
runs: list[list[Hit]], k: int = 60, weights: list[float] | None = None
) -> list[Hit]:
"""Fuse ranked lists by RANK, not by score.
This is the detail that matters. BM25 scores are unbounded and corpus
dependent; cosine similarities sit in [-1, 1]. Normalising and averaging
them is the single most common production RAG bug - it looks principled
and it silently lets whichever retriever happens to have the wider score
range dominate the result. RRF sidesteps it by discarding magnitudes.
k dampens the contribution of top ranks; 60 is the widely used default
and is worth tuning on your own labelled set, not copying blindly.
"""
weights = weights or [1.0] * len(runs)
scores: dict[str, float] = {}
by_id: dict[str, Hit] = {}
for run, w in zip(runs, weights):
for rank, hit in enumerate(run, start=1):
scores[hit.chunk_id] = scores.get(hit.chunk_id, 0.0) + w / (k + rank)
by_id[hit.chunk_id] = hit
ranked = sorted(scores.items(), key=lambda kv: kv[1], reverse=True)
return [by_id[cid] for cid, _ in ranked]
async def retrieve(query: str, as_of: str, top_k: int = 12) -> list[Hit]:
# Both retrievers over-fetch. Fusion needs depth to work with, and the
# reranker is what buys back precision.
sparse, dense = await asyncio.gather(
bm25_search(query, as_of=as_of, limit=top_k * 6),
vector_search(embed(query), as_of=as_of, limit=top_k * 6),
)
# Weight sparse slightly higher on financial corpora. This is not a
# universal truth - it is what our labelled sets keep showing for
# filings, prospectuses and research notes.
fused = reciprocal_rank_fusion([sparse, dense], weights=[1.2, 1.0])
# Cross-encoder reranking: reads (query, candidate) jointly instead of
# comparing two independently-computed vectors. Far more accurate, far
# more expensive - which is why it runs on ~70 candidates, not 70,000.
return await rerank(query, fused[: top_k * 6])[:top_k]Chunking: The Step That Decides Your Ceiling
Retrieval quality is capped by chunking quality, and financial documents punish naive chunking harder than most corpora because so much of the meaning is structural. A fixed 512-token window slid across a 10-K will cut tables in half, separate figures from their units, and detach a number from the period it describes.
- 01Chunk on document structure, not on token count. Filings have sections, notes and numbered items. Split there and let chunk sizes vary; a note that runs to 1,800 tokens is one idea and should stay one chunk.
- 02Never split a table. Extract tables separately, serialise each with its caption and column headers intact, and index it as a unit. A half table retrieves as confidently as a whole one and the model will sum it without hesitation.
- 03Prepend context to every chunk before embedding. Document title, section path, entity, period. Contextual retrieval yields consistent gains where most query-expansion tricks do not, and this is the cheapest version of it.
- 04Index identifiers separately as well as inline. Extract every ISIN, ticker, LEI, CUSIP and section reference into a keyword field so an exact-match query can hit it directly rather than hoping BM25's tokeniser preserved it.
- 05Carry knowledge time on every chunk. Financial facts are bitemporal - a period they describe and a moment they became known - and a retriever that cannot filter on the second will answer a March question with a June restatement.
# Tokenisation decides whether BM25 can see an identifier at all.
# A standard analyser will happily destroy every one of these.
FINANCIAL_ANALYZER = {
"tokenizer": "standard",
"filter": ["lowercase"],
# Keep identifier-shaped tokens whole. Without this, GB00B03MLX29 may be
# split on digit/letter boundaries and the exact-match advantage - the
# entire reason BM25 is in the pipeline - disappears silently.
"char_filter": [{
"type": "pattern_replace",
"pattern": r"\b([A-Z]{2}[A-Z0-9]{9}\d)\b", # ISIN
"replacement": " __isin_$1 ",
}],
}
IDENTIFIER_PATTERNS = {
"isin": r"\b[A-Z]{2}[A-Z0-9]{9}\d\b",
"cusip": r"\b[0-9A-Z]{8}[0-9]\b",
"lei": r"\b[A-Z0-9]{18}[0-9]{2}\b",
"sedol": r"\b[B-DF-HJ-NP-TV-XZ0-9]{6}\d\b",
"ticker": r"\b[A-Z]{1,5}(?:\.[A-Z]{1,2})?\b",
}
def index_chunk(chunk: Chunk) -> dict:
ids = {k: sorted(set(re.findall(p, chunk.text)))
for k, p in IDENTIFIER_PATTERNS.items()}
return {
# Prepended context - this is what makes an embedding of a bare
# table row mean anything at all.
"text": f"{chunk.doc_title} | {chunk.section_path} | "
f"{chunk.entity} | {chunk.period}\n\n{chunk.text}",
"raw": chunk.text,
"identifiers": ids, # keyword fields, exact match, no analysis
"period_start": chunk.period_start,
"period_end": chunk.period_end,
"known_from": chunk.known_from, # bitemporal: as-at filtering
"known_until": chunk.known_until,
}Route The Query, Because Not Every Query Needs Both
A refinement worth adding once the basic pipeline works. Not every query benefits from both retrievers, and the research is explicit that BM25-only is sufficient when queries are exact-match dominated - product catalogues with SKUs, legal case numbers, financial instrument identifiers. Detecting that case is cheap and saves both latency and an embedding call.
def route(query: str) -> Literal["sparse", "dense", "hybrid"]:
"""Cheap, deterministic, and right most of the time. Resist the urge to
use a model here - a regex that runs in microseconds beats a classifier
that adds 200ms and its own failure mode."""
has_identifier = any(re.search(p, query) for p in IDENTIFIER_PATTERNS.values())
has_quoted = '"' in query
word_count = len(query.split())
# "GB00B03MLX29" or 'find "material adverse change" clause'
if (has_identifier or has_quoted) and word_count <= 6:
return "sparse"
# "what did management say about margin pressure in the Americas"
if not has_identifier and word_count > 12:
return "dense"
return "hybrid"What Does Not Help As Much As Advertised
A short list, because retrieval has accumulated a lot of folklore and financial corpora expose which parts of it are load-bearing.
- Query expansion techniques such as HyDE and multi-query provide limited benefit for precise numerical queries. They help on vague conceptual questions and add latency and noise on the exact-match queries that dominate financial workloads.
- Larger embedding models help less than better chunking. The gains from moving up a model tier are usually smaller than the gains from not splitting tables, and the second is free at inference time.
- Bigger top-k is not better recall. Past a point you are feeding the model more distractors, and the lost-in-the-middle effect means the extra context can reduce answer quality. Fix ranking, do not widen the funnel.
- Contextual retrieval is the exception - prepending document and section context to chunks before embedding yields consistent gains, and is one of the few techniques that reliably transfers between corpora.
Measuring It Properly
Retrieval is measured separately from generation, and conflating them is how teams end up tuning a prompt to compensate for a retriever that never returned the right document. Three measures, reported per query class:
- 01Recall@k on the labelled set, split by query class. A pipeline that is excellent on conceptual questions and poor on identifier lookups has a headline number that hides the thing your analysts will actually complain about.
- 02Mean reciprocal rank, to capture whether the right document is first rather than merely present. With reranking in the pipeline this is the measure that moves when the reranker is doing its job.
- 03End-to-end answer accuracy with retrieval held fixed, and again with retrieval varied. The difference tells you whether your next hour is better spent on the retriever or the prompt - a question most teams answer by instinct and get wrong about half the time.
The Bottom Line
The assumption that semantic search supersedes keyword search is wrong on financial documents, and the evidence is now clear enough to design around: BM25 outperforms state-of-the-art dense retrieval on this corpus type because filings are dense with identifiers and exact figures that embeddings are trained to smooth away, while vector search at roughly 78% recall@10 still cannot find an ISIN. The answer is a two-stage pipeline - both retrievers over-fetching independently, fused on ranks with reciprocal rank fusion rather than on incompatible scores, then reranked with a cross-encoder - which reaches Recall@5 of 0.816 on text-and-table financial documents. Chunk on structure and never split a table, prepend context before embedding, index identifiers as exact-match keyword fields, carry knowledge time so a March question never gets a June restatement, route the obvious exact-match queries straight to sparse, and build two hundred labelled pairs before you tune anything. None of that is exotic. All of it is the difference between a research assistant your analysts trust and one they quietly stop using. That is the retrieval engineering we do for financial clients in London.
References & Further Reading
- From BM25 to Corrective RAG: benchmarking retrieval strategies for text-and-table documents (arXiv). arxiv.org/html/2604.01733v1
- Cormack, Clarke & Buettcher - Reciprocal rank fusion outperforms Condorcet and individual rank learning methods. plg.uwaterloo.ca/~gvcormac/cormacksigir09-rrf.pdf
- Denser - Hybrid search for RAG: combining BM25 and dense vector search (2026 guide). denser.ai/blog/hybrid-search-for-rag
- Prem AI - Hybrid search for RAG: BM25, SPLADE and vector search combined. premai.io/blog/hybrid-search-for-rag-bm25-splade-and-vector-search-combined
- Digital Applied - Hybrid search: BM25, vector and reranking reference 2026. digitalapplied.com/blog/hybrid-search-bm25-vector-reranking-reference-2026
- Anthropic - Introducing contextual retrieval. anthropic.com/news/contextual-retrieval
- Robertson & Zaragoza - The probabilistic relevance framework: BM25 and beyond. staff.city.ac.uk/~sbrp622/papers/foundations_bm25_review.pdf
AlchmAI Engineering
Engineering, London
Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.
Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information