retrieval evaluation / run_01
Chunking should be measured, not chosen by vibes.
8 strategies. 1 corpus. 1 labeled question set. The committed BM25 run scores retrieval quality beside the context tokens required to clear 90% recall.
Loading committed benchmark artifacts…
01 / leaderboard
BM25 Retrieval Scorecard
Click a column heading to sort. The highlighted row has the highest nDCG@10, with context cost as the tie-breaker.
n/a means recall never exceeded 90% at k = 1, 3, 5, or 10.
02 / curves
Quality and Context Cost
The leaderboard above is the text alternative for both plots.
03 / chunk inspector
See the Boundaries
Highlighted spans are returned context. Overlap bands show text covered by more than 1 chunk.
Select a strategy and document.
04 / failure explorer
Inspect a Miss
These examples missed the gold answer in the top five retrieved contexts.
Select a recorded failure.
Top retrieved chunk
05 / live BYOK
Run the Same Benchmark on Your Text
Client-side chunking, embedding retrieval, and the same metric formulas. The live semantic strategy uses adjacent-sentence embeddings instead of the committed TF-IDF proxy. No server is involved.
| Strategy | R@1 | R@3 | R@5 | R@10 | MRR@10 | nDCG@10 | Est. context tokens to >90% |
|---|
06 / method
Method
The committed run uses Okapi BM25 with k1 = 1.2 and b = 0.75. Tokens are lowercase alphanumerics split on every other character, with no stemming. A query is a hit when any returned context contains a case-insensitive, whitespace-normalized SQuAD gold answer string. Recall@k is the fraction of queries with at least 1 hit by rank k. MRR@10 averages the reciprocal rank of the first hit. nDCG@10 discounts later relevant chunks and divides by the ideal ranking.
Context cost sums the returned context tokens at the first measured k where aggregate recall exceeds 90%. Sentence-window retrieval indexes 1 sentence and returns 2 neighbors on each side. Parent-document retrieval indexes 250-character children and returns their 1,000-character parent blocks. Semantic (TF-IDF proxy) splits at the 80th percentile of adjacent-sentence TF-IDF cosine distance. The BYOK live mode uses embeddings for that signal.
07 / data
Data
SQuAD v1.1 dev paragraphs are grouped back into source articles. Seed 20260729 selects 40 articles and 800 labeled questions. The offline builder downloads the source file when it is absent and ranks indexed chunks locally with BM25. It uses tiktoken when installed, otherwise the metadata and every displayed token column identify the UTF-8 byte estimate.
08 / limitations
Limitations
- String containment does not recognize semantic equivalents and can over-credit common answer strings.
- 1 factoid QA corpus does not establish a universal winner.
- BM25 measures lexical retrieval. The optional live embedding run answers a different retrieval question and is not directly comparable.
- No reranker, query rewriting, hybrid retrieval, or generation step is included.
- Results do not transfer blindly to another corpus.
- Live context tokens use a visible byte-based estimate because no tokenizer library is shipped to the browser.