retrieval evaluation / run_01

Chunking should be measured, not chosen by vibes.

8 strategies. 1 corpus. 1 labeled question set. The committed BM25 run scores retrieval quality beside the context tokens required to clear 90% recall.

corpus
loading…
sample
loading…
retriever
loading…
seed
loading…

Loading committed benchmark artifacts…

01 / leaderboard

BM25 Retrieval Scorecard

Click a column heading to sort. The highlighted row has the highest nDCG@10, with context cost as the tie-breaker.

n/a means recall never exceeded 90% at k = 1, 3, 5, or 10.

02 / curves

Quality and Context Cost

The leaderboard above is the text alternative for both plots.

Recall@k
nDCG@10 vs context tokens to >90% recall

03 / chunk inspector

See the Boundaries

Highlighted spans are returned context. Overlap bands show text covered by more than 1 chunk.

Select a strategy and document.

04 / failure explorer

Inspect a Miss

These examples missed the gold answer in the top five retrieved contexts.

Select a recorded failure.

Gold span


        

Top retrieved chunk


        

05 / live BYOK

Run the Same Benchmark on Your Text

Client-side chunking, embedding retrieval, and the same metric formulas. The live semantic strategy uses adjacent-sentence embeddings instead of the committed TF-IDF proxy. No server is involved.

Your key stays in sessionStorage for this tab, is sent only to api.openai.com over HTTPS, and is never logged. Use Forget key to clear it.

1 pair per line, separated by a tab. Add 2 to 20 pairs. Answers may use || for accepted variants.

06 / method

Method

The committed run uses Okapi BM25 with k1 = 1.2 and b = 0.75. Tokens are lowercase alphanumerics split on every other character, with no stemming. A query is a hit when any returned context contains a case-insensitive, whitespace-normalized SQuAD gold answer string. Recall@k is the fraction of queries with at least 1 hit by rank k. MRR@10 averages the reciprocal rank of the first hit. nDCG@10 discounts later relevant chunks and divides by the ideal ranking.

Context cost sums the returned context tokens at the first measured k where aggregate recall exceeds 90%. Sentence-window retrieval indexes 1 sentence and returns 2 neighbors on each side. Parent-document retrieval indexes 250-character children and returns their 1,000-character parent blocks. Semantic (TF-IDF proxy) splits at the 80th percentile of adjacent-sentence TF-IDF cosine distance. The BYOK live mode uses embeddings for that signal.

07 / data

Data

SQuAD v1.1 dev paragraphs are grouped back into source articles. Seed 20260729 selects 40 articles and 800 labeled questions. The offline builder downloads the source file when it is absent and ranks indexed chunks locally with BM25. It uses tiktoken when installed, otherwise the metadata and every displayed token column identify the UTF-8 byte estimate.

08 / limitations

Limitations

  • String containment does not recognize semantic equivalents and can over-credit common answer strings.
  • 1 factoid QA corpus does not establish a universal winner.
  • BM25 measures lexical retrieval. The optional live embedding run answers a different retrieval question and is not directly comparable.
  • No reranker, query rewriting, hybrid retrieval, or generation step is included.
  • Results do not transfer blindly to another corpus.
  • Live context tokens use a visible byte-based estimate because no tokenizer library is shipped to the browser.