- 36questions with labelled answer files
- 0.8195% interval 0.67 to 0.92hit rate@5, shipped configuration
- 0.6595% interval 0.52 to 0.78recall@5, shipped configuration
- 0.57795% interval 0.444 to 0.702mean reciprocal rank
- 18configurations scored on the same questions
What is measured
The corpus is four repositories of mine, checked out at fixed commits and listed file by file, with line counts and hashes, in the results directory. The evaluator refuses to run on a corpus that differs from that list, so the numbers here can be regenerated by anyone.
| Project | Repository | Commit | Files | Lines |
|---|---|---|---|---|
difflens | DiffLens-Engine | 3ca9978e05b9 | 60 | 6,640 |
gesturecontrol | GestureControl-Engine | 53c4b04bf767 | 13 | 1,316 |
portfolio | portfolio | 7c361eea5f1d | 10 | 1,458 |
taskforge | TaskForge-Engine | 8c02b8a514bb | 33 | 2,405 |
| total | 116 | 11,819 |
The question set is 36 questions I wrote against these projects, phrased the way someone new to a repository would ask rather than by pasting identifiers, each with the files that answer it. Relevance is judged at file level: a retrieved chunk counts when its source file is one of the labelled answers. I wrote both the questions and the code, and there has been one labeller so far; the labelling protocol for a second round with two independent labellers is in the repository.
The questions are split 50/50 into a dev half and a test half. The only setting chosen on the data, the per-file cap value, is chosen on the dev half; every table reports all 36 questions with a bootstrap interval, because eighteen questions per half say very little on their own.
Metrics
- hit rate@k: 1 when any labelled file appears in the top k chunks.
- recall@k: the share of the labelled files that appear in the top k chunks. For a question with four answer files, one of them in the top 5 is a hit rate of 1 and a recall of 0.25.
- MRR: one over the rank of the first relevant chunk, averaged over questions.
- nDCG@10: discounted gain with one binary gain per relevant file at its first chunk, normalised by the ideal ordering of the labelled set.
Every cell carries a 95% percentile bootstrap interval over questions (10,000 resamples). Every claim that one configuration beats another rests on a paired bootstrap of the per-question differences and an exact sign test on the questions where the two differ. When the interval includes zero the page says so.
Results
The shipped configuration is merged structural chunking, hybrid retrieval, and a per-file cap of 1.
| Configuration | hit rate@1 | hit rate@3 | hit rate@5 | recall@5 | MRR | nDCG@10 |
|---|---|---|---|---|---|---|
| structural, merged / hybrid + per-file cap | 0.42 [0.25, 0.58] | 0.69 [0.56, 0.83] | 0.81 [0.67, 0.92] | 0.65 [0.52, 0.78] | 0.577 [0.444, 0.702] | 0.573 [0.469, 0.674] |
Every configuration
- BM25structural, merged
- densestructural, merged
- hybrid (RRF)structural, merged
- hybrid + per-file capstructural, merged
- hybrid + reranker (text only)structural, merged
- hybrid + reranker (prefixed)structural, merged
- BM25structural, unmerged
- densestructural, unmerged
- hybrid (RRF)structural, unmerged
- hybrid + per-file capstructural, unmerged
- hybrid + reranker (text only)structural, unmerged
- hybrid + reranker (prefixed)structural, unmerged
- BM25window
- densewindow
- hybrid (RRF)window
- hybrid + per-file capwindow
- hybrid + reranker (text only)window
- hybrid + reranker (prefixed)window
| Chunking | Strategy | Chunks | hit rate@1 | hit rate@5 | recall@5 | MRR | nDCG@10 |
|---|---|---|---|---|---|---|---|
| structural, merged | BM25 | 942 | 0.14 [0.03, 0.25] | 0.67 [0.50, 0.81] | 0.53 [0.38, 0.66] | 0.342 [0.245, 0.448] | 0.389 [0.297, 0.486] |
| structural, merged | dense | 942 | 0.39 [0.22, 0.56] | 0.83 [0.69, 0.94] | 0.64 [0.52, 0.75] | 0.532 [0.406, 0.661] | 0.532 [0.421, 0.640] |
| structural, merged | hybrid (RRF) | 942 | 0.42 [0.25, 0.58] | 0.72 [0.58, 0.86] | 0.57 [0.44, 0.71] | 0.547 [0.415, 0.679] | 0.522 [0.414, 0.628] |
| structural, merged | hybrid + per-file cap | 942 | 0.42 [0.25, 0.58] | 0.81 [0.67, 0.92] | 0.65 [0.52, 0.78] | 0.577 [0.444, 0.702] | 0.573 [0.469, 0.674] |
| structural, merged | hybrid + reranker (text only) | 942 | 0.28 [0.14, 0.42] | 0.72 [0.56, 0.86] | 0.55 [0.42, 0.68] | 0.423 [0.302, 0.550] | 0.433 [0.332, 0.536] |
| structural, merged | hybrid + reranker (prefixed) | 942 | 0.25 [0.11, 0.39] | 0.78 [0.64, 0.92] | 0.59 [0.46, 0.72] | 0.455 [0.342, 0.572] | 0.478 [0.383, 0.572] |
| structural, unmerged | BM25 | 1281 | 0.25 [0.11, 0.39] | 0.56 [0.39, 0.72] | 0.45 [0.31, 0.59] | 0.382 [0.264, 0.510] | 0.390 [0.289, 0.492] |
| structural, unmerged | dense | 1281 | 0.31 [0.17, 0.45] | 0.69 [0.53, 0.83] | 0.57 [0.43, 0.71] | 0.474 [0.350, 0.602] | 0.500 [0.393, 0.608] |
| structural, unmerged | hybrid (RRF) | 1281 | 0.39 [0.22, 0.56] | 0.64 [0.47, 0.81] | 0.50 [0.36, 0.63] | 0.521 [0.390, 0.654] | 0.498 [0.396, 0.601] |
| structural, unmerged | hybrid + per-file cap | 1281 | 0.39 [0.22, 0.56] | 0.75 [0.61, 0.89] | 0.60 [0.47, 0.73] | 0.568 [0.447, 0.692] | 0.589 [0.490, 0.682] |
| structural, unmerged | hybrid + reranker (text only) | 1281 | 0.31 [0.17, 0.47] | 0.67 [0.50, 0.81] | 0.50 [0.37, 0.64] | 0.436 [0.313, 0.563] | 0.443 [0.339, 0.548] |
| structural, unmerged | hybrid + reranker (prefixed) | 1281 | 0.28 [0.14, 0.42] | 0.78 [0.64, 0.92] | 0.57 [0.44, 0.70] | 0.461 [0.347, 0.582] | 0.482 [0.385, 0.581] |
| window | BM25 | 919 | 0.22 [0.08, 0.36] | 0.64 [0.47, 0.78] | 0.50 [0.37, 0.64] | 0.385 [0.266, 0.506] | 0.386 [0.284, 0.491] |
| window | dense | 919 | 0.36 [0.22, 0.53] | 0.75 [0.61, 0.89] | 0.59 [0.46, 0.72] | 0.514 [0.387, 0.643] | 0.508 [0.398, 0.620] |
| window | hybrid (RRF) | 919 | 0.39 [0.25, 0.56] | 0.69 [0.53, 0.83] | 0.50 [0.37, 0.62] | 0.515 [0.382, 0.651] | 0.487 [0.377, 0.594] |
| window | hybrid + per-file cap | 919 | 0.39 [0.22, 0.56] | 0.81 [0.67, 0.92] | 0.65 [0.52, 0.78] | 0.567 [0.446, 0.693] | 0.573 [0.470, 0.672] |
| window | hybrid + reranker (text only) | 919 | 0.31 [0.17, 0.44] | 0.67 [0.50, 0.81] | 0.52 [0.39, 0.66] | 0.449 [0.323, 0.578] | 0.427 [0.321, 0.534] |
| window | hybrid + reranker (prefixed) | 919 | 0.28 [0.14, 0.44] | 0.67 [0.50, 0.81] | 0.52 [0.38, 0.66] | 0.442 [0.323, 0.565] | 0.453 [0.359, 0.549] |
What beat what, and what was within noise
Structural boundaries against windows
Merged structural chunking with hybrid retrieval scores hit rate@1 0.42 against 0.39 for windows (difference +0.03, 95% interval -0.06 to +0.11; better on 2 questions, worse on 1, same on 33). The interval includes zero, so this is within noise at this sample size. The same chunking scores MRR 0.547 against 0.515 for windows (difference +0.032, 95% interval -0.027 to +0.089; better on 11 questions, worse on 4, same on 21). The interval includes zero, so this is within noise at this sample size. Both chunkers are packed to the same token budget, so the comparison is about where boundaries fall, not about chunk size.
Hybrid against dense
Hybrid retrieval scores MRR 0.547 against 0.532 for dense retrieval alone (difference +0.015, 95% interval -0.126 to +0.160; better on 14 questions, worse on 11, same on 11). The interval includes zero, so this is within noise at this sample size. Hybrid retrieval scores hit rate@5 0.72 against 0.83 for dense retrieval (difference -0.11, 95% interval -0.25 to +0.00; better on 1 question, worse on 5, same on 30). The interval includes zero, so this is within noise at this sample size. Fusing in the BM25 ranking pushes some files that dense retrieval had in the top 5 below the cut.
Hybrid against BM25
Hybrid retrieval scores MRR 0.547 against 0.342 for BM25 alone (difference +0.205, 95% interval +0.095 to +0.318; better on 20 questions, worse on 3, same on 13). The interval excludes zero.
The per-file cap
With the cap, hybrid retrieval scores recall@5 0.65 against 0.57 for the same ranking without it (difference +0.08, 95% interval +0.01 to +0.17; better on 4 questions, worse on 0, same on 32). The interval excludes zero. The cap scores MRR 0.577 against 0.547 for no cap (difference +0.029, 95% interval +0.011 to +0.052; better on 9 questions, worse on 0, same on 27). The interval excludes zero. Under file-level relevance the cap cannot lower hit rate or MRR, because the first chunk of every file is always kept and only moves up; the metric it can lower is recall, and it did not.
| Per-file cap | dev MRR | dev recall@5 | test MRR | test recall@5 | all MRR | all hit rate@5 | all recall@5 |
|---|---|---|---|---|---|---|---|
| no cap | 0.604 | 0.48 | 0.491 | 0.67 | 0.547 | 0.72 | 0.57 |
| cap 3 | 0.604 | 0.48 | 0.494 | 0.67 | 0.549 | 0.72 | 0.57 |
| cap 2 | 0.607 | 0.51 | 0.498 | 0.67 | 0.553 | 0.75 | 0.59 |
| cap 1 (chosen) | 0.614 | 0.53 | 0.539 | 0.78 | 0.577 | 0.81 | 0.65 |
The shipped configuration against the next best cell
The shipped configuration scores MRR 0.577 against 0.568 for structural, unmerged / hybrid + per-file cap (difference +0.009, 95% interval -0.039 to +0.059; better on 7 questions, worse on 4, same on 25). The interval includes zero, so this is within noise at this sample size.
All comparisons
| Comparison (A vs B) | Metric | A | B | A - B [95% CI] | better / worse / same | sign test p | Reading |
|---|---|---|---|---|---|---|---|
| structural boundaries vs windows, hybrid | hit rate@1 | 0.42 | 0.39 | +0.03 [-0.06, +0.11] | 2 / 1 / 33 | 1.000 | within noise |
| structural boundaries vs windows, hybrid | MRR | 0.547 | 0.515 | +0.032 [-0.027, +0.089] | 11 / 4 / 21 | 0.118 | within noise |
| structural boundaries vs windows, hybrid | nDCG@10 | 0.522 | 0.487 | +0.035 [-0.004, +0.077] | 12 / 6 / 18 | 0.238 | within noise |
| structural boundaries vs windows, dense | hit rate@1 | 0.39 | 0.36 | +0.03 [-0.08, +0.14] | 3 / 2 / 31 | 1.000 | within noise |
| structural boundaries vs windows, dense | MRR | 0.532 | 0.514 | +0.019 [-0.069, +0.101] | 10 / 4 / 22 | 0.180 | within noise |
| merged vs unmerged structural, hybrid | hit rate@1 | 0.42 | 0.39 | +0.03 [-0.06, +0.11] | 2 / 1 / 33 | 1.000 | within noise |
| merged vs unmerged structural, hybrid | MRR | 0.547 | 0.521 | +0.026 [-0.027, +0.085] | 10 / 3 / 23 | 0.092 | within noise |
| hybrid vs dense, merged | hit rate@1 | 0.42 | 0.39 | +0.03 [-0.17, +0.22] | 7 / 6 / 23 | 1.000 | within noise |
| hybrid vs dense, merged | hit rate@5 | 0.72 | 0.83 | -0.11 [-0.25, +0.00] | 1 / 5 / 30 | 0.219 | within noise |
| hybrid vs dense, merged | MRR | 0.547 | 0.532 | +0.015 [-0.126, +0.160] | 14 / 11 / 11 | 0.690 | within noise |
| hybrid vs dense, merged | nDCG@10 | 0.522 | 0.532 | -0.010 [-0.111, +0.087] | 14 / 12 / 10 | 0.845 | within noise |
| hybrid vs dense, window | hit rate@1 | 0.39 | 0.36 | +0.03 [-0.11, +0.17] | 4 / 3 / 29 | 1.000 | within noise |
| hybrid vs dense, window | MRR | 0.515 | 0.514 | +0.001 [-0.109, +0.111] | 12 / 10 / 14 | 0.832 | within noise |
| hybrid vs bm25, merged | hit rate@5 | 0.72 | 0.67 | +0.06 [-0.08, +0.19] | 4 / 2 / 30 | 0.688 | within noise |
| hybrid vs bm25, merged | MRR | 0.547 | 0.342 | +0.205 [+0.095, +0.318] | 20 / 3 / 13 | 0.001 | interval excludes zero |
| per-file cap vs none, merged hybrid | hit rate@3 | 0.69 | 0.67 | +0.03 [+0.00, +0.08] | 1 / 0 / 35 | 1.000 | within noise |
| per-file cap vs none, merged hybrid | hit rate@5 | 0.81 | 0.72 | +0.08 [+0.00, +0.19] | 3 / 0 / 33 | 0.250 | within noise |
| per-file cap vs none, merged hybrid | recall@5 | 0.65 | 0.57 | +0.08 [+0.01, +0.17] | 4 / 0 / 32 | 0.125 | interval excludes zero |
| per-file cap vs none, merged hybrid | MRR | 0.577 | 0.547 | +0.029 [+0.011, +0.052] | 9 / 0 / 27 | 0.004 | interval excludes zero |
| per-file cap vs none, merged hybrid | nDCG@10 | 0.573 | 0.522 | +0.051 [+0.027, +0.078] | 16 / 0 / 20 | 0.000 | interval excludes zero |
| reranker (text only) vs hybrid, merged | hit rate@1 | 0.28 | 0.42 | -0.14 [-0.25, -0.03] | 0 / 5 / 31 | 0.062 | interval excludes zero |
| reranker (text only) vs hybrid, merged | MRR | 0.423 | 0.547 | -0.124 [-0.220, -0.042] | 5 / 12 / 19 | 0.143 | interval excludes zero |
| reranker (prefixed) vs hybrid, merged | hit rate@1 | 0.25 | 0.42 | -0.17 [-0.31, -0.03] | 1 / 7 / 28 | 0.070 | interval excludes zero |
| reranker (prefixed) vs hybrid, merged | MRR | 0.455 | 0.547 | -0.092 [-0.194, +0.008] | 10 / 13 / 13 | 0.678 | within noise |
| reranker prefixed vs text only, merged | hit rate@1 | 0.25 | 0.28 | -0.03 [-0.14, +0.06] | 1 / 2 / 33 | 1.000 | within noise |
| reranker prefixed vs text only, merged | MRR | 0.455 | 0.423 | +0.032 [-0.024, +0.082] | 16 / 3 / 17 | 0.004 | within noise |
| reranker (text only) vs hybrid, window | hit rate@1 | 0.31 | 0.39 | -0.08 [-0.25, +0.08] | 3 / 6 / 27 | 0.508 | within noise |
| reranker (text only) vs hybrid, window | MRR | 0.449 | 0.515 | -0.066 [-0.183, +0.045] | 10 / 10 / 16 | 1.000 | within noise |
| reranker (prefixed) vs hybrid, window | hit rate@1 | 0.28 | 0.39 | -0.11 [-0.28, +0.06] | 3 / 7 / 26 | 0.344 | within noise |
| reranker (prefixed) vs hybrid, window | MRR | 0.442 | 0.515 | -0.074 [-0.192, +0.039] | 11 / 11 / 14 | 1.000 | within noise |
| reranker prefixed vs text only, window | hit rate@1 | 0.28 | 0.31 | -0.03 [-0.08, +0.00] | 0 / 1 / 35 | 1.000 | within noise |
| reranker prefixed vs text only, window | MRR | 0.442 | 0.449 | -0.007 [-0.046, +0.023] | 5 / 4 / 27 | 1.000 | within noise |
| shipped configuration vs best other cell by MRR | hit rate@1 | 0.42 | 0.39 | +0.03 [-0.06, +0.11] | 2 / 1 / 33 | 1.000 | within noise |
| shipped configuration vs best other cell by MRR | hit rate@5 | 0.81 | 0.75 | +0.06 [-0.06, +0.17] | 3 / 1 / 32 | 0.625 | within noise |
| shipped configuration vs best other cell by MRR | MRR | 0.577 | 0.568 | +0.009 [-0.039, +0.059] | 7 / 4 / 25 | 0.549 | within noise |
The cross-encoder reranker
Reranking the top 30 with cross-encoder/ms-marco-MiniLM-L-6-v2 was run two ways: on the chunk text alone, and on the same path :: name payload the embedder sees, so that the reranker is not judged on less information than the first stage had. Every chunk fits the reranker's window, so nothing is truncated in either variant.
Reranking the chunk text alone scores MRR 0.423 against 0.547 for hybrid retrieval without it (difference -0.124, 95% interval -0.220 to -0.042; better on 5 questions, worse on 12, same on 19). The interval excludes zero. Reranking the same path-and-name payload the embedder sees scores MRR 0.455 against 0.547 for hybrid retrieval (difference -0.092, 95% interval -0.194 to +0.008; better on 10 questions, worse on 13, same on 13). The interval includes zero, so this is within noise at this sample size.
The prefixed variant scores MRR 0.455 against 0.423 for the text-only variant (difference +0.032, 95% interval -0.024 to +0.082; better on 16 questions, worse on 3, same on 17). The interval includes zero, so this is within noise at this sample size. On window chunks, the text-only reranker scores MRR 0.449 against 0.515 for hybrid retrieval (difference -0.066, 95% interval -0.183 to +0.045; better on 10 questions, worse on 10, same on 16). The interval includes zero, so this is within noise at this sample size.
What the data supports: in this setup the reranker does not improve ranking, and on merged structural chunks the text-only variant is worse by more than the noise. It does not support a cause. Whether the model's training domain, the small candidate set, or the file-level judgement explains it is not something this experiment tested.
Where it fails
The shipped configuration misses 7 of 36 questions at rank 5: tf03, tf08, tf09, tf12, dl06, dl13, gc05. Two things were measured about what fills the top 5 on those questions.
| Share | On misses | On hits | Difference [95% CI] |
|---|---|---|---|
| share of the top 5 that is documentation | 0.34 | 0.30 | +0.04 [-0.09, +0.17] |
| share of the top 5 from the wrong project | 0.40 | 0.16 | +0.24 [+0.03, +0.46] |
The documentation share differs by +0.04 between misses and hits, and its interval includes zero, so it is within noise. The wrong-project share differs by +0.24, and its interval excludes zero. The corpus holds four separate projects and nothing in the retriever scopes a query to one of them; a project filter is the obvious next fix, and it would be tested the same way.
How it works
files ──► chunker ──┬─► BM25 index ┐
└─► MiniLM embeddings ├─► RRF fusion ─► per-file cap ─► top k
┘
Tree-sitter parses Python and Java. A class becomes a header chunk (its signature, fields and docstring) plus one chunk per method, so no line is indexed twice; imports and module-level code become their own chunks; Markdown splits on headings and carries the heading path into the text. Files without a grammar (JavaScript, YAML) use windows in every chunking. Every chunk is sized to fit the embedder's window: sentence-transformers/all-MiniLM-L6-v2 reads at most 256 word pieces, so the path :: name prefix plus the text is kept within 254 tokens, and a declaration longer than that is split into consecutive parts. The window baseline packs lines to the same budget with a quarter of each window repeated in the next.
| Chunking | Chunks | Median lines | p95 lines | Max lines | Median tokens | Max tokens |
|---|---|---|---|---|---|---|
| structural, merged | 942 | 12 | 25 | 37 | 225 | 254 |
| structural, unmerged | 1281 | 8 | 22 | 37 | 158 | 254 |
| window | 919 | 16 | 29 | 37 | 242 | 254 |
BM25 runs on a tokeniser that splits identifiers (RateLimitFilter is reachable from "rate limit"), drops English stopwords and stems with Snowball. Fusion is Reciprocal Rank Fusion with k = 60 over the top 60 of each list. The per-file cap keeps the first chunk of every file at the head of the list and pushes the rest below every survivor, so nothing is discarded. Both models run on CPU at pinned revisions.
Limitations
- The set is 36 questions. The intervals are the honest width: a hit rate of 0.81 at rank 5 sits in 0.67 to 0.92. Most differences between configurations are inside that width.
- One labeller, who also wrote the code. A second round with two independent labellers, an agreement statistic and an adjudication log is specified but not done.
- File-level relevance is generous: a chunk from the right file counts even when it is the wrong function in that file.
- The dev/test halves are small. The cap value was chosen on eighteen questions; the test half agrees on direction, and that is all it can say.
- Retrieval only: no generation, so no measurement of whether an answer follows from the retrieved context.