Case study

CodeAtlas

Semantic search over four of my own projects, and the evaluation harness that measures it. The numbers on this page are rendered from the committed results files by a script; a rerun on the pinned corpus regenerates them.

What is measured

The corpus is four repositories of mine, checked out at fixed commits and listed file by file, with line counts and hashes, in the results directory. The evaluator refuses to run on a corpus that differs from that list, so the numbers here can be regenerated by anyone.

The four projects, at the commits recorded in results/corpus_manifest.json
ProjectRepositoryCommitFilesLines
difflensDiffLens-Engine3ca9978e05b9606,640
gesturecontrolGestureControl-Engine53c4b04bf767131,316
portfolioportfolio7c361eea5f1d101,458
taskforgeTaskForge-Engine8c02b8a514bb332,405
total11611,819

The question set is 36 questions I wrote against these projects, phrased the way someone new to a repository would ask rather than by pasting identifiers, each with the files that answer it. Relevance is judged at file level: a retrieved chunk counts when its source file is one of the labelled answers. I wrote both the questions and the code, and there has been one labeller so far; the labelling protocol for a second round with two independent labellers is in the repository.

The questions are split 50/50 into a dev half and a test half. The only setting chosen on the data, the per-file cap value, is chosen on the dev half; every table reports all 36 questions with a bootstrap interval, because eighteen questions per half say very little on their own.

Metrics

Every cell carries a 95% percentile bootstrap interval over questions (10,000 resamples). Every claim that one configuration beats another rests on a paired bootstrap of the per-question differences and an exact sign test on the questions where the two differ. When the interval includes zero the page says so.

Results

The shipped configuration is merged structural chunking, hybrid retrieval, and a per-file cap of 1.

Shipped configuration on all 36 questions; value and 95% bootstrap interval
Configurationhit rate@1hit rate@3hit rate@5recall@5MRRnDCG@10
structural, merged / hybrid + per-file cap0.42 [0.25, 0.58]0.69 [0.56, 0.83]0.81 [0.67, 0.92]0.65 [0.52, 0.78]0.577 [0.444, 0.702]0.573 [0.469, 0.674]

Every configuration

Every configuration on the same 36 questions; value and 95% bootstrap interval
ChunkingStrategyChunkshit rate@1hit rate@5recall@5MRRnDCG@10
structural, mergedBM259420.14 [0.03, 0.25]0.67 [0.50, 0.81]0.53 [0.38, 0.66]0.342 [0.245, 0.448]0.389 [0.297, 0.486]
structural, mergeddense9420.39 [0.22, 0.56]0.83 [0.69, 0.94]0.64 [0.52, 0.75]0.532 [0.406, 0.661]0.532 [0.421, 0.640]
structural, mergedhybrid (RRF)9420.42 [0.25, 0.58]0.72 [0.58, 0.86]0.57 [0.44, 0.71]0.547 [0.415, 0.679]0.522 [0.414, 0.628]
structural, mergedhybrid + per-file cap9420.42 [0.25, 0.58]0.81 [0.67, 0.92]0.65 [0.52, 0.78]0.577 [0.444, 0.702]0.573 [0.469, 0.674]
structural, mergedhybrid + reranker (text only)9420.28 [0.14, 0.42]0.72 [0.56, 0.86]0.55 [0.42, 0.68]0.423 [0.302, 0.550]0.433 [0.332, 0.536]
structural, mergedhybrid + reranker (prefixed)9420.25 [0.11, 0.39]0.78 [0.64, 0.92]0.59 [0.46, 0.72]0.455 [0.342, 0.572]0.478 [0.383, 0.572]
structural, unmergedBM2512810.25 [0.11, 0.39]0.56 [0.39, 0.72]0.45 [0.31, 0.59]0.382 [0.264, 0.510]0.390 [0.289, 0.492]
structural, unmergeddense12810.31 [0.17, 0.45]0.69 [0.53, 0.83]0.57 [0.43, 0.71]0.474 [0.350, 0.602]0.500 [0.393, 0.608]
structural, unmergedhybrid (RRF)12810.39 [0.22, 0.56]0.64 [0.47, 0.81]0.50 [0.36, 0.63]0.521 [0.390, 0.654]0.498 [0.396, 0.601]
structural, unmergedhybrid + per-file cap12810.39 [0.22, 0.56]0.75 [0.61, 0.89]0.60 [0.47, 0.73]0.568 [0.447, 0.692]0.589 [0.490, 0.682]
structural, unmergedhybrid + reranker (text only)12810.31 [0.17, 0.47]0.67 [0.50, 0.81]0.50 [0.37, 0.64]0.436 [0.313, 0.563]0.443 [0.339, 0.548]
structural, unmergedhybrid + reranker (prefixed)12810.28 [0.14, 0.42]0.78 [0.64, 0.92]0.57 [0.44, 0.70]0.461 [0.347, 0.582]0.482 [0.385, 0.581]
windowBM259190.22 [0.08, 0.36]0.64 [0.47, 0.78]0.50 [0.37, 0.64]0.385 [0.266, 0.506]0.386 [0.284, 0.491]
windowdense9190.36 [0.22, 0.53]0.75 [0.61, 0.89]0.59 [0.46, 0.72]0.514 [0.387, 0.643]0.508 [0.398, 0.620]
windowhybrid (RRF)9190.39 [0.25, 0.56]0.69 [0.53, 0.83]0.50 [0.37, 0.62]0.515 [0.382, 0.651]0.487 [0.377, 0.594]
windowhybrid + per-file cap9190.39 [0.22, 0.56]0.81 [0.67, 0.92]0.65 [0.52, 0.78]0.567 [0.446, 0.693]0.573 [0.470, 0.672]
windowhybrid + reranker (text only)9190.31 [0.17, 0.44]0.67 [0.50, 0.81]0.52 [0.39, 0.66]0.449 [0.323, 0.578]0.427 [0.321, 0.534]
windowhybrid + reranker (prefixed)9190.28 [0.14, 0.44]0.67 [0.50, 0.81]0.52 [0.38, 0.66]0.442 [0.323, 0.565]0.453 [0.359, 0.549]

What beat what, and what was within noise

Structural boundaries against windows

Merged structural chunking with hybrid retrieval scores hit rate@1 0.42 against 0.39 for windows (difference +0.03, 95% interval -0.06 to +0.11; better on 2 questions, worse on 1, same on 33). The interval includes zero, so this is within noise at this sample size. The same chunking scores MRR 0.547 against 0.515 for windows (difference +0.032, 95% interval -0.027 to +0.089; better on 11 questions, worse on 4, same on 21). The interval includes zero, so this is within noise at this sample size. Both chunkers are packed to the same token budget, so the comparison is about where boundaries fall, not about chunk size.

Hybrid against dense

Hybrid retrieval scores MRR 0.547 against 0.532 for dense retrieval alone (difference +0.015, 95% interval -0.126 to +0.160; better on 14 questions, worse on 11, same on 11). The interval includes zero, so this is within noise at this sample size. Hybrid retrieval scores hit rate@5 0.72 against 0.83 for dense retrieval (difference -0.11, 95% interval -0.25 to +0.00; better on 1 question, worse on 5, same on 30). The interval includes zero, so this is within noise at this sample size. Fusing in the BM25 ranking pushes some files that dense retrieval had in the top 5 below the cut.

Hybrid against BM25

Hybrid retrieval scores MRR 0.547 against 0.342 for BM25 alone (difference +0.205, 95% interval +0.095 to +0.318; better on 20 questions, worse on 3, same on 13). The interval excludes zero.

The per-file cap

With the cap, hybrid retrieval scores recall@5 0.65 against 0.57 for the same ranking without it (difference +0.08, 95% interval +0.01 to +0.17; better on 4 questions, worse on 0, same on 32). The interval excludes zero. The cap scores MRR 0.577 against 0.547 for no cap (difference +0.029, 95% interval +0.011 to +0.052; better on 9 questions, worse on 0, same on 27). The interval excludes zero. Under file-level relevance the cap cannot lower hit rate or MRR, because the first chunk of every file is always kept and only moves up; the metric it can lower is recall, and it did not.

Per-file cap values tried on the dev half (n=18) of the merged structural chunking with hybrid retrieval; the test half (n=18) and all questions shown for reference
Per-file capdev MRRdev recall@5test MRRtest recall@5all MRRall hit rate@5all recall@5
no cap0.6040.480.4910.670.5470.720.57
cap 30.6040.480.4940.670.5490.720.57
cap 20.6070.510.4980.670.5530.750.59
cap 1 (chosen)0.6140.530.5390.780.5770.810.65

The shipped configuration against the next best cell

The shipped configuration scores MRR 0.577 against 0.568 for structural, unmerged / hybrid + per-file cap (difference +0.009, 95% interval -0.039 to +0.059; better on 7 questions, worse on 4, same on 25). The interval includes zero, so this is within noise at this sample size.

All comparisons

Paired comparisons on the same questions: mean difference with a 95% paired-bootstrap interval, question counts and an exact sign test
Comparison (A vs B)MetricABA - B [95% CI]better / worse / samesign test pReading
structural boundaries vs windows, hybridhit rate@10.420.39+0.03 [-0.06, +0.11]2 / 1 / 331.000within noise
structural boundaries vs windows, hybridMRR0.5470.515+0.032 [-0.027, +0.089]11 / 4 / 210.118within noise
structural boundaries vs windows, hybridnDCG@100.5220.487+0.035 [-0.004, +0.077]12 / 6 / 180.238within noise
structural boundaries vs windows, densehit rate@10.390.36+0.03 [-0.08, +0.14]3 / 2 / 311.000within noise
structural boundaries vs windows, denseMRR0.5320.514+0.019 [-0.069, +0.101]10 / 4 / 220.180within noise
merged vs unmerged structural, hybridhit rate@10.420.39+0.03 [-0.06, +0.11]2 / 1 / 331.000within noise
merged vs unmerged structural, hybridMRR0.5470.521+0.026 [-0.027, +0.085]10 / 3 / 230.092within noise
hybrid vs dense, mergedhit rate@10.420.39+0.03 [-0.17, +0.22]7 / 6 / 231.000within noise
hybrid vs dense, mergedhit rate@50.720.83-0.11 [-0.25, +0.00]1 / 5 / 300.219within noise
hybrid vs dense, mergedMRR0.5470.532+0.015 [-0.126, +0.160]14 / 11 / 110.690within noise
hybrid vs dense, mergednDCG@100.5220.532-0.010 [-0.111, +0.087]14 / 12 / 100.845within noise
hybrid vs dense, windowhit rate@10.390.36+0.03 [-0.11, +0.17]4 / 3 / 291.000within noise
hybrid vs dense, windowMRR0.5150.514+0.001 [-0.109, +0.111]12 / 10 / 140.832within noise
hybrid vs bm25, mergedhit rate@50.720.67+0.06 [-0.08, +0.19]4 / 2 / 300.688within noise
hybrid vs bm25, mergedMRR0.5470.342+0.205 [+0.095, +0.318]20 / 3 / 130.001interval excludes zero
per-file cap vs none, merged hybridhit rate@30.690.67+0.03 [+0.00, +0.08]1 / 0 / 351.000within noise
per-file cap vs none, merged hybridhit rate@50.810.72+0.08 [+0.00, +0.19]3 / 0 / 330.250within noise
per-file cap vs none, merged hybridrecall@50.650.57+0.08 [+0.01, +0.17]4 / 0 / 320.125interval excludes zero
per-file cap vs none, merged hybridMRR0.5770.547+0.029 [+0.011, +0.052]9 / 0 / 270.004interval excludes zero
per-file cap vs none, merged hybridnDCG@100.5730.522+0.051 [+0.027, +0.078]16 / 0 / 200.000interval excludes zero
reranker (text only) vs hybrid, mergedhit rate@10.280.42-0.14 [-0.25, -0.03]0 / 5 / 310.062interval excludes zero
reranker (text only) vs hybrid, mergedMRR0.4230.547-0.124 [-0.220, -0.042]5 / 12 / 190.143interval excludes zero
reranker (prefixed) vs hybrid, mergedhit rate@10.250.42-0.17 [-0.31, -0.03]1 / 7 / 280.070interval excludes zero
reranker (prefixed) vs hybrid, mergedMRR0.4550.547-0.092 [-0.194, +0.008]10 / 13 / 130.678within noise
reranker prefixed vs text only, mergedhit rate@10.250.28-0.03 [-0.14, +0.06]1 / 2 / 331.000within noise
reranker prefixed vs text only, mergedMRR0.4550.423+0.032 [-0.024, +0.082]16 / 3 / 170.004within noise
reranker (text only) vs hybrid, windowhit rate@10.310.39-0.08 [-0.25, +0.08]3 / 6 / 270.508within noise
reranker (text only) vs hybrid, windowMRR0.4490.515-0.066 [-0.183, +0.045]10 / 10 / 161.000within noise
reranker (prefixed) vs hybrid, windowhit rate@10.280.39-0.11 [-0.28, +0.06]3 / 7 / 260.344within noise
reranker (prefixed) vs hybrid, windowMRR0.4420.515-0.074 [-0.192, +0.039]11 / 11 / 141.000within noise
reranker prefixed vs text only, windowhit rate@10.280.31-0.03 [-0.08, +0.00]0 / 1 / 351.000within noise
reranker prefixed vs text only, windowMRR0.4420.449-0.007 [-0.046, +0.023]5 / 4 / 271.000within noise
shipped configuration vs best other cell by MRRhit rate@10.420.39+0.03 [-0.06, +0.11]2 / 1 / 331.000within noise
shipped configuration vs best other cell by MRRhit rate@50.810.75+0.06 [-0.06, +0.17]3 / 1 / 320.625within noise
shipped configuration vs best other cell by MRRMRR0.5770.568+0.009 [-0.039, +0.059]7 / 4 / 250.549within noise

The cross-encoder reranker

Reranking the top 30 with cross-encoder/ms-marco-MiniLM-L-6-v2 was run two ways: on the chunk text alone, and on the same path :: name payload the embedder sees, so that the reranker is not judged on less information than the first stage had. Every chunk fits the reranker's window, so nothing is truncated in either variant.

Reranking the chunk text alone scores MRR 0.423 against 0.547 for hybrid retrieval without it (difference -0.124, 95% interval -0.220 to -0.042; better on 5 questions, worse on 12, same on 19). The interval excludes zero. Reranking the same path-and-name payload the embedder sees scores MRR 0.455 against 0.547 for hybrid retrieval (difference -0.092, 95% interval -0.194 to +0.008; better on 10 questions, worse on 13, same on 13). The interval includes zero, so this is within noise at this sample size.

The prefixed variant scores MRR 0.455 against 0.423 for the text-only variant (difference +0.032, 95% interval -0.024 to +0.082; better on 16 questions, worse on 3, same on 17). The interval includes zero, so this is within noise at this sample size. On window chunks, the text-only reranker scores MRR 0.449 against 0.515 for hybrid retrieval (difference -0.066, 95% interval -0.183 to +0.045; better on 10 questions, worse on 10, same on 16). The interval includes zero, so this is within noise at this sample size.

What the data supports: in this setup the reranker does not improve ranking, and on merged structural chunks the text-only variant is worse by more than the noise. It does not support a cause. Whether the model's training domain, the small candidate set, or the file-level judgement explains it is not something this experiment tested.

Where it fails

The shipped configuration misses 7 of 36 questions at rank 5: tf03, tf08, tf09, tf12, dl06, dl13, gc05. Two things were measured about what fills the top 5 on those questions.

What fills the top 5 for the 7 questions the shipped configuration misses at rank 5, against the 29 it hits
ShareOn missesOn hitsDifference [95% CI]
share of the top 5 that is documentation0.340.30+0.04 [-0.09, +0.17]
share of the top 5 from the wrong project0.400.16+0.24 [+0.03, +0.46]

The documentation share differs by +0.04 between misses and hits, and its interval includes zero, so it is within noise. The wrong-project share differs by +0.24, and its interval excludes zero. The corpus holds four separate projects and nothing in the retriever scopes a query to one of them; a project filter is the obvious next fix, and it would be tested the same way.

How it works

files ──► chunker ──┬─► BM25 index          ┐
                    └─► MiniLM embeddings   ├─► RRF fusion ─► per-file cap ─► top k
                                            ┘

Tree-sitter parses Python and Java. A class becomes a header chunk (its signature, fields and docstring) plus one chunk per method, so no line is indexed twice; imports and module-level code become their own chunks; Markdown splits on headings and carries the heading path into the text. Files without a grammar (JavaScript, YAML) use windows in every chunking. Every chunk is sized to fit the embedder's window: sentence-transformers/all-MiniLM-L6-v2 reads at most 256 word pieces, so the path :: name prefix plus the text is kept within 254 tokens, and a declaration longer than that is split into consecutive parts. The window baseline packs lines to the same budget with a quarter of each window repeated in the next.

Chunk statistics; every payload fits the 254-token budget
ChunkingChunksMedian linesp95 linesMax linesMedian tokensMax tokens
structural, merged942122537225254
structural, unmerged128182237158254
window919162937242254

BM25 runs on a tokeniser that splits identifiers (RateLimitFilter is reachable from "rate limit"), drops English stopwords and stems with Snowball. Fusion is Reciprocal Rank Fusion with k = 60 over the top 60 of each list. The per-file cap keeps the first chunk of every file at the head of the list and pushes the rest below every survivor, so nothing is discarded. Both models run on CPU at pinned revisions.

Limitations