I tried to reproduce a paper's headline number on two repos I actually own, on a laptop with no GPU and a free-tier model. It reproduces on the 249-file repo and loses on the 33-file one — and the thing that decides which is not the technique.
Paper: CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents, Yu et al., arXiv:2607.25431v1.
The paper claims 50–87% fewer trajectory tokens than paired grep/read_file, by
changing how repository context reaches a coding agent. I wanted to know whether that
survives contact with code I have on disk.
Four arms, 8 behavioral questions, 2 private repos, 32 runs, gemma-4-31b at temperature 0,
16-turn cap. The questions are phrased behaviorally and never name the file or the function —
"a user types (514) 555-0142 and it's stored as +15145550142, where does that happen" —
so grep has to earn it. Ground truth was verified by reading each target file before the
questions were written.
| arm | what varies | total input tokens | vs control | found |
|---|---|---|---|---|
control (grep + read_file) |
— | 98,153 | 100% | 7/8 |
| codenib mcp (tool swap — not a paper arm) | tool set | 97,659 | 99.5% | 8/8 |
| eager (paper) | prompt history | 70,366 | 71.7% | 8/8 |
| eager + compact (paper) | prompt history | 42,172 | 43.0% | 8/8 |
Split by repository, which is the actual finding:
| arm | Leadpipe, 249 files | SalesRabbit, 33 files |
|---|---|---|
| codenib mcp (tool swap) | 71.8% | 323.0% |
| eager | 44.1% | 294.4% |
| eager + compact | 32.3% | 129.3% |
On the big repo, compact uses 32.3% of grep/read tokens with 5/5 correctness against the
control's 4/5 — inside the paper's claimed band. On the small repo every arm loses,
including the paper's own, which carry no tool-schema overhead at all. So the penalty isn't
the tools and isn't the delivery policy. It's whether the control agent's grep was going to
succeed anyway. When it was, you're paying for retrieval you didn't need.
The one task where grep genuinely failed — lp-dberror, 16 turns, 47k tokens, no answer —
is the one that lands in the paper's regime. There compact used 11.6%.
My first experiment swapped the agent's tools: grep + read_file out, codenib mcp in.
It came back at 99.5% of control and I wrote down that the paper doesn't reproduce.
That was accurate about what I measured and wrong as a claim about CodeNib. The paper never swaps tools. Its arms all share one tool set and differ only in what's already sitting in the prompt before the agent's first move. I'd measured a different experiment and compared it to their number.
Same retrieval engine, same eight tasks, run their way instead: 43.0%. RESULTS.md keeps
both parts in the order I ran them, retraction included.
A wide tool surface changes the model's policy, not just its cost per call. Nine tools
offer nine plausible next actions, and on lp-sms the agent ran four different searches
before answering — fewer turns than control, 2.4× the tokens. Two tools force convergence.
No experiment that holds the tool set fixed can observe this.
Compaction buys tokens with round trips. Against eager it cut cost on 7 of 8 tasks but added a turn on 4 of them. The rewrite discards the injected candidates, so when the first read wasn't the right file, the agent has to go looking again. Cheaper, not faster.
Tool schemas cost ~875 tokens per turn, measured. The paper's design cancels that term by construction, so it can't appear in their accounting at all.
Injected context invented a new failure mode. On lp-decision the agent answered from the
candidates at turn 1 with zero tool calls, citing a file it never opened — against an explicit
system-prompt instruction not to. grep/read_file structurally cannot fail that way.
| file | what it does |
|---|---|
cafe.py |
runs the arms, writes one receipt per (task, arm) |
order_book.py |
the 8 questions and their ground-truth files |
filter_menu.py |
the control tool set — grep and read_file, nothing else |
espresso_menu.py |
the codenib mcp tool set, spoken over stdio |
brew_kit.py |
the agent loop, rate limiter, receipts, and the eager/compact rewrite |
candidate_tray.py |
freezes the top-10 candidates into candidates.json before any run |
bean_mill.py |
indexing, because codenib index won't expose --batch-size |
tasting_notes.py |
aggregates receipts into the summary table |
latte_art.py |
draws the chart above, straight from the receipts |
repo_walk.py |
counts what the chunker would actually count |
Every run's full message transcript is committed under results_*/transcripts/, and every
number in this README comes out of results_*/receipts.json. Nothing here is estimated.
Q3 (static vs live navigation) and Q4 (incremental maintenance) both need baselines that were
out of scope. The structural view (symbol_graph) could not be built at all — it needs
external SCIP binaries that pip install codenib[graph] doesn't ship — so the treatment arm
ran with two of the paper's three views.
Worth holding against everything above: 8 tasks against the paper's 100 snapshots, one model,
one seed, and two truncation caps that flatter different arms in different directions (both
measured and stated in RESULTS.md). Agreement on sign and on which arm gets selected is all
I'd claim.
The two target repos are private and aren't here. Substitute your own — one small, one medium,
because scale is the variable that decides the sign — and rewrite order_book.py.
nibenv\Scripts\python.exe bean_mill.py <repo> --batch-size 4 --threads 4
nibenv\Scripts\python.exe candidate_tray.py --codenib <exe> --codenib-home <dir> --out candidates.json
nibenv\Scripts\python.exe cafe.py --menus control,codenib --model gemma-4-31b --api-key-file <key> --out results_<repo>
nibenv\Scripts\python.exe cafe.py --menus eager,compact --candidates candidates.json --model gemma-4-31b --api-key-file <key> --out results_paperarms
nibenv\Scripts\python.exe tasting_notes.py "results_*/receipts.json" --out RESULTS_ALL_ARMS.md
python latte_art.py "results_*/receipts.json" --out assetsThat nibenv is the throwaway venv REPRODUCE.md builds in step 1 — codenib doesn't work
from system Python here. latte_art.py is the exception: standard library only, so plain
python redraws the chart from whatever receipts you just produced.
Budget roughly 45 minutes of indexing and 1.5–2 hours for 32 runs if your endpoint caps you at 5 requests a minute. The rate limit is the wall, not the compute. Full setup, checkpoints and the Windows-specific traps are in REPRODUCE.md — start there, not here.
| document | read it for |
|---|---|
| RESULTS.md | every measured number, both parts, and how to read them honestly |
| REPRODUCE.md | step-by-step reproduction, expected checkpoints, troubleshooting |
| PAPER_ARMS.md | what the paper's arms actually are, and where the 50–87% comes from |
| BUILD_LOG.md | the running log — segfaults, wrong hypotheses, dead ends, in order |
The paper itself is in CodeNib.pdf.