Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

7 Commits

Folders and files

Repository files navigation

CodeNib-Replication

I tried to reproduce a paper's headline number on two repos I actually own, on a laptop with no GPU and a free-tier model. It reproduces on the 249-file repo and loses on the 33-file one — and the thing that decides which is not the technique.

Paper: CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents, Yu et al., arXiv:2607.25431v1.

The order

The paper claims 50–87% fewer trajectory tokens than paired grep/read_file, by changing how repository context reaches a coding agent. I wanted to know whether that survives contact with code I have on disk.

Four arms, 8 behavioral questions, 2 private repos, 32 runs, gemma-4-31b at temperature 0, 16-turn cap. The questions are phrased behaviorally and never name the file or the function — "a user types (514) 555-0142 and it's stored as +15145550142, where does that happen" — so grep has to earn it. Ground truth was verified by reading each target file before the questions were written.

What came out

Per-task input tokens as a percentage of the grep/read control. On Leadpipe (249 source files) most bars fall left of the 100% control line, with eager+compact reaching 11.6% on lp-dberror. On SalesRabbit (33 source files) every bar for every arm sits right of the line, from 108.6% to 559.7%.
arm what varies total input tokens vs control found
control (grep + read_file) — 98,153 100% 7/8
codenib mcp (tool swap — not a paper arm) tool set 97,659 99.5% 8/8
eager (paper) prompt history 70,366 71.7% 8/8
eager + compact (paper) prompt history 42,172 43.0% 8/8

Split by repository, which is the actual finding:

arm Leadpipe, 249 files SalesRabbit, 33 files
codenib mcp (tool swap) 71.8% 323.0%
eager 44.1% 294.4%
eager + compact 32.3% 129.3%

On the big repo, compact uses 32.3% of grep/read tokens with 5/5 correctness against the control's 4/5 — inside the paper's claimed band. On the small repo every arm loses, including the paper's own, which carry no tool-schema overhead at all. So the penalty isn't the tools and isn't the delivery policy. It's whether the control agent's grep was going to succeed anyway. When it was, you're paying for retrieval you didn't need.

The one task where grep genuinely failed — lp-dberror, 16 turns, 47k tokens, no answer — is the one that lands in the paper's regime. There compact used 11.6%.

I had it backwards for a whole session

My first experiment swapped the agent's tools: grep + read_file out, codenib mcp in. It came back at 99.5% of control and I wrote down that the paper doesn't reproduce.

That was accurate about what I measured and wrong as a claim about CodeNib. The paper never swaps tools. Its arms all share one tool set and differ only in what's already sitting in the prompt before the agent's first move. I'd measured a different experiment and compared it to their number.

Same retrieval engine, same eight tasks, run their way instead: 43.0%. RESULTS.md keeps both parts in the order I ran them, retraction included.

Things I didn't see coming

A wide tool surface changes the model's policy, not just its cost per call. Nine tools offer nine plausible next actions, and on lp-sms the agent ran four different searches before answering — fewer turns than control, 2.4× the tokens. Two tools force convergence. No experiment that holds the tool set fixed can observe this.

Compaction buys tokens with round trips. Against eager it cut cost on 7 of 8 tasks but added a turn on 4 of them. The rewrite discards the injected candidates, so when the first read wasn't the right file, the agent has to go looking again. Cheaper, not faster.

Tool schemas cost ~875 tokens per turn, measured. The paper's design cancels that term by construction, so it can't appear in their accounting at all.

Injected context invented a new failure mode. On lp-decision the agent answered from the candidates at turn 1 with zero tool calls, citing a file it never opened — against an explicit system-prompt instruction not to. grep/read_file structurally cannot fail that way.

The café

file what it does
cafe.py runs the arms, writes one receipt per (task, arm)
order_book.py the 8 questions and their ground-truth files
filter_menu.py the control tool set — grep and read_file, nothing else
espresso_menu.py the codenib mcp tool set, spoken over stdio
brew_kit.py the agent loop, rate limiter, receipts, and the eager/compact rewrite
candidate_tray.py freezes the top-10 candidates into candidates.json before any run
bean_mill.py indexing, because codenib index won't expose --batch-size
tasting_notes.py aggregates receipts into the summary table
latte_art.py draws the chart above, straight from the receipts
repo_walk.py counts what the chunker would actually count

Every run's full message transcript is committed under results_*/transcripts/, and every number in this README comes out of results_*/receipts.json. Nothing here is estimated.

What I didn't reproduce

Q3 (static vs live navigation) and Q4 (incremental maintenance) both need baselines that were out of scope. The structural view (symbol_graph) could not be built at all — it needs external SCIP binaries that pip install codenib[graph] doesn't ship — so the treatment arm ran with two of the paper's three views.

Worth holding against everything above: 8 tasks against the paper's 100 snapshots, one model, one seed, and two truncation caps that flatter different arms in different directions (both measured and stated in RESULTS.md). Agreement on sign and on which arm gets selected is all I'd claim.

Running it

The two target repos are private and aren't here. Substitute your own — one small, one medium, because scale is the variable that decides the sign — and rewrite order_book.py.

nibenv\Scripts\python.exe bean_mill.py <repo> --batch-size 4 --threads 4
nibenv\Scripts\python.exe candidate_tray.py --codenib <exe> --codenib-home <dir> --out candidates.json
nibenv\Scripts\python.exe cafe.py --menus control,codenib --model gemma-4-31b --api-key-file <key> --out results_<repo>
nibenv\Scripts\python.exe cafe.py --menus eager,compact --candidates candidates.json --model gemma-4-31b --api-key-file <key> --out results_paperarms
nibenv\Scripts\python.exe tasting_notes.py "results_*/receipts.json" --out RESULTS_ALL_ARMS.md
python latte_art.py "results_*/receipts.json" --out assets

That nibenv is the throwaway venv REPRODUCE.md builds in step 1 — codenib doesn't work from system Python here. latte_art.py is the exception: standard library only, so plain python redraws the chart from whatever receipts you just produced.

Budget roughly 45 minutes of indexing and 1.5–2 hours for 32 runs if your endpoint caps you at 5 requests a minute. The rate limit is the wall, not the compute. Full setup, checkpoints and the Windows-specific traps are in REPRODUCE.md — start there, not here.

The rest of the write-up

document read it for
RESULTS.md every measured number, both parts, and how to read them honestly
REPRODUCE.md step-by-step reproduction, expected checkpoints, troubleshooting
PAPER_ARMS.md what the paper's arms actually are, and where the 50–87% comes from
BUILD_LOG.md the running log — segfaults, wrong hypotheses, dead ends, in order

The paper itself is in CodeNib.pdf.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages