Kin Retrieval Proof
Citable within scopePre-registered symmetric protocol v2, n=26
Kin ties a grep-driven agent baseline on file localization. On symbol, span and line localization the point estimate is higher for Kin; we did not test significance on those axes, and under our own pre-registered rule the verdict on all three is a tie. Cost is at parity, and the Kin arm was bit-reproducible in this measured run. The grep-driven agent baseline has the higher file-F1 point estimate; the pre-registered significance rule still returns a tie.
File localization
TIE
Pooled micro-F1 is 0.4148 for Kin vs 0.5325 for the grep-driven agent baseline. The separate paired decision uses mean per-task F1: 0.5257 vs 0.5963, delta -0.0706, 95% CI [-0.1595, +0.0090], McNemar p = 0.125. We claim no win on this axis.
Structural localization
.318 vs .289
On symbol F1 (.318 vs .289), span F1 (.326 vs .297) and line F1 (.307 vs .266), Kin's point estimate is higher; we did not test significance on these axes. Under our own pre-registered rule the verdict on all three is a tie, as it is on file F1.
Cost
1.00x tokens
Cost was at or above parity: Kin averaged 111,015 tokens to answer against the grep-driven agent baseline's 111,345, a ratio of 0.997, at 0.95x turns. The Kin mean is over 22 declared tasks and the baseline mean over roughly 25.7, so the ratio is biased toward Kin. We claim no token saving.
Determinism
26 / 26
Given the same 26 tasks, the same seed, the same machine and temperature 0, Kin returned bit-identical output on 26 of 26 tasks across three passes. The grep-driven agent baseline was bit-identical on 3 of 26. The denominator is tasks, not runs.
Pooled micro-F1 and mean per-task F1 are different aggregations. The pooled values describe magnitude; the paired per-task values govern the pre-registered statistical verdict.
Method
- Suite: 26 Multi-SWE-Bench tasks from the gh CLI repository.
- Symmetry: same model, seed, temperature, context limit, agent loop, and tool-call cap. Only the retrieval surface changed.
- Runtime: temperature 0, seed 0, context 32768, max 24 tool calls, parallelism 1.
- Decision rule: a task-paired 10,000-sample bootstrap over per-task file F1 and an exact McNemar test over file hits. A win requires the confidence interval to exclude zero and p < 0.05.
- Kin gate: the Kin arm had to be bit-identical across three complete passes or the result was not citable.
What this does not prove
- It is not an end-to-end patch-passes-tests or live-agent success result.
- It does not prove lower token cost; the measured result is token parity.
- It does not establish superiority on the file axis; the grep-driven agent baseline has the higher point estimate and the statistical verdict is a tie.
- It does not generalize beyond this suite, language, model, or runtime without another governed run.
Provenance and Artifacts
The run pins Kin at 2508da69, kin-bench at2c9875a, and scorer version3.0.0-gold-denominator. The compact governing verdict and paired decision are published below with SHA-256 hashes. These figures were measured on that commit, not on the current release.
One label correction, stated here rather than edited into the evidence. The comparator arm is keyedlexical in verdict.v2.json. That key is a misnomer: the arm that ran is the grep-driven agent loop described above, not a tuned BM25 retriever. The artifact's own numbers show it, because a deterministic lexical retriever would be bit-identical on 26 of 26 by construction and this arm was bit-identical on 3 of 26. The paired decision artifact names the same armgrep. Both files are published byte-for-byte as produced, so the hashes below still verify.
Disclosure: the full benchmark harness repository and raw run directory are not public today. This page publishes the governing compact artifacts and their exact hashes, but does not claim full third-party reproducibility yet. The complete evidence bundle is available for investor diligence while the public harness boundary is prepared.