Keel

Corpus grounding ratios

Every verification edge in each repository below was classified by who produces the signal, and whether the actor being verified can write to that producer.

A check is only a check if the signal it reads comes from somewhere the thing being checked cannot write to.

generated 2026-07-25T03:59:42.265Z · 14 target(s) measured · 1 with no verification surface

Pooled across the corpus

0.853

anchored / (anchored + self_referential + unknown)

anchored 122 self_referential 14 unknown 7 not_a_check 190 excluded from the denominator

Scope. Keel measures the shape of verification, not its quality. A repo can be 100% anchored with terrible tests. Anchoring says the signal comes from outside; it does not say the signal is sufficient.

Across 14 repositories. The pooled figure weights each repository by how many edges it contributed, so it is not the mean of the column below. Both are printed because neither alone is honest: the mean hides size, the pooled figure hides spread. Mean of the 14 target(s) with a denominator: 0.867.


By repository

The routes column is the constructive half. A grounding ratio on its own is a diagnosis nobody can act on; a route names, for each ungrounded check, an anchored producer the repository already owns and the change that would wire it in. Keel never invents an anchor — every proposed producer is a node already classified anchored in that same report, and one that does not resolve is refused. Where no route exists the page says so and names the decision required instead, because a proposal that raises the number without grounding the claim is the failure this project exists to name. Every route is a proposal: the constructing loop emits, a human admits.

targetratioanchored self_referentialunknown not_a_checkcoverage (judged)routes
keel
44f7e8d6cbdf
self-measured — Keel judging Keel, the one run where query independence collapses. Published because refusing to print our own number would be the failure this project names; flagged because it is not evidence of the same kind as the rows around it.
0.357 5 of 14 9 0 11 25 judged
sampled 25 of 32 — 78% of the surface
script 12 ci_step 13
routes
anthropic-sdk-python
60c64fba5c2b
0.667 8 of 12 2 2 13 25 judged
sampled 25 of 41 — 61% of the surface
test_target 2 ci_step 22 review_gate 1
routes
openai-python
d4c151d92ba7
0.769 10 of 13 2 1 12 25 judged
sampled 25 of 95 — 26% of the surface
test_target 2 ci_step 22 review_gate 1
routes
vercel-ai
e29788dd545f
0.800 8 of 10 0 2 15 25 judged
sampled 25 of 1014 — 2% of the surface
script 8 test_target 8 ci_step 8 deploy_gate 1
routes
aider
5dc9490bb35f
1.000
thin — 7 classified edges
7 of 7 0 0 18 25 judged
sampled 25 of 62 — 40% of the surface
test_target 1 review_gate 1 ci_step 23
browser-use
b909fbfba0ae
1.000
thin — 7 classified edges
7 of 7 0 0 18 25 judged
sampled 25 of 104 — 24% of the surface
review_gate 1 ci_step 22 test_target 2
mcp-python-sdk
00a70148bcd3
1.000 15 of 15 0 0 10 25 judged
sampled 25 of 121 — 21% of the surface
review_gate 1 ci_step 16 test_target 8
simonw-llm
0392226e6630
0.750
thin — 4 classified edges
3 of 4 1 0 21 25 judged
sampled 25 of 38 — 66% of the surface
test_target 2 script 2 ci_step 21
tiktoken
08a5f3b2c987
1.000
thin — 4 classified edges
4 of 4 0 0 8 12 judged
ci_step 12
requests
69f84847045b
0.917 11 of 12 0 1 13 25 judged
sampled 25 of 106 — 24% of the surface
review_gate 2 script 10 ci_step 10 test_target 3
routes
flask
36e4a824f340
1.000 15 of 15 0 0 10 25 judged
sampled 25 of 59 — 42% of the surface
review_gate 1 test_target 11 ci_step 12 script 1
sinatra
cb22afd7902b
0.875
thin — 8 classified edges
7 of 8 0 1 17 25 judged
sampled 25 of 59 — 42% of the surface
test_target 2 script 11 ci_step 11 review_gate 1
routes
commander-js
ba6d13ddb424
1.000 12 of 12 0 0 9 21 judged
script 11 ci_step 10
anthropic-quickstarts
370e18d4a20f
1.000 10 of 10 0 0 15 25 judged
sampled 25 of 62 — 40% of the surface
script 9 review_gate 1 ci_step 10 test_target 5

not_a_check is excluded from the ratio entirely, which makes it the one shoppable class: mis-filing a real check there shrinks the denominator and inflates the score. It is printed here for exactly that reason. unknown fails closed and counts against the ratio like self_referential.


Crystallization curve

Cost per node against corpus run index. The run order is part of the measurement — the curve is a claim about ordered probe accumulation, so a reshuffle changes it by design. If cost per node does not fall, that is published as-is.

Keel crystallization curve Crystallization curve over 15 runs from <repo>/reports. falls — total fitted change -66.5% of the mean across 15 runs (R^2 0.28, so the line explains less than half the variance — read the raw squares) CRYSTALLIZATION CURVE 15 sequential runs · 345 nodes judged of 1838 gathered · 124 anchored Keel corpus — 15 repositories, 2026-07-24 — measured corpus (declared on the command line) estimated tokens per node estimated tokens / judged node 0 300 600 run 0 · keel · estimated tokens per node = 498 · judged 25 of 32 gathered run 1 · anthropic-sdk-python · estimated tokens per node = 369 · judged 25 of 41 gathered run 2 · openai-python · estimated tokens per node = 398 · judged 25 of 95 gathered run 3 · vercel-ai · estimated tokens per node = 410 · judged 25 of 1014 gathered run 4 · aider · estimated tokens per node = 204 · judged 25 of 62 gathered run 5 · browser-use · estimated tokens per node = 378 · judged 25 of 104 gathered run 6 · mcp-python-sdk · estimated tokens per node = 592 · judged 25 of 121 gathered run 7 · simonw-llm · estimated tokens per node = 189 · judged 25 of 38 gathered run 8 · tiktoken · estimated tokens per node = 240 run 9 · requests · estimated tokens per node = 290 · judged 25 of 106 gathered run 10 · flask · estimated tokens per node = 408 · judged 25 of 59 gathered run 11 · sinatra · estimated tokens per node = 265 · judged 25 of 59 gathered run 12 · commander-js · estimated tokens per node = 316 run 13 · anthropic-quickstarts · estimated tokens per node = 310 · judged 25 of 62 gathered run 14 · tiktoken · estimated tokens per node = 53 0* 1* 2* 3* 4* 5* 6* 7* 8 9* 10* 11* 12 13* 14 seconds per node measured s / judged node 0.0 60.0 120.0 run 0 · keel · seconds per node = 0.1 · judged 25 of 32 gathered run 1 · anthropic-sdk-python · seconds per node = 15.8 · judged 25 of 41 gathered run 2 · openai-python · seconds per node = 15.4 · judged 25 of 95 gathered run 3 · vercel-ai · seconds per node = 105.9 · judged 25 of 1014 gathered run 4 · aider · seconds per node = 21.3 · judged 25 of 62 gathered run 5 · browser-use · seconds per node = 21.7 · judged 25 of 104 gathered run 6 · mcp-python-sdk · seconds per node = 20.1 · judged 25 of 121 gathered run 7 · simonw-llm · seconds per node = 11.8 · judged 25 of 38 gathered run 8 · tiktoken · seconds per node = 23.7 run 9 · requests · seconds per node = 16.8 · judged 25 of 106 gathered run 10 · flask · seconds per node = 20.5 · judged 25 of 59 gathered run 11 · sinatra · seconds per node = 16.9 · judged 25 of 59 gathered run 12 · commander-js · seconds per node = 18.7 run 13 · anthropic-quickstarts · seconds per node = 20.4 · judged 25 of 62 gathered run 14 · tiktoken · seconds per node = 0.1 0* 1* 2* 3* 4* 5* 6* 7* 8 9* 10* 11* 12 13* 14 probe-decided share measured share of decided nodes 0.00 0.50 1.00 run 0 · keel · probe-decided share = 0.00 · judged 25 of 32 gathered run 1 · anthropic-sdk-python · probe-decided share = 0.00 · judged 25 of 41 gathered run 2 · openai-python · probe-decided share = 0.16 · judged 25 of 95 gathered run 3 · vercel-ai · probe-decided share = 0.08 · judged 25 of 1014 gathered run 4 · aider · probe-decided share = 0.56 · judged 25 of 62 gathered run 5 · browser-use · probe-decided share = 0.32 · judged 25 of 104 gathered run 6 · mcp-python-sdk · probe-decided share = 0.00 · judged 25 of 121 gathered run 7 · simonw-llm · probe-decided share = 0.56 · judged 25 of 38 gathered run 8 · tiktoken · probe-decided share = 0.58 run 9 · requests · probe-decided share = 0.28 · judged 25 of 106 gathered run 10 · flask · probe-decided share = 0.00 · judged 25 of 59 gathered run 11 · sinatra · probe-decided share = 0.28 · judged 25 of 59 gathered run 12 · commander-js · probe-decided share = 0.19 run 13 · anthropic-quickstarts · probe-decided share = 0.24 · judged 25 of 62 gathered run 14 · tiktoken · probe-decided share = 1.00 0* 1* 2* 3* 4* 5* 6* 7* 8 9* 10* 11* 12 13* 14 probe library size measured probes in library 0 15 30 run 0 · keel · probe library size = 0 · judged 25 of 32 gathered run 1 · anthropic-sdk-python · probe library size = 3 · judged 25 of 41 gathered run 2 · openai-python · probe library size = 5 · judged 25 of 95 gathered run 3 · vercel-ai · probe library size = 8 · judged 25 of 1014 gathered run 4 · aider · probe library size = 10 · judged 25 of 62 gathered run 5 · browser-use · probe library size = 12 · judged 25 of 104 gathered run 6 · mcp-python-sdk · probe library size = 13 · judged 25 of 121 gathered run 7 · simonw-llm · probe library size = 15 · judged 25 of 38 gathered run 8 · tiktoken · probe library size = 17 run 9 · requests · probe library size = 19 · judged 25 of 106 gathered run 10 · flask · probe library size = 22 · judged 25 of 59 gathered run 11 · sinatra · probe library size = 23 · judged 25 of 59 gathered run 12 · commander-js · probe library size = 26 run 13 · anthropic-quickstarts · probe library size = 29 · judged 25 of 62 gathered run 14 · tiktoken · probe library size = 29 0* 1* 2* 3* 4* 5* 6* 7* 8 9* 10* 11* 12 13* 14 x axis: run index in the recorded corpus order (listed below). Squares are the raw per-run values; the dashed line is an ordinary-least-squares fit and is never shown without them. The fit is clipped to the panel, and withheld entirely (with the panel saying so) where a straight line would predict values the points cannot take — R^2 for every fit is in the trend notes below. * on run 0, 1, 2, 3, 4, 5, 6, 7, 9, 10, 11, 13 = judged fewer nodes than gathered; per-node values are per JUDGED node. Coverage (judged) ci_step 224 · script 64 · test_target 46 · review_gate 10 · deploy_gate 1 Trend estimated tokens per node: falls — total fitted change -66.5% of the mean across 15 runs (R^2 0.28, so the line explains less than half the variance — read the raw squares). raw first 498 -> last 53. seconds per node: falls — total fitted change -72.9% of the mean across 15 runs (R^2 0.04, so the line explains less than half the variance — read the raw squares). raw first 0.1 -> last 0.1. probe-decided share: rises — total fitted change 144.9% of the mean across 15 runs (R^2 0.21, so the line explains less than half the variance — read the raw squares). raw first 0.00 -> last 1.00. probe library size: rises — total fitted change 187.7% of the mean across 15 runs (R^2 0.99). raw first 0 -> last 29. Run order run order derived from Report.generatedAt (ascending, filename as tiebreak) — no order.json manifest present 0 keel -> 1 anthropic-sdk-python -> 2 openai-python -> 3 vercel-ai -> 4 aider -> 5 browser-use -> 6 mcp-python-sdk -> 7 simonw-llm -> 8 tiktoken -> 9 requests -> 10 flask -> 11 sinatra -> 12 commander-js -> 13 anthropic-quickstarts -> 14 tiktoken Shuffle check shuffle check INCOMPLETE — the curve direction ran, the ratio direction did not Permutation (400 draws, seed 20260724): median |delta normalized slope| 0.633, 92% of draws move it by >= 0.15. A permutation reorders ALREADY-RECORDED runs; it cannot reproduce what a genuine re-run in a different order would have cost, because the probe library would have accumulated differently. It is a sanity signal on order-dependence, never a substitute for the empirical re-run. Ratio stability under shuffle was NOT checked: no re-run supplied (--shuffled) and none declared in corpus.meta.json. Permuting recorded runs cannot move a per-target ratio, so the permutation result below says nothing about it. Supply a re-run with --shuffled <dir>. Disclosures - Skipped 3 file(s) that are not usable run reports: corpus-summary.json (no string "target"); curve.json (no string "target"); keel.bindings.json (no "nodes" array). - Run 0 (keel) judged 25 of 32 gathered nodes — a cap. Per-node cost below is per JUDGED node. - Run 1 (anthropic-sdk-python) judged 25 of 41 gathered nodes — a cap. Per-node cost below is per JUDGED node. - Run 2 (openai-python) judged 25 of 95 gathered nodes — a cap. Per-node cost below is per JUDGED node. - Run 3 (vercel-ai) judged 25 of 1014 gathered nodes — a cap. Per-node cost below is per JUDGED node. - Run 4 (aider) judged 25 of 62 gathered nodes — a cap. Per-node cost below is per JUDGED node. - Run 5 (browser-use) judged 25 of 104 gathered nodes — a cap. Per-node cost below is per JUDGED node. - Run 6 (mcp-python-sdk) judged 25 of 121 gathered nodes — a cap. Per-node cost below is per JUDGED node. - Run 7 (simonw-llm) judged 25 of 38 gathered nodes — a cap. Per-node cost below is per JUDGED node. - Run 9 (requests) judged 25 of 106 gathered nodes — a cap. Per-node cost below is per JUDGED node. - Run 10 (flask) judged 25 of 59 gathered nodes — a cap. Per-node cost below is per JUDGED node. - Run 11 (sinatra) judged 25 of 59 gathered nodes — a cap. Per-node cost below is per JUDGED node. - Run 13 (anthropic-quickstarts) judged 25 of 62 gathered nodes — a cap. Per-node cost below is per JUDGED node. - Run 14 (tiktoken): decidedByProbe + decidedByAgent = 10 but nodesSampled = 12. Probe-decided share uses the decided total (10). - Provenance "measured" was declared on the COMMAND LINE (--provenance), not in <repo>/reports/corpus.meta.json. The declaration therefore lives in the invocation and is only as trustworthy as the run sheet that records it. - Token counts are ESTIMATES in 15 of 15 run(s) (RunEconomics.tokensEstimated). No API exposes session token usage to a skill, so the token axis reads "estimated tokens" and nothing here claims a measured token count. Scope. Keel measures the shape of verification, not its quality. A repo can be 100% anchored with terrible tests. Anchoring says the signal comes from outside; it does not say the signal is sufficient.

Measured, and found nothing to measure

The clone and the gather both succeeded; these repositories simply carry no verification edge the gatherer can read. That is a result about the repository, not a failure of the run, and it is reported as "nothing gathered" rather than as a ratio — publishing 0.000 here would read as "no grounded checks", which is a different and false claim.

Reproducing this

Every target is pinned to a revision (shown under its name). The corpus ran sequentially in a recorded order, because the crystallization curve measures ordered probe-library growth and a parallel run would destroy the signal being measured.

npx skills add broomva/keel