Keel
Corpus grounding ratios
Every verification edge in each repository below was classified by who
produces the signal, and whether the actor being verified can write to that producer.
A check is only a check if the signal it reads comes from
somewhere the thing being checked cannot write to.
generated 2026-07-25T03:59:42.265Z ·
14 target(s) measured · 1 with no verification surface
Pooled across the corpus
0.853
anchored / (anchored + self_referential + unknown)
anchored 122
self_referential 14
unknown 7
not_a_check 190
excluded from the denominator
Scope. Keel measures the shape of verification, not its
quality. A repo can be 100% anchored with terrible tests. Anchoring says the signal comes
from outside; it does not say the signal is sufficient.
Across 14 repositories. The pooled figure weights each
repository by how many edges it contributed, so it is not the mean of the column below.
Both are printed because neither alone is honest: the mean hides size, the pooled figure
hides spread. Mean of the 14 target(s) with a denominator:
0.867 .
By repository
The routes column is the constructive half. A grounding
ratio on its own is a diagnosis nobody can act on; a route names, for each ungrounded
check, an anchored producer the repository already owns and the change that would
wire it in. Keel never invents an anchor — every proposed producer is a node already
classified anchored in that same report,
and one that does not resolve is refused. Where no route exists the page says so and names
the decision required instead, because a proposal that raises the number without grounding
the claim is the failure this project exists to name. Every route is a
proposal : the constructing loop emits, a human admits.
target ratio anchored
self_referential unknown
not_a_check coverage (judged) routes
keel
44f7e8d6cbdf
self-measured — Keel judging Keel, the one run where query
independence collapses. Published because refusing to print our own number
would be the failure this project names; flagged because it is not evidence
of the same kind as the rows around it.
0.357
5 of 14
9
0
11
25 judged sampled 25 of 32 — 78% of the surface
script 12 ci_step 13
routes
anthropic-sdk-python
60c64fba5c2b
0.667
8 of 12
2
2
13
25 judged sampled 25 of 41 — 61% of the surface
test_target 2 ci_step 22 review_gate 1
routes
openai-python
d4c151d92ba7
0.769
10 of 13
2
1
12
25 judged sampled 25 of 95 — 26% of the surface
test_target 2 ci_step 22 review_gate 1
routes
vercel-ai
e29788dd545f
0.800
8 of 10
0
2
15
25 judged sampled 25 of 1014 — 2% of the surface
script 8 test_target 8 ci_step 8 deploy_gate 1
routes
aider
5dc9490bb35f
1.000 thin — 7 classified edges
7 of 7
0
0
18
25 judged sampled 25 of 62 — 40% of the surface
test_target 1 review_gate 1 ci_step 23
—
browser-use
b909fbfba0ae
1.000 thin — 7 classified edges
7 of 7
0
0
18
25 judged sampled 25 of 104 — 24% of the surface
review_gate 1 ci_step 22 test_target 2
—
mcp-python-sdk
00a70148bcd3
1.000
15 of 15
0
0
10
25 judged sampled 25 of 121 — 21% of the surface
review_gate 1 ci_step 16 test_target 8
—
simonw-llm
0392226e6630
0.750 thin — 4 classified edges
3 of 4
1
0
21
25 judged sampled 25 of 38 — 66% of the surface
test_target 2 script 2 ci_step 21
—
tiktoken
08a5f3b2c987
1.000 thin — 4 classified edges
4 of 4
0
0
8
12 judged ci_step 12
—
requests
69f84847045b
0.917
11 of 12
0
1
13
25 judged sampled 25 of 106 — 24% of the surface
review_gate 2 script 10 ci_step 10 test_target 3
routes
flask
36e4a824f340
1.000
15 of 15
0
0
10
25 judged sampled 25 of 59 — 42% of the surface
review_gate 1 test_target 11 ci_step 12 script 1
—
sinatra
cb22afd7902b
0.875 thin — 8 classified edges
7 of 8
0
1
17
25 judged sampled 25 of 59 — 42% of the surface
test_target 2 script 11 ci_step 11 review_gate 1
routes
commander-js
ba6d13ddb424
1.000
12 of 12
0
0
9
21 judged script 11 ci_step 10
—
anthropic-quickstarts
370e18d4a20f
1.000
10 of 10
0
0
15
25 judged sampled 25 of 62 — 40% of the surface
script 9 review_gate 1 ci_step 10 test_target 5
—
not_a_check is excluded from the ratio
entirely, which makes it the one shoppable class: mis-filing a real check there
shrinks the denominator and inflates the score. It is printed here for exactly that reason.
unknown fails closed and counts against the ratio like
self_referential.
Crystallization curve
Cost per node against corpus run index. The run order is part of the
measurement — the curve is a claim about ordered probe accumulation, so a reshuffle
changes it by design. If cost per node does not fall, that is published as-is.
Keel crystallization curve
Crystallization curve over 15 runs from <repo>/reports. falls — total fitted change -66.5% of the mean across 15 runs (R^2 0.28, so the line explains less than half the variance — read the raw squares)
CRYSTALLIZATION CURVE
15 sequential runs · 345 nodes judged of 1838 gathered · 124 anchored
Keel corpus — 15 repositories, 2026-07-24 — measured corpus (declared on the command line)
estimated tokens per node
estimated tokens / judged node
0
300
600
run 0 · keel · estimated tokens per node = 498 · judged 25 of 32 gathered
run 1 · anthropic-sdk-python · estimated tokens per node = 369 · judged 25 of 41 gathered
run 2 · openai-python · estimated tokens per node = 398 · judged 25 of 95 gathered
run 3 · vercel-ai · estimated tokens per node = 410 · judged 25 of 1014 gathered
run 4 · aider · estimated tokens per node = 204 · judged 25 of 62 gathered
run 5 · browser-use · estimated tokens per node = 378 · judged 25 of 104 gathered
run 6 · mcp-python-sdk · estimated tokens per node = 592 · judged 25 of 121 gathered
run 7 · simonw-llm · estimated tokens per node = 189 · judged 25 of 38 gathered
run 8 · tiktoken · estimated tokens per node = 240
run 9 · requests · estimated tokens per node = 290 · judged 25 of 106 gathered
run 10 · flask · estimated tokens per node = 408 · judged 25 of 59 gathered
run 11 · sinatra · estimated tokens per node = 265 · judged 25 of 59 gathered
run 12 · commander-js · estimated tokens per node = 316
run 13 · anthropic-quickstarts · estimated tokens per node = 310 · judged 25 of 62 gathered
run 14 · tiktoken · estimated tokens per node = 53
0*
1*
2*
3*
4*
5*
6*
7*
8
9*
10*
11*
12
13*
14
seconds per node
measured s / judged node
0.0
60.0
120.0
run 0 · keel · seconds per node = 0.1 · judged 25 of 32 gathered
run 1 · anthropic-sdk-python · seconds per node = 15.8 · judged 25 of 41 gathered
run 2 · openai-python · seconds per node = 15.4 · judged 25 of 95 gathered
run 3 · vercel-ai · seconds per node = 105.9 · judged 25 of 1014 gathered
run 4 · aider · seconds per node = 21.3 · judged 25 of 62 gathered
run 5 · browser-use · seconds per node = 21.7 · judged 25 of 104 gathered
run 6 · mcp-python-sdk · seconds per node = 20.1 · judged 25 of 121 gathered
run 7 · simonw-llm · seconds per node = 11.8 · judged 25 of 38 gathered
run 8 · tiktoken · seconds per node = 23.7
run 9 · requests · seconds per node = 16.8 · judged 25 of 106 gathered
run 10 · flask · seconds per node = 20.5 · judged 25 of 59 gathered
run 11 · sinatra · seconds per node = 16.9 · judged 25 of 59 gathered
run 12 · commander-js · seconds per node = 18.7
run 13 · anthropic-quickstarts · seconds per node = 20.4 · judged 25 of 62 gathered
run 14 · tiktoken · seconds per node = 0.1
0*
1*
2*
3*
4*
5*
6*
7*
8
9*
10*
11*
12
13*
14
probe-decided share
measured share of decided nodes
0.00
0.50
1.00
run 0 · keel · probe-decided share = 0.00 · judged 25 of 32 gathered
run 1 · anthropic-sdk-python · probe-decided share = 0.00 · judged 25 of 41 gathered
run 2 · openai-python · probe-decided share = 0.16 · judged 25 of 95 gathered
run 3 · vercel-ai · probe-decided share = 0.08 · judged 25 of 1014 gathered
run 4 · aider · probe-decided share = 0.56 · judged 25 of 62 gathered
run 5 · browser-use · probe-decided share = 0.32 · judged 25 of 104 gathered
run 6 · mcp-python-sdk · probe-decided share = 0.00 · judged 25 of 121 gathered
run 7 · simonw-llm · probe-decided share = 0.56 · judged 25 of 38 gathered
run 8 · tiktoken · probe-decided share = 0.58
run 9 · requests · probe-decided share = 0.28 · judged 25 of 106 gathered
run 10 · flask · probe-decided share = 0.00 · judged 25 of 59 gathered
run 11 · sinatra · probe-decided share = 0.28 · judged 25 of 59 gathered
run 12 · commander-js · probe-decided share = 0.19
run 13 · anthropic-quickstarts · probe-decided share = 0.24 · judged 25 of 62 gathered
run 14 · tiktoken · probe-decided share = 1.00
0*
1*
2*
3*
4*
5*
6*
7*
8
9*
10*
11*
12
13*
14
probe library size
measured probes in library
0
15
30
run 0 · keel · probe library size = 0 · judged 25 of 32 gathered
run 1 · anthropic-sdk-python · probe library size = 3 · judged 25 of 41 gathered
run 2 · openai-python · probe library size = 5 · judged 25 of 95 gathered
run 3 · vercel-ai · probe library size = 8 · judged 25 of 1014 gathered
run 4 · aider · probe library size = 10 · judged 25 of 62 gathered
run 5 · browser-use · probe library size = 12 · judged 25 of 104 gathered
run 6 · mcp-python-sdk · probe library size = 13 · judged 25 of 121 gathered
run 7 · simonw-llm · probe library size = 15 · judged 25 of 38 gathered
run 8 · tiktoken · probe library size = 17
run 9 · requests · probe library size = 19 · judged 25 of 106 gathered
run 10 · flask · probe library size = 22 · judged 25 of 59 gathered
run 11 · sinatra · probe library size = 23 · judged 25 of 59 gathered
run 12 · commander-js · probe library size = 26
run 13 · anthropic-quickstarts · probe library size = 29 · judged 25 of 62 gathered
run 14 · tiktoken · probe library size = 29
0*
1*
2*
3*
4*
5*
6*
7*
8
9*
10*
11*
12
13*
14
x axis: run index in the recorded corpus order (listed below). Squares are the raw per-run values; the dashed line is an
ordinary-least-squares fit and is never shown without them. The fit is clipped to the panel, and withheld entirely (with the
panel saying so) where a straight line would predict values the points cannot take — R^2 for every fit is in the trend notes
below. * on run 0, 1, 2, 3, 4, 5, 6, 7, 9, 10, 11, 13 = judged fewer nodes than gathered; per-node values are per JUDGED node.
Coverage (judged)
ci_step 224 · script 64 · test_target 46 · review_gate 10 · deploy_gate 1
Trend
estimated tokens per node: falls — total fitted change -66.5% of the mean across 15 runs (R^2 0.28, so the line explains less
than half the variance — read the raw squares). raw first 498 -> last 53.
seconds per node: falls — total fitted change -72.9% of the mean across 15 runs (R^2 0.04, so the line explains less than half
the variance — read the raw squares). raw first 0.1 -> last 0.1.
probe-decided share: rises — total fitted change 144.9% of the mean across 15 runs (R^2 0.21, so the line explains less than
half the variance — read the raw squares). raw first 0.00 -> last 1.00.
probe library size: rises — total fitted change 187.7% of the mean across 15 runs (R^2 0.99). raw first 0 -> last 29.
Run order
run order derived from Report.generatedAt (ascending, filename as tiebreak) — no order.json manifest present
0 keel -> 1 anthropic-sdk-python -> 2 openai-python -> 3 vercel-ai -> 4 aider -> 5 browser-use -> 6
mcp-python-sdk -> 7 simonw-llm -> 8 tiktoken -> 9 requests -> 10 flask -> 11 sinatra -> 12 commander-js -> 13
anthropic-quickstarts -> 14 tiktoken
Shuffle check
shuffle check INCOMPLETE — the curve direction ran, the ratio direction did not
Permutation (400 draws, seed 20260724): median |delta normalized slope| 0.633, 92% of draws move it by >= 0.15. A permutation
reorders ALREADY-RECORDED runs; it cannot reproduce what a genuine re-run in a different order would have cost, because the
probe library would have accumulated differently. It is a sanity signal on order-dependence, never a substitute for the
empirical re-run.
Ratio stability under shuffle was NOT checked: no re-run supplied (--shuffled) and none declared in corpus.meta.json. Permuting
recorded runs cannot move a per-target ratio, so the permutation result below says nothing about it. Supply a re-run with
--shuffled <dir>.
Disclosures
- Skipped 3 file(s) that are not usable run reports: corpus-summary.json (no string "target"); curve.json (no string "target");
keel.bindings.json (no "nodes" array).
- Run 0 (keel) judged 25 of 32 gathered nodes — a cap. Per-node cost below is per JUDGED node.
- Run 1 (anthropic-sdk-python) judged 25 of 41 gathered nodes — a cap. Per-node cost below is per JUDGED node.
- Run 2 (openai-python) judged 25 of 95 gathered nodes — a cap. Per-node cost below is per JUDGED node.
- Run 3 (vercel-ai) judged 25 of 1014 gathered nodes — a cap. Per-node cost below is per JUDGED node.
- Run 4 (aider) judged 25 of 62 gathered nodes — a cap. Per-node cost below is per JUDGED node.
- Run 5 (browser-use) judged 25 of 104 gathered nodes — a cap. Per-node cost below is per JUDGED node.
- Run 6 (mcp-python-sdk) judged 25 of 121 gathered nodes — a cap. Per-node cost below is per JUDGED node.
- Run 7 (simonw-llm) judged 25 of 38 gathered nodes — a cap. Per-node cost below is per JUDGED node.
- Run 9 (requests) judged 25 of 106 gathered nodes — a cap. Per-node cost below is per JUDGED node.
- Run 10 (flask) judged 25 of 59 gathered nodes — a cap. Per-node cost below is per JUDGED node.
- Run 11 (sinatra) judged 25 of 59 gathered nodes — a cap. Per-node cost below is per JUDGED node.
- Run 13 (anthropic-quickstarts) judged 25 of 62 gathered nodes — a cap. Per-node cost below is per JUDGED node.
- Run 14 (tiktoken): decidedByProbe + decidedByAgent = 10 but nodesSampled = 12. Probe-decided share uses the decided total
(10).
- Provenance "measured" was declared on the COMMAND LINE (--provenance), not in <repo>/reports/corpus.meta.json.
The declaration therefore lives in the invocation and is only as trustworthy as the run sheet that records it.
- Token counts are ESTIMATES in 15 of 15 run(s) (RunEconomics.tokensEstimated). No API exposes session token usage to a skill,
so the token axis reads "estimated tokens" and nothing here claims a measured token count.
Scope. Keel measures the shape of verification, not its quality. A repo can be 100% anchored with terrible tests. Anchoring says
the signal comes from outside; it does not say the signal is sufficient.
Measured, and found nothing to measure
The clone and the gather both succeeded; these repositories simply carry
no verification edge the gatherer can read. That is a result about the repository, not a
failure of the run, and it is reported as "nothing gathered" rather than as a ratio —
publishing 0.000 here would read as "no grounded checks", which
is a different and false claim.
anthropic-courses — nothing gathered (0 edges)
Reproducing this
Every target is pinned to a revision (shown under its name). The corpus
ran sequentially in a recorded order, because the crystallization curve measures ordered
probe-library growth and a parallel run would destroy the signal being measured.
npx skills add broomva/keel