1
Canonical protocols and admission
Source URLs, provider launch pages, aggregators, captures and publication hashes are provenance—not benchmark identity. The reviewed map keeps incompatible versions, tracks and evaluation protocols separate, including MRCR context regimes and distinct agent harnesses.
identity = owner + benchmark + exact version + metric/unit
+ genuine track/tool mode + evaluation protocol
≥5 models: full calibration · 2–4: pairwise only · 1: display only
2
Frozen protocol normalization
Methodology 2.1 preserves the reviewed 2.0 canonical protocols and their direction-aware normalization as immutable history. Ordinary ranking and data refreshes cannot relearn either protocol calibration or family offsets.
p = (count below + ½ count tied) ÷ n
P = 15 + 70p
z = direction(x − median) ÷ max(1.4826·MAD, minimum spread)
M = 50 + 35·tanh(z ÷ 2)
BenchmarkScore = clamp(0.55P + 0.45M, 0, 100)
3
Deterministic benchmark information quality
Discrimination uses normalized p90–p10 spread; saturation measures metric-endpoint clustering; support uses log-scaled canonical-model count. Version dates—not check dates—drive freshness, while aliases and mirrors reduce independence. Quality factors are normalized within a pillar and also weight pairwise edges.
Q = discrimination^.30 · saturation^.20 · support^.15
· freshness^.10 · source quality^.20 · independence^.05
freshness half-life: 4 years · missing bounds/dates: neutral 0.5
4
Frozen cross-family offsets
Only broad direct-evidence cohort rows with four pillars, two source clusters and owner or independent support enter the offline learner. Fewer than three overlaps shrink to zero; three, four and five-plus overlaps use 25%, 50% and 75%. The reviewed artifact is versioned and hash-pinned, so adding one model cannot recalibrate the leaderboard.
residual = family score − weighted median(≥2 other families)
offset = clamp(median residual × overlap shrinkage, −15, +15)
calibrated family = family score − frozen offset
5
Equal pillars and normal relative evidence
Related versions remain separate but share family, duplicate, mirror, configuration and source-cluster caps. Existing pairwise, small-table anchored and one-hop successor routes run first and reserve the exact semantic comparison edges they consume. Missing pillars remain missing rather than becoming zero.
PillarScore = source-capped quality-weighted mean(calibrated families)
Category PillarScore = mean(available pillar scores)
normal relative routes = peer + anchored + successor/shared
6
Shared-benchmark peer model
Full and pairwise-only canonical protocols create quality-weighted edges. Larger margins matter more, correlated families are capped, and a win against stronger opposition carries greater fitted value.
outcome = 0.5 + 0.5·tanh(score margin ÷ 8)
all canonical-protocol pairs · regularized Bradley–Terry fit
PeerScore = 0.55 percentile score + 0.45 rating-margin score
7
Successor evidence and generic prior
Exact compatible protocols are combined inside pillars, then across at least two pillars. A two-pillar category-local delta is preferred; otherwise a global lineage delta may inform Score but never category eligibility. Positive and negative deltas are symmetric, one predecessor hop is allowed, and the frozen fallback prior stays a separate component.
SuccessorDelta = equal-pillar reliability-weighted median(shared deltas)
SuccessorScore = predecessor category + clamp(delta, −8, +8)
max successor weight: 25% × [1, .8, .5, .2, 0] for 0,1,2,3,4+ direct pillars
8
Evidence-depth base category blend
If PeerScore is unavailable, its weight moves to PillarScore—never to the prior. Successor evidence is reported separately, its weight decays with direct breadth, and neither successor nor anchored evidence is counted as a full-direct pillar.
4+ direct: .65 Pillar + .35 Peer + 0 Prior
2–3 direct: .65 Pillar + .30 Peer + .05 Prior
1 direct: .70 Pillar + .20 Peer + .10 Prior
0 direct: generic Prior only · Confidence D · no ordinal rank
final = (1 − successor weight)·base + successor weight·SuccessorScore
9
Controlled launch-relative evidence
A raw result can still contribute normally to its absolute family and also supply new relative information. Deduplication occurs at exact benchmark/protocol/configuration/evaluation-system comparison-edge level. A launch edge already used by peer, anchored, successor/shared or another relative route is rejected. Mirrors collapse inside one release event, and launch evidence creates no pillars, eligibility, connectivity or Confidence.
margin = clamp(direction·(candidate − anchor) ÷ max(p90−p10, floor), −1, 1)
estimate = clamp(anchor pre-launch category + 8·margin, 0, 100)
launch weight: 20% at 1–2 direct pillars · 10% at 3 · 0% at 0 or 4+
10
Sparse family-fragility safeguard
The safeguard is symmetric and does nothing merely because a category is sparse. It switches off at three direct pillars, requires an actual comparison replacement signal and never reuses a launch margin numerically.
trigger: ≤2 direct pillars + ≥2 families + leave-one-family range > 8
replacement: real cross-family peer signal backed by ≥2 non-zero-overlap families
movement = clamp(.15 × clamp(replacement − score, −8, +8), −1.2, +1.2)
11
Overall, Confidence and uncertainty
Unallocated weight uses nominally combined global/peer category estimates bounded within ±6 of the model prior and no higher than one point below the weakest observed category. Zero-direct-category models retain a diagnostic prior but receive no competitive Overall Score or rank. Eligibility, Confidence and nominal weights remain unchanged.
C = observed nominal category coverage
effective observed weight = nominal × min(1/C, 1.67)
p = .15 × (1 − C)
Overall = (1 − p) × (.85 fixed composite + .15 observed median) + p × bounded model prior