A closed, ordered vocabulary of evidence classes. An auditor (model or human) assigns a rung; the percent is the rung's code — a fixed integer, never re-estimated per claim. Tier is chosen by a decidable test, the number derived from the tier, so ratings are idempotent (byte-identical on re-audit of an unchanged claim) and comparable across raters and models.
What the number means
T-value = the confidence you are entitled to place in a well-formed, specifically scoped claim when THIS is the only support offered, before any independent verification.
It is an entitlement ceiling, not a probability estimate of the claim:
- A low rung means "nearly all trust must still be earned elsewhere," not "the claim is probably false." (T10 hearsay is weak evidence, not anti-evidence.)
- Priors still dominate extremes: an extraordinary claim at T60 remains unsettled; a near-tautological claim needs no rung at all. The rung caps entitlement; it never creates it.
- Only two rungs escape the empirical middle: T100 (true by meaning within a stated formal system) and T0 (demonstrated false). Everything about the world sits strictly between — Cromwell's rule, quantized.
The Ladder
The table scrolls in two directions on smaller screens. Focus it, then use the arrow keys or Shift + mouse wheel to explore.
Grades give the coarse 6-level ordinal used for gap arithmetic; they align with the warrant scale: W5=97, W4=88, W3=68, W2=43, W1=20, W0=5 each fall inside grades 5, 4, 3, 2, 1, 0 respectively.
| T | Grade | Rung | Assignment test | Example |
|---|
44 rungs. All integers unique; unused integers (e.g., 87, 45, 35) are reserved slots for future rungs — extendable toward the 100-item ideal without renumbering anything.
Assignment Protocol
- Scope the claim as a single specific proposition. Re-scoping is re-rating: the same record is T88 about its subject and T22 as a generalization.
- Identify the best support actually offered — not the best conceivable.
- Walk top-down; take the FIRST rung whose test the support provably passes. Tests are ordered strictly harder upward, so first-pass is unique.
- On doubt between adjacent rungs, take the lower and flag for review. An honest T48 beats a T65 overclaim — the exact analog of "an honest W1/forced beats a W4/clean overclaim."
- Weakest link. A claim resting on a chain of dependencies takes the minimum rung across links.
- Aggregation is a new classification event, never addition. Ten independent anecdotes are not a meta-analysis; but three independent labs do pass the T92 test. Re-run the walk on the combined support as its own object.
- Emit rung + one-clause test-pass, mirroring the Warrant output format:
T65: "objective endpoint, n=1,200, data on OSF — no replication yet". The clause is what a second auditor checks; agreement on the clause implies agreement on the number.
Integration Notes (for the split/audit pipeline)
- Store the rung integer as the canonical value (like
warrant.tier); render the percent from it, never the reverse. Idempotency requirement carries over: an unchanged claim re-audited must yield a byte-identical rung. - Gap arithmetic runs on GRADES, not rungs.
confidenceGapcounts levels; one level on this 44-rung ladder is ~7× finer than one warrant tier, so a rung-based gap would silently retune every threshold. Grade(local) − Grade(floor) is directly comparable with warrant-tier gaps because the grade bands contain the warrant midpoints (97→5, 88→4, 68→3, 43→2, 20→1, 5→0). - Grade 0 auto-flags for human review, matching W0 behavior.
- Adding a rung: claim an unused integer, write its one-clause test so it is strictly harder than the rung below and easier than the rung above, assign the grade of its band, bump the version constant. Never renumber, never reuse a retired integer.
Changelog vs. v1 (the 25-row banded table)
- Bands → fixed integers at (roughly) band midpoints; conditional bands split into condition-keyed rungs: eyewitness 60/40/25, isolated paper 65/48, tertiary 80/50, journalism 70/38, patent 83/30, intuition 53/8, heuristic 42/12, preprint 55/44, case report 88/22.
- Endpoints re-anchored: T100 analytic only; T0 demonstrated-false only (rumor moved to T10 — untraceable, not "likely fabricated"); adversarial source added at T2.
- "Conservation of mass" example dropped (it was overturned — the framework's own lesson); replaced by second law of thermodynamics under convergent core theory.
- Added common claim classes previously missing: regulatory review (86), primary document (72), official statistics (78), demonstrated prototype (83).