Home

Reliability and corrections

Our methodologies promise audits: reproducibility checks, verification outcomes, and a corrections register. A methodology that promises an audit is not the same object as one that reports it, so this page reports it, with real counts, starting from the first measured round on July 24, 2026. Numbers here get updated as new rounds run, and the ugly ones stay up.

Verdict reproducibility, round one

28published logic-test verdicts re-coded blind
0.89raw agreement (25 of 28)
0.77Cohen's kappa across the four verdict categories
3disagreements, all between adjacent categories, listed below

Design: every logic-test verdict published to date (the fifteen Fault Line Report cards and the thirteen tech and AI special report cards) was re-coded by an independent model coder shown the claim, both stated readings, and each prediction with its observation, blinded to the published holds, verdict, plain reading, and confidence. This is model-to-model agreement, disclosed as such; it measures whether the recorded evidence compels the recorded verdict, not whether a human would design the same test. A human-anchored round is one we intend to run when the funding or volunteer co-coders exist to run it properly; until then, model-based rounds are the honest ceiling of this operation, and the gold-set false-positive audit likewise remains open until that set is built.

CardPublished verdictBlind re-code
New York assisted dying (pol-ny-maid)not testablemixed evidence
NDAA blockade (pol-ndaa-blockade)mixed evidencecontradicted by checks
Jordan deaths (pol-jordan-deaths)mixed evidencecontradicted by checks

We read the pattern in the disagreements as the blind coder weighing failed strongest-reading checks more heavily than the published codings did. That is exactly the boundary the steelman discipline in the polarization methodology exists to police, and these three cards are queued for third-coder review.

Verification-pass outcomes

Every report runs two-stage verification: authors verify each kept item at its source, then independent adversarial checkers re-fetch quotes, links, and dates and attack the test designs. What that machinery actually caught:

PassVolumeCaught
Fault Line Report, July 14 to 20138 candidates screened, 62 dropped; 133 adversarial checks on the survivors1 misattributed quote, 1 unverifiable carrier (card restructured and rescored), 1 misquote, 1 wrong-source quote, 1 unsupported claim (verdict confidence lowered)
Tech and AI special report, Jan to Jul622 items screened, 40 passed the gate, 13 cards; 179 quote checks160 verified verbatim; 10 mismatches, 4 wrong dates, 4 unreachable, 1 wrong attribution, all corrected or cut before scoring; 2 ignition statuses downgraded; 1 carrier's official role corrected mid-window across 3 cards
Religious hate report, Jan to Jul925 accounts, 38 lanes, 87 candidates28 candidates dropped in verification; 59 survived to publication
Roster expansion, Canada UK Europe126 verified candidates across 10 lanes2 rejected, 5 tier downgrades, 16 duplicates caught across two dedup layers; roughly 140 discovery near-misses logged unpublished

Corrections register

Public corrections to published surfaces, dated. Silent fixes are not made; anything that changed a published number or verdict lands here.

DateCorrection
2026-07-22Fault Line Report headline recounted from four to three contradicted narratives after an external review of our test designs; the political-terrorism ministerial card's verdict moved from a fails-own-logic coding to mixed evidence under a rescoped check, and the verdict vocabulary was renamed to describe what we did rather than a property of the narrative.
2026-07-22Fault Line Report NDAA card restructured after one leadership quote could not be verified at its cited source; subsequent first-party verification restored the mirror frame with different, verified quotes and the card was rescored.
2026-07-22Religious hate report: one verified item recoded as targeting both Jewish and Christian people, moving the anti-Christian account-level count from zero to one.
2026-07-22Weekly report window corrections before republication: one speech excluded as predating the window; two items included after full-text recodes.

Staffing, stated plainly

This is a very small operation, and the methodologies specify more reliability apparatus than the current staffing can run: second-coder reproducibility at scale, an ethics reviewer with sign-off authority, and blind re-adjudication of silent drops are designed and partially exercised, not fully operating. The model-based rounds above are what is achievable today, labeled as what they are. If you are a researcher or graduate student interested in co-coding against a documented instrument with published anchors and a stated formula, that collaboration solves the biggest open reliability gap on this page: jerm@hobocode.net.

Version register

InstrumentVersion
Hate methodologyv1.2, amended 2026-07-24 (§16 geography and converged narratives)
Polarization methodologyv1.0, adopted 2026-07-22; amended 2026-07-24 (§16; §15 publication-selection bias named)

Scoring constants (tier cutoffs, weighting exponents, marker thresholds) are conventions under the parameter-provenance discipline in both documents: tunable, dated on change, and any change lands in the corrections register above.