Hobocode Counter-hate Monitor
How we decide what counts as hate, how bad it is, and who we watch for it
This is the operating manual for a public instrument: the tests a post has to pass before we call it hate, the model we use to follow a narrative from first spark to decay, and the rules that govern which accounts we watch and why.
This section animates one worked example in eight steps: the post, the target, the implied claim, the technique, the decision, the score, the bridge, and the verdict. Each step's full text is in the captions that follow.
The method in practice
Follow one post through the method
The post, the account, and the match below are invented, and no real person is quoted. Every step applies the method exactly as the full text specifies it.
The post
Minutes after a match ends, this lands in a monitored stream. None of its words is a slur, and each clause on its own could pass for ordinary post-match anger. This is the region of language the method exists for: the place where legitimate criticism and coded hostility look identical on the surface (§5).
Typical. Every single one of them cheats, it’s who they are. Send them all back.
The target
The first question concerns the target. Criticism of a team’s conduct, however harsh, is protected and stays that way. Here, though, “them” resolves to the losing side’s nationality, which is a protected basis. From this point the burden of proof sits on the hate call (§5.1).
The implied claim
Before we flag anything, the implied claim has to be statable in plain form: they are dishonest by nature. If producing that sentence required adding hostility the text fails to supply, the item would be classed as criticism and dropped. In this case the sentence assembles from the words on the page (§5.1).
The technique
We name specific, testable techniques. This item carries three: a dispositional “they” that turns behavior into essence, collective attribution that expands one accusation to a whole people, and a policy warrant that converts the trait into a demand. That is the recode ladder climbing from behavior to trait to nationality to policy (§5.3, §4).
The decision
A classifier surfaced this item, and we treat its verdict as advisory, setting its self-reported confidence aside entirely. On a representative day of real traffic, automated hate detection averages 9.4% precision against the 87% that curated benchmarks suggest, so a human reads the item in full context and sets the confidence judgment (§5.5).
The score
We score severity on five anchored factors, reach, toxicity, coordination, mainstreaming risk, and harm potential, and combine them with a geometric mean, so a narrative has to be bad on several axes at once to rank high. This item scores low on reach and coordination, and its base score is 1.9. The harm weighting engages because the item cleared Track A, and on numbers this mild it barely moves the score, so the item logs at P4: recorded and watched passively, kept out of the daily queue by design (§8).
The bridge
Overnight, two similar posts appear. Their timing is scattered, they share no coined term or fabricated detail, and no account links them, so this is convergence: independent people reacting to the same event on their own. A coordination verdict requires a confirmed bridge account (§9).
The verdict
We publish the aggregate: one more item in a tracked narrative, currently at the framing-contest stage. We reserve naming for public figures and organized networks, so this small private account stays out of print. The subject of the monitoring is the narrative (§1, §12).
That was one pass through the whole pipeline. The full methodology holds every rule we used and the evidence behind each one.
Section §01
Standing principles
Seven rules hold on every pass, without exception, and everything else in this document is downstream of them.
We drop non-hits without a write-up, and we log genuinely borderline exclusions with sourcing, for transparency. The rule governs publication and leaves measurement intact: every pass keeps an aggregate screened count, and a sampled drop audit (§13) estimates what the silence hides.
Every hit needs a real, independently fetchable source with a URL and a confirmed date, read in full context. We have been burned once already by a hallucinated summary, and it will happen again if we skip this step.
The burden of proof sits on the hate call. The line is drawn at the target: hostility toward a group on a protected basis, or a confirmed proxy, versus conduct, ideas, or power.
We keep no capability to query "everything by or about group X" as a population; we query only "content exhibiting hallmark Y." We aggregate over spotlight, and we never reproduce a slur past the clinical minimum the lexicon needs.
We run no sock puppets and no bait accounts. A tool that shapes the behavior it records has stopped observing and started manufacturing its own evidence.
Public figures and organized networks get named, and that is accountability. A private individual's home address, employer, or family stays out of print, because publishing it would be doxxing.
Every naming decision passes through a human. A classifier's confidence, human or machine, stays advisory on the way to that decision; §5.5 shows why this matters most for the machines.
Section §02
Vocabulary discipline
We police our own vocabulary here: the terms we reach for, the bars they must clear, and the upgrade that retired "genuine vs. opportunistic" for good.
Loose language is how a monitoring project quietly drifts into advocacy. We use three families of terms consistently everywhere, in every hunt prompt and every published sentence:
2.1Units, movement, and measurement
actor · account · content item · narrative · frame · discourse · ideological ecosystem · network (reserved for a demonstrable connection; "several people said similar things" falls short) · community/milieu.
diffusion · cascade · amplification (name the mechanism: algorithmic, influencer, press, paid, coordinated) · organic reach (a claim about payment alone) · mainstreaming vs. normalization · narrative laundering.
prevalence · volume ≠ reach ≠ exposure ≠ impact · engagement (state which actions count) · virality (it takes more than a view count to establish) · the base-rate problem.
2.2Terms that need to clear a bar before publication
| Term | Bar to meet |
|---|---|
| "Both sides" | Say specifically what was symmetrical and what wasn't. |
| "Organic" | Means unpaid distribution, and only that. |
| "Coordinated" | Requires evidence of actual arrangement; similar posts landing near each other fall short. |
| "Radicalized" | Don't collapse belief, identity, participation, and violence-preparation into one word. |
| "Extremist" | Needs a stated definition and a shown connection between the actor and the behavior. |
| "Hate group" | Name whose designation methodology produced the label. It's never self-evident. |
| "Dog whistle" | Needs contextual evidence of an encoded meaning and an intended audience, both. |
| "Went viral" | Replace with the actual observed scale and the actual observed time. |
| "Led to" / "sparked" | Reserved for downstream harm that an outtake_evidence value backs (§4.1). "Was followed by" covers everything else. |
pre_existing → activated_by_event → adapted_to_event → weakly_event_linked → event_generated
fabricated → unsupported → mixed → well_supported
Section §03
The Ember Model: a narrative's lifecycle
A hate narrative is a living object. It sparks, it spreads, it decays, and what it leaves behind lowers the ignition energy for whatever comes next.
Three things about that lifecycle are load-bearing enough to shape everything downstream. The unit of analysis is the frame: a compressed claim about what an event means. One event spawns several competing frames at once, and posts stand as evidence of a frame's health. Stages form a state machine: a narrative can skip, stall, loop, or die at the framing contest, which is where most of them end, and triage is the model's first job. And decay leaves embers: a burned-out narrative deposits a reusable lexicon and a dormant network that lower the activation energy for the next matching spark, which is why a narrative can be grafted: recruited into a larger host narrative and scored at the host's own downstream stage. That inheritance has an evidence bar: two of three among lexicon overlap with the host's term set, carrier overlap with the host's amplifiers, and explicit bridging content, recorded on the card. Suspicion alone never inherits the host's posture.
Eight stages, one trace
click a stage for its signals, its measurable indicator, and its intervention window
3.2Confirming a cross-platform jump takes more than one crossover
Stage S5 requires plural, independently-acting carriers reinforcing the same lexicon in the same rolling window (72 hours as a working default); a single amplifier hopping platforms falls short of it. That's a well-corroborated finding: Centola & Macy's complex-contagion "wide bridges" result1 holds up on real political-hashtag data2 and real far-right radicalization case data3 alike. We don't hard-code the count itself: the plurality requirement is evidenced while any exact count would be a guess, and an analyst keeps an override path for one ultra-high-reach account that has unambiguously adopted the frame on its own. A single bridge with no plural reinforcement sets an interim S4_to_S5_watch status. That status answers a genuinely different question from the astroturf checklist's "confirmed bridge" signal in §9: one bridge is enough to call covert coordination, and it takes plural carriers to call genuine cross-platform uptake.
3.3Amplifier tracking doesn't default to follower count
A frame carried by a critical mass of ordinary mid-tier accounts is often a stronger durability signal than one large account's momentary spike, which decays faster4. Every narrative card codes a carrier_pattern: single_large_account distributed_mid_tier mixed.
linked_narratives populated when a later spark grafted onto it, and did the recorded reignition latency match the true time-to-reignition?
Section §04
Coding schema
We code fourteen dimensions on every item that clears either detection track.
Reading this as one flat JSON object hides what it's actually doing, so we grouped the fields here by the job each one does.
narrative_familytargetrhetorical_operation: sixteen named narrative families (replacement/invasion, hidden hand, collective guilt, conspiracy fusion, and eleven more) and five possible targets, from individual up through ideological opponent.
coding_methoddeniability_mechanism: explicit through distributed; ten named deniability mechanisms from innocent-literal-meaning to criticism-as-censorship, so a coder has to name the exact escape hatch being used.
evidentiary_groundingevent_relationship: the two-axis replacement from §2.2, required on every item, and kept as two axes.
harm_levelevidence_typeconfidenceclassifier_provenance: six harm levels from biased to calling-for-violence; primary item through inferred association; low/moderate/high, always human-adjudicated; plus model, prompt version, and date on every machine-assisted classification.
outtake_evidenceobserved_outcome: an item cannot claim any outcome past exposure or engagement unless outtake evidence is populated with something real. See below, this is the field that closes an actual hole.
distribution_stagerecode_ladder_stagegraft_operation: the six-point distribution ladder, the five-move recode ladder (behavior → trait → nationality → immigration category → policy warrant), and whether this item is itself a graft attempt onto a host narrative.
4.1What the grounding and outcome fields close
outtake_evidence gates observed_outcome because a record could otherwise jump straight from "someone saw it" to "it caused harassment" with no checkpoint asking whether the audience noticed, recalled, or repeated the narrative first. We modeled the gate on AMEC's Integrated Evaluation Framework, which inserts exactly that stage between Output and Outcome6. recode_ladder_stage and graft_operation turn the project's best piece of analytical work, the five-move recode ladder worked out case by case, into a field that gets coded every time instead of re-derived once per case study. Two further fields close the loop: classifier_provenance gives §5.5's "log which model produced this" mandate a field to live in. And the outtake gate carries a stated cost: with degraded search access, real outcomes go unrecorded for visibility reasons alone, so outcome-frequency claims and cross-narrative outcome comparisons carry an undercount caveat.
observed_outcome value beyond exposure or engagement unless outtake_evidence holds something other than engagement_only_no_registration_confirmed. "Went viral, so it must have caused X" is exactly the inference this field exists to block.
Section §05
Detection method
The signal that matters lives exactly where legitimate criticism, in-group banter, and coded hate are lexically indistinguishable. No single classifier survives that region, so nothing here is one.
We run both tracks on every item, plus an independent coded-language pass that reads past whatever either gate finds on the surface text. A "no" on Track A means "check the other track"; extremism overall remains an open question at that point.
5.1Track A: protected-group hate
Track A gate
- Dehumanization / inferiority
- Whole-group generalization
- Threat or incitement
- Coded term with a real cue
- Grievance / victim-reversal
- "Just asking questions" scaffolding
- Target is policy, institution, or power
- Identity terms present, stance supportive/neutral
- Counter-speech quoting to condemn
- In-group or reclaimed use
- Specific-conduct claim only
- Satire aimed upward at power
The articulability test is the single highest-precision tool here: the classifier has to state the implied hateful claim in <target> {are/do/commit} <predicate> form. If that sentence can only be built by adding hostility the text and context fail to supply, it's criticism and it stops there.
5.2Track B: political and ideological-outgroup extremism
Gate: is the target a political party, ideological movement, or a proxy for one? If so, check for a dehumanization metaphor (disease, vermin, parasite, "root out," predator/prey; §6 covers the animalistic/mechanistic split), a collective-attribution synecdoche with a criminal or subhuman category swap, an incitement construction that restructures a threat as a prediction, or an eliminationist register whose natural end-state is removal or destruction, past any normal political contest. A trip classifies as political_outgroup_dehumanization or incitement_or_violence_normalization and gets reported under the extremism label. A political-idiom guard: the metaphor check fires on dehumanization applied to people, and stock idiom applied to conduct passes through: "root out corruption" targets a practice and stays protected; "root out these people" targets humans and trips. The gold set carries hard-protected political exemplars so this boundary stays measured (§13).
5.3The coded-language pass, independent of both gates
This pass runs regardless of whether either track's surface-text gate trips. A "coded" classification stands without an explicit predicate or a named group; that is definitionally what coded means.
| Technique | Detection handle |
|---|---|
| Dogwhistle | Watchlist only; fires only paired with a target reference plus one other technique. |
| In-group euphemism | Evasive referent: demonstrative pronoun, no antecedent, evaluative predicate. |
| Dispositional "they" | Universal quantifier plus trait predicate applied to a nationality. |
| Dehumanization metaphor | Metaphor-domain lexicon applied to a human group. Highest precision of the ten. |
| Numeric / emoji / leet codes | Exact-match list plus emoji sequence plus numeric handles. |
| Ironic deniability humor | Joke-frame wrapped around a slur; "can't take a joke" as the follow-move. |
| Algospeak / char substitution | Normalize (unicode-fold, deleet, dehomoglyph) before fuzzy matching. |
| Adjacent-topic laundering | Topic shift: the premise is sports, the conclusion is immigration, same target throughout. |
| Collective attribution / synecdoche | Actor-set expansion: "some fans" silently becomes "a nation." |
| Pseudo-civility framing | Dual-use; codes as hate only once fused with dehumanization or collectivization. |
A technique family is necessary, and a hostile predicate completes it. The predicate can be carried by construction: a known replacement-rhetoric template, a self-avowed motive recycled approvingly, a documented authorial pattern. A satire register is not an exemption if the piece names protected classes and frames their protections as illegitimate.
5.4The dangerous-speech supplement
For anything scored glorifying_violence or calling_for_violence, add Susan Benesch's Dangerous Speech Project hallmarks on top7: a narrower question, "does this raise real mass-violence risk," used only at the top of the harm scale: dehumanization, accusation in a mirror (a real, sourced term: the anonymous propaganda manual found in Butare, Rwanda after the 1994 genocide instructing "the party which is using terror will accuse the enemy of using terror," in Des Forges' rendering of the manual8), a threat to group purity, an assertion of attack on women or children, and questioning in-group loyalty.
5.5LLM-classifier discipline
- Real-world performance runs far below benchmark performance: in the HateDay study, models averaged 9.4% precision on a representative day of real Twitter traffic, against roughly 40% on curated academic benchmarks and 87% on HateCheck9. A human always makes the final classification.
- Identity-term over-flagging persists in LLMs beyond the older tools like Perspective API: current models still misclassify benign statements that mention protected groups as hate1011. Bigger, multimodal models genuinely do read context better, but demographic and lexical biases persist, especially in smaller models; context-sensitivity and fairness remain separate claims12.
- A model's self-reported confidence does not track its accuracy. Treat it as decorative. A separately human-adjudicated confidence judgment stays authoritative.
- Persona and role framing in a prompt measurably shifts classifier output, and the more heavily safety-aligned a model is, the better it resists that steering13, so we keep prompts minimal and test them periodically for sensitivity to small wording changes.
- The same model API can silently change behavior when a vendor updates policy. Log which model/version produced each classification; re-score a fixed set of past items against the current model periodically. The schema's
classifier_provenancefield is where that log lives. - The gold-set agreement baseline covers the verdict alone. A double-coded sample across all fourteen schema dimensions and stage assignment, repeated after any schema change, is what tells us whether two coders can operate this instrument the same way; a field that resists consistent coding becomes a collapse candidate.
Section §06
Sharpening dehumanization detection
Track B's dehumanization check used to treat dehumanization as one bucket. It's actually two, with different psychological correlates.
Denies uniquely-human traits. Vermin, insect, disease, and bestial language sit here: the register that strips a target of refinement, culture, and self-restraint.
Denies human-nature traits instead. Object, machine, and automaton language sits here: the register that strips a target of warmth, emotion, and agency.
Haslam's 2006 integrative review is where this split comes from14, and it's coded every time a dehumanization flag fires. Kteily et al.'s validated 0–100 "Ascent of Man" scale15 can inform an ad hoc toxicity judgment. It was built and validated as a survey self-report instrument, so any adapted use on rhetorical artifacts is a heuristic.
Section §07
Reach & virality measurement
We try the tiers in strict order and use the first that produces a real, citable number. Tiers stay separate, and we label a low tier as exactly what it is: a ceiling.
Tiered lookup, try in order
X/Twitter and YouTube expose public view counts directly. We cite the post URL and the date checked, a floor as of that date. Labeled platform-reported, with the platform's own definition of a "view" attached where known (X counts roughly two seconds of dwell): a floor over time; the count of unique humans reached stays unknown.
A Nielsen-sourced ratings roundup for the broadcast window. Labeled explicitly as the show's average; the segment's real audience remains unmeasured.
Tranco domain rank16, via tools/reach_lookup.py. Anything outside roughly the top 500,000 global rank is flagged as upper-bound-leaning.
An upper bound on potential audience, only when tiers 1–3 are empty, always labeled as a ceiling, and kept free of engagement-rate multipliers that would manufacture a synthetic reach figure.
We say so directly and leave the estimate blank.
7.3Trajectory as a secondary, corroborating signal
tools/reach_lookup.py computes a trajectory ratio from GDELT's narrative timeline17 (rate-limited to one request per five seconds): two adjacent windows, recent divided by prior. Output always carries a confidence string, literally "provisional (2-point ratio)" at two observations and "trend (n>=3 windows)" beyond that, and the string travels with every printed value. This corroborates the EPS trajectory multiplier in §8; it doesn't replace the primary signal, since GDELT tracks external media coverage, which most early-stage or fringe-only narratives never register on at all.
Section §08
Severity scoring: the Ember Priority Score
We combine five factors with a geometric mean, which damps compensation without removing it (a 5 still pulls up a 2): a narrative has to be bad on more than one axis to rank high, and a narrative bad on one axis alone can be suppressed, which is why the floor exists.
Five factors, illustrative reading
| Factor | 1 | 3 | 5 |
|---|---|---|---|
| Reach | <1k, one small community | ~100k, single platform | Multi-platform, millions, media pickup |
| Toxicity, read as intent | Heated, non-dehumanizing | Sustained othering / coded terms | Explicit dehumanization or a call for harm |
| Coordination | Organic, scattered | Timing clusters, shared hashtag | Confirmed network, bots, cross-platform ops |
| Mainstreaming-risk | Confined to fringe | Active fringe-to-mainstream bridges | Already adopted by mainstream figures |
| Harm-potential | Discourse only | Identifiable target, doxxing risk | Concrete target + time/place + mobilization |
Coordination also carries a defined anchor at 4, documented co-platforming: a named event, shared stage, or common organizing hub transmitting the same narrative, without evidence of a persistent cross-platform operation. The factor measures degree of organization; whether the organization was hidden belongs to the coordination verdict in §9, and overt organization scores as high as covert.
The weighted variant weights harm and mainstreaming up; concretely it is (R·T·C·M1.3·H1.5)1/5.8, the normalized weighted geometric mean. It auto-triggers the moment any item in the narrative has actually cleared Track A or B, a deliberate, stated carve-out that closes a real inter-analyst gap where the choice used to be pure discretion. scoring.weighting_applied gets recorded on the card so a tier assignment stays auditable after the fact.
8.3Priority tiers
Immediate. Continuous monitoring, escalate to response owners.
Active. Worked daily, card kept current.
Watch. Reviewed weekly, alert thresholds set.
Log only. Recorded, passively monitored.
Ordinary rivalry, satire, and heated-but-non-dehumanizing discourse are supposed to land on P4 and stay logged, resting outside the daily queue; the rubric exists to keep those out of the queue as much as to flag real danger. A grafting narrative scores at its host's downstream posture once it clears the graft evidence bar (§3), which is exactly what pushes it up the queue before its own metrics would justify that on their own.
Two rules close a gap the base formula leaves open. A small-account item carrying a concrete, explicit threat (Reach 1, Toxicity 5, Coordination 1, Mainstreaming 1, Harm 5) computes to roughly 1.9 base, 2.0 weighted, and 1.2 under a decaying τ, which lands on "log only" for a post naming a target, a time, and a place. So: harm-potential 5 now sets a P2 floor regardless of the computed score, recorded as scoring.floor_applied so the quarterly distribution audit can see how often formula and floor disagree. And the §12 escalation ladder runs on its own thresholds, independent of the EPS tier in both directions: the EPS ranks the monitoring queue, and emergencies travel the ladder.
8.2The trajectory multiplier τ
Computed primarily from the narrative card's own already-tracked rate-of-change field, zero cost, always available: above +15% per window is accelerating (τ = 1.3), between −15% and +15% is stable (τ = 1.0), below −15% is decaying (τ = 0.6): illustrative starting bands, tunable after real weeks of data. A low-confidence two-window accelerating reading needs corroboration from at least one other flipped indicator before it can push a narrative into P1. τ applies as a genuine post-hoc multiplier on the already-computed base score, kept outside the geometric mean itself, the same separation cybersecurity keeps between CVSS severity and EPSS exploitation-probability1819, on purpose. One caveat: the rate-of-change field is computed over the open-platform footprint only (§11), so a narrative accelerating on X, TikTok, or Meta can read as stable here. τ is footprint-bounded: divergence against a tier-1 lookup or the GDELT check gets recorded for judgment, and τ computed across a surge-mode collection boundary (§11) always carries a flag.
Section §09
Coordination & astroturf detection
Coordination is neither hate nor inauthenticity by itself. Activists, fandoms, and breaking-news crowds converge on shared language all the time, with nothing hidden about it.
9.1Signals, checked in this order
- Named event, willing co-participants. Overt coordination, as a documented fact. Name the actual hub; score it as one event-narrative card, since nothing was hidden.
- The same triggering event, at finer grain than a shared topic. A mixed-trigger cluster gets split explicitly; scoring it as one is a methodology error.
- A shared semantic payload, checked at the embedding level: the same novel claim, coined term, fabricated detail, or unusual framing sequence recurring across accounts, with shared errors counting most of all. Lexical near-duplication still counts when it appears, and a language model makes rephrasing free, so the absence of matching phrasing now carries almost no exculpatory weight.
- Synchronized timing. A tight co-timed window, hours to a couple of days, is a real signal. A multi-week scatter is not, by itself.
- A bridge account that actually cross-links or reposts, the single highest-value check of the five. We always run it before finalizing an "organic parallel" verdict.
| Pattern found | Verdict | Card type |
|---|---|---|
| Named event, willing co-participants, hub identifiable | Overt coordination | One event-narrative card; name the hub |
| Same trigger, tight window, shared semantic payload, bridge confirmed | Covert coordination | Coordination-network card; name the bridge |
| Same topic, scattered timing, idiosyncratic phrasing, no cross-link | Organic parallel | Individual hits on the highest-value item |
| Same trope, unrelated messengers, no cross-link | Notable convergence | Convergence card, explicitly labeled as such |
Cross-lane convergence is a real, distinct finding on its own, and it becomes coordination only when signal 5 independently turns it up. A vendor's coordination or bot-probability score, however high it reads, serves as supplementary data; a confirmed named bridge remains the required evidence.
Section §10
Standing rejections
We have rejected these marketing and growth-hacking techniques in writing, each one tempting precisely because it looks like harmless self-measurement.
No hook-rate, retention-curve, or K-factor score may advance distribution_stage or substitute for Reach, Coordination, or Mainstreaming-risk. These metrics carry no harm/polarity term: a well-produced hateful post and a well-produced benign post register identically as "high-performing content."
These are undisclosed black boxes with no published accuracy figures. One such tool sits in the project's own tool stack and must never touch coded-language triage, an EPS factor, or a reach estimate; its output may describe production craft only.
Virality science treats ambiguity and novelty as accelerants to maximize, exactly the property dogwhistles exploit to preserve deniability. Using one to triage the other would silently invert the coded-language pass's whole purpose.
Not even during a live spark/framing-contest event. A tight reaction window is evidence that a single skilled actor moved fast, and on inspection it cuts the wrong way in the checklist's timing signal if used as a calibration anchor.
No hook rate, no retention engineering, no Growth-Loops design, in any form, including reader-retention analytics framed as innocent self-measurement. A de-amplification-first publication chasing its own virality inverts the project's own operating principles.
Permitted only over public amplifier accounts that already independently meet the naming bar. The audience and the targeted community's own network stay out of scope, permanently, as does any general-purpose version of the capability.
Section §11
Data collection architecture
The open-API era is over. Every source here is partial, biased, or sampled, and the honest posture is a portfolio with disclosed gaps.
Reshare-tree data and hook-rate/retention data are both, independent of desirability, largely uncomputable on this exact stack; Bluesky is the one platform where the former is realistically free. Reference pipeline: per-source collectors write append-only raw payloads to object storage, normalization feeds Postgres + pgvector, enrichment runs lexicon match / toxicity / embeddings / coordination signals, and analysis surfaces narrative clustering, actor graphs, spread velocity, and cross-platform hops into a dashboard.
Scale-up path: Junkipedia27 (Truth Social, Gab, GETTR, Rumble, BitChute, VK, with no engineering needed) → a university or nonprofit partner for TikTok Research API / Meta Content Library → funding for a commercial listening seat or metered X → EU vetted-researcher status under DSA Art. 4028 → image-hash matching for cross-platform media tracing.
Surge mode: when a live spark pushes the graft check to hourly (§13), collection surges on a pre-stated plan instead of improvised polling: a reserved share of each rate-limited budget, event collectors first, then P1/P2 cards, with roster-wide baseline polling throttled explicitly and on the record. GDELT stays at its one-request-per-five-seconds limit regardless; it's a corroborating signal on its own timescale. Surge start and stop get logged on the affected cards, because a volume jump at a surge boundary reflects the instrument's own movement.
Section §12
Ethics, legal & safety guardrails
These are functional requirements, load-bearing from the start. Nothing we run ships without them.
Eight operating principles: watch the hate, not the hated · minimize by default · observe, don't provoke · aggregate over spotlight · punch up, not down · assume the tool will be attacked or misused · reversibility is a feature · independent oversight gates anything that names a person.
12.2The escalation ladder
The lowest sufficient channel, tried first, always.
DSA Art. 22 routes28, where the platform recognizes one.
Domain experts equipped for a specific harm type.
Never download or store. Preserve pointers only, report to NCMEC.
Publication ethic: strategic silence and de-amplification by default; quantify and contextualize, with reproduction held to the clinical minimum; strip PII; no target lists, no evasion playbook, no manifesto text; delay real-time findings that could steer a live campaign; coordinate disclosure with affected communities.
Researcher wellbeing: hard exposure limits and time-boxing on the worst queues, rotation, no lone-working on the worst material, aggregate-first views, blur/grayscale rendering, trauma-informed training.
Section §13
Cadence, roles & KPIs
Even a two-person team can run this, as long as the roles and the rhythm are written down and explicit.
Owns the queue, triage, and the daily graft-tripwire check.
Collectors, enrichment pipeline, uptime.
Human-confirms every candidate coded term before it enters the lexicon.
Sign-off authority over any naming or escalation decision. A two-person team borrows this role in.
13.2Rhythm
- Daily: we work the P1/P2 cards and run the graft-tripwire check hourly instead of daily during a live spark/framing-contest event, per our own Euro 2020 finding that hate-narrative latency runs in minutes.
- Weekly: we review the P3 watch list, re-score, and prune P4.
- Per-event: we post-mortem the stage map and the reignition-latency estimate.
- Quarterly: we run the loop-back audit (§3), the EPS score-distribution audit that checks for unnatural clustering at tier boundaries, and the LLM-vs-human agreement disclosure.
The one first-class KPI: false-positive rate on the hard-protected gold set. We reject outright any recall gained by regressing here, and this metric outranks every other on the list on the list, which also includes time-to-detection, lexicon precision per target group, share of escalations meeting the full threshold, and analyst-exposure hours against wellbeing limits.
A false-negative drop audit: every pass records an aggregate screened count, and a quarterly random sample of silent drops gets re-adjudicated blind, with an estimated miss rate published next to the false-positive figure and the sampled content deleted after adjudication. Without this audit, the FP-first KPI plus silent dropping would make conservative drift invisible by construction. A gold-set composition rule: the hard-protected set must include heated-but-protected political speech, so the first-class metric exercises Track B's boundary too, and the distribution of Track B flags across lanes gets disclosed periodically. And a degradation order: when the queue exceeds measured capacity, P3 review slips first, then discovery, then P4 logging; P1/P2 work, escalation checks, and the human gates hold their place at every load level. A parameter provenance register marks every working constant as evidenced, convention, or operational, with a dated changelog on any change.
Section §14
Open questions & limits
We carry these forward deliberately, on the record.
- Full coverage is structurally impossible: the three highest-value surfaces, X, TikTok, and Meta, are also the three hardest and most expensive to access.
- Cluster granularity and stance on isolated posts remain unsolved knobs; target resolution and irony detection genuinely fail without conversational context.
- Lexicons rot within days under adversarial drift. Techniques and disambiguation signals are the durable feature; human confirmation stays non-negotiable.
- Academic and regulatory access (TikTok's Research API, Meta's Content Library, DSA Art. 40) is slow and enclave-bound: useful for rigor and history, with real-time alerting beyond its reach.
- The hardest ethical tension has no resolution in general: the same capability that catches a coded campaign early is the capability a bad actor most wants. The mitigations reduce it, and a residue remains.
- We flagged these for future research and left them unfilled from memory: the Rabat Plan of Action's six-part incitement threshold, the most on-point legal framework for
harm_levelnot yet incorporated; SPLC, CCDH, HOPE not Hate, and NCRI as operational analogs; the UK Online Safety Act 2023; the algospeak evasion literature underlying the char-substitution technique this schema already names.
Section §15
Roster construction & the sampling frame
Everything above governs how we code an item once we collect it. This section governs which accounts we watch in the first place.
The roster's own metadata said the quiet part plainly: a seed list, parsed and deduped, no rigor applied to why any specific name was chosen. The project's own cautionary tale for this exact failure mode is CCDH's "Disinformation Dozen" report: NPR reported Facebook's objection that "it was not clear what criteria the group used" to build the list32, and Meta's rebuttal put the twelve accounts at about 0.05% of vaccine-content views on its platforms33.
15.2The gate every entry has to clear
A specific, dated, independently verifiable instance that would itself pass Track A or B: a citable post, statement, or documented affiliation, with a date and a source.
A checkable measure of real standing: a stated follower threshold, a named leadership role, or recurring citation by other already-vetted accounts.
Independent citation · a confirmed network tie · unambiguous category fit · inside the 24-month lookback window.
15.3Discovery: models supply leads, and primary sources decide
- We anchor on a real, dated event: a rally, a filing, a deplatforming, a viral incident with press coverage.
- We snowball from primary sources: livestream credits, press bylines, court filings, designated-org lists, and the public network of already-vetted accounts.
- We triangulate across at least two independent sources before a name becomes a candidate at all; a model counts as at most one lead.
- If we use a model at all, it supplies a lead to verify independently: we query it about a specific dated event, keep open "list accounts in category X" prompts out of the workflow, and treat non-convergent answers across providers as extra-scrutiny candidates.
- We actively search for the tail. Popularity-driven discovery, model or otherwise, systematically favors the already-famous; if every candidate we find is someone already recognized, the pass reproduced head-bias and found nothing new.
15.6Two additional categories
Both sit in Track B: the target class is an institution or ideological system and its human proxies. Neither gets a symmetric paired lane, the same way Anti-Semitic & Syncretic Conspiracy Accounts has no paired opposite already.
Techno-pessimist / anti-civilization. Kaczynski-lineage ideology, insurrectionary anti-tech sabotage networks, the AI-doomer "Zizian" violent splinter (tied to six deaths across 2022–202534), right-coded "Great Reset" conspiracism.
Explicitly protected, outside this lane's scope: mainstream AI-safety advocacy (PauseAI, Stop AI, both explicitly nonviolent by their own stated commitments3536), AI regulation debate, automation/job-loss criticism, environmental objections to data centers, e/acc and neoreaction commentary.
Scoped deliberately narrow to the medical-tyranny/antigovernment strain, which carries essentially all the documented violence: the August 2025 CDC headquarters shooting37, the Whitmer kidnapping plot's lockdown grievance38, armed intimidation of health boards.
Explicitly protected, outside this lane's scope: ordinary vaccine hesitancy, informed-consent advocacy, mandate-policy criticism.
15.9The circularity problem, and the two-tier fix
Gate 1 requires a documented, rubric-clearing act as a condition of inclusion. That's the right design for a monitoring watchlist, because a watchlist should contain known offenders. It is the wrong design for comparing prevalence across ideological lanes: an account qualifies for the roster because it already produced a qualifying example, so reporting what share of roster accounts produce one in a given period reports a recidivism rate in a population pre-selected on the outcome being measured; an estimate of how much hate exists in that community would require a differently built sample. This isn't hypothetical: the religion-narratives special report's 73% and 78% antisemitism/Islamophobia hit rates were checked against the original ungoverned seed list, while its 2.6% anti-Christian figure came from a hand-supplied roster selected on identity alone. That gap may reflect a real difference between lanes. It may also just reflect that one roster was built the right way for this comparison and the other wasn't; both explanations can be true at once, and the old construction method couldn't tell which was doing how much.
Gate 1 exactly as written. Valid for what a watchlist is actually for: tracing a narrative's movement, tracking recidivism and escalation, feeding the EPS and coding-schema work this project runs every week.
For any "X% of this lane does this" or "lane A is worse than lane B" claim it stays out of bounds on its own, because inclusion already required doing it once.
Gate 1 replaced with verified lane membership (a real, checkable reason the account belongs to the ideological current the category names), with no requirement that it has ever cleared Track A/B. Gate 2 stays unchanged.
Valid for an actual, if still bounded, prevalence estimate, since whether the account ever clears the rubric is now a genuinely measured outcome.
Appendix
References
Every externally sourced claim on this page carries a numbered marker that resolves here.
- Centola, D. & Macy, M. (2007). "Complex Contagions and the Weakness of Long Ties." American Journal of Sociology 113(3), 702–734. journals.uchicago.edu/doi/10.1086/521848
- Romero, D. M., Meeder, B. & Kleinberg, J. (2011). "Differences in the Mechanics of Information Diffusion Across Topics: Idioms, Political Hashtags, and Complex Contagion on Twitter." Proceedings of WWW 2011, 695–704. archives.iw3c2.org (PDF)
- Youngblood, M. (2020). "Extremist ideology as a complex contagion: the spread of far-right radicalization in the United States between 2005 and 2017." Humanities and Social Sciences Communications 7, art. 49. nature.com/articles/s41599-020-00546-3
- Watts, D. J. & Dodds, P. S. (2007). "Influentials, Networks, and Public Opinion Formation." Journal of Consumer Research 34(4), 441–458. academic.oup.com
- Balfour, B. (2018). "Growth Loops are the New Funnels." Reforge. reforge.com/blog/growth-loops
- AMEC. "Integrated Evaluation Framework." amecorg.com/amecframework (accessed 2026-07-21)
- Dangerous Speech Project (Benesch, S., et al.). "Dangerous Speech: A Practical Guide." First published 2018; last revised 2024-08-20. dangerousspeech.org/libraries/guide
- Marcus, K. L. (2012). "Accusation in a Mirror." Loyola University Chicago Law Journal 43(2), 357–393. lawecommons.luc.edu (the manual's English wording follows Des Forges, Leave None to Tell the Story, 1999, as quoted by Marcus)
- Tonneau, M., Liu, D., Malhotra, N., Hale, S. A., Fraiberger, S. P., Orozco-Olvera, V. & Röttger, P. (2025). "HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter." Proceedings of ACL 2025, 2297–2321. aclanthology.org/2025.acl-long.115
- Dixon, L., Li, J., Sorensen, J., Thain, N. & Vasserman, L. (2018). "Measuring and Mitigating Unintended Bias in Text Classification." Proceedings of AIES '18, 67–73. dl.acm.org/doi/10.1145/3278721.3278729
- Zhang, M., He, J., Ji, T. & Lu, C.-T. (2024). "Don't Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech Detection." Proceedings of ACL 2024. aclanthology.org/2024.acl-long.652
- Davidson, T. (2026). "Multimodal large language models can make context-sensitive hate speech evaluations aligned with human judgement." Nature Human Behaviour 10, 514–530. nature.com/articles/s41562-025-02360-w
- Selvaganapathy, S. & Nasim, M. (2025). "Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection." arXiv:2509.00673. arxiv.org/abs/2509.00673
- Haslam, N. (2006). "Dehumanization: An Integrative Review." Personality and Social Psychology Review 10(3), 252–264. journals.sagepub.com
- Kteily, N., Bruneau, E., Waytz, A. & Cotterill, S. (2015). "The Ascent of Man: Theoretical and Empirical Evidence for Blatant Dehumanization." Journal of Personality and Social Psychology 109(5), 901–931. pubmed.ncbi.nlm.nih.gov/26121523
- Le Pochat, V., Van Goethem, T., Tajalizadehkhoob, S., Korczyński, M. & Joosen, W. (2019). "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation." Proceedings of NDSS 2019. tranco-list.eu
- The GDELT Project. "GDELT DOC 2.0 API" (timeline volume of global news coverage). blog.gdeltproject.org
- FIRST.org (2023). "Common Vulnerability Scoring System v4.0: Specification Document." first.org/cvss/v4.0
- FIRST.org. "Exploit Prediction Scoring System (EPSS)." first.org/epss
- CISA. "National Cyber Incident Scoring System (NCISS)." cisa.gov
- World Health Organization (2024). Emergency Response Framework, ed. 2.1. who.int/publications/i/item/9789240058064
- Reed, C., Biggerstaff, M., Finelli, L., et al. (2013). "Novel Framework for Assessing Epidemiologic Effects of Influenza Epidemics and Pandemics." Emerging Infectious Diseases 19(1), 85–91. wwwnc.cdc.gov
- Wired (2023). "Twitter Data API Prices Out Nearly Everyone." wired.com (corroborated by TechCrunch, 2023-03-29)
- TechTarget (2023). "Reddit pricing: API charge explained" ($0.24 per 1,000 calls, effective 2023-07-01). techtarget.com (the $12k/yr figure in the table is a derived illustration at ~50M calls/yr)
- TikTok for Developers. "Research API." developers.tiktok.com (accessed 2026-07-21)
- ICPSR / Social Media Archive (SOMAR), University of Michigan. "Meta Content Library." icpsr.umich.edu/sites/somar (accessed 2026-07-21; the ~$371/mo fee is SOMAR's Virtual Data Enclave charge per research team from January 2026; UI access and Meta's own secure research environment carry no fee)
- Junkipedia (Algorithmic Transparency Institute / National Conference on Citizenship). junkipedia.org (accessed 2026-07-21)
- Regulation (EU) 2022/2065 (Digital Services Act), OJ L 277, 27.10.2022; Art. 22 (trusted flaggers), Art. 40 (data access and scrutiny). eur-lex.europa.eu
- Counterman v. Colorado, 600 U.S. 66 (2023). supremecourt.gov (PDF)
- Brandenburg v. Ohio, 395 U.S. 444 (1969). law.cornell.edu
- Bellingcat & Global Legal Action Network, Justice & Accountability Unit (2022). Online investigations methodology manual (updated 2022-12-14), "Algorithmic Effects" annex. bellingcat.com (PDF) (the unit wound down in July 2025; the work continues at GLAN as DRAGNET)
- Bond, S. (2021). "Just 12 People Are Behind Most Vaccine Hoaxes On Social Media, Research Shows." NPR, 2021-05-13. npr.org
- Bickert, M. (2021). "How We're Taking Action Against Vaccine Misinformation Superspreaders." Meta Newsroom, 2021-08-18. about.fb.com
- Associated Press (2025). "A timeline of cultlike 'Zizian' group tied to killing of a Border Patrol agent in Vermont." Via PBS NewsHour, 2025-02-18. pbs.org/newshour
- PauseAI. "Values" ("We will never use, encourage, or tolerate violence"). pauseai.info/values (accessed 2026-07-21)
- Stop AI ("democratic and non-violent methods"). stopai.info (accessed 2026-07-21)
- Associated Press (2025). "Shooter attacked CDC headquarters to protest COVID-19 vaccines, authorities say." Via PBS NewsHour, 2025-08-12. pbs.org/newshour
- US Department of Justice (2020). "Six Arrested on Federal Charge of Conspiracy to Kidnap the Governor of Michigan," 2020-10-08. justice.gov; lockdown-grievance motive per NPR/AP trial coverage, 2022-04-08. npr.org
- US Department of Homeland Security (2020). Homeland Threat Assessment, October 2020. dhs.gov (PDF)
- FBI & DHS (2021). Strategic Intelligence Assessment and Data on Domestic Terrorism, May 2021. fbi.gov (PDF)
Fuller working bibliographies (ISD, Moonshot, ADL, HateLab, Latent Hatred, Silent Signals, HateCheck, the Santa Clara Principles, the Menlo Report, and the complete Mexico–England case sourcing) live in the project's internal research files.
Hobocode Counter-hate Monitor. We revise this methodology as our detection, scoring, and reach-measurement practice changes; it is the same standard we apply to every Weather Report bulletin and special report.