ChaosWatch
New here? Start here →
Evidence inspector

North America · 2026-07-04

The receipts. This is the exact material every model was shown for this day and region, followed by what each of them made of it — in its own words, next to the number it gave.

It is worth reading once even if you skip everything else: it is where confident-sounding reasoning and a score that turns out to be no better than chance sit side by side.

Why the evidence is identical, and how you can check

Evidence is gathered once per (date, region, depth_profile), hashed, and reused for every model in the run. A disagreement below is therefore a disagreement about identical input, not two models getting different search results. The evidence hash shown on this page is what makes that checkable.

Evidence hash
1645cb5a28c50a53562ef143b3fe818e130f295cf15eb2d03353c2e19fedcfb7

SHA-256 over the canonically serialised payload minus the hash itself, so it is stable across processes and machines. It is stored on the run and on every score row below — provenance is checkable, not asserted. Every score row below carries this same hash.

Source material · 10

shallow

These headlines, in this order, are the entire world the models were given. Nothing else about North America on 2026-07-04 was in the prompt.

03

The 10 Largest Cities in North America

Which is the biggest city in North America? Not New York. Explore the 10 largest, from Mexico City's sprawl to the metros challenging the old guard.

worldatlas.com

Extracts: not gathered at shallow depth — Stage 7 populates extracts only for standard and deep profiles. Zero here means "not requested", not "requested and empty". Full article texts: not gathered at shallow depth — only the deep profile fetches them.

The shallow profile requests 20 headlines and 10 are stored. That is a known provider cap, not a clipped payload: the Serper news endpoint returns at most 10 items per request regardless of the count asked for, and pagination has not been added. The three profiles currently differ in enrichment, not in headline breadth.

Model rationales

prompt p1

Same headlines, same rendered prompt, different readings. Where two models diverge, the disagreement is about interpretation, because the input was byte-identical and the hash proves it. The scored verdicts here span 17.0 points of the 0-100 scale.

claude-haiku-4-5
anthropic
32.0
okElevated

The evidence shows mixed conditions with concerning environmental and economic signals: extreme heat is identified as a serious public health threat with inadequate preparedness, and US hiring has slowed sharply in June, suggesting economic momentum is weakening. However, these challenges are partially offset by improved farmer sentiment and ongoing normal civic activities, preventing a higher instability classification on this Independence Day snapshot.

evidence 1645cb5a28c5…prompt p1
claude-opus-5
anthropic
22.0
okElevated

The evidence shows a largely routine, celebratory US news day around the 250th Independence Day anniversary, with the main stressors being a prolonged extreme heat dome and a sharp slowdown in June hiring that dented Fed rate expectations. No civil unrest, security incidents, or institutional breakdown appears in the record, so risk stays modest and confined to weather and economic softening.

evidence 1645cb5a28c5…prompt p1
gemini-3.1-pro-preview
google
20.0
okStable

The United States is experiencing a period of relative stability marked by peaceful nationwide celebrations for its 250th Independence Day and optimistic agricultural economic indicators. However, a severe and prolonged heat dome presents a notable environmental hazard, while a slight slowdown in hiring indicates mild macroeconomic cooling.

evidence 1645cb5a28c5…prompt p1
llama-3.3-70b-versatile
groq
20.0
okElevated

The evidence suggests a sense of normalcy and celebration in the US on its 250th anniversary, with various articles discussing travel, economy, and cultural events, but also touches on concerns such as extreme heat and its potential unpreparedness, contributing to a slightly elevated chaos score. Overall, the tone remains more positive and festive than alarming, hence the moderate score.

evidence 1645cb5a28c5…prompt p1
deepseek-reasoner
deepseek
15.0
okStable

The headlines reflect routine economic reports, holiday festivities, and weather warnings, with no indications of civil unrest or political breakdown. While extreme heat poses a public health risk, its impact is not reflected as chaos, keeping the overall assessment stable.

evidence 1645cb5a28c5…prompt p1

Every model returned a parseable verdict for this evidence, so no failure card is shown. Non-ok rows are rendered by the same component — with error_text where the rationale would be — and are covered by tests/rationale_card_test.tsx.