ChaosWatch
New here? Start here →
Evidence inspector

North America · 2026-06-27

The receipts. This is the exact material every model was shown for this day and region, followed by what each of them made of it — in its own words, next to the number it gave.

It is worth reading once even if you skip everything else: it is where confident-sounding reasoning and a score that turns out to be no better than chance sit side by side.

Why the evidence is identical, and how you can check

Evidence is gathered once per (date, region, depth_profile), hashed, and reused for every model in the run. A disagreement below is therefore a disagreement about identical input, not two models getting different search results. The evidence hash shown on this page is what makes that checkable.

Evidence hash
39d4198b0d4edfe3ed7e1038d14dc5b379b4deab2781fbaafc60f6b9100d34b7

SHA-256 over the canonically serialised payload minus the hash itself, so it is stable across processes and machines. It is stored on the run and on every score row below — provenance is checkable, not asserted. Every score row below carries this same hash.

Source material · 10

shallow

These headlines, in this order, are the entire world the models were given. Nothing else about North America on 2026-06-27 was in the prompt.

Extracts: not gathered at shallow depth — Stage 7 populates extracts only for standard and deep profiles. Zero here means "not requested", not "requested and empty". Full article texts: not gathered at shallow depth — only the deep profile fetches them.

The shallow profile requests 20 headlines and 10 are stored. That is a known provider cap, not a clipped payload: the Serper news endpoint returns at most 10 items per request regardless of the count asked for, and pagination has not been added. The three profiles currently differ in enrichment, not in headline breadth.

Model rationales

prompt p1

Same headlines, same rendered prompt, different readings. Where two models diverge, the disagreement is about interpretation, because the input was byte-identical and the hash proves it. The scored verdicts here span 17.0 points of the 0-100 scale.

claude-haiku-4-5
anthropic
32.0
okElevated

The evidence shows mixed signals with some economic dynamism (solar project rush, manufacturing growth) but underlying structural tensions including political polarization, climate-related challenges in Europe, and AI-driven political manipulation entering U.S. discourse. While no acute crisis is evident, these factors suggest elevated instability rather than baseline stability in the American political and social landscape as of mid-2026.

evidence 39d4198b0d4e…prompt p1
deepseek-reasoner
deepseek
25.0
okElevated

Evidence points to policy-driven market shifts and ongoing political polarization, but no acute violent conflict or institutional collapse is indicated. The headlines reflect normal democratic processes and economic adjustments, warranting a low-to-moderate chaos score.

evidence 39d4198b0d4e…prompt p1
claude-opus-5
anthropic
24.0
okElevated

The evidence is dominated by energy policy churn — a rush to beat the July 4 clean energy tax credit cutoff, AI-flooded public comment processes, and mining/energy expansion — indicating regulatory disruption but no violence, institutional breakdown, or emergency. References to American polarization and unease among trans travelers point to persistent social tension, though the coverage is reflective rather than reporting active unrest.

evidence 39d4198b0d4e…prompt p1
llama-3.3-70b-versatile
groq
20.0
okElevated

The provided headlines indicate a mix of climate concerns, economic activities, and social issues in the US, but none of them suggest immediate or severe chaos. The overall tone is one of awareness and preparation for potential challenges, such as climate change and economic shifts, rather than outright crisis or conflict.

evidence 39d4198b0d4e…prompt p1
gemini-3.1-pro-preview
google
15.0
okStable

The provided evidence indicates a generally stable environment characterized by routine economic activity, scientific research, and standard political discourse regarding energy policy. While opinion pieces highlight underlying societal polarization, the absence of acute crises or civil unrest keeps the immediate threat level low.

evidence 39d4198b0d4e…prompt p1

Every model returned a parseable verdict for this evidence, so no failure card is shown. Non-ok rows are rendered by the same component — with error_text where the rationale would be — and are covered by tests/rationale_card_test.tsx.