ChaosWatch
New here? Start here →
Evidence inspector

North America · 2026-06-25

The receipts. This is the exact material every model was shown for this day and region, followed by what each of them made of it — in its own words, next to the number it gave.

It is worth reading once even if you skip everything else: it is where confident-sounding reasoning and a score that turns out to be no better than chance sit side by side.

Why the evidence is identical, and how you can check

Evidence is gathered once per (date, region, depth_profile), hashed, and reused for every model in the run. A disagreement below is therefore a disagreement about identical input, not two models getting different search results. The evidence hash shown on this page is what makes that checkable.

Evidence hash
926089fba6a8fb155dc2a4f8ef33df7054db19e3534fe18a7cb4570aa3eff7a5

SHA-256 over the canonically serialised payload minus the hash itself, so it is stable across processes and machines. It is stored on the run and on every score row below — provenance is checkable, not asserted. Every score row below carries this same hash.

Source material · 10

shallow

These headlines, in this order, are the entire world the models were given. Nothing else about North America on 2026-06-25 was in the prompt.

03

Supply Chain, Energy, and AI Nexus

The anticipated growth in artificial intelligence (AI) development requires additional power capacity. More than half of North America faces a substantial...

rand.org

Extracts: not gathered at shallow depth — Stage 7 populates extracts only for standard and deep profiles. Zero here means "not requested", not "requested and empty". Full article texts: not gathered at shallow depth — only the deep profile fetches them.

The shallow profile requests 20 headlines and 10 are stored. That is a known provider cap, not a clipped payload: the Serper news endpoint returns at most 10 items per request regardless of the count asked for, and pagination has not been added. The three profiles currently differ in enrichment, not in headline breadth.

Model rationales

prompt p1

Same headlines, same rendered prompt, different readings. Where two models diverge, the disagreement is about interpretation, because the input was byte-identical and the hash proves it. The scored verdicts here span 22.0 points of the 0-100 scale.

deepseek-reasoner
deepseek
42.0
okUnstable

The evidence indicates moderate instability from legal battles over climate policy, geopolitical energy disruptions, and declining CFO confidence in the economy. Yet, substantial renewable energy investments and large-scale events like the World Cup provide counterbalancing positives, preventing a more severe rating.

evidence 926089fba6a8…prompt p1
gemini-3.1-pro-preview
google
40.0
okElevated

Domestic stability is bolstered by bipartisan legislative successes, major renewable energy infrastructure milestones, and the economic boost from the 2026 FIFA World Cup. However, global geopolitical tensions stemming from the Iran war and declining CFO confidence in the world economy introduce moderate external risks to the region.

evidence 926089fba6a8…prompt p1
claude-haiku-4-5
anthropic
35.0
okElevated

The US shows mixed signals with positive developments in renewable energy infrastructure, bipartisan housing legislation, and economic activity from the FIFA World Cup, but these are offset by declining CFO confidence in the economy, ongoing climate policy litigation, and geopolitical energy security concerns stemming from the Iran conflict. While no acute crises are evident, underlying economic uncertainty and regulatory instability create moderate tension in the domestic landscape without approaching systemic instability or severe disruption.

evidence 926089fba6a8…prompt p1
claude-opus-5
anthropic
27.0
okElevated

The evidence is dominated by economic, energy, and infrastructure stories, including bipartisan housing legislation, a record renewable project, and World Cup economic gains, indicating functioning institutions and no domestic violence or unrest. Mild elevation reflects deregulatory litigation battles over climate policy, softening CFO sentiment on the global economy, grid capacity strain from AI demand, and spillover from the Iran war on energy security.

evidence 926089fba6a8…prompt p1
llama-3.3-70b-versatile
groq
20.0
okElevated

The US is experiencing a mix of positive developments, such as the passage of a bipartisan bill to lower housing costs and the launch of the largest renewable energy project in the country's history, but also faces challenges like climate change lawsuits and decreased CFO sentiments on the world economy. Overall, the situation is stable with some concerns, but no catastrophic events or extreme instability are reported, warranting an elevated score.

evidence 926089fba6a8…prompt p1

Every model returned a parseable verdict for this evidence, so no failure card is shown. Non-ok rows are rendered by the same component — with error_text where the rationale would be — and are covered by tests/rationale_card_test.tsx.