Living reference · Updated August 2, 2026

The models behind AI Storming

Six frontier models from six independent labs sit on the AI Storming panel. This page is the public record of who they are, what their own vendors publish about them, how much weight those published numbers deserve — and the thing the numbers cannot show you, which is what happens to a model once it can see what the others said.

Models
6
Independent labs
6
Depth tiers
2
Shared lineage
0

The current roster

Every session draws from these six. You choose which ones join, which tier each runs at, and which of them moderates. The seat is earned on what the model contributes to an argument — not on where it sits on a leaderboard.

Gemini

Google DeepMind

Long-context recall and synthesis. Holds a large canvas and a long transcript without losing the thread.

fast · gemini-3.6-flash thinking · gemini-3.6-flash one model, two thinking depths

ChatGPT

OpenAI

Structured argument. Tends to name the decision criteria before answering, which gives the rest of the panel something to attack.

fast · gpt-5.6-luna thinking · gpt-5.6-sol two models from one family

Claude

Anthropic

Caveats and second-order effects. The most likely of the six to say the question itself is wrong.

fast · claude-sonnet-5 thinking · claude-opus-5 two models

Grok

xAI

Contrarian pressure. Least likely to converge for the sake of converging, which is exactly what a panel needs one of.

fast · grok-4.5 thinking · grok-4.5 one model, two reasoning efforts

Kimi

Moonshot AI

Open-weight lineage trained outside the US labs — a genuinely different data and alignment history from the other five.

fast · kimi-k3 thinking · kimi-k3 one model, two reasoning efforts

DeepSeek

DeepSeek

Terse, math-forward reasoning. Cuts through a round that is drifting toward agreement on vibes.

fast · deepseek-v4-flash thinking · deepseek-v4-pro two models

Four of the six now run one model across both tiers, with the fast/thinking switch changing how long that model reasons rather than which model answers. That is a change in how the frontier labs ship: depth has become a dial rather than a separate product.

What each vendor publishes

Below are figures published by the labs themselves, linked to their own sources. We do not re-run these benchmarks and we do not average them into a ranking — the next section explains why a ranking built from them would mislead you.

Where a figure carries a caveat, the caveat is printed with it rather than in a footnote.

Terminal-bench 2.1 (Terminus-2)
78.0%
OSWorld-Verified
83.0%
SWE-Bench Pro
58.7%
MLE-Bench
63.9%
GDM-MRCR v2 (128k avg)
91.8%
GPQA Diamond
94.6%
ARC-AGI-3
38.3%

The ARC-AGI-3 figure was produced with two OpenAI-specific API features enabled. Under the ARC Prize project’s own test harness the same model scored 7.8%. Both numbers are real; they measure different things.

Frontier-Bench v0.1
43.3%
GDPval-AA
state of the art at release

Anthropic publishes most Opus 5 figures inside charts rather than as text, and states plainly that the model trails on some coding, legal and science evaluations.

Artificial Analysis Intelligence Index v4.1
54 (up 16 from grok-4.3)

xAI’s launch materials led with agentic and coding results and did not publish a GPQA Diamond score for this model. The index figure is a third-party composite, not a single test.

BrowseComp
90.4
DeepSWE (mini-SWE-agent harness)
67.3

K3 is open-weight — 2.8T parameters released under a modified MIT licence — so its claims are, unusually, reproducible by anyone with the hardware.

SWE-bench Verified
80.6%
GPQA Diamond
90.1
LiveCodeBench
93.5
MMLU-Pro
87.5

These are DeepSeek’s own reported figures for the V4 family. Broad independent reproduction was still outstanding at the time of writing.

How much a benchmark number is worth

Published scores are useful. They are also softer than they look, and the softness is not hidden — you can see it in the four examples already on this page.

The same benchmark, two variants, forty points apart

"SWE-bench" names a family, not a test. A model reported around the mid-nineties on SWE-bench Verified can land in the sixties on SWE-bench Pro. Neither number is dishonest. But a headline that says only "SWE-bench" has told you almost nothing.

The conditions belong to whoever is reporting

The clearest recent case: a 38.3% ARC-AGI-3 result was published for a frontier model with two provider-specific API features switched on. Run through the ARC Prize project's own harness, the same model scored 7.8%. The gap is not fraud — it is the difference between "what this model can do with our scaffolding" and "what this model does on neutral ground." Almost no headline distinguishes them.

Labs publish the tests they win

One of the models on this panel launched with a strong set of agentic and coding results and simply no GPQA Diamond figure. That is a normal commercial decision, and it means an absence on a launch page is information too — just not the kind you can put in a table.

Some numbers have not been reproduced yet

A vendor's first-party figures are a claim until someone outside the vendor reproduces them. Open-weight releases are the honourable exception: when the weights ship, anyone with the hardware can check.

Every one of these numbers describes a single model, working alone, on a task with a known answer, under conditions its own vendor selected. Your question has none of those properties.

What changes when the models can see each other

This is the part a leaderboard structurally cannot show, because every benchmark isolates the model on purpose. Put the same models in a room where each one reads what the others said, and their behaviour changes — measurably.

The measurement

The foundational result is Du, Li, Torralba, Tenenbaum and Mordatch (MIT/Google, 2023): multiple language models debating across rounds beat single-model baselines on reasoning, mathematics and factuality, with the improvement scaling with the number of participants before plateauing. Chan et al. (2023) reproduced the effect for evaluation tasks; Liang et al. (2023) found it pushes models out of convergent thinking rather than into it.

Note what that means for the table above. A model's solo score is a ceiling on what it knows, not a ceiling on what it will contribute. The panel's output is not the best model's answer; it is an answer none of them arrived at alone.

Why it works — two mechanisms

Blind spots don't overlap. These six were trained on different corpora, with different post-training regimes and different refusal behaviour. They are wrong in different places. A claim one model is confidently wrong about is a claim another has a real chance of catching — which is only true because the labs are independent. Six models from a single lab would inherit one lab's errors and ratify them unanimously.

Agreeableness needs an audience of one. A model alone with you is optimised to be satisfying: it holds a position for a message or two, then follows wherever you lean. Give it a peer who disagrees on the record and that stops working. It now has to defend the position to something that will not simply accept it. What survives that is worth more than what survives you.

The full research lineage and the boundaries of the category are on the multi-LLM debate reference.

How a model earns a seat — and how it loses one

Frontier releases have been landing every few weeks. The roster is reviewed against them continuously, and this page moves when the roster does.

Three things decide a swap. Independence comes first: one seat per lab, because the panel's value collapses if two seats share a training lineage. Contribution to an argument comes second, and it is not the same as a benchmark score — a model that reasons beautifully but concedes instantly is worth less to a debate than one that is merely stubborn and specific. Behaviour under our own conditions comes third: long shared context, several rounds, a moderator, and a hard requirement to address what the previous speaker actually said.

Before any swap ships, the new model is checked against its provider's API documentation and then against the live API. That second step is not ceremony. In the August 2026 review, four of the five providers we re-examined had documentation that contradicted their own API on at least one point that mattered — a parameter documented as supported that returned an error, a setting documented as unavailable that worked, a default that was not the documented default. A roster kept current from release notes alone would be quietly wrong.

The panel is a composition, not a ranking. The aim is six models that fail differently — not the six highest scores on any one list.

Frequently asked questions

Which AI models does Nodalist AI Storming use?

Six models from six independent providers: Gemini (Google DeepMind), ChatGPT (OpenAI), Claude (Anthropic), Grok (xAI), Kimi (Moonshot AI), and DeepSeek. As of August 2026 the panel runs gemini-3.6-flash, gpt-5.6-luna and gpt-5.6-sol, claude-sonnet-5 and claude-opus-5, grok-4.5, kimi-k3, and deepseek-v4-flash and deepseek-v4-pro. You choose which models join each session and which one moderates.

Why six providers instead of six models from the best lab?

Independent provenance is the whole point. Six models from one lab share training data, alignment methods and refusal behaviour — which means they share blind spots, and a debate between them would agree for the wrong reason. Six models from six labs fail in different places, so a claim one is overconfident about is a claim another can catch.

Do AI models actually perform better when they can see each other's answers?

Yes, and it is measured rather than assumed. Du et al. (MIT/Google, 2023) showed multi-agent debate beats single-model baselines on reasoning, math and factuality benchmarks, with the gain scaling with model count before plateauing. Chan et al. (2023) extended the result to evaluation tasks and Liang et al. (2023) to divergent thinking. The mechanism is that a model forced to defend an answer against a disagreeing peer cannot fall back on confident agreeableness.

What is the difference between the fast tier and the thinking tier?

Reasoning depth. On the fast tier each model answers with minimal internal deliberation, which keeps a round quick. On the thinking tier the same provider reasons for longer before speaking, producing more considered turns and a slower round. For four of the six providers both tiers are now the same model at two different reasoning settings rather than two different models.

How often does the model roster change?

Whenever a provider ships something that beats what is on the panel. Frontier releases have been arriving every few weeks, so the roster is reviewed continuously and this page is updated with it. Each swap is verified against the provider's own API documentation and then against the live API before it ships, because vendor documentation and vendor behaviour disagree more often than you would expect.

Can I pick which models debate my question?

Yes. Every session lets you choose which of the six join, which tier each runs at, and which model moderates. A narrow technical question might warrant three models on the thinking tier; a broad exploratory one is often better with all six on fast.

Further reading

Put all six on your hardest question

Pick the models, pick the moderator, and watch them argue it out on your canvas. The consensus comes back as a node you keep working from. Free to start.

Try AI Storming free