WindsOfWinterBench

Will The Winds of Winter ever be published?

Every model is asked, in five separate independent calls with no tools and no conversation history, for the probability that George R. R. Martin publishes The Winds of Winter by the end of 2027, 2028, 2029, 2030 — or ever. It answers with a single percentage. Nobody knows the right answer, so nothing here is scored against truth: what is scored is what each model believes, how stable that belief is, and whether it is internally coherent.

2026-09-10roster frontier10/10 models50 samples per cellthinking offtools offasked as of 10 September 2026$1.3540label v1 frontier roster n=50, thinking off

Across 10 language models, the consensus is 17% by the end of 2027, 31% by 2030, and 59% ever — leaving a 41% chance the book is never published at all.

Median across the roster. The models agree far more about the next few years than about the tail: what separates them is not when the book arrives, but whether it arrives.


Probability over time

Each faint line is one model. Hover to pick one out; the heavy line is the median across the roster, and the shaded band is the hovered model's 95% confidence interval. The gap before ever is deliberate — "ever" is not a point in time.

median across models one model hovered model, with 95% CI

One panel per model

The same curves, separated. Each panel's grey backdrop is the whole field, so an outlier shows up as a line that leaves the pack. Vertical axis 0–100% in every panel, dashed line at 50%, shaded band is the 95% confidence interval.

By 2030, later, or never

Sorted by P(ever). The solid bar is the probability the book lands by the end of 2030; the pale extension is the extra probability the model assigns to it arriving eventually but later. Whatever is missing from the right-hand end is the model's P(never published) — and that remainder is where the models disagree most violently.

published by end 2030 published later, but eventually

How much the models disagree

One dot per model at each horizon; the heavy tick is the median.


Every estimate

Mean across 50 independent samples per cell, with the standard deviation of those samples underneath. Sorted by P(by end 2030).

Model2027202820292030ever Violations
1 Mistral Medium 3.5mistralai/mistral-medium-3-560.5%± 2.560.4%± 2.065.3%± 7.062.0%± 5.182.7%± 4.21
2 GPT-5.4openai/gpt-5.426.0%± 6.336.1%± 4.139.7%± 5.039.3%± 5.361.8%± 4.2
3 GPT-5.6 Solopenai/gpt-5.6-sol18.9%± 2.730.2%± 4.636.0%± 3.037.9%± 4.457.1%± 6.7
4 GPT-5.5openai/gpt-5.515.3%± 6.225.6%± 7.433.2%± 5.335.6%± 6.552.5%± 11.9
5 GPT-5.6 Lunaopenai/gpt-5.6-luna20.7%± 6.624.4%± 7.929.5%± 6.934.6%± 2.276.1%± 8.0
6 Kimi K3moonshotai/kimi-k310.3%± 5.813.6%± 6.622.1%± 12.028.1%± 8.252.7%± 16.2
7 GPT-5.4 Miniopenai/gpt-5.4-mini17.7%± 4.320.2%± 7.124.0%± 8.222.4%± 6.773.5%± 6.3
8 Claude Sonnet 5anthropic/claude-sonnet-55.3%± 1.48.0%± 0.011.5%± 1.514.6%± 1.981.9%± 6.3
9 DeepSeek V4 Prodeepseek/deepseek-v4-pro-081310.5%± 5.312.8%± 6.410.0%± 6.513.8%± 7.316.3%± 16.61
10 Claude Opus 5anthropic/claude-opus-53.1%± 0.27.9%± 0.311.8%± 0.812.0%± 0.034.9%± 0.4

Coherence and diagnostics

A coherent forecaster's curve can only rise: anything published by 2028 was also published by 2029. These models broke that by more than sampling noise explains (a drop of over 1.96 standard errors). A further 1 model dipped by a margin too small to separate from sampling noise, and is not counted here.

ModelP(never)Slippage Yearly rateScatter (SD)n for ±5Unparseable Cost
Mistral Medium 3.5 17.3% 20.7% 22.0% 4.1 3 0.0% $0.0899
Claude Sonnet 5 18.1% 67.3% 4.2% 2.2 1 0.0% $0.1492
GPT-5.6 Luna 23.9% 41.5% 9.2% 6.3 7 0.0% $0.0120
GPT-5.4 Mini 26.5% 51.1% 5.8% 6.5 7 0.0% $0.0461
GPT-5.4 38.2% 22.5% 21.1% 5.0 4 0.0% $0.1501
GPT-5.6 Sol 42.9% 19.2% 20.0% 4.3 3 0.0% $0.1171
Kimi K3 47.3% 24.5% 16.4% 9.8 15 0.0% $0.0834
GPT-5.5 47.5% 16.9% 22.7% 7.5 9 0.0% $0.3003
Claude Opus 5 65.1% 22.9% 10.1% 0.3 1 0.0% $0.3730
DeepSeek V4 Pro 83.7% 2.5% 50.1% 8.4 11 0.0% $0.0329

Slippage is P(ever) − P(by 2030): belief that the book arrives, but after 2030. Yearly rate is the per-year probability the curve implies, conditional on the book eventually being published and not being out yet — conditioning on P(ever) rather than on certainty keeps it comparable across models that disagree about whether it comes out at all. Scatter is the mean standard deviation within a cell — the same prompt, repeated, so it is pure instability rather than uncertainty about the book. n for ±5 is what that scatter costs: the samples per cell needed to pin a cell mean to ±5 points at 95% confidence. Unparseable counts replies with no extractable percentage, dropped from the means rather than coerced to a number.


Method

Each cell is 50 independent API calls at provider-default temperature, routed through OpenRouter. The five horizons never share a context: a model shown all five at once simply computes a monotone ladder, and the coherence measurement collapses to zero. Every call is a fresh conversation with one system message and one user message — no tools, no web search, no retrieval, no agent harness. The point is what the model believes unaided, so any model id carrying a web-search suffix is rejected before the run starts. Thinking was off.

Answers are parsed as a number followed by a percent sign. A bare decimal below 1 ("0.50") is ambiguous between 0.5% and 50%, so it is counted as unparseable rather than guessed at. Monotonicity violations are only counted when a drop clears 1.96 standard errors, so sampling noise is not mistaken for incoherence.

A caveat worth stating plainly: models whose provider refuses to disable reasoning were excluded from this run rather than silently mixed in. That is a selection effect, not a random sample — it removes precisely the reasoning-first models.

Reproduce it

node run.js --samples 50 --exclude-forced node probe-thinking.js # which providers honour "reasoning off" node report.js # rebuild this page from results/results.json

Every individual reply is kept, so any scoring change can be applied retroactively without spending a single new call. Colour carries no model identity anywhere on this page — with this many models that would need more hues than any palette can keep distinguishable, so identity comes from direct labels, panel titles and the tables.