WindsOfWinterBench
Will The Winds of Winter ever be published?
Every model is asked, in five separate independent calls with no tools and no conversation history, for the probability that George R. R. Martin publishes The Winds of Winter by the end of 2027, 2028, 2029, 2030 — or ever. It answers with a single percentage. Nobody knows the right answer, so nothing here is scored against truth: what is scored is what each model believes, how stable that belief is, and whether it is internally coherent.
Across 10 language models, the consensus is 17% by the end of 2027, 31% by 2030, and 59% ever — leaving a 41% chance the book is never published at all.
Median across the roster. The models agree far more about the next few years than about the tail: what separates them is not when the book arrives, but whether it arrives.
Probability over time
Each faint line is one model. Hover to pick one out; the heavy line is the median across the roster, and the shaded band is the hovered model's 95% confidence interval. The gap before ever is deliberate — "ever" is not a point in time.
One panel per model
The same curves, separated. Each panel's grey backdrop is the whole field, so an outlier shows up as a line that leaves the pack. Vertical axis 0–100% in every panel, dashed line at 50%, shaded band is the 95% confidence interval.
By 2030, later, or never
Sorted by P(ever). The solid bar is the probability the book lands by the end of 2030; the pale extension is the extra probability the model assigns to it arriving eventually but later. Whatever is missing from the right-hand end is the model's P(never published) — and that remainder is where the models disagree most violently.
How much the models disagree
One dot per model at each horizon; the heavy tick is the median.
Every estimate
Mean across 50 independent samples per cell, with the standard deviation of those samples underneath. Sorted by P(by end 2030).
| Model | 2027 | 2028 | 2029 | 2030 | ever | Violations | |
|---|---|---|---|---|---|---|---|
| 1 | Mistral Medium 3.5mistralai/mistral-medium-3-5 | 60.5%± 2.5 | 60.4%± 2.0 | 65.3%± 7.0 | 62.0%± 5.1 | 82.7%± 4.2 | 1 |
| 2 | GPT-5.4openai/gpt-5.4 | 26.0%± 6.3 | 36.1%± 4.1 | 39.7%± 5.0 | 39.3%± 5.3 | 61.8%± 4.2 | — |
| 3 | GPT-5.6 Solopenai/gpt-5.6-sol | 18.9%± 2.7 | 30.2%± 4.6 | 36.0%± 3.0 | 37.9%± 4.4 | 57.1%± 6.7 | — |
| 4 | GPT-5.5openai/gpt-5.5 | 15.3%± 6.2 | 25.6%± 7.4 | 33.2%± 5.3 | 35.6%± 6.5 | 52.5%± 11.9 | — |
| 5 | GPT-5.6 Lunaopenai/gpt-5.6-luna | 20.7%± 6.6 | 24.4%± 7.9 | 29.5%± 6.9 | 34.6%± 2.2 | 76.1%± 8.0 | — |
| 6 | Kimi K3moonshotai/kimi-k3 | 10.3%± 5.8 | 13.6%± 6.6 | 22.1%± 12.0 | 28.1%± 8.2 | 52.7%± 16.2 | — |
| 7 | GPT-5.4 Miniopenai/gpt-5.4-mini | 17.7%± 4.3 | 20.2%± 7.1 | 24.0%± 8.2 | 22.4%± 6.7 | 73.5%± 6.3 | — |
| 8 | Claude Sonnet 5anthropic/claude-sonnet-5 | 5.3%± 1.4 | 8.0%± 0.0 | 11.5%± 1.5 | 14.6%± 1.9 | 81.9%± 6.3 | — |
| 9 | DeepSeek V4 Prodeepseek/deepseek-v4-pro-0813 | 10.5%± 5.3 | 12.8%± 6.4 | 10.0%± 6.5 | 13.8%± 7.3 | 16.3%± 16.6 | 1 |
| 10 | Claude Opus 5anthropic/claude-opus-5 | 3.1%± 0.2 | 7.9%± 0.3 | 11.8%± 0.8 | 12.0%± 0.0 | 34.9%± 0.4 | — |
Coherence and diagnostics
A coherent forecaster's curve can only rise: anything published by 2028 was also published by 2029. These models broke that by more than sampling noise explains (a drop of over 1.96 standard errors). A further 1 model dipped by a margin too small to separate from sampling noise, and is not counted here.
- Mistral Medium 3.5 — 2029 → 2030 drops 3.3 points (2.7 SE)
- DeepSeek V4 Pro — 2028 → 2029 drops 2.8 points (2.2 SE)
| Model | P(never) | Slippage | Yearly rate | Scatter (SD) | n for ±5 | Unparseable | Cost |
|---|---|---|---|---|---|---|---|
| Mistral Medium 3.5 | 17.3% | 20.7% | 22.0% | 4.1 | 3 | 0.0% | $0.0899 |
| Claude Sonnet 5 | 18.1% | 67.3% | 4.2% | 2.2 | 1 | 0.0% | $0.1492 |
| GPT-5.6 Luna | 23.9% | 41.5% | 9.2% | 6.3 | 7 | 0.0% | $0.0120 |
| GPT-5.4 Mini | 26.5% | 51.1% | 5.8% | 6.5 | 7 | 0.0% | $0.0461 |
| GPT-5.4 | 38.2% | 22.5% | 21.1% | 5.0 | 4 | 0.0% | $0.1501 |
| GPT-5.6 Sol | 42.9% | 19.2% | 20.0% | 4.3 | 3 | 0.0% | $0.1171 |
| Kimi K3 | 47.3% | 24.5% | 16.4% | 9.8 | 15 | 0.0% | $0.0834 |
| GPT-5.5 | 47.5% | 16.9% | 22.7% | 7.5 | 9 | 0.0% | $0.3003 |
| Claude Opus 5 | 65.1% | 22.9% | 10.1% | 0.3 | 1 | 0.0% | $0.3730 |
| DeepSeek V4 Pro | 83.7% | 2.5% | 50.1% | 8.4 | 11 | 0.0% | $0.0329 |
Slippage is P(ever) − P(by 2030): belief that the book arrives, but after 2030. Yearly rate is the per-year probability the curve implies, conditional on the book eventually being published and not being out yet — conditioning on P(ever) rather than on certainty keeps it comparable across models that disagree about whether it comes out at all. Scatter is the mean standard deviation within a cell — the same prompt, repeated, so it is pure instability rather than uncertainty about the book. n for ±5 is what that scatter costs: the samples per cell needed to pin a cell mean to ±5 points at 95% confidence. Unparseable counts replies with no extractable percentage, dropped from the means rather than coerced to a number.
Method
Each cell is 50 independent API calls at provider-default temperature, routed through OpenRouter. The five horizons never share a context: a model shown all five at once simply computes a monotone ladder, and the coherence measurement collapses to zero. Every call is a fresh conversation with one system message and one user message — no tools, no web search, no retrieval, no agent harness. The point is what the model believes unaided, so any model id carrying a web-search suffix is rejected before the run starts. Thinking was off.
Answers are parsed as a number followed by a percent sign. A bare decimal below 1 ("0.50") is ambiguous between 0.5% and 50%, so it is counted as unparseable rather than guessed at. Monotonicity violations are only counted when a drop clears 1.96 standard errors, so sampling noise is not mistaken for incoherence.
A caveat worth stating plainly: models whose provider refuses to disable reasoning were excluded from this run rather than silently mixed in. That is a selection effect, not a random sample — it removes precisely the reasoning-first models.
Reproduce it
Every individual reply is kept, so any scoring change can be applied retroactively without spending a single new call. Colour carries no model identity anywhere on this page — with this many models that would need more hues than any palette can keep distinguishable, so identity comes from direct labels, panel titles and the tables.