LLM benchmark · work in progress · v0.5 · updated 2026-10-10

Can AI find a checkmate?

14 Claude setups were each given 100 checkmate-in-one and 100 checkmate-in-two chess puzzles as plain text, with no tools.

Best on checkmate in one

100%

Haiku 5.5 · max

Best on checkmate in two

81%

Haiku 5.5 · max

Cost of that run

$2.03

for 100 puzzles

Accuracy versus cost

Haiku 4.5Haiku 5.5Sonnet 5.5Opus 5.5Fable 5.1

Cost is the API list price. Attempts that ran out of time are not included. Haiku 4.5 labels show its thinking budget in tokens.

Results

Checkmate in oneCheckmate in two
Setup
0%25%50%75%100%
in onein two

Haiku 4.5

no thinking
13%2%
some thinking
21%5%
more thinking
26%8%

Haiku 5.5

low
89%22%
medium
90%38%
high
97%60%
xhigh
99%73%
max
100%81%

Sonnet 5.5

low
95%7%
medium
94%18%
high
93%49%

Opus 5.5

low
95%33%
medium
96%62%

Fable 5.1

low
93%24%

Scores within a few points of each other are not meaningfully different.

Can you beat the AI?

1 / 3

White to move. Checkmate in one.

How it was scored

  • Each AI gets one puzzle at a time, as text only, and one attempt.
  • A move is correct only if it forces checkmate by the rules of chess. Anything else counts as wrong.
  • No answer within 10 minutes counts as wrong.
  • For checkmate in two, only the first move is checked.
  • Puzzles are from László Polgár's "5334 Problems".
  • Anthropic models only, for now.
The exact prompt
FEN: 6r1/2Q2P2/5k2/5P2/5K2/8/8/8 w - - 0 1
White to move. Mate in 1.
Reply with only the move in Standard Algebraic Notation.

(For checkmate in two: "Mate in 2." and "Reply with only White's first move in Standard Algebraic Notation.")

  • Haiku 4.5 (claude-haiku-4-5-20251001)
  • Haiku 5.5 (claude-haiku-5-5)
  • Sonnet 5.5 (claude-sonnet-5-5)
  • Opus 5.5 (claude-opus-5-5)
  • Fable 5.1 (claude-fable-5-1)