Chess rating test
Real positions, each rated by the thousands of players who have tried it. You get a range and a confidence, not a made-up number.
Loading
Setting out the pieces
The test needs JavaScript. If nothing appears in a moment, the puzzle bank did not load — reload, or open the course instead.
- 12 positions at most
- 10,440 rated puzzles behind it
- No account, email or clock
- Free, and it always will be
What is my chess rating?
There is no single number, and that is the first thing every other tool on this search gets wrong. A rating does not describe you. It describes your results inside one pool. You have one rating on Chess.com, a different one on Lichess, a different one again over the board if you play rated tournaments, and a puzzle rating that is different from all three. None of them is the real one.
So a test that promises to tell you "your chess rating" has to answer an obvious question: which one? This one estimates your Lichess puzzle rating, and it can say that with a straight face for a specific reason. Every position in the bank comes from the Lichess open puzzle database, and every one arrives with the Glicko rating it earned from the real attempts of everyone who has tried it. The scale is not invented here. It is the scale the positions already sit on, and all this page does is find where on it you belong.
What this number is not
It is not a playing rating. Solving a puzzle means finding a move that you have been told exists, with no clock and nobody trying to stop you. Playing a game means building a position worth having, not blundering for forty moves, and converting the ending. Both of those are chess and they are not the same skill. The two correlate, and people who solve well generally do play well, but there is no conversion factor between them that survives contact with an individual player, so this page does not print one.
It is not a Chess.com rating. Chess.com and Lichess rate separate pools with different formulas and different starting points, and the same player routinely carries different numbers on each. The size of that gap is not a constant, it is not the same for blitz and rapid, and it is not something this site has measured. So we are not going to publish a conversion table that looks authoritative and is guesswork. If what you actually want is your Chess.com rating, the only honest way to get it is to play rated games on Chess.com.
It is not a certificate. The result is an estimate from at most twelve observations, and the page tells you how wide the uncertainty is because the width is part of the answer. Twelve questions cannot do better than about a hundred points of standard error under this model. Anyone showing you a clean integer after ten puzzles is showing you a number they cannot support.
How the test works
It is a Bayesian adaptive test, which sounds grander than it is. Before the first question the test holds a wide, deliberately uninformative belief about how strong you might be. After each answer it multiplies that belief by how likely the answer was at every possible strength, and renormalises. What is left is a probability distribution over your rating, and the range on the result card is a genuine interval read straight off it rather than a designer's guess at how wide a range should look.
The formula it uses to score an answer is this one:
P(solve) = 1 / (1 + 10(puzzle − you) / 400)
which is the Elo expected-score curve, and that is not a modelling convenience. It is the model the puzzle ratings themselves were produced under. Using anything else would mean estimating on a scale the difficulties were never calibrated for. Two adjustments sit on top of it. The probability is squeezed into a band that never quite reaches zero or one, because a strong player still hangs a piece occasionally and a weak one still finds a brilliancy occasionally, and without that floor a single fluke would drag the whole estimate a band out of place. And an answer given more than three minutes after the position appeared counts for half, because someone who left the tab open and came back has not told us anything.
Each question is the most useful question available. Under this model an item tells you the most when its difficulty matches your strength, so the next position is drawn from the ten unused ones nearest the current estimate, preferring a motif you have not been shown yet. That is why the test converges in eight to twelve questions instead of thirty, and why it is not the same twelve positions twice.
There is no clock, and speed is not in the score. Lichess puzzle ratings are earned with no time limit, so a speed term would be estimating something the difficulties never measured. Your times are recorded and reported back to you as their own line on the result card, because "you find these, but slowly" is real information and deserves to be read as itself rather than smuggled into a rating.
How well does it actually do?
Two thousand simulated solvers per rating level were run through the shipped estimator. Against a solver who behaves exactly as the model assumes, the estimate lands within about 20 points of the truth between 1000 and 2000, drifts about 35 points low by 2200, and the stated 80% interval contained the true value 84% of the time — slightly wider than advertised, which is the safe direction to be wrong in. Against a sharper solver, one who reliably gets everything below their level and reliably misses everything above it, the estimate lands within about 20 points and the interval is comfortably conservative. Against a careless solver who fumbles one in five puzzles they should get, the estimate reads low, which is the correct behaviour: on the day, that is how they solved.
Why the test will not guess below 800
Because it cannot measure down there, and it says so instead. Here is the whole corpus, counted:
| Rating | Positions | Your moves | Ends in mate | Commonest named motif |
|---|---|---|---|---|
| 400-599 | 24 | 2 | 71% | back rank mate |
| 600-799 | 133 | 2 | 60% | back rank mate |
| 800-999 | 829 | 2 | 48% | back rank mate |
| 1000-1199 | 513 | 2 | 43% | fork |
| 1200-1399 | 1,198 | 2 | 36% | fork |
| 1400-1599 | 2,287 | 2 | 25% | fork |
| 1600-1799 | 1,715 | 2 | 23% | fork |
| 1800-1999 | 2,096 | 2 | 18% | fork |
| 2000-2199 | 1,111 | 3 | 13% | fork |
| 2200-2399 | 442 | 3 | 14% | fork |
| 2400-2599 | 82 | 3 | 1% | a quiet move |
| 2600+ | 10 | 3 | 0% | en passant |
| Total | 10,440 | |||
157 positions out of 10,440 are rated under 800. Twelve questions drawn from a pool that thin cannot separate 500 from 750, so the test stops early, says under 800, and points you at the beginning of the course. The same applies at the top: there are 92 positions above 2400, and past that the test says 2400 or above rather than guessing at 2600. Every tool that hands you a confident three-digit number at either end made it up.
The table is worth reading for its own sake, because it shows what actually changes as puzzle ratings rise, and it is not what people expect. The lines do not get much longer — the median stays at two of your own moves almost the whole way up. What collapses is mate. Seven in ten of the easiest positions end in checkmate, and by 2400 it is one in a hundred. Easy puzzles ask you to finish something. Hard ones ask you to see that something is there at all.
Chess rating bands and what they mean
The numbers themselves come from Arpad Elo, a physicist and strong amateur who built the system FIDE adopted in 1970. The property that makes it useful is the one on the curve above: rating differences translate directly into expected results. A 100-point gap is roughly a 64% score for the stronger player, a 200-point gap is roughly 76%, and a 400-point gap is roughly ten games in eleven. That is the entire content of a rating. It is a prediction of results against other rated people, not a description of how much chess you know.
Lichess, Chess.com and FIDE all use variants of the same idea — Lichess runs Glicko-2, which adds an explicit uncertainty to every rating for exactly the reason this page reports a range. But they run it over separate populations, which is why the numbers do not transfer. A pool is a currency. The exchange rate is not published because nobody can honestly publish it.
Here is what the bands mean on this scale, described by what the positions contain rather than by what kind of player you are:
- Under 800. One-move problems on quiet boards. Mate in one, or a piece sitting undefended in the open. More than half of them end in checkmate.
- 800 to 1200. The basic double attacks. A fork or a back-rank idea that is visible once you look at the whole board rather than the part of it you were already thinking about.
- 1200 to 1600. Two of your own moves, and the first one is usually a check or a capture. The skill being tested is following a line to its end instead of stopping at the first move that looks active.
- 1600 to 2000. The first move stops being the obvious one. A defender has to be removed, or a piece has to be given up for something that only pays two moves later.
- 2000 to 2400. Quiet moves, and lines where nothing is hanging and no check exists. The mate share has dropped to about one in eight; most of these are about material or a position you have to want before you can find it.
- Above 2400. Ninety-two positions in the whole corpus. The test cannot resolve this band and does not pretend to.
Questions people ask
How accurate is this chess rating test?
Within about 20 points of the truth on average between 1000 and 2000, with a run-to-run spread of roughly 130 points, measured over 2,000 simulated solvers per level. The interval printed on the result card contained the true value 84% of the time against a model-consistent solver, which is slightly better than the 80% it claims. The limit is arithmetic, not effort: twelve questions cannot buy much more precision than that.
Can I take it more than once?
Yes, and you will get different positions. The test picks from the ten unused puzzles nearest your current estimate rather than the single nearest one, and it avoids repeating a motif inside a run, so two attempts are not the same twelve questions. If two runs disagree by more than a couple of hundred points, the honest reading is that your true rating is somewhere between them.
Does the test use a timer?
No, and this is deliberate. The puzzle ratings it scores you against were earned on a site with no puzzle clock, so adding time pressure here would be measuring something the difficulties were never calibrated for. Your times are recorded and reported separately. The only place time touches the score is a data-quality rule: an answer given more than three minutes after the position appeared is counted at half weight, because a tab left open is not evidence.
What happens if I guess?
Very little. There is no multiple choice to guess from — you have to play a legal move on a real board, and in most of these positions there are thirty of them and one is right. The model also assumes a small rate of lucky finds at every level, so one improbable solve moves the estimate less than an answer at your own level does.
Is this a chess Elo calculator?
Not in the usual sense. An Elo calculator takes two known ratings and a result and tells you the points that change hands. This does the harder thing: it starts with nothing, asks questions whose difficulty is known, and works backwards to the rating that best explains your answers.
What do I do with the result?
Use it as a starting line. The result card names a skill in the free Reflex Chess course and the level inside it whose positions have the median rating closest to your estimate, which is the part of the recommendation that is measured rather than guessed. Then stop thinking about the number. A rating is a lagging indicator of work you already did.
Keep going
The motifs the test is built out of, each with worked positions and the answers collapsed so you can try them first: