Reflex Chess

Chess rating test

Real positions, each rated by the thousands of players who have tried it. You get a range and a confidence, not a made-up number.

Loading

Setting out the pieces

The test needs JavaScript. If nothing appears in a moment, the puzzle bank did not load — reload, or open the course instead.

What is my chess rating?

There is no single number, and that is the first thing every other tool on this search gets wrong. A rating does not describe you. It describes your results inside one pool. You have one rating on Chess.com, a different one on Lichess, a different one again over the board if you play rated tournaments, and a puzzle rating that is different from all three. None of them is the real one.

So a test that promises to tell you "your chess rating" has to answer an obvious question: which one? This one estimates your Lichess puzzle rating, and it can say that with a straight face for a specific reason. Every position in the bank comes from the Lichess open puzzle database, and every one arrives with the Glicko rating it earned from the real attempts of everyone who has tried it. The scale is not invented here. It is the scale the positions already sit on, and all this page does is find where on it you belong.

What this number is not

It is not a playing rating. Solving a puzzle means finding a move that you have been told exists, with no clock and nobody trying to stop you. Playing a game means building a position worth having, not blundering for forty moves, and converting the ending. Both of those are chess and they are not the same skill. The two correlate, and people who solve well generally do play well, but there is no conversion factor between them that survives contact with an individual player, so this page does not print one.

It is not a Chess.com rating. Chess.com and Lichess rate separate pools with different formulas and different starting points, and the same player routinely carries different numbers on each. The size of that gap is not a constant, it is not the same for blitz and rapid, and it is not something this site has measured. So we are not going to publish a conversion table that looks authoritative and is guesswork. If what you actually want is your Chess.com rating, the only honest way to get it is to play rated games on Chess.com.

It is not a certificate. The result is an estimate from at most twelve observations, and the page tells you how wide the uncertainty is because the width is part of the answer. Twelve questions cannot do better than about a hundred points of standard error under this model. Anyone showing you a clean integer after ten puzzles is showing you a number they cannot support.

How the test works

It is a Bayesian adaptive test, which sounds grander than it is. Before the first question the test holds a wide, deliberately uninformative belief about how strong you might be. After each answer it multiplies that belief by how likely the answer was at every possible strength, and renormalises. What is left is a probability distribution over your rating, and the range on the result card is a genuine interval read straight off it rather than a designer's guess at how wide a range should look.

The formula it uses to score an answer is this one:

P(solve) = 1 / (1 + 10(puzzle − you) / 400)

which is the Elo expected-score curve, and that is not a modelling convenience. It is the model the puzzle ratings themselves were produced under. Using anything else would mean estimating on a scale the difficulties were never calibrated for. Two adjustments sit on top of it. The probability is squeezed into a band that never quite reaches zero or one, because a strong player still hangs a piece occasionally and a weak one still finds a brilliancy occasionally, and without that floor a single fluke would drag the whole estimate a band out of place. And an answer given more than three minutes after the position appeared counts for half, because someone who left the tab open and came back has not told us anything.

Each question is the most useful question available. Under this model an item tells you the most when its difficulty matches your strength, so the next position is drawn from the ten unused ones nearest the current estimate, preferring a motif you have not been shown yet. That is why the test converges in eight to twelve questions instead of thirty, and why it is not the same twelve positions twice.

There is no clock, and speed is not in the score. Lichess puzzle ratings are earned with no time limit, so a speed term would be estimating something the difficulties never measured. Your times are recorded and reported back to you as their own line on the result card, because "you find these, but slowly" is real information and deserves to be read as itself rather than smuggled into a rating.

How well does it actually do?

Two thousand simulated solvers per rating level were run through the shipped estimator. Against a solver who behaves exactly as the model assumes, the estimate lands within about 20 points of the truth between 1000 and 2000, drifts about 35 points low by 2200, and the stated 80% interval contained the true value 84% of the time — slightly wider than advertised, which is the safe direction to be wrong in. Against a sharper solver, one who reliably gets everything below their level and reliably misses everything above it, the estimate lands within about 20 points and the interval is comfortably conservative. Against a careless solver who fumbles one in five puzzles they should get, the estimate reads low, which is the correct behaviour: on the day, that is how they solved.

Why the test will not guess below 800

Because it cannot measure down there, and it says so instead. Here is the whole corpus, counted:

Every position in Reflex Chess by Lichess puzzle rating. “Your moves” is the median number of moves the solver has to find, not counting the opponent’s replies.
RatingPositionsYour movesEnds in mateCommonest named motif
400-59924271%back rank mate
600-799133260%back rank mate
800-999829248%back rank mate
1000-1199513243%fork
1200-13991,198236%fork
1400-15992,287225%fork
1600-17991,715223%fork
1800-19992,096218%fork
2000-21991,111313%fork
2200-2399442314%fork
2400-25998231%a quiet move
2600+1030%en passant
Total10,440

157 positions out of 10,440 are rated under 800. Twelve questions drawn from a pool that thin cannot separate 500 from 750, so the test stops early, says under 800, and points you at the beginning of the course. The same applies at the top: there are 92 positions above 2400, and past that the test says 2400 or above rather than guessing at 2600. Every tool that hands you a confident three-digit number at either end made it up.

The table is worth reading for its own sake, because it shows what actually changes as puzzle ratings rise, and it is not what people expect. The lines do not get much longer — the median stays at two of your own moves almost the whole way up. What collapses is mate. Seven in ten of the easiest positions end in checkmate, and by 2400 it is one in a hundred. Easy puzzles ask you to finish something. Hard ones ask you to see that something is there at all.

Chess rating bands and what they mean

The numbers themselves come from Arpad Elo, a physicist and strong amateur who built the system FIDE adopted in 1970. The property that makes it useful is the one on the curve above: rating differences translate directly into expected results. A 100-point gap is roughly a 64% score for the stronger player, a 200-point gap is roughly 76%, and a 400-point gap is roughly ten games in eleven. That is the entire content of a rating. It is a prediction of results against other rated people, not a description of how much chess you know.

Lichess, Chess.com and FIDE all use variants of the same idea — Lichess runs Glicko-2, which adds an explicit uncertainty to every rating for exactly the reason this page reports a range. But they run it over separate populations, which is why the numbers do not transfer. A pool is a currency. The exchange rate is not published because nobody can honestly publish it.

Here is what the bands mean on this scale, described by what the positions contain rather than by what kind of player you are:

Questions people ask

How accurate is this chess rating test?

Within about 20 points of the truth on average between 1000 and 2000, with a run-to-run spread of roughly 130 points, measured over 2,000 simulated solvers per level. The interval printed on the result card contained the true value 84% of the time against a model-consistent solver, which is slightly better than the 80% it claims. The limit is arithmetic, not effort: twelve questions cannot buy much more precision than that.

Can I take it more than once?

Yes, and you will get different positions. The test picks from the ten unused puzzles nearest your current estimate rather than the single nearest one, and it avoids repeating a motif inside a run, so two attempts are not the same twelve questions. If two runs disagree by more than a couple of hundred points, the honest reading is that your true rating is somewhere between them.

Does the test use a timer?

No, and this is deliberate. The puzzle ratings it scores you against were earned on a site with no puzzle clock, so adding time pressure here would be measuring something the difficulties were never calibrated for. Your times are recorded and reported separately. The only place time touches the score is a data-quality rule: an answer given more than three minutes after the position appeared is counted at half weight, because a tab left open is not evidence.

What happens if I guess?

Very little. There is no multiple choice to guess from — you have to play a legal move on a real board, and in most of these positions there are thirty of them and one is right. The model also assumes a small rate of lucky finds at every level, so one improbable solve moves the estimate less than an answer at your own level does.

Is this a chess Elo calculator?

Not in the usual sense. An Elo calculator takes two known ratings and a result and tells you the points that change hands. This does the harder thing: it starts with nothing, asks questions whose difficulty is known, and works backwards to the rating that best explains your answers.

What do I do with the result?

Use it as a starting line. The result card names a skill in the free Reflex Chess course and the level inside it whose positions have the median rating closest to your estimate, which is the part of the recommendation that is measured rather than guessed. Then stop thinking about the number. A rating is a lagging indicator of work you already did.

The course the test sends you to 87 skills and 10,440 rated positions, in a few minutes a day. Free, no account needed to start, and it works offline. Open Reflex Chess

Keep going

The motifs the test is built out of, each with worked positions and the answers collapsed so you can try them first: