Measuring chess strength against an opponent that never blunders — an experiment, and an invitation
Your chess rating is measured with a strange instrument: other people, whose own ratings are wrong to varying degrees and whose play swings wildly from game to game. We want to test a different instrument — a full-strength engine giving fixed material odds — and find out how accurately, and how quickly, it can locate a player's real strength. Neither instrument is perfect, and this article is honest about the flaws of both. At the end, we invite you to help us run the measurement.
First, an uncomfortable truth about ratings
A chess rating feels like a personal property, like height or weight. It isn't. A rating is a prediction: it says what score you are expected to make against the other rated players in one specific pool. "1200 on lichess" means "scores about 50% against lichess 1200s" — and nothing more.
You can see this immediately in a fact most online players have noticed: the same person usually carries noticeably different ratings on chess.com and on lichess, often hundreds of points apart through much of the range. It is tempting to conclude that at least one of the sites is measuring badly. The correct conclusion is stranger: there is no site-independent "true Elo" for either of them to get right. Each site's number is defined relative to its own population, its own starting values, its own rating floor. Both sites measure their own quantity perfectly well; the quantities are simply different. Any attempt to measure chess strength — including ours — has to pick a reference pool and state its results in that pool's units.
So the honest question is never "what is this player's true rating?" It is: given a reference pool, how accurately and how quickly can an instrument estimate the number that pool would eventually assign?
The standard instrument, and why it is noisy at low ratings
The standard instrument is human opposition: play many rated games, let the rating system average the results. At high rating levels this works well — the opponents are stable, their ratings are converged, and the pool is dense. At low rating levels, where most improving players live, the instrument is remarkably noisy, for reasons that compound each other.
Every game is noisy on both sides of the board. A game result mixes your fluctuations with your opponent's. Low-rated players are enormously variable — the same person plays one game at the level of a 700 and the next at the level of a 300 — so a win or a loss says as much about which opponent showed up that day as about you.
The opponents' labels are unreliable. Low-rated pools are crowded with brand-new accounts on provisional ratings, fast-improving juniors whose number lags weeks behind their strength, returning adults, and the occasional sandbagger. When you are crushed by a "350" and then crush a "500," you have mostly learned that those two labels were inaccurate — the modern Glicko systems even quantify this with a rating deviation, which is largest exactly where the low-rated pool is.
Averaging is slow. Rating systems do overcome this noise, but only across many games. Small samples are dominated by luck of the pairing — which is why a run of 10 or 20 online games can move a low rating by an amount that has little to do with any change in skill.
An analogy we like, with its limits stated honestly: estimating your strength from low-rated human games is like estimating your height by comparing yourself against a crowd of people whose own recorded heights are wrong by varying amounts. The analogy is imperfect in one important way — height exists independently of any comparison, while a chess rating is constituted by comparisons. But if you grant that some underlying playing strength exists, of which every rating is a noisy estimate, the picture holds: the standard instrument is honest, but it is noisy, and noisiest at the bottom of the scale.
The alternative instrument: a full-strength engine giving material odds
Now consider a very different opponent. Take an engine searching at full strength — twenty-four plies deep or more, no artificial errors, no injected noise — and remove pieces from its side before move one: a queen, a rook and a knight, various combinations, always from the same squares for a given configuration. The engine then plays essentially perfect chess with what remains. It never hangs a piece, never misses a tactic you allow, never has a good day or a bad day. Its entire weakness is fixed before the game begins.
(This is not a novelty, by the way — it is a revival. Before rating systems existed, graded handicaps were how chess expressed strength differences: nineteenth-century masters played serious games at "pawn and move," knight odds and rook odds, and a player's class was described by the handicap a master needed to give them. We have written about the deep difference between these odds engines and ordinary "weak bots" in a separate article.)
As a measuring instrument, such an opponent has one outstanding property: it contributes almost no noise of its own. With fixed settings and varied openings, essentially all the variance left in your results is yours. Players who try a ladder of these configurations typically discover something step-like: one handicap they beat almost every time, a slightly smaller handicap they almost never beat. The outcome is governed by a threshold — whether the material surplus is large enough to survive that player's typical error rate against perfect punishment — rather than by the mutual, fluctuating blundering that decides human games. That threshold is a stable, personal quantity. The question is what it corresponds to on a human rating scale.
The quirks of the bot ladder — stated as plainly as the other side's
It would be easy to stop here and declare the odds engine the superior instrument. That would be exactly the kind of one-sided argument this article is trying to avoid. The ladder has its own imperfections, and an honest calibration experiment has to design around every one of them.
The rungs are elastic
A removed piece subtracts a fixed amount of material, but material and rating have no fixed exchange rate — the exchange rate is set by the receiving player's conversion skill. A strong club player given knight odds knows the winning recipe cold: trade relentlessly, avoid complications, steer into a won endgame. Against them, the handicap is worth an enormous amount. A novice given the identical position cannot execute that recipe — they fail to trade when trading wins, allow counterplay, leak material back until the odds have evaporated, and then face a full-strength engine on level terms. In the limiting case, a complete beginner given queen-and-rook odds can still lose most games, simply by donating more than a queen and rook through their own blunders. The consequence: the ladder's rungs are physically fixed but elastically spaced in rating terms — stretched apart at some strength levels, bunched together at others. The mapping from configuration to rating must be measured empirically across the whole range; it can never be assumed linear, and it will not even be evenly spaced.
Which square matters, not just which piece
A handicap is a position, not a point count. Removing the f-pawn half-opens the king's shelter and changes the game from move one; removing the a-pawn is a quiet deficit that may go unfelt for thirty moves. A missing king's knight reshapes the opening; a missing queen's rook may not matter until the middlegame. Two configurations of identical material value can differ substantially in real difficulty — so configurations have to be ordered by measured difficulty, not by adding up piece values.
A deterministic opponent can be memorized
An engine at fixed depth from a fixed position plays close to deterministically. Without countermeasures, a player can eventually pass a rung the way one beats a video-game boss — by trial and error against the same replies — which demonstrates memory, not chess. The fix is simple but mandatory: randomize the openings, so every attempt is a fresh test.
Odds chess has its own learning curve
The first games anyone plays at material odds partly measure unfamiliarity with an unusual situation ("I have a queen for nothing — now what?") rather than chess strength. A serious protocol discards each player's first few odds games as practice.
The instrument goes blind to one whole skill — and this caps its range
This is the most important limitation, and it deserves to be stated without softening. A large part of practical chess strength — arguably the decisive part in human games below master level — is the ability to exploit the opponent's mistakes: noticing the hung piece, spotting the mate your opponent just allowed, swindling from a lost position, pressing a nervous opponent on the clock. A full-strength odds engine never errs, so it tests this ability exactly zero. It measures, with beautiful precision, a different family of skills: keeping your own pieces safe, spotting threats before moving, simplifying when ahead, converting material without allowing counterplay.
How much this matters depends on rating level. Below roughly 1000–1200, avoiding blunders and converting material explain most of the variation in rating — the skill the ladder measures is, to a good approximation, the skill that ratings at that level reflect, which is why we expect the instrument to work best there. Higher up, style, preparation, practical judgment against imperfect resistance and clock handling carry an increasing share of the rating, and two equally rated players can score very differently against the same configuration because they earn their rating through different skill mixtures. We therefore expect the ladder's accuracy to be good at low ratings and to degrade measurably as ratings rise — and finding where that ceiling sits is one of the explicit goals of the experiment, not an embarrassment to be hidden from it.
The experiment
Here is what we want to measure, and how.
Volunteers with established online ratings play a series of games against full-strength engines at graded material handicaps, with randomized openings, inside ChessMend. The protocol is adaptive: rather than grinding out losses at configurations far too hard, the ladder moves with your results — win, and the next game is a slightly harder configuration; lose, and it is slightly easier. The sequence oscillates around the configuration where your winning chances are close to even, and that crossing point is your ladder threshold.
A game outcome carries the most information when the result is genuinely uncertain. Establishing that you lose 90% rather than 80% of games at some rung takes a large number of games, because both look like "mostly losses" in a small sample. Homing in on the even-chances point extracts far more information per game — which is why this staircase design appears throughout measurement science: it is how hearing tests find the softest tone you can detect, and how computerized adaptive tests such as the GRE locate a test-taker's level in a fraction of the questions a fixed test needs. Each item is placed where it is maximally informative. We are applying the same logic to chess, with the added advantage that our "test items" — the engine's games — contain essentially no noise of their own.
Alongside each volunteer's threshold we record their established rating (site, time control, number of games and rating deviation — we need ratings that are actually converged, in a time control comparable to the odds games). With enough volunteers across the whole strength range, we fit the curve that maps ladder thresholds to ratings in a chosen reference pool.
And here is the elegant part: the experiment measures the very number that decides whether the idea works. Once the curve is fitted, the scatter of real players around it — the residual spread — is the accuracy of the instrument. If the spread comes out tight in the lower rating bands, then a short adaptive ladder session genuinely estimates a low-rated player's strength about as well as a long and noisy run of human games, and we will have shown it with data. If style differences push the spread wide, the idea fails — and the data will show precisely where and why, including the rating level at which the blind spot described above begins to dominate. Either outcome is a real result. That is the difference between an experiment and a sales pitch.
Two expectations we are stating in advance, so we can be held to them: we expect the fitted curve to be visibly nonlinear (the elastic-rungs effect), and we expect accuracy to be best at low ratings and worse at high ratings (the blind-spot effect). If the data contradicts either expectation, that is interesting too.
What this is not
To keep expectations honest: this experiment will not produce a "truer" rating than chess.com or lichess — we have argued above that no such thing exists, and our calibration will itself be stated in one reference pool's units. It will not replace rated human play, which remains the only arena where the full blend of chess skills, including exploiting human error, is tested at once. And a ladder threshold will never be a substitute for a converged rating built on hundreds of games. The claim under test is narrower and more useful: that for lower-rated players, a short, structured session against a noise-free opponent can estimate strength faster and more repeatably than a similar number of human games — and can serve afterward as an unusually clear progress meter, because an opponent that never has a bad day is an opponent you cannot get lucky against.
Volunteer for the experiment
We need players of every strength — from a few hundred to as high as we can get. Weaker players are not less useful to this study; they are the heart of it. What participation involves:
- An established chess.com or lichess rating (roughly 50+ games in one time control, so the number has converged).
- Around 20–40 games against handicapped full-strength engines in ChessMend's Play mode, at your own pace — the adaptive ladder does the scheduling, and the first few games count as practice, not data.
- Sharing your results together with your rating, site and time control. Data is used in anonymized, aggregate form for the calibration study only.
In return you get an early read of your own ladder threshold, a genuinely luck-proof baseline to measure your future improvement against — and a hand in building the first properly calibrated odds-engine rating scale we know of.
Volunteer via the contact form →Or email us — the address is in the site footer. Tell us your rating, site and preferred time control, and we will send the protocol.