An IQ score is a rank, not a quantity
This is the single idea that makes everything else on the page follow. Nobody possesses 118 units of intelligence. A score of 118 says that a performance beat roughly 88% of the reference distribution on one particular set of tasks, on one particular day.
Two consequences fall straight out. Scores cannot be averaged across different tests, because each is a rank within its own distribution. And the same person's number moves when the comparison group changes — which is usually the explanation when a school score from decades ago looks better than a result today.
The scale, and where the numbers come from
Modern IQ uses a deviation scale: the population mean is fixed at 100 and the standard deviation at 15. Percentiles then follow from the normal curve rather than from anything about the questions.
| Score | Percentile | Conventional band | Out of 100 adults |
|---|---|---|---|
| 130 and up | 98th+ | Very superior | About 2 people |
| 120–129 | 91st–97th | Superior | About 7 people |
| 110–119 | 75th–90th | High average | About 16 people |
| 90–109 | 25th–74th | Average | About 50 people |
| 80–89 | 9th–24th | Low average | About 16 people |
| 70–79 | 2nd–8th | Borderline | About 7 people |
| Below 70 | Under 2nd | Extremely low | About 2 people |
The band labels come from clinical convention and carry no moral content whatsoever. Two thirds of all scores fall between 85 and 115. A score of 95 sits in the widest, most crowded part of the distribution alongside half the adult population, and describes a person about as usefully as a shoe size does.
Why the scale is not evenly spaced
This catches almost everyone, and it is worth ten seconds of attention. Percentile ranks are packed tightly in the middle of the curve and stretched thin at the edges:
- The ten points from 100 to 110 cover 25 percentile ranks.
- The ten points from 130 to 140 cover barely two.
So a ten-point gap near the average separates two genuinely different positions in the population, while the same gap at the top separates almost nobody. Differences at the high end are far smaller in human terms than they look written down — which is the opposite of how they are usually discussed.
It also means the percentile is least stable exactly where most people land. The curve is steepest at the middle, so a few points of measurement error sweep across a wide span of ranks. That is why our result page prints a percentile range rather than a single figure.
Read the band before you read the number
Every psychological measurement carries error, and the quantity describing it is the standard error of measurement. It comes from the scale's standard deviation and the test's reliability coefficient:
SEM = SD × √(1 − reliability) · 95% interval = 1.96 × SEM
| Reliability | Standard error | 95% interval |
|---|---|---|
| .98 — WAIS-IV Full-Scale IQ | 2.1 points | ±4 |
| .90 | 4.7 points | ±9 |
| .81 — the 16-item structure used here | 6.5 points | ±13 |
| .70 | 8.2 points | ±16 |
A result of 112 on this test therefore means a true score most likely between 99 and 125 — a range that spans “ordinary” and “well above average”. A two-point difference between you and a friend carries no information at all.
Before trusting any test's number, ask what its reliability is. Reviewing the free IQ tests ranking for this subject in July 2026, we could not find one that publishes a reliability coefficient or a confidence interval for its own scale, and several quote the WAIS-IV's ±4 alongside their own results — which is the precision of a different instrument entirely. A test that cannot answer the question has not been calibrated.
Why boundaries matter less than people think
People chase 130 because of what it unlocks socially, then read 128 as a failure and 131 as an achievement. With a ±13 band those two results are indistinguishable; even on a supervised instrument with ±4 they are barely separable. Any threshold applied to a measurement with error will misclassify people at the edge in both directions, and no amount of confidence in the number changes that.
Reading a score from this test specifically
Ours is sixteen items across four reasoning domains, six options each, untimed. That produces a few specific reading rules.
The total is worth more than the breakdown
Each domain rests on four items. One lucky guess moves a domain score by a quarter. We report domain performance as a count out of four rather than a per-domain IQ, and flag a difference only when it reaches three items or more. Below that, the domains are not reliably distinguishable — which is stated on the page rather than left for you to infer.
A very low score may be measuring nothing
With six options, someone answering entirely at random reaches five correct or better about eleven times in a hundred. A raw score of five or under is inside what guessing produces, so the result page says so instead of presenting a figure. Rushing, guessing, or reading the verbal items in a second language all land in the same place.
A perfect run returns 132, not 145
Sixteen items cannot separate people at the very top of the distribution. The scale stops where the evidence stops. If you clear all sixteen, the honest reading is “above the ceiling of this test”, and a supervised assessment is the only way to put a number on it.
The methodology page gives the full conversion table and the arithmetic behind it.
Why two tests give you different numbers
Because each scale has its own standard deviation, its own norming sample and its own error band. A thirteen-point spread between two short online tests is entirely expected. Neither result is the true one; the defensible reading is wherever their confidence intervals overlap.
The standard-deviation problem is worth an example, because it produces numbers that look wildly different and are not:
| Scale | SD | 98th percentile |
|---|---|---|
| Wechsler (WAIS, WISC), Stanford–Binet 5 | 15 | 130 |
| Older Stanford–Binet forms | 16 | 132 |
| Cattell III B | 24 | 148 |
Those are one threshold written three ways, not three levels of difficulty. Comparing a Cattell figure with a Wechsler figure without converting is the most common error in this area, and it inflates the Cattell result by eighteen points.
What a retake does to your score
A second attempt raises scores by about four IQ points on average, with no change whatsoever in underlying ability. That comes from a meta-analysis of 50 studies covering 107 samples and 134,436 participants (Hausknecht et al., 2007). What matters most is which items you see again: repeating identical items produces close to six points, while a parallel form with fresh items produces about three. By a third sitting the accumulated effect is larger still.
Your first honest attempt is the most informative score you will get. Retake it if you like, subtract about five points before believing the improvement, and leave several months so item memory fades.
What the number does not tell you
Cognitive ability predicts job performance considerably less strongly than the figure usually quoted. The familiar correlation near .51, from a 1998 research summary, was re-examined in 2022 and found to rest on an overcorrection for range restriction; the revised estimate is .31 (Sackett et al., 2022), which accounts for under ten per cent of the variation in performance. A meta-analysis restricted to this century reported a mean observed validity of .16.
For one individual the translation is blunt: your score forecasts very little about your career, and an error band of ±13 makes the forecast weaker still. An IQ figure says nothing about judgement, originality, persistence, or how someone holds up during a hard month — and it must never be used to self-diagnose anything. If you are worried about a real cognitive change, that is a question for a clinician rather than a web page.
Where this sits
This page explains how to read any IQ score. The methodology page covers how ours specifically is produced. The terms page lists the purposes a result from here cannot serve, and the about page explains what we refuse to publish and why.
Sources
- Condon, D. M. & Revelle, W. (2014). The International Cognitive Ability Resource. Intelligence, 43, 52–64.
- Young, S. R. & Keith, T. Z. (2020). Convergent validity of the ICAR16 and WAIS-IV. Journal of Psychoeducational Assessment, 38(8), 1052–1059.
- Hausknecht, J. P. et al. (2007). Retesting in selection: a meta-analysis of coaching and practice effects. Journal of Applied Psychology, 92(2), 373–385.
- Sackett, P. R., Zhang, C., Berry, C. M. & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection. Journal of Applied Psychology, 107(11), 2040–2068.
- Mensa International — published qualifying criteria and accepted-test list. Checked 25 July 2026.
Checked against the sources above on 25 July 2026. Found something wrong? Tell us and it will be corrected with a note.