The specification in one table
| Property | Value |
|---|---|
| Items | 16 |
| Domains | 4, with 4 items each |
| Options per item | 6 |
| Time limit | None |
| Typical completion | About ten minutes |
| Scale | Mean 100, standard deviation 15 |
| Reported range | 55–145 |
| Published reliability | .81 (inherited — see below) |
| Standard error of measurement | 6.54 points |
| 95% confidence interval | ±13 points |
| Score by guessing alone | About 2.7 of 16 |
| Minimum age | 16 |
| Where scoring happens | In your browser |
Where the item structure comes from
The test follows the structure of the ICAR16 — a sixteen-item public-domain measure built by David Condon and William Revelle, published in Intelligence in 2014 and used in research since. Four item types, four items apiece: letter and number series, matrix reasoning, verbal reasoning, and three-dimensional rotation.
A distinction worth stating plainly, because most sites in this niche blur it. The structure is theirs. The items are ours, written to the same formats. We do not reproduce the ICAR item bank, and that has a consequence we return to below: every psychometric figure published for the ICAR16 describes their items, not ours, and is inherited rather than earned.
What each domain asks
Letter and number series
A short sequence, and you supply the next term. The rule may be a growing difference, a term built from the one before it, or two sequences interleaved. When the sixteen-item set was validated against the WAIS-IV in adults, this was the only one of the four domains that aligned strongly with fluid reasoning as theory predicts (Gf, r = .70) — which makes it the domain with the clearest interpretation of the four.
Matrix reasoning
A three-by-three grid of figures with the last cell missing, in the tradition of Raven's Progressive Matrices. Rules run along rows and down columns, and the harder items require both to hold at once. Matrix reasoning is widely described as the purest available marker of fluid intelligence; the validation of this item structure complicates that claim, which we take up under what the evidence actually says.
Verbal reasoning
Meaning and inference: distinguishing near-synonyms, holding a quantifier steady through an argument, recognising the relation behind an analogy. This domain leans on language exposure more than the other three, and a result here is worth discounting for anyone testing in a second language.
Three-dimensional rotation
A cube carrying a different mark on each face is shown from one angle, and you decide which option shows the same cube turned in space. This is the only one of the four whose population scores rose across the 2006–2018 period in which the other three fell.
How the rotation items are constructed
These are the easiest items in the format to get wrong, and the error is invisible in a screenshot, so the construction is worth setting out.
You can see three faces of the target. You therefore cannot know what is on the other three. If an option shows a mark you have never seen, in a position nothing constrains, the item is a coin toss rather than a question — and it would still look like a rotation puzzle.
The rule we enforce instead rests on a fact about cubes: any two visible faces in known positions fix the whole object, so the third face is not a free choice. From that:
- The correct option shows the target's own three faces, cycled one position around the corner they share — the same three marks, visibly turned.
- Three distractors show those same three marks running the other way around the corner. Reversing that cycle requires a reflection, not a rotation, so they are impossible.
- Two distractors keep a valid pair of familiar faces and place an unseen mark in the slot that pair has already determined.
Every distractor is therefore rejectable from the target alone, and exactly one option is a genuine rotation. The generator checks all of this when the page loads — six distinct options, exactly one valid rotation, no second valid orientation leaking in as a hidden correct answer, and every distractor deducible — and reports an error rather than serving a broken item.
Why six options
With six options, answering all sixteen at random returns about 2.7 correct. That places a blind run near the floor of the scale rather than in the middle of it, which is the point: on a four-option test the same behaviour lands closer to average and the score flatters the test-taker.
It also sets a threshold worth publishing. Under a binomial model with sixteen items and a one-in-six chance of a correct guess, someone answering entirely at random reaches five correct or better about eleven times in a hundred, and six or better under four times in a hundred. A raw score of five or under cannot be distinguished from guessing, and the result page says so in place of treating the number as a measurement.
Why there is no clock
Speed carries real information about cognitive ability. In an unsupervised browser it also carries information about your connection, your device and whether someone knocked on the door, and the two cannot be separated afterwards. A time limit in this setting scores the circumstances as much as the reasoning.
The researchers who built this item structure reached the same conclusion, noting that unsupervised online testing has to allow for interruptions and technical problems. Response times are still recorded and used for one purpose only: the result page describes how you approached the items — quickly and accurately, quickly and carelessly, slowly and at the edge of difficulty — without any of it touching the score.
How a raw score becomes an IQ figure
Your raw score is the count of correct answers, nothing more. There is no weighting by item difficulty and no item-response model. It is placed on the deviation scale directly:
z = (raw − 8.6) ÷ 3.5
IQ = 100 + 15z, clamped to the range 55–145
Which produces, end to end:
| Raw | IQ | Raw | IQ | Raw | IQ |
|---|---|---|---|---|---|
| 0 | 63 | 6 | 89 | 12 | 115 |
| 2 | 72 | 8 | 97 | 14 | 123 |
| 4 | 80 | 10 | 106 | 16 | 132 |
Two things follow that are easy to miss. A perfect run returns 132, not 145: sixteen items cannot separate people at the very top of the distribution, and the scale stops where the evidence does. And IQ here is a rank rather than a quantity — 118 does not mean 118 units of anything, it means this performance beat roughly 88% of the reference distribution on these particular tasks.
The two constants deserve their own line. 8.6 and 3.5 are assumed values for this item structure, not estimates from a sample we collected. They are the single largest source of uncertainty on the page and the reason the percentile is labelled an estimate.
The error band, and the arithmetic behind it
The published internal consistency for this sixteen-item structure is .81. The standard error of measurement follows from the scale's standard deviation and that coefficient:
SEM = 15 × √(1 − .81) = 6.54 points
95% interval = 1.96 × 6.54 = ±13 points
So a result of 112 means a true score most likely between 99 and 125 — a range spanning “ordinary” and “well above average”. A supervised WAIS-IV, with full-scale reliability near .98, carries about ±4 by the same arithmetic. The gap between those two figures is the honest measure of what a browser test can and cannot tell you.
Note that the band is the same width for everyone. Reliability under this model is a property of the test rather than of the person, so we do not print individually calculated intervals. An earlier version of this site did, derived from an item-response model; that model was removed along with the items it was fitted to, and printing a person-specific error from assumed parameters was more precision than the instrument supports.
Why domain scores are reported out of four
Each domain rests on four items. Four items cannot support a confident conclusion about anything: one lucky guess moves a domain score by a quarter, and one misread instruction moves it the other way by the same amount.
So the result page reports domain performance as a count out of four rather than converting it into a per-domain IQ figure. Converting would invite exactly the over-reading the number cannot bear. A difference between your strongest and weakest domain is called out only when it reaches three items or more; below that the page states, in place of a ranking, that the domains are not reliably distinguishable on this attempt.
This inverts the usual arrangement, and the result page says so: on a sixteen-item test the total is the trustworthy part and the breakdown is not.
What the evidence actually says
The strongest evidence for this item structure is a head-to-head study against the clinical standard: total scores correlated .81 with WAIS-IV Full-Scale IQ, and the latent general factors of the two instruments correlated .94 (Young & Keith, 2020). The original development paper reported .75 against Raven's Advanced Progressive Matrices after correction for range restriction (Condon & Revelle, 2014).
Those figures need their caveats attached rather than dropped. The WAIS-IV comparison used 97 participants — 67 university volunteers and 30 people tested at a university assessment centre — and corrected for range restriction and unreliability, so it describes a relationship between constructs rather than agreement between two score reports. The authors called for replication in larger samples. None has been published.
The same study produced a result its authors flagged as surprising: only letter and number series aligned strongly with fluid reasoning, while matrix reasoning, verbal reasoning and three-dimensional rotation all correlated most strongly with visual-spatial ability (Gv, r = .35–.75) — inconsistent with what CHC theory predicts. A 2025 validation of a children's version reached the opposite conclusion for matrix reasoning, finding it loaded on Gf as theory expects.
Two well-conducted studies, two different answers, different age groups. We report the domain profile as a description of performance on four specific tasks, not as a map of anyone's cognitive architecture, because that is as far as the evidence reaches.
What this test has not established
Every free IQ test sits roughly here. Few say so, and the difference is the saying.
| Criterion | Status |
|---|---|
| Standardisation sample of our own | Not established |
| Reliability measured on our own item forms | Inherited from the published structure, not measured |
| Factor analysis confirming four distinct domains | Not established |
| Correlation against a supervised instrument, our forms | Not established |
| Item difficulties estimated from responses | Not established — items are assigned by design |
There is a second uncertainty underneath the ±13, and almost nobody states it. Measurement error describes how much a score bounces on a retake. Norm error describes how far the whole scale might be displaced, and it exists whenever a test is anchored to published population parameters rather than to a sample the publisher collected. Ours is anchored that way. The interval we print is a floor on the true uncertainty rather than the whole of it, and we cannot yet say by how much.
What the result page shows, in order of confidence
- Solid. Your total out of sixteen, and the broad band it places you in.
- Reasonable. The shape of your profile — which of your own domains came out stronger. You are your own comparison group, so this does not depend on norms; but each domain rests on four items, so a one-item difference is noise.
- Provisional. The exact IQ figure and the percentile, both resting on assumed population parameters.
The page also prints the worked reasoning for all sixteen items — what you chose, what was correct, and why. Sixteen items cannot measure anyone accurately; sixteen worked explanations are the part of the output with unambiguous value.
Where scoring happens, and what leaves your device
The scoring rules and the norms ship with the page. Your answers are read by JavaScript on your own device and the result is drawn on the same page, so nothing has to be transmitted anywhere for the arithmetic to run. There is no email field, no card field and no paid tier anywhere on this site, so nothing is withheld and nothing is asked for.
Two values are written to your browser's local storage: your last score and its timestamp, used only to tell you whether a change on a retake clears the reliable-change threshold. Sharing a result encodes the score, the raw total and the four domain counts into the link itself — there is no database and no server-side record. The privacy page itemises all of it, and the cookie page names every stored key.
What changed, and when
| Date | Change |
|---|---|
| 25 July 2026 | Rebuilt on the ICAR16 structure: 16 items, four domains, six options, no clock. Working memory and processing speed removed. Item-response scoring replaced with a raw-score conversion and a fixed error band. Worked explanations for every item added to the result page. |
| 22 July 2026 | Previous rebuild: 21 items in three domains with two timed tasks and a five-index composite. Superseded. |
Removing two measured indices is a reduction in what the test claims, and it was the right direction. A browser can present a memory task at a controlled rate, so the earlier decision to add one was defensible; what it could not do was norm the result against anything. Reporting fewer things with a stated error beats reporting more things with an implied one.
Sources
- Condon, D. M. & Revelle, W. (2014). The International Cognitive Ability Resource: development and initial validation of a public-domain measure. Intelligence, 43, 52–64. doi:10.1016/j.intell.2014.01.004
- Young, S. R. & Keith, T. Z. (2020). An examination of the convergent validity of the ICAR16 and WAIS-IV. Journal of Psychoeducational Assessment, 38(8), 1052–1059.
- Dworak, E. M., Revelle, W. & Condon, D. M. (2023). Looking for Flynn effects in a recent online U.S. adult sample. Intelligence, 98, 101734. doi:10.1016/j.intell.2023.101734
- Validation of a children's form of the same item structure, Behavior Research Methods, 2025.
- Pearson Clinical Assessment — Wechsler Adult Intelligence Scale documentation, for the reliability and error figures quoted. Checked 25 July 2026.
Every figure on this page was checked against the source above on 25 July 2026. If something here is wrong or out of date, tell us and it will be corrected with a note.