BrainRank IQ TestIQ TestBrainRankSign in

What Is an AI's IQ? Why a Score Above 140 Can't Be Compared to a Human's

What Is an AI's IQ? Why a Score Above 140 Can't Be Compared to a Human's

In brief: The IQ figures reported for AI are produced by giving models an IQ test standardized on humans and converting the result against a human reference group. Some models exceed 140 on publicly available tests, but scores drop by 20 to 40 points when the same format is swapped for unseen items. IQ expresses relative standing within a human population, so applying it to AI does not carry the same meaning.

The IQ figures reported for AI are produced by giving models a test built for humans and converting the result against a human reference group. Scores in the 140s have been reported on publicly available tests, but swap the same format for unseen items and the scores fall substantially. This article separates what those numbers show from what they do not.

How AI IQ scores are produced

The best-known effort is the ongoing measurement at Tracking AI. Each model is given the 35 visual pattern items that Mensa Norway publishes openly, and the number correct is converted into an IQ-equivalent figure using human norming data. The format belongs to the same family as Raven's Progressive Matrices: a language-independent measure of fluid reasoning.

As of August 2026, several models have been reported at 140 or above. Mapped onto the human distribution, an IQ of 140 sits around the top 0.4% (see what percentage has an IQ of 140). That the phrase 'top X%' presupposes a human population is exactly the point that matters later.

557085100115130145
Where IQ 140+ sits on the distribution (mean 100, SD 15): roughly the top 0.38%.

A 20-to-40-point gap between public and private tests

The same site also measures those models on a private item set — one written by a Mensa member that has never appeared on the public internet. There, even the leading models land in the 120s, a gap of 20 to 40 points against the public test.

The most natural explanation for the gap is exposure to training data. Published items and their worked solutions may sit inside a model's training corpus, and a high score under those conditions is not evidence that unseen problems were solved by reasoning. Benchmarks leaking into training data is known as contamination, and it is a general problem in AI evaluation.

Human IQ testing works the same way. Someone who has already worked through the same items will score higher, but intelligence did not rise — the measurement broke. We cover this in how accurate are free IQ tests.

Three reasons IQ does not transfer to AI

  • There is no reference group: IQ is designed to express relative position within a human age group (see what is IQ). An AI is not part of that group, so a converted figure means no more than 'a human here would sit at this position.'
  • The profile has a different shape: in humans, scores across domains correlate, and what they share is summarized as the g factor. AI models show gaps between strengths and weaknesses that look nothing like the human pattern, so the premise for collapsing them into one number does not hold.
  • Only one format is being measured: 35 visual items capture one facet of fluid reasoning. Clinical batteries such as the WAIS combine many subtests precisely because a single format cannot stand in for the whole.

These three mirror the reasons simple cross-country IQ comparisons are difficult (see how to read average IQ by country). The moment the assumptions behind standardization break, the number stops meaning what it appears to mean.

The trend line is still informative

That comparison with humans fails does not make the measurement worthless. Run the same item set through the same procedure and you can track differences between models and changes over time. The gap between public and private tests is itself a useful estimate of how much contamination is in play.

The issue is how the number is read. 'This AI has an IQ of 140' is wrong as a comparison with people; 'this model answered this many of 35 items in this format' is a legitimate fact. What the number measured has to travel with the number.

An AI's score does not change what a human's score means

That an AI can solve pattern items does not reduce the point of a human solving them. An IQ test is a tool for locating human cognition relative to other people, not a contest against machines.

The practical question is whether you can verify what an AI produces. We treat that in AI and cognitive offloading and in how to build critical thinking.

Check your own number against a human baseline

BrainRank's free IQ test scores 20 questions across spatial reasoning, pattern recognition, logical reasoning, and classification (about 10 minutes) using item response theory (IRT), returning an estimated IQ on the mean-100, SD-15 scale along with domain tendencies. Because the reference group is human, none of the conversion problems in this article apply. Results are estimates for entertainment and self-understanding, not a medical or diagnostic assessment.

Frequently asked questions

What is ChatGPT's IQ?
Scores in the 140s have been reported on the public Mensa Norway test. That figure is a conversion of how many of 35 visual pattern items were answered correctly, scaled against human norms, and it falls sharply when the items are replaced with unseen ones. It does not mean the model's intelligence is 140.
If AI scores higher than humans, is AI smarter?
The comparison does not hold. IQ is designed to express relative standing inside a human age group, and an AI is not a member of that reference group. Two numbers sitting side by side were not produced on the same ruler.
Why do scores drop on private tests?
Publicly posted items and their solutions may be present in a model's training data. Benchmarks leaking into training data is called contamination, and it is a shared problem across AI evaluation.

Share this article

Related articles

Editorial note & disclaimer

BrainRank Editorial Team

This article was written and edited by the BrainRank Editorial Team with reference to academic literature on psychometrics, including CHC theory and Item Response Theory (IRT). Statistics and percentages are calculated from a normal distribution model with a mean of 100 and a standard deviation of 15.

The tests on this site provide estimates for entertainment and self-understanding purposes only. They are not medical or clinical assessments, nor official psychological (intelligence) tests.

Sourcing standards and correction policy →