Is AI smarter than humans?
On specific tasks, yes — frontier models now meet or exceed human baselines on PhD-level science, competition math (Gemini Deep Think won IMO gold), and coding, where SWE-bench Verified went from ~60% to near 100% in a year. Yet the same top model reads analog clocks correctly 50.1% of the time. AI intelligence is jagged, not uniformly above or below ours.
Why — the first-principles explanation
The question assumes intelligence is one dial — a single quantity you could turn up until the AI number passes the human number. That assumption is the reason the question feels unanswerable. It isn't one dial. Calculators have crushed humans at arithmetic since the 1960s, and nobody concluded calculators were smarter than people. We just quietly stopped counting arithmetic as intelligence.
What we actually measure are specific competencies, and there the record is lopsided in AI's favor and getting more so. Stanford's 2026 AI Index reports frontier models meeting or exceeding human baselines on PhD-level science questions, multimodal reasoning, and competition mathematics; Gemini Deep Think earned an International Mathematical Olympiad gold medal. SWE-bench Verified went from about 60% to near 100% in a single year. Any one of these, seen in a human, would mark them as exceptional.
Then comes the fact that detonates the whole framing. That same top model reads analog clocks correctly just 50.1% of the time — a coin flip on a skill most humans master around age seven. This is not a small footnote. It reveals that these systems aren't climbing a ladder that runs from "dumb" through "average human" to "genius." They're building a shape that's spiky: mountainous in some directions, at ground level in others, with no obvious rule for which is which. François Chollet's ARC-AGI benchmark was built around exactly this, testing skill-acquisition efficiency on unknown tasks rather than accumulated knowledge — because, as ARC's framing puts it, intelligence is marked by generalization "rather than skill itself."
The mechanism explains the shape. These models learned by finding statistical structure in enormous amounts of human-produced data. Where humanity wrote down a great deal — code, math proofs, scientific text — the model has rich structure to absorb, and it can exceed nearly any individual because no individual read all of it. Where humans learned from living in bodies in a physical world — reading a clock face, knowing a glass will spill, judging whether someone means it — the training signal is thin, because we never wrote it down. We just knew. AI is superhuman precisely where humans documented themselves, and childlike where we didn't. So "smarter" is the wrong word. Ask instead: at what, compared to whom, and does it know when it's wrong? On that last one, humans still hold a large lead.
An example that makes it click
Imagine meeting someone who has read every book ever written but has never once left the library.
Ask about the fall of Rome, protein folding, or Portuguese grammar and you get a better answer than any professor could give — he's read more than any professor ever could. You'd swear he's the smartest person alive.
Then you hand him a wall clock and ask the time, and he squints and guesses. You ask if the coffee is too hot to drink and he has no idea. Not because he's stupid — because nobody ever wrote a book called What Time It Looks Like When the Little Hand Is There. We all just learned that by looking at clocks. He never got to look at anything.
That's exactly the shape of AI intelligence. Superhuman where humanity took notes. Toddler-level where we didn't bother, because we assumed everyone would just live it.
Key facts
- Frontier models now meet or exceed human baselines on PhD-level science questions, multimodal reasoning, and competition mathematics, per Stanford HAI's 2026 AI Index.
- Gemini Deep Think earned a gold medal at the International Mathematical Olympiad.
- SWE-bench Verified (autonomous software engineering) performance rose from approximately 60% to close to 100% within a single year.
- The same top-performing model reads analog clocks correctly just 50.1% of the time — a skill most humans acquire around age seven.
- AI agents complete roughly 66% of real computer tasks on the OSWorld benchmark, up from 12% previously, still failing about one task in three.
- ARC-AGI, created by François Chollet, deliberately measures 'skill-acquisition efficiency on unknown tasks' rather than crystallized knowledge, on the principle that intelligence is marked by generalization 'rather than skill itself.'
- 73% of AI experts expect positive job impacts versus 23% of the general public — a 50-point perception gap on what these capabilities mean.
▶ The 60-second explainer (script)
Is AI smarter than humans? On specific tasks, yes — and it's not close. Frontier models now meet or beat human baselines on PhD-level science, multimodal reasoning, and competition math. Gemini Deep Think won a gold medal at the International Math Olympiad. Coding went from about sixty percent to near a hundred on SWE-bench in one year. Any one of those in a person would make them exceptional. And then: that same top model reads an analog clock correctly fifty point one percent of the time. A coin flip. On a skill most kids get by age seven. That fact breaks the whole question. Because the question assumes intelligence is one dial you turn up until AI passes us. It isn't. Calculators crushed us at arithmetic in the nineteen sixties and nobody said calculators were smarter than people — we just quietly stopped counting arithmetic as intelligence. What's really happening is a shape, not a ladder. Spiky. Mountains in some directions, ground level in others, and no obvious rule for which. Here's the mechanism. These models learned by finding structure in enormous amounts of human-produced text. Where humanity wrote a lot down — code, proofs, scientific papers — there's rich structure to absorb, and the model beats almost any individual, because no individual read all of it. But where humans learned by living in bodies — reading a clock face, knowing a glass will spill, telling whether someone means it — we never wrote it down. We just knew. So there's no training signal. AI is superhuman exactly where we documented ourselves, and childlike where we didn't. Imagine meeting someone who read every book ever written but never left the library. Rome, protein folding, Portuguese grammar — better than any professor. Then you point at the wall clock and ask the time, and he squints and guesses. Not stupid. Nobody ever wrote a book called 'What Time It Looks Like When The Little Hand Is There.' We all just looked at clocks. He never got to look at anything. So stop asking 'is it smarter.' Ask: at what, compared to whom, and does it know when it's wrong? On that last one, we're still well ahead.
What authoritative sources say
People also ask
Has AI passed human intelligence overall?
There's no coherent way to measure 'overall.' It exceeds human baselines on PhD-level science, competition math, and coding, while failing tasks a seven-year-old handles — like reading an analog clock, at 50.1% accuracy.
Why is AI great at math but bad at simple physical things?
Because it learned from what humans wrote down. We wrote extensively about math and code; we never wrote down how a clock face looks or that a glass will spill. We just lived it, so there's little training signal.
Does beating a benchmark mean it's intelligent?
Not necessarily. That's why François Chollet built ARC-AGI to test skill acquisition on unknown tasks instead of accumulated knowledge — measuring generalization rather than memorized skill.
What do humans still clearly do better?
Knowing when we're wrong, learning from a couple of examples, judging what matters, and acting in the physical world. AI agents still fail about one in three real computer tasks.
Is IQ a useful way to compare AI and humans?
No. IQ is calibrated on the human distribution of correlated abilities. AI abilities aren't correlated the same way — a system can win an IMO gold medal and fail at clock reading, which no human IQ score can describe.