Here is the clearest definition of artificial general intelligence in use anywhere in 2026. It does not come from a lab bench, a textbook, or a cognitive scientist. It comes, reportedly, from a commercial contract: Microsoft and OpenAI are said to have agreed that AGI is reached when OpenAI builds a system capable of generating $100 billion in profit (New Scientist, 2025). Read that again. The milestone we are told will remake civilization within a few short years has been operationally defined, by two of the companies closest to it, as a revenue number. When "human-level intelligence" is measured in dollars, the conversation stopped being about science a while ago.

I build with these tools every day. I have watched AI agents do in an afternoon what once took a team a week, and I do not say that with a shrug — it changed what I believe about who gets to create. So I am not here to tell you the machines are dumb. I am here to tell you that "AGI" has become a word we point at instead of a thing we can measure, and that the gap between the two is where all the hype lives. Separate the science from the marketing and a soberer, more honest picture appears — one that is, if anything, more interesting than the myth.

The Definition Nobody Can Agree On

Start with the uncomfortable foundation: there is no agreed definition of AGI. Not a fuzzy-around-the-edges one — no shared operational test at all. Google DeepMind frames it as a system that can outperform all humans on a battery of cognitive tests. François Chollet, who built the field's toughest reasoning benchmark, defines intelligence as efficient skill acquisition — how well a system learns genuinely new things, not how much it has already stored. Huawei researchers have suggested AGI requires embodiment, a body to act in the world. Microsoft and OpenAI, reportedly, went with the hundred billion dollars. And as IEEE Spectrum bluntly noted in 2025, some people simply define AGI "by vibes."

This is not a trivia problem. It is the whole story. Because the finish line is unsettled, any headline that says "AGI is three years away" is unfalsifiable — you can slide the goalposts in either direction and never be wrong. An undefined target is one you can always claim to be approaching. That is not a scientific forecast. It is rhetoric wearing a lab coat.

What the Science Actually Measures

Underneath the noise, real researchers are doing real measurement — and the results are humbling. Consider ARC-AGI, a family of puzzles built to test fluid intelligence: the ability to look at a few examples, abstract a brand-new rule, and apply it. This is different from crystallized intelligence, the vast recall of facts and patterns that large language models are extraordinary at. Chollet's insight is that recall is not reasoning. "To solve any problem," he has said, "you need some knowledge, and then you're going to recombine that knowledge on the fly." The recombining is the hard part.

On ARC-AGI-2, launched in March 2025, the best AI systems score about 16 percent. The average human scores 60 percent (IEEE Spectrum, 2025). That is the widest frontier-AI-versus-ordinary-person gap of any major benchmark — a reality check you can hold in your hand. And when a model did crack the original ARC-AGI, the asterisk was enormous: OpenAI's o3 reportedly hit 88 percent at an estimated $20,000 of compute per puzzle, on a model that was never publicly released. A human needs a cup of coffee. As Melanie Mitchell of the Santa Fe Institute put it, even a strong benchmark score "doesn't capture what people mean when they say general intelligence." Benchmarks are proxies, not proof.

An undefined finish line is one you can always claim to be approaching. That is not a forecast. It is rhetoric wearing a lab coat.

The Insiders and the Field

Now the honest counterargument, because there is one. The leaders of OpenAI, Anthropic, and Google DeepMind have each said recently that they expect AGI within a few years. These people see the internal, unreleased models. Their optimism cannot simply be waved away as ignorance — they are as close to the frontier as anyone alive.

But proximity to the frontier is also proximity to the people who profit from the story. AGI narratives drive fundraising, valuations, and recruiting; the insiders most bullish on the timeline are the ones whose balance sheets bend to it. And when you step outside the labs, the independent research community disagrees at scale. In a survey of 475 researchers compiled by the Association for the Advancement of Artificial Intelligence, 76 percent said scaling up current approaches is "unlikely" or "very unlikely" to achieve AGI (AAAI, via New Scientist, 2025). Eighty percent said public perception of AI overshoots reality. As Thomas Dietterich, who contributed to that report, observed, systems "proclaimed to be matching human performance… still make bone-headed mistakes." Even the industry's pivot to "reasoning" and inference-time scaling — spending more compute at the moment of answering — is, as Princeton's Arvind Narayanan noted, "unlikely to be a silver bullet." The pivot itself is a quiet admission that pure scaling stalled.

76%of surveyed AI researchers doubt current methods scale to AGI (AAAI, 2025)
80%say public perception of AI capability exceeds reality (AAAI, 2025)
16% vs 60%best AI vs average human on ARC-AGI-2 (IEEE Spectrum, 2025)
2047median expert forecast for a 50% chance machines outperform humans at every task (Grace et al., 2023)

Why the Hype Keeps Working

If the science is this sober, why do the headlines run so hot? Partly incentives, which we have covered. Partly the unfalsifiable definition, which lets everyone claim progress toward a line no one has drawn. And partly because we have been here before and keep forgetting it. In 1970, Marvin Minsky — a founding father of the field — told Life magazine that "in three to eight years we will have a machine with the general intelligence of an average human being," one that could read Shakespeare, grease a car, play office politics, and tell a joke. That was fifty-five years ago. The office politics and the physical grease-work remain unsolved. Confident near-term AGI predictions from credible insiders are not new. Their track record is just bad.

The measured researcher view holds the uncertainty honestly. A survey of 2,778 authors at top AI venues put the median 50 percent chance of machines outperforming humans at every task at 2047 — yet the 50 percent mark for all occupations becoming fully automatable sat all the way out at 2116 (Grace et al., 2023). Nearly a century of daylight between "can do every task" and "can do every job." That gap is where the human work lives.

What to Do With an Undefined Finish Line

Geoffrey Hinton, who won a Nobel for the foundations of this technology, likes to say, "We're building alien beings." I think that is exactly right, and exactly why the AGI question is the wrong one to organize your life around. These systems are not lesser humans creeping toward parity on some single dial. They are strange instruments with jagged strengths — superhuman at recall, childlike at novelty — and the "when is it human-level" frame flattens all of that into a number a marketer can sell.

So here is my one concrete step, the thing to actually do: stop waiting for AGI and start measuring the tool in front of you against the work in front of you. Take one real task this week — a draft, a dataset, a design — and test what today's models genuinely do and where they fall on their face. You will learn more about the future of intelligence from that hour than from a year of timeline headlines. The finish line is undefined on purpose. The tool on your desk is not. Build with the one you can measure.