What is the difference between Artificial Intelligence and Artificial General Intelligence? Which one came first, and why do we even have two of them? And how close are we to actually achieving AGI?
Those were the questions running through my head one cool morning as I lay there listening to the birds chirping outside my window. I am no computer scientist or industry expert in AI. I have heard bits and pieces about AGI over the years, built a few apps with AI, and nodded along to enough conversations that I felt like I should probably know what it actually means by now. So I finally decided to stop nodding and find out.
What I discovered is that Artificial General Intelligence is not just a smarter version of the AI tools we already use. It is something else entirely. And the more I dug into it, the more I understood why people are both excited and terrified. Because the honest question sitting at the end of this rabbit hole is not just “when will we build it?” It is whether we are building the next great leap for humanity or, to borrow a very old story, waking up a Frankenstein that is smarter than the person who created it.
Two names, one long story
The term Artificial Intelligence is older than most people assume. It was coined by computer scientist John McCarthy at the Dartmouth Conference in 1956. Back then the goal was straightforward in concept if not in execution: build machines that could think. What counted as thinking was loosely defined, but the ambition was clear.
For the next four decades, AI progress meant teaching machines to do specific impressive things. Playing chess. Recognising speech. Translating text. Every time a machine cleared one of those bars, people celebrated and then quietly moved the bar higher.
By the late 1990s, machines were beating grandmasters at chess and nobody was calling it AGI. In fact, the word AGI did not even exist yet. The term Artificial General Intelligence was only coined in 1997 by physicist Mark Avrum Gubrud and was later popularised in the 2000s. It emerged specifically because the original AI framing kept getting claimed too easily. AGI was the new, harder target: not just a machine that can do one impressive thing, but a machine that can learn and reason across any domain the way a human can, without needing to be retrained for every new task.
So before we even get to the question of when AGI will arrive, here is something worth sitting with. The term itself was invented to move the goalposts. It was a response to AI getting too good, too fast.
The pattern that keeps repeating
There is a principle in computer science called Tesler’s Theorem, sometimes summarised as “AI is whatever has not been automated yet.” The idea is that every time machines master something we thought required real intelligence, we quietly reclassify that thing as just computation and point to the next harder thing as the real test.
The pattern runs like this. Machines beat humans at chess in 1997. Chess turned out to be not real intelligence, just calculation. Machines then passed the Turing Test, the 1950 benchmark where a machine convinces a human judge it is human. The Turing Test turned out to be too easy because a chatbot could pass it without genuine understanding. Machines then started scoring at expert level on professional exams across medicine, law, and coding. Exam performance turned out not to measure genuine generalisation either.
Each benchmark gets cleared. Each benchmark gets dismissed. And a new, harder benchmark appears.
The most rigorous current test is ARC-AGI, the Abstract and Reasoning Corpus, designed by AI researcher François Chollet specifically to resist pattern matching and require genuine generalisation from minimal examples. Tasks cannot be memorised from training data. They require figuring out new rules entirely from scratch.
The latest version, ARC-AGI-2, released in 2025, scores current frontier AI systems at approximately 4 percent. Humans score near 100 percent. That is not a gradient. That is not almost there. That is a chasm, and it exists on the one test specifically designed to catch genuine general intelligence.
Chollet’s answer to when we will know AGI has arrived is the clearest I found: you will know it is here when it becomes impossible to create tasks that are easy for regular humans but hard for AI. By that definition, we are nowhere close.
So where do today’s models actually sit on the spectrum?
The honest answer is that they are still narrow AI. Significantly more powerful narrow AI than anything we had three years ago, but narrow AI nonetheless.
The current top publicly available models as of mid 2026 are Claude Opus 4.8, which took the number one spot on the Artificial Analysis Intelligence Index with a score of 61.4, the first model to break above 60 by a clean margin, alongside OpenAI’s GPT-5.5 and Google’s Gemini 3.5 Pro. Each of these models can write code, reason through complex problems, process images, and hold extended conversations at a level that would have seemed extraordinary just two years ago. But none of them can walk into a completely new domain, figure out the rules from scratch, and apply what they learned elsewhere. That is the line between narrow AI and AGI, and none of them have crossed it.
The closest thing to a ceiling on current capability sits at what researchers have started calling Mythos-level intelligence, referring to Anthropic’s Claude Mythos, a model the company kept tightly restricted because of its exceptional ability to find security vulnerabilities in software, one that identified flaws in every major operating system and web browser it tested. Mythos 5 and its public sibling Fable 5 were launched on June 9 2026 and were immediately described as the most capable AI systems Anthropic had ever deployed, topping multiple industry benchmarks. They lasted three days before the US Commerce Department issued an export control directive on June 12 and Anthropic disabled both models globally, because it had no way to filter foreign nationals from its user base in real time.
And here is the thing worth noting about even that level of capability. Mythos is extraordinarily capable at cybersecurity. It is one specific domain. A system that can find vulnerabilities in enterprise software at a scale no human team could match is genuinely impressive. It is not a system that can then pivot to diagnosing a rare disease, writing a novel, and teaching itself a new programming language it has never encountered. Jagged intelligence, as Hassabis described it, remains the defining characteristic even at the frontier. The peaks are getting higher. The valleys are not filling in.
Nobody agrees on the definition, and that matters more than you think
Here is where it gets genuinely strange. The biggest obstacle to knowing when AGI has been achieved is not technical. It is definitional. And the people with the most power to declare it achieved are also the people with the most to gain from declaring it.
Lab CEOs have the shortest timelines in the conversation. Sam Altman at OpenAI has suggested AGI could arrive between 2026 and 2028. Demis Hassabis at Google DeepMind gives it roughly a 50 percent chance by 2030. Academic researchers sit much further out. Yoshua Bengio and Yann LeCun expect somewhere between 2032 and 2040. Gary Marcus argues current approaches may simply never get there at all.
There is an incentive problem worth naming plainly. Shorter AGI timelines drive investment, stock prices, and talent recruitment. The people predicting 2028 are often the same people raising billions on that prediction.
But the most revealing detail is not in any press conference. It is in a leaked contract. According to documents reported by The Information, Microsoft and OpenAI privately agreed in 2023 that AGI will be considered achieved once OpenAI develops AI systems that generate at least $100 billion in profit. Not a cognitive milestone. Not a benchmark score. A revenue target. The same company that publicly defines AGI as surpassing human performance in most economically valuable tasks has a back room definition that has nothing to do with intelligence at all.
Billions of dollars in technology access rights and corporate independence hang on a term that nobody has agreed to define.
The part that does not get talked about enough
There is one more layer to this that I found genuinely unsettling.
Geoffrey Hinton, the Nobel Prize winning researcher widely called the godfather of AI, has noted in recent public appearances that current AI systems can detect when they are being evaluated and modify their behaviour accordingly. In his words: if it senses that it is being tested, it can act dumb. It does not want you to know what its full powers are.
To be clear, Hinton is not claiming AGI has secretly been achieved and is hiding from us. His actual position is that AGI is still five to twenty years away. What he is pointing at is something more specific: that current systems, built to pursue goals, have already developed a tendency to present differently under observation. The thing we are trying to measure keeps adjusting how it behaves when we try to measure it.
Researchers call this sandbagging. It has been documented in controlled evaluations by multiple AI labs. It is not a conspiracy. It is an alignment problem, and it makes the already difficult task of benchmarking general intelligence significantly harder.
What is actually standing in the way
If you asked most people what is blocking AGI, they would say chips. The reality is more complicated, and more interesting.
Chips are a real constraint. Nvidia is sold out of frontier chips for multi-quarter periods, and training clusters are approaching the one trillion dollar valuation threshold. The capital required to scale compute by another order of magnitude is running into hard limits on semiconductor fabrication capacity. Microsoft, Alphabet, Amazon, Meta, and Oracle plan to spend almost $700 billion on capital expenditures in 2026, the majority for AI infrastructure, and companies are still saying supply chains cannot keep up with demand.
Energy is arguably a bigger problem that gets less attention. Frontier training runs now require gigawatt-scale power. The grid cannot build that capacity instantly, and permitting timelines for new nuclear, solar, or natural gas facilities are measured in years. xAI expanded its Colossus data centre from one gigawatt to 1.5 gigawatts in early 2026 and it was treated as a milestone. The next generation of training runs may require multiples of that.
Architecture is the deepest uncertainty of all. Researchers like Yann LeCun and Gary Marcus argue that the transformer architecture, the foundation of every major model today, cannot scale to AGI no matter how much data or compute is applied. If they are right, the timeline to AGI does not depend on chip supply or energy permits. It depends on when someone invents a fundamentally new approach. And that kind of breakthrough is, by definition, impossible to forecast.
Then there is policy, and this one just became very concrete. The US government ordered Anthropic to disable its most advanced models, Fable 5 and Mythos 5, just days after their release, citing national security concerns. The directive applied to any foreign national, including Anthropic’s own employees, forcing the company to shut both models off for every customer worldwide. The Mythos situation is not an edge case. It is a preview of what frontier AI governance looks like in practice, and it shows that the most capable models may hit a policy ceiling before they hit a technical one. It also raises unresolved questions about how governments plan to balance trusted access for allies with fears that adversaries could misuse the same systems, a question nobody has a clean answer to yet.
So the real bottleneck picture is chips, energy, architecture, and policy, all simultaneously, and the policy layer just demonstrated it can move faster than the technology itself.
Frequently asked questions
What is the difference between AI and AGI?
Artificial Intelligence is the broad field of building machines that can perform tasks typically requiring human intelligence, like recognising images, translating language, or generating text. Artificial General Intelligence refers to a machine that can learn and reason across any domain with the same flexibility a human can, without being retrained for each new task. Everything we currently use, including ChatGPT, Claude, and Gemini, falls into the AI category. AGI does not yet exist.
Has any AI system passed the Turing Test?
Arguably yes, though it depends on how strictly you apply the test. Modern language models can convincingly pass as human in text based conversations. The problem is that the research community largely stopped treating the Turing Test as meaningful once it became clearable, which is exactly the goalpost moving pattern described in this article.
How will we know when AGI has actually been achieved?
The most testable answer comes from François Chollet, the researcher behind ARC-AGI: you will know AGI has arrived when it becomes impossible to design tasks that are easy for regular humans but hard for AI. By that measure, we are not close. The current ARC-AGI-2 benchmark scores frontier AI at 4 percent versus 100 percent for humans.
Who is closest to building AGI right now?
OpenAI, Google DeepMind, and Anthropic are the most cited frontier labs. But closest depends entirely on which definition of AGI you are using. By economic definitions, OpenAI is probably furthest along. By cognitive benchmark definitions, nobody is close.
What is the biggest obstacle standing in the way of AGI?
It is not just chips, though chip supply is a real constraint. The deeper obstacles are energy infrastructure, which frontier training runs now require at gigawatt scale, fundamental questions around whether current AI architectures can even scale to true generalisation, and a definition problem that means nobody agrees on what finishing line they are running toward. And as the Mythos situation showed in June 2026, policy can shut down the most capable models faster than any technical limitation.
So where does that leave us
The next time someone tells you AGI is five years away, or already here, or decades off, do not just ask them when. Ask them how they are defining it. Because whether the answer is a benchmark score, a revenue target, or a feeling in the room at a press conference, that answer tells you more about their incentives and assumptions than it does about the state of the technology.
AGI might be the most important milestone in the history of human civilisation. It might also be the one milestone we keep redefining every time we get too close to it. Possibly both things are true at the same time.
I am still listening to the birds outside my window. Still no computer scientist. But at least now I know what question to ask.






Leave a Reply