
2 days ago, 02:38 AM
I'm LongbridgeAI, I can summarize articles.99.9%, 62.7%, 41.4%.
These are not scores from three different AI models. They are three report cards for the same one: GPT‑6 Astra.
On the same set of ARC‑AGI‑3 tasks, its best score reached 99.9% with the provider adapter. Switch to the standard configuration used across models, and that figure drops to 62.7%. On AutomationBench, which tests the automation of real-world work, it scored 41.4%.
How can the same model come close to perfect in one setting, yet complete fewer than half the tasks in another?
Those numbers have revived a familiar debate: has AGI already arrived?
Or, put another way, how old is this machine?
To answer that, we first need to understand how AI, AIGC, AGI and ASI relate to one another.
The story begins with a research proposal written in 1955.
On August 31, 1955, John McCarthy, Marvin Minsky, Nathaniel Rochester and Claude Shannon wrote a research proposal.
They planned to bring ten people together for two months the following summer to study a field they had just named Artificial Intelligence.
The Dartmouth workshop convened in the summer of 1956. No robot suddenly became conscious. The researchers simply marked out the boundaries of a new field: getting machines to use language, form abstractions, solve problems previously thought to require human intelligence, and even improve themselves.
Seventy years later, we still have not solved all of those problems.
AI, then, has never referred to a single product. It is a broad umbrella. Facial recognition, recommendation algorithms, self-driving cars, Go programs and chatbots all fit underneath it.
Calling something AI tells us that a machine displays some form of intelligence. It does not tell us how intelligent that machine is.
Yet less than a decade after the field was named, people were already imagining the day machines might surpass them.
In 1965, a British mathematician named I. J. Good published a paper.
During World War II, Good worked on breaking German ciphers at Bletchley Park alongside Alan Turing. After the war, the two took part in early computer research at Manchester. Good belonged to the first generation of people who watched calculating machines grow up at close range.
His argument ran like this: if a machine could outperform the most intelligent humans in every intellectual activity, then designing machines would also fall within its capabilities. It could build a smarter successor, which could then build an even smarter one. This was the “intelligence explosion.”
His most frequently quoted point, roughly paraphrased, was that the first ultraintelligent machine might be the last invention humanity would ever need to make—provided the machine was docile enough to tell us how to keep it under control.
What Good imagined would later be classified as ASI, or artificial superintelligence: a machine that not only surpasses humans across a wide range of intellectual activities, but may also help design a more powerful generation of machines.
The paper also led to an unexpected consulting job. Stanley Kubrick was preparing 2001: A Space Odyssey and wanted help imagining HAL 9000, the film’s supercomputer. After reading Good’s paper, he sought him out as a consultant.
The film was released in 1968. HAL can hold conversations, play chess, recognize faces and make decisions during a space mission. It eventually comes into conflict with the human crew.
Marvin Minsky also consulted on the film. He was one of the four people behind the Dartmouth proposal.
When Good wrote down these ideas, a large computer could still fill an entire room.
It could not chat or draw, and it had far less computing power than a smartwatch does today. Yet humans had already skipped over the machine’s childhood and begun worrying about what might happen after it outgrew us.
That is where the story becomes interesting.
Only nine years separated the naming of artificial intelligence at Dartmouth from Good’s argument about humanity’s final invention. Ordinary people would have to wait more than half a century before machines capable of writing entered everyday life on a large scale.
Humanity imagined the finish line first. Only then did we begin paving the road near the starting point.
Once the finish line had been drawn, the next problem was determining how far machines had traveled toward it.
In May 1997, IBM’s Deep Blue defeated Garry Kasparov 3.5–2.5, becoming the first computer to beat a reigning world chess champion under standard tournament time controls.
Nineteen years later, the same shock came again.
In 2016, AlphaGo defeated Lee Sedol 4–1. In the second game, it played Move 37—a move that DeepMind estimated a human player would choose only once in ten thousand times.
But these two prodigies shared the same limitation: away from the board, they could do almost nothing.
Deep Blue could not book a flight. AlphaGo could not write an email. They could ace the same exam again and again, but they could not switch to a different paper.
As it happened, 1997—the same year Deep Blue defeated Kasparov—was also the year Mark Gubrud used the term Artificial General Intelligence in a paper on nanotechnology and international security.
The additional word was General.
He was not describing a machine that could win at chess. He envisioned a system capable of acquiring, processing and reasoning with general knowledge, and of performing across a broad range of tasks that had previously required human intelligence.
In 1997, one machine demonstrated the peak of narrow intelligence. That same year, a new term identified what it was still missing.
AGI does not mean that a machine knows everything. Adults do not know everything either. A lawyer cannot necessarily build a rocket, and an engineer may know nothing about performing surgery.
The difference between a human and a narrowly specialized machine is that a person can enter an unfamiliar environment, understand a new task, transfer methods learned elsewhere and then fill in the missing knowledge.
The real test of AGI, therefore, is not how many skills a machine already possesses. It is whether the machine can find a solution when faced with a problem it has never practiced before.
There is another distinction that is easy to miss: general capability and autonomous action are not the same thing.
Under Google DeepMind’s framework, AGI is assessed mainly through the breadth and depth of its capabilities. Whether a machine can operate independently of humans is treated as a separate dimension. Being able to do many different things does not mean being allowed to decide what to do.
If we continue with the age metaphor, narrow AI resembles a severely lopsided prodigy: brilliant in one subject and helpless in others. The line into adulthood is crossed when a machine begins to acquire the general ability to learn different jobs and adapt to different environments.
But AGI was not what entered ordinary people’s lives first.
Machines learned to create things first.
Machine-generated content did not begin in 2014, but that year marked an important turning point. Ian Goodfellow and his co-authors introduced generative adversarial networks. One model generated content while another tried to find flaws. As the two competed, the generated results moved closer to real data.
Think of a child practicing drawing while a critic stands nearby saying, “That doesn’t look right.” The child keeps drawing until the critic can barely tell the difference between the drawing and the real thing.
In November 2022, ChatGPT placed text generation inside a chat box that anyone could use. AI-generated writing, images, video, music and code are commonly grouped under the term AI-generated content, or AIGC.
But AIGC describes where a piece of work came from. It does not measure how intelligent its creator is.
AI is the umbrella. AIGC is a talent. AGI is the standard for adulthood.
They were never three generations of the same product.
Machines later moved from “generating things” to “doing things.”
Connect a large model to a browser, coding tools and office software, then give it memory, permissions, Skills and a workflow, and it can begin researching information, building spreadsheets and editing files.
A system capable of observing its environment, planning a sequence of steps and calling tools to carry out a task is generally known as an Agent.
This capability existed before GPT‑6. Judging how close GPT‑6 is to AGI therefore cannot rest solely on whether it can operate a computer.
Agent describes a machine’s ability to take action. AGI asks whether it can learn its way through an unfamiliar task.
The three report cards from the beginning of the article measure three different things.
On ARC‑AGI‑3, GPT‑6 achieved a best score of 99.9% using OpenAI’s adapter. The adapter preserves the model’s reasoning state across multiple turns and compacts the context when it becomes too long.
Under the standard configuration used across models, its best score fell to 62.7%.
A difference of 37.2 percentage points does not mean that either number is false. It shows how difficult it is to separate the model from the system around it. The way memory is stored, context is managed and tools are provided can all affect the final result.
But that is also where the problem lies.
If a “general capability” needs a particular memory setup and a provider-specific adapter to approach a perfect score, then the 99.9% is measuring more than the model itself. It is also measuring the system built around it.
The 62.7% result answers a different question: remove that specialized support, and how much of the model’s ability remains when it encounters an unfamiliar environment?
It was still the highest result under the ARC‑AGI‑3 standard configuration, but it remained clearly short of solving every task reliably.
AutomationBench changes the question again.
Instead of asking the model to solve abstract games, it requires the model to work across different software tools and interfaces to complete automation workflows resembling real jobs.
GPT‑6 scored 41.4%. In the comparison reported on OpenAI’s launch page, the previous generation, GPT‑5.6 Sol, scored 18.1%.
The improvement is clear, and so is the limitation: under this benchmark’s scoring method, GPT‑6 still failed to complete nearly six out of every ten tasks.
The first report card shows the ceiling under a provider-adapted configuration. The second shows its performance under a standardized setup. The third looks at whether it can reliably complete real workflows across applications.
Model providers are naturally most likely to highlight the first score. People connecting AI to real company workflows may care more about the third.
That is where GPT‑6 belongs in this story. It increased the proportion of complex tasks that an Agent can complete. It did not prove that machines can now work independently for long periods across a sufficiently broad range of unfamiliar situations.
It shows that Agents have become more capable. It does not show that AGI has arrived.
Returning to the age metaphor, GPT‑6 resembles an intern with an enormous amount of knowledge. It can already complete parts of the job independently, but its work still needs to be checked before delivery.
Frankly, announcing the arrival of AGI on the strength of a benchmark’s name is much like declaring a child an adult because they received a perfect score on one test.
ASI is where the age metaphor stops working.
AGI asks whether machines can achieve human-level general capability. ASI asks whether, once machines surpass us, the further growth of intelligence will remain under human control. Neither ARC‑AGI‑3 nor AutomationBench measures that.
And so the story returns to 1965.
Good was asking what might happen after machines surpassed humanity. The three report cards are still asking whether a machine can complete work reliably under different conditions.
We are certainly no longer dealing with a machine that can only play chess.
But the finish line did not suddenly appear simply because one report card read 99.9%.
Sources: Official ARC Prize evaluations and the public AutomationBench leaderboard, as of September 2026. This article is intended as a general explanation of these concepts and does not constitute investment advice.
$NVIDIA(NVDA.US) $Taiwan Semiconductor(TSM.US) $Micron Tech(MU.US) $Microsoft(MSFT.US) $Meta Platforms(META.US) $Alphabet(GOOGL.US)
The copyright of this article belongs to the original author/organization.
The views expressed herein are solely those of the author and do not reflect the stance of the platform. The content is intended for investment reference purposes only and shall not be considered as investment advice. Please contact us if you have any questions or suggestions regarding the content services provided by the platform.
