{"id":79771,"date":"2026-09-04T08:04:35","date_gmt":"2026-09-04T12:04:35","guid":{"rendered":"https:\/\/overcentral.com\/en\/?p=79771"},"modified":"2026-09-04T08:04:35","modified_gmt":"2026-09-04T12:04:35","slug":"gpt-6-astra-arc-agi-3-79771","status":"publish","type":"post","link":"https:\/\/overcentral.com\/en\/gpt-6-astra-arc-agi-3-79771\/","title":{"rendered":"GPT-6 Astra pulls Chollet&#8217;s AGI forecast forward on ARC-AGI-3"},"content":{"rendered":"<article>\n<p><a href=\"https:\/\/overcentral.com\/en\/openai-gpt-6-astra-agi-era-begins-wait-the-slug-should-be-3-6-words-let-me-adjust-openai-gpt6-astra-agi-4-words-no-hyphens-for-numbers-better-openai-gpt-6-astra-agi-5-words-but-earlier-i-sa-79647\/\" title=\"OpenAI Launches GPT-6 Astra as AGI Era Begins\" data-iacss-internal=\"1\">GPT-6 Astra<\/a> has pulled Fran\u00e7ois Chollet&#8217;s AGI forecast forward on ARC-AGI-3, posting a 62.7 percent score on the benchmark in a standard configuration\u2014more than eight times GPT-5.6 Sol&#8217;s 7.8 percent and more than double Claude Opus 5&#8217;s 30.2 percent. The result pushed Chollet, the <a href=\"https:\/\/arcprize.org\/\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">ARC Prize<\/a> co-founder, to say that his earlier forecast of AGI by 2030 should be moved sooner. But the score is only the visible surface of a more complicated story about cost, harness design, mathematical proof, and a model that is learning to write its own symbolic notation.<\/p>\n<h2>What is ARC-AGI-3?<\/h2>\n<p>ARC-AGI-3 is a benchmark that drops an AI into a series of unfamiliar game worlds without explaining the rules or the goals. The system has to determine what is happening, what it is supposed to do, and how to do it through trial and error. It was designed to test exploration under uncertainty, adaptation without guidance, and causal world modeling from sparse data\u2014the qualitative properties that general intelligence should exhibit. Chollet and ARC Prize have repeatedly emphasized that solving the benchmark is not proof of AGI, but that did not stop Astra&#8217;s result from shifting their timeline.<\/p>\n<p>The benchmark is small in scale compared with the real world, with deterministic mechanics, closed-ended goals, and no representation of the open world. It is also beginning to fall. When ARC-AGI-3 was released about six months ago, Chollet estimated it would take roughly a year to saturate. Astra arrived about twice as fast.<\/p>\n<h2>How GPT-6 Astra Pulled Chollet&#8217;s AGI Forecast Forward on ARC-AGI-3<\/h2>\n<p>Across the broader evaluation landscape, Astra&#8217;s profile is not uniformly dominant. It is the strongest model on several key measures, while its main rival holds the lead on coding. But the ARC-AGI-3 jump is unmistakable.<\/p>\n<table>\n<caption>Scores from <a href=\"https:\/\/epochai.org\/\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">Epoch AI<\/a>, Artificial Analysis, and ARC Prize. Astra&#8217;s ARC-AGI-1 figure carries an asterisk because it depends on reasoning effort: 98.5 percent at xhigh, 97.5 percent at max. ARC-AGI-1 is now widely treated as saturated.<\/caption>\n<thead>\n<tr>\n<th>Benchmark<\/th>\n<th>Astra<\/th>\n<th>Sol<\/th>\n<th>Fable 5.1<\/th>\n<th>Opus 5<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Epoch ECI, overall score<\/td>\n<td>169<\/td>\n<td>162<\/td>\n<td>163<\/td>\n<td>162<\/td>\n<\/tr>\n<tr>\n<td>AA Intelligence Index<\/td>\n<td>61<\/td>\n<td>61<\/td>\n<td>66<\/td>\n<td>63<\/td>\n<\/tr>\n<tr>\n<td>ARC-AGI-3, unfamiliar game worlds<\/td>\n<td>62.7%<\/td>\n<td>7.8%<\/td>\n<td>no data<\/td>\n<td>30.2%<\/td>\n<\/tr>\n<tr>\n<td>ARC-AGI-2, abstract visual puzzles<\/td>\n<td>95.0%<\/td>\n<td>92.5%<\/td>\n<td>90.0%<\/td>\n<td>90.4%<\/td>\n<\/tr>\n<tr>\n<td>ARC-AGI-1, older version<\/td>\n<td>98.5%*<\/td>\n<td>97.5%<\/td>\n<td>97.5%<\/td>\n<td>97.5%<\/td>\n<\/tr>\n<tr>\n<td>FrontierMath Erd\u0151s, open math<\/td>\n<td>3%<\/td>\n<td>0%<\/td>\n<td>0%<\/td>\n<td>no data<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Astra leads the group on Epoch ECI and is the only model with an ARC-AGI-3 score in double digits. It also posts the best ARC-AGI-2 result: 95.0 percent, against 92.5 percent for Sol and 90.4 for Opus 5. On FrontierMath Erd\u0151s it is the only model to score at all. Fable 5.1, meanwhile, leads the AA Intelligence Index at 66, and Fable 5 and Fable 5.1 have not yet been run on ARC-AGI-3.<\/p>\n<p>Read carefully, the table shows something more interesting than a new leader. Astra&#8217;s strengths are concentrated in math, knowledge, and puzzles. Coding is where the established order still holds.<\/p>\n<h2>The Economics of Astra: Pricier Than Sol, Cheaper Than Claude<\/h2>\n<p>Astra is clearly more expensive than its own predecessor. OpenAI charges two and a half times as much per unit of processed text, yet a typical task costs roughly 75 percent more than it did with Sol because Astra uses so few tokens. Against Anthropic, the picture flips. Artificial Analysis places Astra at the same coding score as Claude Fable 5 while Astra&#8217;s per-task cost is less than half Claude Fable 5&#8217;s. The reason is how sparing the model is. Astra needs only a third of the compute steps Sol uses and a fifth of what Opus 5 uses.<\/p>\n<p>That kind of efficiency is what allows a more expensive model to be cheaper in practice. On the Coding Agent Index, Astra reaches 67 points at roughly a third of Sol&#8217;s <a href=\"https:\/\/overcentral.com\/en\/agentic-token-usage-openrouter-77457\/\" title=\"Agentic Token Usage Shows 14x Jump on OpenRouter\" data-iacss-internal=\"1\">token usage<\/a>, while Fable 5.1 leads with 70. Fable 5.1 holds the top marks on nearly every coding test. On AA-Omniscience, the hallucination rate drops from 92 percent to 51 percent. The same model loses about 80 Elo points on GDPval-AA v2 and slips on banking support, SciCode, and long-context reasoning tasks.<\/p>\n<p>The qualification is that Astra&#8217;s gains are not universal. Epoch has recorded only one coding score for Astra so far, and that score came from a run at medium reasoning effort. Coding remains the clearest gap between Astra and Fable 5.1, as well as the area with the least complete evidence.<\/p>\n<h2>Two Lean-Verified Proofs on FrontierMath Erd\u0151s<\/h2>\n<p>One of the most striking numbers in the table is the small one: 3 percent on FrontierMath Erd\u0151s. That represents two of 68 open Erd\u0151s problems solved with Lean-verified proofs, and Epoch AI reports that <a href=\"https:\/\/overcentral.com\/en\/openai-gpt-6-astra-agi-79684\/\" title=\"OpenAI launches GPT-6 Astra, founder declares AGI is here\" data-iacss-internal=\"1\">GPT-6<\/a> Astra was the only model to solve any of them on the official budget of $300 per attempt. The model was not simply producing conjecture; it was generating machine-checkable mathematics.<\/p>\n<p>Three more solutions emerged in non-standardized extra runs that burned through more than $220,000 in compute. Epoch says those do not count toward the score. That distinction matters: the official result shows what Astra can do within a fixed cost envelope, while the unofficial runs show how far a brute-force version of the same model can still go when money stops being a constraint.<\/p>\n<h2>Why Did OpenAI Report 99.9 Percent on ARC-AGI-3?<\/h2>\n<p>OpenAI&#8217;s promotional figure for GPT-6 Astra on ARC-AGI-3 was 99.9 percent, not 62.7 percent. Both numbers are real; they come from different harnesses. The OpenAI-built harness preserves reasoning chains between individual requests and automatically summarizes long runs. ARC Prize measured that setup as 3.66 times faster and using 49 percent fewer tokens than the internal ARC harness, across 167 game-reasoning pairs solved by both configurations.<\/p>\n<p>ARC Prize has been cautious about vendor harnesses since the Sol era, when the same disagreement surfaced. Only the 62.7 percent run on the internal harness allows a fair comparison between vendors, the organization says, and it plans to publish vendor-harness numbers in the future as well. The score, in other words, is a property of a system that includes the model and the scaffold.<\/p>\n<h2>When More Reasoning Lowers the Bill<\/h2>\n<p>The cost relationship on ARC-AGI-3 is unusual. Standard logic says a higher reasoning level should make a test run more expensive because the model does more work. Astra inverts that. On the standard ARC scaffold, cost drops from $49,791 with no reasoning to $26,098 at maximum reasoning, while the score climbs from 35.2 percent to 62.7 percent. ARC Prize attributes this to Astra solving the games in fewer moves, which means fewer model calls and fewer tokens.<\/p>\n<p>Then there is the outlier. The &#8220;low&#8221; reasoning level scores 17.5 percent, worse than running with no reasoning at all. ARC Prize does not comment on it, but Astra&#8217;s behavior on other benchmarks suggests a plausible explanation: the new architecture presumably loops processing internally before generating its first token, which lets it solve some longer-horizon tasks without an external reasoning trace. Low reasoning may interrupt that internal process in a way that no reasoning does not.<\/p>\n<p>For context, human testers earned $115 per 90-minute session plus $5 for each solved game; with about nine attempts in a session, the base compensation works out to roughly $12.78 per attempted game. Count only the metabolic energy of the brain as electricity, and ARC Prize arrives at 0.067 cents per game. The human advantage was never about direct energy cost, though; it was about how little information a person needs to understand a new game.<\/p>\n<p>That measure is where Astra comes closest to crossing the human line. Before launch, ARC Prize had about 500 testers play ARC-AGI-3 with no pre-screening and recorded the median number of moves among those who solved each level. On the OpenAI-scaffold run, Astra cleared 96 percent of levels in fewer moves than that median, and on average used a little over half as many moves. This metric does not track compute. It tracks how much experience with an environment the model needed before it mastered it.<\/p>\n<p>This is the territory where ARC Prize expected humans to keep a lasting edge. Brute-force search still belongs to humans. But with top models, the pattern is becoming almost binary: once the model has figured out the mechanics, its execution lands in the human efficiency range.<\/p>\n<h2>Astra&#8217;s Self-Invented Algebraic Shorthand<\/h2>\n<p>The key to that efficiency is not just architecture. Astra keeps its own notes, and those notes take the form of a self-invented, algebra-like shorthand. It records objects, coordinates, rules, and open plans in compressed strings: &#8220;extend8 to3; retract10 to2&#8221; for an ordered sequence of moves, or &#8220;Turn 5: P=(24,20), empty, facing west&#8221; as a state note. In environment s5i5, for example, it noted &#8220;L8: hub q2 (8\u2193). Lengths: 14=1\u2026&#8221; and mapped operations to exact controls. Other models have tried similar note-taking, but ARC Prize singles out Astra for its precision and information density.<\/p>\n<p>On the standard harness, this matters because anything the model does not save to its own visible notes is lost. Chollet describes the behavior as &#8220;highly efficient, on-the-fly symbolic world modeling for each game and level.&#8221; Astra effectively develops its own shorthand DSL to represent in-game situations\u2014what Chollet calls &#8220;essentially a game-specific algebraic notation.&#8221; The deeper observation is that &#8220;Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses, so harness capabilities are increasingly shifting into the model itself.&#8221;<\/p>\n<h2>Inside the PRO-LONG Sandbox: Astra Builds Its Own Tools<\/h2>\n<p>ARC Prize also used a third-party agent framework called PRO-LONG as an early red-teaming partner for ARC-AGI-3. In that setup, Astra could execute code inside a sandbox. It took full advantage, writing small program libraries for each game: parsers for the game board, state models, search algorithms, and planners. In a maze game with guards, it built a pathfinder, a combat rules module, a patrol movement model, and a script that continuously compared predictions against observations.<\/p>\n<p>ARC Prize observed no attempts to escape the sandbox. These runs cannot be compared to human test conditions, because the human testers had neither a code interpreter nor a notepad. What they show is Astra&#8217;s combined performance with self-built tools\u2014and the direction where frontier models are heading.<\/p>\n<h2>Chollet: Faster Than Expected, Still Not Proof of AGI<\/h2>\n<p>ARC Prize explicitly does not treat Astra&#8217;s benchmark scores as evidence of general intelligence. &#8220;All we know about the system so far are its benchmark scores,&#8221; Chollet said. When ARC-AGI-3 was released, the group emphasized the same point in every presentation: &#8220;Solving it is not proof of AGI. It is not intended as a finish line.&#8221;<\/p>\n<p>The benchmark does test the qualitative properties expected of an AGI system, but it does so &#8220;on a small scale.&#8221; The games run on time scales orders of magnitude shorter than real-world tasks, with less data, less modeling complexity, and less on-the-fly learning. Still, the pace of change caught even the designer unprepared. Asked about saturation, Chollet had said &#8220;about a year&#8221; depending on how focused the approach was. Astra arrived roughly twice as fast.<\/p>\n<p>His newest forecast is less a prediction than a warning: &#8220;I believe the pace of progress will surprise many people, and what the new models are capable of will challenge the perception of AI that people have formed based on earlier generations of models.&#8221; When a user asked whether his earlier AGI forecast for 2030 still held, his answer was succinct: &#8220;Sooner, because progress is happening faster than I expected.&#8221;<\/p>\n<h2>ARC-AGI-4 Is Already in Development<\/h2>\n<p>If Astra has pulled the arc forward, the next stage is already underway. ARC-AGI-4 has been in development since the release of ARC-AGI-3 and is scheduled for the first quarter of 2027. Chollet considers ARC-AGI-3 itself quite limited: deterministic mechanics, closed-ended goals, and no representation of the open real world. The next generation is intended to explore recursive self-improvement and open innovation.<\/p>\n<p>The immediate consequence of Astra&#8217;s ARC-AGI-3 run is not that AGI has arrived. It is that the reasoning scaffold is becoming the model. The harness capabilities ARC Prize once had to build externally are now appearing inside a single system, and the gap between &#8220;model&#8221; and &#8220;agent&#8221; is collapsing. Chollet&#8217;s revised timeline may still be too conservative; the benchmark cycle is already half over.<\/p>\n<\/article>\n","protected":false},"excerpt":{"rendered":"<p>GPT-6 Astra has pulled Fran\u00e7ois Chollet&#8217;s AGI forecast forward on ARC-AGI-3, posting a 62.7 percent score on the benchmark in a standard configuration\u2014more than eight times GPT-5.6 Sol&#8217;s 7.8 percent and more than double Claude Opus 5&#8217;s 30.2 percent. The result pushed Chollet, the ARC Prize co-founder, to say that his earlier forecast of AGI [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":82954,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/79771.png","fifu_image_alt":"GPT-6 Astra pulls Chollet's AGI forecast forward on ARC-AGI-3","footnotes":""},"categories":[31],"tags":[],"class_list":["post-79771","post","type-post","status-publish","format-standard","has-post-thumbnail","category-technology"],"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/79771.png","fifu_image_alt":"GPT-6 Astra pulls Chollet's AGI forecast forward on ARC-AGI-3","_links":{"self":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/79771","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/comments?post=79771"}],"version-history":[{"count":0,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/79771\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media\/82954"}],"wp:attachment":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media?parent=79771"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/categories?post=79771"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/tags?post=79771"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}