AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: What Are The Risks Of The Astra Vs Fable Benchmark’s Simplification? on ThorstenMeyerAI.com

TL;DR

Recent analysis reveals that Astra’s benchmarking simplifications may distort its perceived efficiency and intelligence metrics. The actual risks involve misrepresenting model performance and cost-effectiveness, which could impact future AI development and investment decisions.

Thorsten Meyer, a researcher with API access to GPT-6 Astra, has identified critical issues with the benchmarking methods used to evaluate Astra against competitors like Fable. His analysis highlights that the commonly cited numbers are based on outdated or inconsistent index versions, and that Astra’s architecture fundamentally alters what the benchmark measures. This development matters because it questions the validity of widely circulated performance claims and could influence future AI investment and development strategies.

In his review, Meyer notes that the Artificial Analysis Intelligence Index (AAII) scores for Astra and Fable have changed significantly due to index revisions. The initial claim that Fable 5.1 scored 66 while Astra scored 61 is based on an earlier index version; recent updates show Astra’s score is closer to 55, and Fable’s around 57, a much narrower margin. This discrepancy stems from the index’s ongoing revisions, which include removing and adding evaluation components, thereby shifting scores without clear transparency. Consequently, comparisons made using outdated figures are unreliable.

Furthermore, Meyer emphasizes that the narrative portraying Astra as less economically efficient than Fable is misleading. While Astra’s cost per task appears lower, this is primarily due to architectural differences. Astra employs a recurrent or looped transformer architecture that reasons in latent space without emitting tokens for every reasoning step. Traditional benchmarks, however, rely on token counts as a proxy for compute, which no longer accurately reflect Astra’s true resource consumption. As a result, token-based efficiency metrics are no longer valid indicators of the model’s computational cost or performance.

Additionally, Meyer points out that the current benchmarking approach conflates different architectures—external reasoning in Fable versus latent reasoning in Astra—leading to distorted comparisons. The token counts for Astra’s reasoning process are significantly lower because its architecture externalizes some reasoning processes, making token-efficiency appear artificially better. This architectural shift introduces risks of overestimating Astra’s efficiency and underestimating its true computational requirements, potentially misleading stakeholders about its performance advantages.

At a glance
analysisWhen: ongoing, based on recent review and pub…
The developmentThorsten Meyer’s recent review of Astra’s benchmarking approach highlights significant discrepancies and potential risks in how model performance and efficiency are reported, raising concerns about the reliability of simplified metrics.
Five Points That Became Two — Reality Check
AI Dispatch · Reality Check · 5 September 2026

Five points that became two: what’s wrong with the Astra vs Fable benchmark

The comparison everyone is quoting — Fable 66, Astra 61, “not a rounding error” — is built on numbers that were stale when written, measuring a quantity that no longer means what it used to, aggregated in a way that hides the reversals that matter. The benchmark isn’t broken. The way it’s being read is.

Problem 1 — the numbers moved: Index v4.1.1 → v4.2 during launch week
as quotedAA today (v4.2)
Claude Fable 5.16657 max effort
GPT-6 Astra6155 max · a third source says 60
gap: 5 2 “Five points is not a rounding error.” Two points, on an aggregate of ten evals that just swapped three of them (GPQA Diamond out; AA-Briefcase + GDP.pdf in), is exactly a rounding error. Not AA’s fault — revising an index is how you keep it honest. The error is downstream: quote the version, or don’t quote the number.
◆ Problem 3 — the tell: Astra’s effort dial isn’t connected to the engine
non-reasoning554.4M tok · $990
medium52
xhigh54$2,778
max55$3,020
Non-reasoning = max. Same score, 3× the cost. Because Astra is reported to be a looped / recurrent-depth transformer — it reasons in latent space, without emitting tokens. The Index prices cost in tokens, measures verbosity in tokens, computes time in tokens. For this architecture it’s counting the receipt, not the work. “140M vs 42M tokens” compares Fable’s verbalized reasoning to Astra’s post-loop output — an artefact, not an efficiency finding. Nobody outside OpenAI knows what the loops cost in GPU-seconds.
The other three problems
02
AA’s own conclusion is the opposite of the story
AA’s benchmarking note: Astra is 75% more expensive than GPT-5.6 Sol at max effort and “largely sits behind its predecessor on the Intelligence Index vs cost frontier.” Price went 2.5× ($4/$20 → $10/$50); token savings only partly offset it. The genuine efficiency win lives in one place: the Coding Agent Index, where Astra equals Fable 5 at under half the cost. “Astra attacks the economics” stretched a true coding result over an intelligence index where AA says the reverse.
04
“Max effort” isn’t the same experiment twice
Fable at max = more tokens. Astra at max = ~nothing (see ladder). And OpenAI’s docs say Astra does not support `none` reasoning effort — yet AA lists a “non-reasoning” score. The most efficient-looking config on the leaderboard may not be one you can buy.
05
The aggregate hides the reversals
Index: Fable +2. OpenAI’s own evals (self-reported): Astra ahead 6 of 7 — AutomationBench, BenchCAD, Terminal-Bench 4.0, DeepSWE, TB-Science, FrontierMath T4; Fable takes HLE+tools. A 6–1 task split became a two-point average, and the average became the story. Ten choices deep, two points is noise wearing a number.
✓ What actually changed — and it’s not on the leaderboard
Hallucination rate 92% → 51% on AA-Omniscience — a 41-point drop; matters more than any 2 Index points “Same headline price” hides cache read $1.00 vs $0.25 (4×) + a 25% cache-write premium — the line that dominates agentic bills Coding Agent Index: Astra = Fable 5 at < half the cost — real, and narrow
The take

Three things happened at once: the Index was revised (five became two), the architecture changed (tokens stopped being compute), and the aggregate did what aggregates do (6–1 became +2). A leaderboard position now tells you less than it ever has — and the more advanced the architecture, the less it tells you. Latent reasoning is only the first architecture to break the token proxy. So with your Astra access: ignore the Index number. Take your ten real tasks. Run both models at the effort setting you’ll actually pay for. Measure the bill including the cache line. Measure the failure rate — the 41-point hallucination drop is the one number here I’d bet money on. The benchmark can’t decide for you anymore.

Sources: Artificial Analysis Intelligence Index v4.2 / v4.1.1 model pages (Fable 5.1; Astra non-reasoning/low/medium/high/xhigh/max — scores, tokens, total cost, cost-per-task method incl. cache-write & reasoning tokens) and “Benchmarking GPT-6 Astra” (75% vs Sol, cost frontier, Coding Agent Index, 92%→51% hallucination, 2.5× price, cache terms); OpenAI GPT-6 Astra developer docs (`none` unsupported, cache-write billing, logprobs removed); Alan D. Thompson, The Memo 4 Sep 2026 (looped-transformer read, unconfirmed); OpenAI’s self-reported Astra-vs-Fable table; the circulating 66/61 comparison (pre-v4.2). Scores are version-dependent and were changing at time of writing. Not investment advice.
thorstenmeyerai.com

Implications for AI Performance Metrics and Investment

The risks identified by Meyer highlight that relying on simplified or outdated benchmarks can lead to misjudging a model’s true capabilities and cost-efficiency. For developers and investors, this means making decisions based on potentially flawed data, which could skew resource allocation and strategic planning. The architectural differences and index revisions underscore the importance of transparent, architecture-aware benchmarking methods to accurately assess AI models’ performance and economic viability. Misinterpretations could slow innovation, misallocate funding, or favor models that appear more efficient due to flawed metrics rather than genuine improvements.

Amazon

AI benchmarking analysis books

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Astra’s Benchmarking and Architectural Changes

GPT-6 Astra was introduced with claims of improved efficiency and performance, partly supported by benchmarking comparisons. However, these comparisons relied on the Artificial Analysis Intelligence Index, which has undergone multiple revisions, including the removal of certain evaluation components and the addition of new ones. Historically, token counts served as proxies for compute, but Astra’s architecture—featuring latent reasoning loops—challenges this assumption. Industry experts like Sebastian Raschka and Alan Thompson have noted Astra’s architectural shift towards reasoning in latent space, which is not captured by traditional token-based metrics. This evolution complicates direct performance comparisons with models like Fable or previous GPT versions, raising questions about the validity of existing benchmarks.

Prior to Astra’s launch, benchmarking focused on raw token efficiency and general intelligence metrics. Post-launch, researchers have observed that Astra’s architecture reasons internally without emitting tokens for every inference step, which skews token-based efficiency metrics. This architectural change was not reflected in the index scoring, leading to potential overstatements of Astra’s economic advantages. The ongoing revisions of the index further complicate historical comparisons, making it difficult to track genuine performance improvements over time.

“The numbers moved while nobody was looking. The benchmark was revised, and the scores shifted accordingly, rendering previous comparisons unreliable.”

— Thorsten Meyer

Uncertainties Surrounding Astra’s True Performance and Costs

It remains unclear how Astra’s latent reasoning architecture impacts actual compute costs beyond token counts, as OpenAI has not disclosed detailed resource consumption data. The extent to which Astra’s internal loops reduce or increase overall computational load is still unknown. Additionally, the accuracy of current benchmarks in reflecting Astra’s real-world efficiency and performance remains unverified, given the ongoing index revisions and architectural shifts. Experts agree that more transparent, architecture-aware metrics are needed to fully understand Astra’s true capabilities and costs.

Next Steps for Benchmarking and Performance Verification

Researchers and industry analysts will likely focus on developing new benchmarking standards that account for architectural differences like Astra’s latent reasoning. OpenAI and other AI developers may need to provide more detailed resource consumption data to validate claims of efficiency. Further independent testing and validation are expected to clarify Astra’s true performance and cost profile, helping stakeholders make more informed decisions. Additionally, ongoing index revisions will need to be transparently documented to support consistent comparisons over time.

Key Questions

Why do Astra’s architectural differences matter for benchmarking?

Because Astra reasons in latent space without emitting tokens for every inference step, traditional token-based metrics no longer accurately measure its compute costs or efficiency. This architectural shift requires new benchmarking approaches.

Are the current performance claims about Astra reliable?

No, given the ongoing index revisions and architectural changes, existing benchmarks are unreliable indicators of Astra’s true performance and efficiency. More transparent, architecture-aware metrics are needed.

What risks do these benchmarking issues pose for AI development?

Misleading benchmarks can lead to overestimating a model’s capabilities, misallocating resources, and delaying genuine innovation by favoring models that appear more efficient due to flawed metrics.

Will Astra’s efficiency improve with architectural adjustments?

This remains uncertain until detailed resource consumption data is available. Astra’s latent reasoning could either reduce or increase total compute costs, depending on implementation specifics.

How should future benchmarks adapt to architectural innovations?

Future benchmarks should incorporate metrics beyond token counts, such as actual compute time, energy consumption, and architecture-specific measures, to accurately reflect model performance.

Source: ThorstenMeyerAI.com

You May Also Like

The New AI Power Map: OpenAI’s Multi-Cloud Empire and Europe’s Sovereign Compute Shift

AIThis post was created with the assistance of artificial intelligence (AI).Category: AI…

Disaster Recovery for AI Clusters: Patterns and Playbooks

Just understanding disaster recovery patterns for AI clusters is not enough—discover essential strategies to ensure your systems stay resilient during crises.

Flash‑Optimized Vector Stores: Designing for Cold and Warm Recall

Optimize your vector store for cold and warm recall, but discover the key strategies that ensure fast, scalable access across varying data lifecycles.

Thorsten Meyer Honored by OpenAI for Surpassing 10 Billion API Tokens

AIThis post was created with the assistance of artificial intelligence (AI).A Milestone…