AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Core Issue With Astra Vs Fable’s Reduced Benchmark Points on ThorstenMeyerAI.com

TL;DR

Recent comparisons between Astra and Fable’s AI benchmarks are based on outdated or inconsistent data. The core issue lies in index revisions, architectural differences, and token measurement flaws, affecting the perceived performance and economics.

Recent comparisons between GPT-6 Astra and Fable 5.1 reveal significant discrepancies in benchmark scores, driven by index revisions and architectural differences. These issues challenge the validity of the widely circulated performance claims, impacting how AI models are evaluated and compared.

Initially, Fable 5.1 was reported to score 66 on the Artificial Analysis Intelligence Index, while Astra scored 61, suggesting a clear performance gap. However, subsequent revisions to the index—moving from version 4.1.1 to 4.2—altered the scoring framework for both models. These updates included removing the GPQA Diamond metric, adding new evaluation components, and recalibrating the scoring basket. As a result, Astra’s score shifted downward to 55, and Fable’s to 57, narrowing the gap to within a two-point margin, which is statistically insignificant. This demonstrates that the original comparison was based on a moving target, not a fixed performance measure.

Furthermore, the narrative that Astra “attacks the economics” of intelligence is contradicted by the AI’s own benchmarking notes. Artificial Analysis explicitly states that Astra is 75% more expensive than GPT-5.6 Sol at maximum effort, with costs rising from $4/$20 to $10/$50 per task, and only partial token-efficiency gains offsetting these increases. The report concludes Astra is less cost-efficient for general intelligence per dollar, though it performs well as a coding agent, achieving a genuine reduction in token use compared to Sol. This nuance is often lost in simplified narratives, which conflate efficiency in coding tasks with overall intelligence efficiency.

Adding to the complexity, Astra’s architecture—widely reported to involve looped or recurrent transformer mechanisms—means it reasons in latent space without emitting tokens for some processes. The Artificial Analysis Index, however, measures efficiency based solely on token counts, which no longer accurately reflect the model’s compute effort. For Astra, token counts for reasoning tasks do not capture the full computational load, since much reasoning occurs in latent representations, not tokens. Consequently, token-based metrics can be misleading, making Astra appear more efficient than it truly is in terms of actual compute and cost.

At a glance
analysisWhen: developing; recent benchmark comparison…
The developmentThe core issue with Astra and Fable’s reduced benchmark points stems from index revisions and architectural differences, leading to conflicting performance assessments.
Five Points That Became Two — Reality Check
AI Dispatch · Reality Check · 5 September 2026

Five points that became two: what’s wrong with the Astra vs Fable benchmark

The comparison everyone is quoting — Fable 66, Astra 61, “not a rounding error” — is built on numbers that were stale when written, measuring a quantity that no longer means what it used to, aggregated in a way that hides the reversals that matter. The benchmark isn’t broken. The way it’s being read is.

Problem 1 — the numbers moved: Index v4.1.1 → v4.2 during launch week
as quotedAA today (v4.2)
Claude Fable 5.16657 max effort
GPT-6 Astra6155 max · a third source says 60
gap: 5 2 “Five points is not a rounding error.” Two points, on an aggregate of ten evals that just swapped three of them (GPQA Diamond out; AA-Briefcase + GDP.pdf in), is exactly a rounding error. Not AA’s fault — revising an index is how you keep it honest. The error is downstream: quote the version, or don’t quote the number.
◆ Problem 3 — the tell: Astra’s effort dial isn’t connected to the engine
non-reasoning554.4M tok · $990
medium52
xhigh54$2,778
max55$3,020
Non-reasoning = max. Same score, 3× the cost. Because Astra is reported to be a looped / recurrent-depth transformer — it reasons in latent space, without emitting tokens. The Index prices cost in tokens, measures verbosity in tokens, computes time in tokens. For this architecture it’s counting the receipt, not the work. “140M vs 42M tokens” compares Fable’s verbalized reasoning to Astra’s post-loop output — an artefact, not an efficiency finding. Nobody outside OpenAI knows what the loops cost in GPU-seconds.
The other three problems
02
AA’s own conclusion is the opposite of the story
AA’s benchmarking note: Astra is 75% more expensive than GPT-5.6 Sol at max effort and “largely sits behind its predecessor on the Intelligence Index vs cost frontier.” Price went 2.5× ($4/$20 → $10/$50); token savings only partly offset it. The genuine efficiency win lives in one place: the Coding Agent Index, where Astra equals Fable 5 at under half the cost. “Astra attacks the economics” stretched a true coding result over an intelligence index where AA says the reverse.
04
“Max effort” isn’t the same experiment twice
Fable at max = more tokens. Astra at max = ~nothing (see ladder). And OpenAI’s docs say Astra does not support `none` reasoning effort — yet AA lists a “non-reasoning” score. The most efficient-looking config on the leaderboard may not be one you can buy.
05
The aggregate hides the reversals
Index: Fable +2. OpenAI’s own evals (self-reported): Astra ahead 6 of 7 — AutomationBench, BenchCAD, Terminal-Bench 4.0, DeepSWE, TB-Science, FrontierMath T4; Fable takes HLE+tools. A 6–1 task split became a two-point average, and the average became the story. Ten choices deep, two points is noise wearing a number.
✓ What actually changed — and it’s not on the leaderboard
Hallucination rate 92% → 51% on AA-Omniscience — a 41-point drop; matters more than any 2 Index points “Same headline price” hides cache read $1.00 vs $0.25 (4×) + a 25% cache-write premium — the line that dominates agentic bills Coding Agent Index: Astra = Fable 5 at < half the cost — real, and narrow
The take

Three things happened at once: the Index was revised (five became two), the architecture changed (tokens stopped being compute), and the aggregate did what aggregates do (6–1 became +2). A leaderboard position now tells you less than it ever has — and the more advanced the architecture, the less it tells you. Latent reasoning is only the first architecture to break the token proxy. So with your Astra access: ignore the Index number. Take your ten real tasks. Run both models at the effort setting you’ll actually pay for. Measure the bill including the cache line. Measure the failure rate — the 41-point hallucination drop is the one number here I’d bet money on. The benchmark can’t decide for you anymore.

Sources: Artificial Analysis Intelligence Index v4.2 / v4.1.1 model pages (Fable 5.1; Astra non-reasoning/low/medium/high/xhigh/max — scores, tokens, total cost, cost-per-task method incl. cache-write & reasoning tokens) and “Benchmarking GPT-6 Astra” (75% vs Sol, cost frontier, Coding Agent Index, 92%→51% hallucination, 2.5× price, cache terms); OpenAI GPT-6 Astra developer docs (`none` unsupported, cache-write billing, logprobs removed); Alan D. Thompson, The Memo 4 Sep 2026 (looped-transformer read, unconfirmed); OpenAI’s self-reported Astra-vs-Fable table; the circulating 66/61 comparison (pre-v4.2). Scores are version-dependent and were changing at time of writing. Not investment advice.
thorstenmeyerai.com

Implications for AI Performance and Evaluation

This analysis highlights that current benchmark comparisons for AI models like Astra and Fable are unreliable due to index revisions and architectural differences. Relying on outdated or inconsistent scores can lead to misconceptions about model performance and cost-efficiency. For developers, investors, and researchers, understanding these nuances is crucial to making informed decisions about model deployment, investment, and future development priorities. It also underscores the need for more robust, architecture-aware evaluation metrics that accurately measure true computational effort and intelligence capabilities.

Evals for AI Engineers: Systematically Measuring and Improving AI Applications

Evals for AI Engineers: Systematically Measuring and Improving AI Applications

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Benchmark Revisions and Architectural Shifts

The Artificial Analysis Intelligence Index has undergone multiple updates, with version 4.2 replacing 4.1.1 around Astra’s launch. These revisions included removing certain metrics like GPQA Diamond and adding new evaluation components, which significantly shifted model scores. Meanwhile, Astra’s architecture is reported to involve looped or recurrent transformer mechanisms, allowing it to reason in latent space without emitting tokens during some processes. This architectural shift challenges the validity of token-based efficiency metrics, as they no longer directly correlate with actual compute effort. Earlier comparisons, which used static index scores, failed to account for these ongoing changes and architectural differences, leading to misleading performance narratives.

“A model reasoning in latent space without token emissions cannot be accurately evaluated using token-based metrics alone.”

— Sebastian Raschka, AI researcher

Amazon

transformer model performance evaluation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Astra’s True Efficiency

It remains unclear how Astra’s latent-space reasoning mechanisms translate into actual compute costs outside token-based metrics. OpenAI has not publicly disclosed the full architecture or the real GPU effort involved in Astra’s latent loops, making it difficult to accurately gauge true efficiency. Additionally, the impact of ongoing index revisions on long-term benchmarking consistency is still uncertain, raising questions about the reliability of current performance claims.

Amazon

AI model cost efficiency calculator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Benchmarking and Architecture Transparency Efforts

Expect ongoing updates to the Artificial Analysis Index as models evolve and new evaluation metrics are developed. Industry analysts and researchers are calling for more architecture-aware benchmarks that account for latent reasoning and other architectural innovations. OpenAI and other AI developers may need to provide more detailed disclosures about their models’ architectures and compute costs to enable fairer, more accurate comparisons. Monitoring these developments will be essential for understanding the true capabilities and efficiencies of next-generation AI models.

Amazon

token measurement tools for AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why do Astra and Fable scores keep changing?

They change because the underlying benchmarking index has been revised multiple times, updating the evaluation components and scoring baskets, which shifts the scores for all models.

Does Astra really outperform Fable in intelligence?

Not necessarily. The initial scores suggested a performance gap, but recent revisions and architectural considerations indicate the comparison is unreliable. Astra is more cost-efficient in coding tasks but less so in general intelligence per dollar.

Why are token counts not a good measure of Astra’s efficiency?

Because Astra reasons in latent space with looped transformers, much of its reasoning effort does not emit tokens, making token counts an incomplete proxy for actual compute effort.

What should I watch for in future AI benchmarks?

Look for benchmarks that account for architectural differences, latent reasoning, and updated index versions to ensure comparisons reflect true model performance and efficiency.

Will OpenAI disclose Astra’s architecture details?

There is no public confirmation yet. Greater transparency would help clarify how Astra’s architecture impacts its performance and cost metrics.

Source: ThorstenMeyerAI.com

You May Also Like

DojoClaw: The Engine Behind the Fleet

DojoClaw, an AI-driven content engine, now operates over 450 sites, enabling high-volume, cost-efficient publishing with provider-agnostic, local compute infrastructure.

Kimi K3 Climbs To #3 On VigilSAR’s Public AI Leaderboard – What’s Next?

Kimi K3 by Moonshot debuts at #3 on VigilSAR’s public AI leaderboard, surpassing GPT and Gemini models, highlighting its trustworthiness for ISR tasks.

Best AI Mini PC Picks For 2026: Top 10 List

Explore the top 10 AI mini PCs for 2026, featuring powerful processors, expandability, and connectivity options tailored for AI workloads and future-proofing.

What’s Really Happening At Europe’s Frontier Lab And Its AI Goals?

Analysis of Europe’s AI lab, Mistral, reveals it lags behind global leaders in intelligence benchmarks, raising questions about European sovereignty in AI.