GPT-6 Astra vs Claude Fable 5.1: The Frontier AI Race Moves Beyond Chatbots
By Suad Seferi · Sep 4, 2026

OpenAI and Anthropic have released two of their most capable AI models within days of each other, and the competition between them is starting to look very different from the chatbot race of the past few years. Anthropic introduced Claude Fable 5.1 on September 1, followed two days later by OpenAI's release of GPT-6 Astra. Both companies are positioning the models for demanding professional work rather than everyday prompting: software development, research, computer use, long-running tasks and workflows that can continue with less human supervision. There are similarities, but the models are not being deployed in quite the same way. OpenAI is pushing Astra strongly toward end-to-end computer work and agentic execution. Anthropic is presenting Fable 5.1 primarily around coding, knowledge work and long-horizon research, while separating some of its more sensitive capabilities into Claude Mythos 5.1. That difference may be more important than the benchmark tables. Astra is built to do more on the computer OpenAI describes GPT-6 Astra as its most capable model for complex reasoning, coding, computer use, research and document creation. The model has a context window of just over one million tokens and supports up to 128,000 output tokens through the API. OpenAI lists API pricing at $10 per million input tokens and $50 per million output tokens. Computer use is one of the areas OpenAI is emphasising most heavily. On AutomationBench, which tests professional computer workflows, OpenAI reports Astra scoring 41.4%, compared with 31.4% for Claude Fable 5.1. On BenchCAD, Astra scored 95.9%, while Fable 5.1 recorded 84.3%. OpenAI also reports a 72.6% partial score for Astra on OSWorld 2.0, a benchmark designed around computer interaction. These numbers should be read carefully. Benchmark configurations, tools and evaluation conditions can differ between companies, and vendor-published comparisons are not substitutes for independent testing. Still, OpenAI's direction is clear. Astra is being designed as a model that does not simply answer questions or generate code, but can increasingly carry a task through several stages itself. That includes working across software, browsing, files and professional tools. Fable 5.1 focuses heavily on sustained work Anthropic's pitch for Claude Fable 5.1 looks slightly different. The company describes it as its most capable generally available model for coding and knowledge work, with improvements in long-running problem solving and scientific research. Anthropic reports a substantial improvement over Fable 5 on several of its own tests. On Terminal-Bench 4.0, Fable 5.1 scored 55.8%. On CursorBench 3.2 it reached 73.4%, and on the company's version of Humanity's Last Exam it scored 60.9% without tools and 65% with tools. Fable also recorded 31.4% on AutomationBench, almost twice the 17.1% Anthropic reports for Fable 5. But some of the more interesting examples are outside the leaderboard. Anthropic says early users have run Fable 5.1 on coding and research tasks lasting many hours. Shopify reported workflows that continued for extended periods without losing track of the objective, while Ramp described an unattended 38-hour machine-learning investigation. Those examples come from Anthropic's launch partners rather than independent studies, but they show what the company wants Fable to become: a model that can remain useful when a task stops looking like a conversation and starts looking like actual work. The biggest difference may be security Astra also arrives with something unusual attached to its launch. OpenAI says GPT-6 Astra is its first broadly deployed model to reach the Critical level for cybersecurity capability under its Preparedness Framework. AI Balkans reported in August that OpenAI had slowed Astra's release while it evaluated exactly this risk. The company now says testing shows Astra can, with appropriate tools and access, identify previously unknown vulnerabilities and develop methods for exploiting protected systems without requiring a human to direct every step. OpenAI says additional protections have therefore been introduced around deployment, monitoring and internal model security. Anthropic is handling the same problem differently. Claude Fable 5.1 and Claude Mythos 5.1 are technically the same underlying model, according to Anthropic, but operate with different safeguards. Fable is generally available and can assist with vulnerability discovery, but Anthropic says its safeguards prevent it from developing exploits. Mythos 5.1 opens more advanced cybersecurity and biology capabilities and is restricted to vetted organisations through trusted-access programmes. So the distinction is becoming less about whether advanced models possess sensitive capabilities. Increasingly, it is about who gets access to those capabilities and under what conditions. Pricing is surprisingly similar For developers using the models directly, the headline token prices are also close. OpenAI lists GPT-6 Astra at $10 per million input tokens and $50 per million output tokens. Anthropic lists Fable 5.1 at the same $10 input and $50 output rate. Anthropic has, however, reduced the price of cache reads to $0.25 per million tokens. The company estimates this makes typical Fable 5.1 workloads around 25% cheaper than Fable 5 and some highly agentic workloads as much as 45% cheaper. The practical cost of running either model will depend heavily on how much context, reasoning and tool use a task requires. Token price alone will not tell companies which system is cheaper to operate. So which one is better? There is no clean answer yet. OpenAI's published results suggest Astra has an advantage in several computer-use and professional automation tasks. Anthropic's own evaluations and customer testing suggest Fable 5.1 is particularly strong in coding, long-running research and work requiring sustained reasoning over many steps. Independent evaluations will matter more once both…