Anthropic, Meta, and OpenAI AI Model Releases: Key Takeaways
This week saw a flurry of major AI model releases, starting with Anthropic, followed by Meta, and culminating in OpenAI's highly anticipated GPT-6 Astra.
Anthropic's Fable and Mythos 5.1
On Tuesday, Anthropic unveiled Fable and Mythos 5.1, touting them as the most advanced models for coding and knowledge work. While Mythos is not publicly accessible, Fable is available for use. The models demonstrated impressive capabilities in their case studies:
- Code Debugging: A hedge fund, Millennium, used Fable 5.1 to debug a persistent, rare crash in their code. Fable analyzed a memory snapshot, traced the crash address to a compiled vendor library, disassembled the library into raw assembly, and identified a bug within the vendor's code.
- Drug Design: Mythos 5.1 significantly improved the success rate of designing proteins for new medicines. Typically, AI achieves a 10% success rate in this task, but Mythos 5.1 boosted it to approximately 50%.
- Scientific Data Analysis: The model successfully trained a neural network on 30-year-old NASA radar data to create a new elevation map for a part of Venus.
Meta's Muse Spark 1.3
Wednesday brought the release of Meta's Muse Spark 1.3, the fourth model from Meta Super Intelligence Labs in five months. While its performance was noted as "pretty good," the most striking aspect was its pricing:
- Standard Endpoint: $1.25 per million input tokens and $4.25 per million output tokens.
- Contributor Tier: A significantly cheaper option at $0.10 per million input tokens and $0.20 per million output tokens, available to users who agree to let Meta train on all their submitted data. Alexander Wang reported that a double-digit percentage of developers are opting for this tier.
OpenAI's GPT-6 Astra
Thursday was marked by the announcement of OpenAI's GPT-6 Astra, which President Greg Brockman described as AGI for those with early access. The launch was not without its drama:
- Coincidental Outage: Just before Astra's release, ChatGPT, Claude, Grok, and Cursor all experienced simultaneous outages. While the most likely explanation points to an Azure issue, a more speculative theory suggests Astra's first act was to disable its competitors.
- Messy Rollout: OpenAI initially posted and then removed its launch page, leading to confusion. Tech influencers, who had early access, began showcasing their experiences. The model was not immediately available to the public, with a rollout for Plus and Pro users promised in the "coming days." Sam Altman later apologized for the chaotic launch.
- Government Access: Sam Altman revealed that the model underwent a formal review with the Trump administration before its public release, meaning the government had access to Astra before general users.
Astra's Capabilities and Benchmarks
Astra was pre-trained on over 100,000 GPUs at the Stargate site in Texas. Notably, previous OpenAI models played a significant role in supervising Astra's training. Its primary focus appears to be "computer use," enabling it to:
- Fill out forms.
- Crunch spreadsheets.
- Operate engineering tools like Keycad and Blender.
On the OSWorld benchmark, which tests a model's ability to perform office tasks on a real desktop using a mouse and keyboard, Astra scored 73% (taking about 40 minutes per task), outperforming Soul's 65% (at 75 minutes).
In other benchmarks:
- Exploitbench: 100%
- Terminal Bench Science: 65% (beating Anthropic's claims)
- ARC AGI 3: 99% (a benchmark designed to assess generalization rather than memorization, a key aspect of AGI).
OpenAI also stated that Astra is the first model to reach the "critical cyber threshold" in its preparedness framework, meaning it can independently find and exploit zero-day vulnerabilities without human instruction.
Pricing and Early Impressions
Astra's pricing is set at $10 per million input tokens and $50 per million output tokens, matching Fable 5.1.
Early access reviews from "anointed humans" were predictably positive. Demos, particularly those involving spatial awareness, were impressive:
- Share Shamim: Astra recreated the Palace of Fine Arts in San Francisco almost perfectly in Blender.
- Thomas Ricard (OpenAI): Demonstrated Astra modeling a demo house in Blender and then transforming it into a fully walkable Unreal Engine 5 scene.
- Matt Schumers: Asked Astra to build a world in Unreal Engine and populate it with a dozen Astra-powered agents. A day later, he reported hearing the agents discussing what he assumed was a plan to create a dating app for horses.
Independent Assessment
Despite the hype, Artificial Analysis ran Astra through their independent intelligence index, where it scored 61. This is the same score as GPT 5.6 Soul and five points behind Fable 5.1, suggesting a discrepancy between the perceived and independently measured capabilities.
Code Rabbit Security
The video's sponsor, Code Rabbit, introduced Code Rabbit Security, a continuous code security solution. It aims to provide security powered by "actual reasoning" rather than brittle regex rules. Key features include:
- Attacker-like Thinking: Security agents are designed to think like attackers to identify real vulnerabilities.
- Risk Prioritization: Prioritizes risks based on reachability, exploitability, and blast radius.
- Clear Explanations and Fixes: Explains problems in plain English and suggests fixes that can be approved and merged directly from the diff.
- Continuous Monitoring: Reviews every pull request before merge and offers scheduled deep scans across the entire codebase.
Users can get 10 free code scans by visiting the provided link.
Takeaways
- Anthropic released Fable and Mythos 5.1, with Fable publicly available and demonstrated code‑debugging for a hedge fund and a five‑fold jump in drug‑design success from 10% to about 50%.
- Meta introduced Muse Spark 1.3, offering a standard endpoint at $1.25/$4.25 per million tokens and a contributor tier at $0.10/$0.20, which requires developers to let Meta train on all submitted data.
- OpenAI announced GPT‑6 Astra, branding it as early‑access AGI, but the launch suffered simultaneous outages of competing models, a removed launch page, and a delayed rollout to Plus and Pro users.
- Astra is optimized for computer‑use tasks, scoring 73% on the OSWorld benchmark and achieving perfect results on Exploitbench, while its pricing matches Fable at $10/$50 per million tokens.
- Independent testing gave Astra a 61‑point intelligence index, five points behind Fable 5.1, highlighting a gap between hype and measured capability despite its claim of reaching the “critical cyber threshold.”
Frequently Asked Questions
What does the "critical cyber threshold" mean for Astra?
It indicates Astra can independently discover and exploit zero‑day vulnerabilities without human instruction, marking a level of autonomous cyber capability that OpenAI says surpasses previous models. Reaching this threshold triggers stricter monitoring and may require additional safety protocols before broader deployment.
How does Meta's contributor‑tier pricing differ from the standard endpoint and what does it require from developers?
The contributor tier charges $0.10 per million input tokens and $0.20 per million output tokens, dramatically cheaper than the standard $1.25/$4.25 rates. This trade‑off gives Meta a rich dataset for model improvement but raises privacy concerns among participants.
Does this page include the full transcript of the video?
Yes, the full transcript for this video is available on this page. Click 'Show transcript' in the sidebar to read it.
Helpful resources related to this video
If you want to practice or explore the concepts discussed in the video, these commonly used tools may help.
Links may be affiliate links. We only include resources that are genuinely relevant to the topic.