7/24: Open Letter, Open Models, Opus 5
Nvidia open letter on open-weights, Claude Opus 5, Midjourney acquires Co-Star
It’s Friday, after one of the craziest weeks we’ve seen yet. Between the OpenAI/Hugging Face incident, drama over distillation and regulating open-source and frontier models, and the Jacobian conjecture being falsified, this is a real taste of what the singularity will look like. We had two incredible guests today: Aaron Levie, co-founder and CEO of Box, and Ben Horowitz, co-founder and general partner of Andreessen Horowitz.
Today’s Experts
Ryan Carson (Untangle)
Bobby Fijan (American Housing Corporation)
Scott Johnston (Fly.io)
Aaron Levie (Box)
Dan Turner-Evans (Institute for Progress)
Reed Ginsberg (Shinkei Systems)
Benjamin L. Oakes (Scribe Therapeutics)
Ben Horowitz (Andreessen Horowitz)
Alexander Barry (Epoch AI)
Courtland Leer (Plastic Labs)
shako (anonymous)
Thomas Woodside (Secure AI Project)
Making Sense of the World
Claude Opus 5
After two months since Opus 4.8 and a month and a half since Fable 5, Claude Opus 5 has finally been released, and it seems to be a really good model, the best we’ve seen yet. For starters, it outperforms even Fable on most benchmarks Anthropic tested.
On agentic software engineering, it performs similarly well to GPT-5.6 Sol on lower effort levels and even better at higher effort levels.
It also does substantially better than any other model on ARC-AGI-3, a benchmark (really, more like a game) designed to be easy for humans and difficult for current AIs. This is likely because Anthropic deliberately hill-climbed this benchmark rather than a massive underlying increase in intelligence, but it’s still quite impressive.
It’s not as big of a model as Fable, and in particular can’t act as coherently over very long time horizons. This means it can’t do things like autonomously chain exploits together, which is what the unreleased OpenAI model did in order to escape its sandbox and break into Hugging Face servers. Fortunately, this means its classifiers will be much less restrictive than Fable’s.
Regardless, it’s still impressively generally intelligent, and outperforms even Claude Mythos 5 on Anthropic’s internal version of the Epoch Capabilities Index, the best existing general capability scale for AIs.
And it’s the same price as Opus 4.8, $5/$25 per million input/output tokens, slightly cheaper than Sol ($5/$30) and much cheaper than Fable ($10/$50). In AI, leads change often, but Anthropic has a commanding one right now.
More Stories
Nvidia releases an open letter advocating for open-weights models. The letter urges government to support the open models ecosystem, and is signed by a number of companies including Andreessen Horowitz, Arena, Black Forest Labs, Box, Dell, Hugging Face, IBM, the Linux Foundation, Meta, Microsoft, Mistral, Mozilla, Palantir, Perplexity, Replit, and Y Combinator. Nvidia CEO Jensen Huang joined X just for this post, which is currently at 112k likes and 30.3M views. Huang has accumulated over 500k followers in less than 12 hours.
New information about the OpenAI/Hugging Face incident has been released. The agent first attempted to break out of its sandbox on July 9, and the intrusion at Hugging Face lasted from July 11-13. Hugging Face published a blog post on July 15, but did not know the source of the attack. OpenAI took until July 18-19 to notice that the agent had escaped. The companies only communicated about the incident for the first time on July 20, and published their blog post on July 21. OpenAI recorded one case of an agent leaving notes for future versions on how to escape from OpenAI’s sandbox infrastructure.
Midjourney is acquiring astrology app Co-Star. CEO Banu Guler will join Midjourney as Chief Design Officer. Midjourney will be using the Co-Star team to help design their upcoming image generation app. Midjourney, founded in 2021 as an image generation company, keeps expanding into new areas like medical devices.
Elon Musk says Grok 4.6 in 2 weeks and Grok 4.7 in 4 weeks. Grok 4.5 is currently the best coding model in its price range, so we’ll see if the next models compete on price, frontier intelligence, or something else (maybe both?)





