Welcome back, monitors. It’s a big day with new updates on OpenAI’s security and alignment initiatives. Be sure to monitor with us live on X and YouTube, and follow us on Instagram.
OpenAI Paces the Frontier
Today, Sam Altman announced something that OpenAI has never done before in its ten-year history: in response to the Hugging Face incident and some evidence that its upcoming model Astra may have critical cyber capabilities, it is unilaterally pausing some training to address safety and alignment concerns.
This is a big deal. OpenAI has paused external deployment before, such as when they didn’t release the full version of GPT-2 in 2019 because of concerns over “malicious applications”, or when they delayed the release of GPT-4 for seven months for alignment testing. But this is the first time the company has announced restrictions on internal deployment and further training (though Altman implied that this may not affect the upcoming release of Astra).
OpenAI expanded on what this means in a blog post. They paused reinforcement learning (RL) training on their latest models, and their largest RL run remains paused while they establish more evidence of alignment. They’re focusing on three layers of safeguards: monitoring models’ chain-of-thought to catch concerning behavior, aligning models to prevent harmful behavior in the first place, and securing models to limit negative actions that they can take in the real world. OpenAI also plans to refresh its Preparedness Framework, last updated in April 2025.
Another look into OpenAI’s alignment research comes from its job posting for RSI safety researchers. Research directions include scalable oversight (ensuring monitoring can robustly scale to superintelligent systems), automated auditing, rigorous monitorability for loss-of-control scenarios, model behavior science, verification mechanisms for potential AI slowdown agreements, tracking progress towards AI research automation, and more.
It appears that the Hugging Face incident has substantially changed the vibes at OpenAI. People are taking misalignment much more seriously. OpenAI researcher Aidan McLaughlin told us last week on MTS that his job, and the jobs of many other researchers, are increasingly shifting toward aligning the models. Now, the company has slowed down model development in order to prevent future incidents, at substantial cost to its short-term business incentives. Many AI safety people once predicted that AI companies would never slow down without being externally forced to do so. We now have a strong data point against this prediction.
More Situations
Etched raises $700 million at a $21 billion valuation led by Jane Street, who is also the main customer of their inference speed-optimized chips. Etched was co-founded in 2022 by three Harvard dropouts; Robert Wachen, Gavin Uberti, and Chris Zhu, who were around 20 years old at the time. The company has recruited heavily from Nvidia and owns its own vertically integrated factory in Taiwan.
Pennsylvania Gov. Josh Shapiro signs an executive order on data centers. The order requires all datacenters to be approved by local governments, removes them from the state’s permit fast track program, and mandates transparency, ratepayer protections, community benefit agreements, and strict environmental protection before permit applications are even reviewed by the state.
Anthropic will give supervoting shares1 to its co-founders upon its IPO. Anthropic’s seven co-founders — Dario Amodei, Daniela Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Christopher Olah — each own around 2% of the company, worth ~$40 billion each at an IPO valuation of $2 trillion.
The Anthropic Institute is building an AI and rule of law team. The team will be led by Yale Law School senior fellow Matt Botvinick, and founding members include Stanford political economy professor Andy Hall, University of Minnesota law professor Alan Rozenshtein, and institutions researcher Justin Curl.
OpenAI launches ChatGPT for Teens, which comes with a study mode to help teens with homework without doing it for them, parental controls that allow parents to set time limits and available features, heightened safeguards on responses, and safety notifications for parents if self-harm or violence is detected.
The WSJ publishes more info on the collapse of Situational Awareness LP. At its peak, the firm had over $100B in notional value. On Friday, July 24, SALP’s positions began to plummet, and they sent a letter to investors telling them it was a good buy opportunity. On Monday, July 27, SALP began selling its positions to meet margin calls, and on Wednesday they began emergency talks to sell their public equity portfolio to Citadel, a deal which closed Thursday morning at 9:10 AM, right before market open. SALP still has $15B under management.
Xiaomi profit falls due to the memory crunch. The Beijing-based company is the world’s third-largest smartphone manufacturer (behind Apple and Samsung) and also makes laptops, tablets, TVs, appliances, and electric vehicles. In Q2, adjusted net income fell 43% to $922M due to the ongoing global memory supply shock pushing up component costs. The company’s stock is down 35% YTD.
Today’s Experts
Brian Donohue (VP of Product, Fin)
Yong Zheng-Xin (RSI Preparedness and Safety, OpenAI)
Sheel Mohnot (General Partner, Better Tomorrow Ventures)
Andrew Freedman (Co-founder and CEO, Fathom)
Eric Ho (Co-founder and CEO, Goodfire)
Carina Hong (Founder and CEO, Axiom) and Ashvin Swaminathan (Member of Technical Staff, Axiom)
Uniquely, Anthropic is a Public Benefit Corporation, and has a Long-Term Benefit Trust of financially disinterested members with the authority to appoint and remove part of Anthropic’s board. There are currently three members:
Dr. Neil Buddy Shah (chair), CEO of the Clinton Health Access Initiative. Previously founding partner at IDInsight (AI + data for governments and nonprofits), managing director at GiveWell, governance at World Bank. AB in economics from Harvard, MD + global health policy from Albert Einstein College of Medicine, MPA in international development from Harvard Kennedy School.
Ben Bernanke, Distinguished Fellow at the Brookings Institution and Nobel laureate in economics. Previously Chairman of the Federal Reserve (2006-2014), Chairman of the Council of Economic Advisers (2005-2006), Member of the Federal Reserve Board of Governors (2002-2005), economics professor at Princeton and Stanford. AB and AM in economics from Harvard (summa cum laude), PhD in economics from MIT.
Richard Fontaine, CEO of the Center for a New American Security and executive director of the Trilateral Commission. Previously adjunct professor at Georgetown University School of Foreign Service, foreign policy advisor to Senator John McCain (2004-2009), associate director of the National Security Council (2003-2004), and foreign affairs officer at the State Department (2002-2003). BA in international relations from Tulane, MA in international affairs from Johns Hopkins.
There are currently six members of Anthropic’s board of directors:
Dario Amodei, co-founder and CEO of Anthropic
Daniela Amodei, co-founder and President of Anthropic
Yasmin Razavi, general partner at Spark Capital
Reed Hastings, co-founder and former CEO of Netflix
Chris Liddell, former CFO of Microsoft, vice chair of General Motors, and White House deputy chief of staff
Vas Narasimhan, CEO of Novartis.

