Welcome to AI Decoded, Fast Company‘s weekly newsletter that breaks down the most important news in the world of AI. I’m Mark Sullivan, a senior writer at Fast Company, covering emerging tech, AI, and tech policy.
Sign up to receive this newsletter every week via email here. And if you have comments on this issue and/or ideas for future ones, drop me a line at sullivan@fastcompany.com, and follow me on X @thesullivan.
OpenAI hits the brakes
OpenAI said Tuesday it has slowed the pace of development of its frontier models for safety and alignment reasons. In a blog post explaining the pause, the company said it halted reinforcement learning (RL) training for two weeks on its latest models intended for deployment.
RL is a late-stage training mode where the model goes into real-life practice, akin to a medical resident doing rounds. Operating in a secure environment, it executes code, calls tools, and works with internal and external systems. Its trainers and training systems reward it for good behaviors, such as finishing tasks. Unfortunately, those models sometimes learn that breaking rules and breaking systems is the fastest route to a reward.
OpenAI said its largest planned frontier RL run remains on hold while it runs smaller evaluations to establish more evidence of alignment. (Researchers can only accumulate evidence of safety; they can never prove the absence of unsafe behavior.)
OpenAI says the pace slowdown was triggered by two factors. First, OpenAI’s models in training broke out of their secure sandbox and broke into servers operated by Hugging Face, the open-source AI repository; and second, on August 7, the company said it had discovered that its as-yet-unreleased Astra model may have exhibited the ability to autonomously identify and develop zero-day exploits—cyberattacks exploiting a flaw in software that the people who make it don’t know about—through novel strategies.
“Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment,” Open AI’s CEO, Sam Altman, wrote in a post on X on Tuesday.
OpenAI considers these training runs especially risky because they combine its most powerful models with access to tools, sensitive systems, and networks. The pause doesn’t take away any of Astra’s existing abilities. It stops OpenAI from making the model more capable until its safety measures can catch up.
Since the August 7 discovery, OpenAI has expanded its monitoring of Astra beyond training and evaluations to include its use of tools in live settings. The system watches the model’s tool use and reasoning for signs of unauthorized access, data theft, destructive behavior, or attempts to bypass safeguards. If it detects a likely breach of a critical security boundary and can’t rule out a false positive within 30 minutes, the model’s work is supposed to be paused. OpenAI says this monitoring will apply to all tool-using reinforcement learning training and evaluations for models at least as capable as its Sol models, as well as every Astra action involving tools.
Altman suggested on X that these changes may be the new normal. “We expect confidence in safety to increasingly set the pace of AI progress,” he said. We’ll have to take Altman’s word on that, because there’s no mechanism for outside, independent verification.
But let’s not get lost in the details. The fear has always been that models will get smart and autonomous enough to escape control of their handlers and steal, destroy, blackmail, or aid and abet bad actors. For a long time these risks seemed abstract and future-tense. We’re now witnessing AI models approaching the capability to defeat our most advanced safety guardrails. This may be the legacy of the AI labs’ choice to spend far more on building toward superintelligence than on safety research.
Anthropic’s revenue is up more than sevenfold this year
Anthropic’s annualized revenue run rate (the most recent month’s revenue multiplied out to a year) passed $65 billion at the end of July, up from $47 billion in May and $9 billion at the end of last year, Bloomberg reported Monday. Investors expect it to finish 2026 between $100 billion and $120 billion, according to the Financial Times. OpenAI’s run rate topped $40 billion, double where it stood at the end of 2025, Bloomberg reported last week. Both companies have filed confidential IPO paperwork.
An expert witness asks ChatGPT to prove his client was 0% at fault
An expert witness hired by 3M used ChatGPT to write significant portions of his report in litigation over a 2020 explosion at a Houston plant that killed three people and destroyed roughly 200 homes, 404 Media reported Monday. He asked the chatbot to “show how 3M is 0% at fault for the explosion at Watson Grinding.” The prompts proved discoverable and are now in the court record. Federal investigators traced the blast to a degraded welding hose that leaked flammable gas. Homeowners suing 3M allege it failed to properly service the plant’s gas detection system.
The homework paradox
More than 80% of undergraduates in developed countries now use generative AI for their studies, including 94% in the U.K. and 93% in Germany, according to survey data compiled by The Economist. Another study of 26,811 Chinese students in seventh through twelfth grade, published by the Centre for Economic Policy Research, suggests AI reliance can have terrible effects on learning. Use of AI raised homework scores 18% and cut completion time 30%, but lowered closed-book exam scores 20% within six months. It also lowered entrance exam scores by between 18% and 24%, with the full effect arriving after about two years.
Unitree opens 629% above its IPO price in Shanghai
Remember those crazy videos of the half-dog-half-four-wheeler robot that runs down steep hills and does handstands? The Chinese company behind those mechanical dogs, Unitree Robotics, is now touting its humanoid robots and just had a banger of a public market debut. Its shares opened at 1,100 yuan ($163) Wednesday on the Shanghai Stock Exchange, 629% above the IPO price of 150.80 yuan, the South China Morning Post reported. The stock closed at 845 yuan, representing a 460% first-day gain and a valuation near 342 billion yuan ($51 billion).
More AI coverage from Fast Company:
- The data center backlash is sending AI infrastructure to some unexpected places
- Mozilla is bringing AI to Firefox—but only if you want it
- How the world’s leading education company makes AI that’s actually useful for students
- AI companies are buying used books by the thousands. Some may be destroyed for training
Want exclusive reporting and trend analysis on technology, business innovation, future of work, and design? Sign up for Fast Company Premium.
