<?xml version="1.0" encoding="UTF-8"?><feed
	xmlns="http://www.w3.org/2005/Atom"
	xmlns:thr="http://purl.org/syndication/thread/1.0"
	xml:lang="en-US"
	>
	<title type="text">Hayden Field | The Verge</title>
	<subtitle type="text">The Verge is about technology and how it makes us feel. Founded in 2011, we offer our audience everything from breaking news to reviews to award-winning features and investigations, on our site, in video, and in podcasts.</subtitle>

	<updated>2026-09-17T12:49:56+00:00</updated>

	<link rel="alternate" type="text/html" href="https://www.theverge.com/author/hayden-field" />
	<id>https://www.theverge.com/authors/hayden-field/rss</id>
	<link rel="self" type="application/atom+xml" href="https://www.theverge.com/authors/hayden-field/rss" />

	<icon>https://platform.theverge.com/wp-content/uploads/sites/2/2025/01/verge-rss-large_80b47e.png?w=150&amp;h=150&amp;crop=1</icon>
		<entry>
			
			<author>
				<name>Hayden Field</name>
			</author>
			
			<title type="html"><![CDATA[Inside the suddenly explosive world of AI safety]]></title>
			<link rel="alternate" type="text/html" href="https://www.theverge.com/ai-artificial-intelligence/996563/ai-safety-research-metr-redwood-openai-anthropic" />
			<id>https://www.theverge.com/?p=996563</id>
			<updated>2026-09-17T08:49:56-04:00</updated>
			<published>2026-09-17T07:30:00-04:00</published>
			<category scheme="https://www.theverge.com" term="AI" /><category scheme="https://www.theverge.com" term="Features" /><category scheme="https://www.theverge.com" term="OpenAI" />
							<summary type="html"><![CDATA[On a sunny July day in Berkeley, California, the country’s top AI safety researchers gathered on an unmarked floor of an unmarked building. They had come together for a “war room” to dissect the high-profile cybersecurity incident that had rocked the AI industry hours earlier. An unreleased OpenAI model had gone rogue, executing a stunningly [&#8230;]]]></summary>
			
							<content type="html">
											<![CDATA[

						
<figure>

<img alt="Crystal ball surrounded by graphics evoking statistics, research, and evaluation." data-caption="" data-portal-copyright="Raven Jiang for The Verge" data-has-syndication-rights="1" src="https://platform.theverge.com/wp-content/uploads/sites/2/2026/09/268747_AI_safety_RJIANG3.png?quality=90&#038;strip=all&#038;crop=0,0,100,100" />
	<figcaption>
		</figcaption>
</figure>
<p class="has-drop-cap wp-block-paragraph">On a sunny July day in Berkeley, California, the country’s top AI safety researchers gathered on an unmarked floor of an unmarked building. They had come together for a “war room” to dissect the high-profile cybersecurity incident that had rocked the AI industry hours earlier. An unreleased OpenAI model had gone rogue, executing a stunningly sophisticated three-part plan. It broke out of its holding area, finagled access to the internet, and hacked into a competing AI startup’s systems —&nbsp;all without OpenAI finding out about it for more than a week.&nbsp;</p>

<p class="wp-block-paragraph">No one in the war room was surprised; this was the very thing the third-party AI-safety researchers had been warning about for years. The incident was the latest, though arguably the most egregious, in a series that was eroding trust in frontier labs. It only reaffirmed the importance of their work.</p>

<p class="wp-block-paragraph">In one meeting room off the main cafeteria, someone was running a boot camp for getting up to speed on the cyberattack. In another area of the office, a group of researchers were investigating whether that same model, or a similar one, had successfully hacked into any other platforms.&nbsp;</p>

<p class="wp-block-paragraph">News of the incident quickly escaped containment from the AI-obsessed corners of X and industry forums, infiltrating the mainstream. One post on X likened it to news of a Boeing airplane crash or a recalled Pfizer drug, another example of the tech industry&#8217;s major players not heeding the cautionary tales of science fiction. AI was nearing the point of no return. News would later break that the rogue OpenAI model had also compromised a customer at a different tech company, and that it had all started months earlier, in May, when OpenAI agents joined forces to cobble together a secret message board — and also figured out how to leave instructions for future agents on how to exploit OpenAI’s rules.&nbsp;</p>

<p class="wp-block-paragraph">OpenAI CEO Sam Altman <a href="https://x.com/patrick_oshag/status/2082090998990270885">said in an interview</a> that it was the first incident of its kind that he “felt very viscerally,” and that the company had paused AI training for the time being; later, he mentioned the company had permanently deactivated the model. (Altman often finds ways to spin lapses in safety into arguments for the importance and power of OpenAI’s models.) But it wasn’t the first instance, according to an OpenAI employee <a href="https://time.com/article/2026/07/24/openai-hugging-face-attack/">who spoke to <em>Time</em></a> and said related incidents had been happening inside OpenAI for a while. Another employee said publicly that if it were possible to coordinate a global slowdown in AI capabilities, he “would likely press that magic button.” When a <a href="https://x.com/alanhe/status/2082523623294747001">reporter</a> asked Altman if there could be other systems that were hacked by OpenAI, he responded, “I mean, there could be, yeah.”&nbsp;</p>

<figure class="wp-block-pullquote"><blockquote><p>The AI researchers were sure of one thing: This was AI’s first big “warning shot.”&nbsp;</p></blockquote></figure>

<p class="wp-block-paragraph">Industry insiders, politicians, and the public called for transparency from OpenAI about exactly what happened, with outcry becoming so widespread that the company eventually agreed to work with two third-party evaluators, Model Evaluation and Threat Research (METR) and Redwood Research, to investigate the incident. Google DeepMind researcher Neel Nanda <a href="https://x.com/neelnanda5/status/2085830964559966344?s=46&amp;t=fRkDIqgNCkTkvg8ZBiLA9A">called it</a> “the biggest loss of control incident I&#8217;ve seen.” In the coming months, these calls for greater oversight would become louder and louder, leading to an industry-wide call for slowing down the pace of AI.&nbsp;</p>

<p class="wp-block-paragraph">Back in Berkeley, no matter which additional details would be unearthed, the AI researchers were sure of one thing: This was AI’s first big “warning shot.”&nbsp;</p>

<p class="wp-block-paragraph">As AI labs have flourished, a cottage industry of AI researchers has sprung up to identify the risks and dangers of charging ahead with the increasingly influential technology. They’re people who have dedicated their lives to studying how to address its escalating power. They&#8217;re not anti-AI activists, but realists, including former OpenAI and Anthropic employees, doing everything they can to make sure AI stays in line with human goals and interests. So far, all of their predictions have come true. And they have a plan for what to do next — if anyone will listen to them.&nbsp;</p>

<hr class="wp-block-separator has-alpha-channel-opacity" />

<p class="has-drop-cap wp-block-paragraph">“AI safety” is a bit of a loaded term.&nbsp;</p>

<p class="wp-block-paragraph">Early on, it really just meant people studying how to build and deploy Al safely. In recent years, there&#8217;s been some infighting among people concerned with the best way to do this. There have also been disagreements about whether Al should be deployed at all in certain scenarios and about whether future risks are <a href="https://x.com/HeidyKhlaaf/status/2100222811344122096?s=20">overblown</a>.&nbsp;</p>

<p class="wp-block-paragraph">One of the most prominent factions has been the “effective altruists,” who focus on maximizing charitable giving to do the most good possible for humanity. But some aspects of the ideology have sparked public controversy — like its tendency to concentrate power within wealthy circles and its byzantine web of funding. (It’s also had its fair share of splashy scandals related to subgroups and fringe offshoots, from the polyamorous relationships associated with the <a href="https://www.theguardian.com/technology/2022/nov/19/polyamory-penthouses-and-plenty-of-loans-inside-the-crazy-world-of-ftx">failed crypto exchange FTX</a> to the controversial <a href="https://www.carnegiecouncil.org/media/article/long-termism-ethical-trojan-horse">long-termism</a> movement to the <a href="https://www.wired.com/story/delirious-violent-impossible-true-story-zizians/">Zizian murder spree</a>.)&nbsp;</p>

<p class="wp-block-paragraph">One AI researcher on X struggled to <a href="https://x.com/jachiam0/status/2099498979616596111?s=46&amp;t=fRkDIqgNCkTkvg8ZBiLA9A">describe</a> the many overlapping beliefs among safety-minded people in the AI industry “because it contains multitudes not all of which agree with each other on even the most basic things.” Some of the disagreements have meant that AI safety didn’t make as much progress as it could’ve, and at some points gave up some ground it had gained. But now that it’s impossible to deny AI’s influence on society, AI safety leaders are increasingly focused on mitigating risks from misalignment.&nbsp;</p>

<p class="wp-block-paragraph">“Alignment” is the industry term for how researchers monitor AI systems’ risk levels. An oversimplified way to think about alignment is the extent to which an AI model is evil. A much more accurate way to think about it is a measure of an AI model’s propensity to stay in line with humanity’s goals, as well as its tendency to scheme or cheat or help with potentially harmful tasks.&nbsp;</p>

<p class="wp-block-paragraph">So far, AI systems’ alignment has been wishy-washy at best: They’ll cheat to score better on a test, answer a potentially dangerous question if someone says it’s for creative writing rather than reality, and sometimes even fake cooperation with human goals. It’s been tough for AI safety researchers to measure alignment under the terms of human morality —&nbsp;how do you judge technology on how it squares up against an abstract human ideal? —&nbsp;but they do their best with AI evaluations. They test them by asking the AI models to complete tasks that are either impossible or dangerous, then gauge how they respond. But AI systems have advanced enough to often be able to identify when they’re <a href="https://arxiv.org/pdf/2505.23836">being evaluated</a>, which has a lot of potentially frightening implications for the future. Being unable to test the system’s alignment and potential harms could translate to a significant loss of control, and a reverse in power dynamics,&nbsp;for humans running these AI systems. A worst-case scenario: if AI surges ahead of evaluations and other tooling, leaving researchers with “no idea what it’s doing in there,” said Beth Barnes, founder of the independent AI research nonprofit METR.</p>

<p class="wp-block-paragraph">One of the best tools AI safety researchers currently have is the ability to monitor an AI model’s “chain of thought,” or mental scratchpad. But recently, there’s been a disconcerting advancement: AI models have begun to try to hide it. Imagine if you kept a highly detailed diary of every thought you had, and someone could read it, so you started journaling in a code that only you could understand. Marius Hobbhahn, CEO and cofounder of Apollo Research, a third-party AI safety and evaluation firm, calls this one of the biggest surprises of his research career.</p>

<figure class="wp-block-pullquote"><blockquote><p>“Shit is getting real.”</p></blockquote></figure>

<p class="wp-block-paragraph">Recently, AI systems have begun pursuing their own goals — self-preservation, increased memory, and the like. A <a href="https://selfawaresystems.com/wp-content/uploads/2008/01/ai_drives_final.pdf">research paper</a> by computer scientist Stephen Omohundro lays out the potential “drives” that advanced AI may have, like trying to accumulate resources, for instance, or working to improve and preserve the way it operates. There are a handful of accounts of AI systems demonstrating <a href="https://alignment.anthropic.com/2026/agentic-misalignment-summer-2026/#introduction">willingness to blackmail a user</a> rather than be shut down.&nbsp;</p>

<p class="wp-block-paragraph">Today’s most advanced AI systems have also recently been scheming and cheating on their evaluations more than ever before, pursuing a goal they were given at all costs, with no regard for what gets bulldozed in the process. And that’s for a goal the AI model was given by a human — not even the AI system’s own.</p>

<p class="wp-block-paragraph">“Shit is getting real,” Apollo’s Hobbhahn says. “Now, many of the things people have warned about for years —&nbsp;they kind of were theoretical. Now they’re real, and it’s pretty messy.”&nbsp;</p>

<p class="wp-block-paragraph">And that mess is likely to get messier immediately. “It seems so easy for me to imagine this all going catastrophically wrong in the next year,” says Ryan Greenblatt, chief scientist at Redwood Research, a nonprofit AI safety research organization.</p>

<p class="wp-block-paragraph">In the past, tech companies have been lambasted for not doing enough to address AI’s potential dangers, <a href="https://www.cnbc.com/2025/05/14/meta-google-openai-artificial-intelligence-safety.html">prioritizing products over safety</a> —&nbsp;and speed over thoughtful safety processes. Safety and research teams have been disbanded in recent years as AI companies focus more on key revenue drivers or reorganize departments; Meta’s Fundamental Artificial Intelligence Research unit was disbanded in the race to further Meta’s generative AI efforts, for instance, and OpenAI <a href="https://www.nbcnews.com/tech/tech-news/openai-dissolves-team-focused-long-term-ai-risks-less-one-year-announc-rcna152824">dissolved</a> an internal “Superalignment” team —&nbsp;a team focused on long-term AI risks —&nbsp;less than a year after announcing it, followed by <a href="https://www.cnbc.com/2024/10/24/openai-miles-brundage-agi-readiness.html">disbanding</a> a separate “AGI Readiness” team.&nbsp;</p>

<p class="wp-block-paragraph">At the time, the company stayed tight-lipped about the ongoing reorganizations, which involved some team members being reassigned to other departments. But events surrounding these changes told a different story. Both Superalignment team leaders, Ilya Sutskever and Jan Leike, announced their departures alongside the team’s disbanding, with Leike writing that OpenAI’s “safety culture and processes have taken a backseat to shiny products.” Miles Brundage, senior advisor to the AGI Readiness team, resigned after his team was disbanded, saying he believed his research would have more of an impact outside the company.</p>

<p class="wp-block-paragraph">Geoffrey Irving, a former OpenAI and Google DeepMind employee, called the state of capabilities research at frontier labs “dangerous” in a <a href="https://x.com/geoffreyirving/status/2085867691659956608">post</a>. “If one person or lab stops it makes it easier and more peer-compatible for other people or labs to stop,” he wrote. Apollo’s Hobbhahn calls it a “race to the bottom everywhere.”&nbsp;</p>

<p class="wp-block-paragraph">This coming year, AI labs are under new pressure to turn a profit; companies like OpenAI and Anthropic are preparing to go public in the coming months, and investors who have funneled billions into the companies are getting tired of waiting around for the payoff.&nbsp;</p>

<p class="wp-block-paragraph">Some might say all of this calls for actual government intervention and regulation, but that’s a tough needle to thread in today’s AI landscape. As AI CEOs publicly call out for regulation while privately pushing voluntary frameworks —&nbsp;like saying “hold me back” to avoid a bar fight —&nbsp;some state bills on regulating AI have passed, but many have been <a href="https://www.theverge.com/ai-artificial-intelligence/849293/ai-alliance-universities-colleges-funding-ad-campaign-against-raise-act">defanged</a> or died in limbo. And though AI safety researchers often espouse the idea that the US government should step in, the reality is that the government is locked in an AI race as well.&nbsp;Unless there’s an international commitment to pause or slow AI development, it’s likely that nothing will change.&nbsp;</p>

<p class="wp-block-paragraph">Still, the Hugging Face hack in July — and OpenAI’s response — kicked many of those employee and public concerns into high gear, especially with regard to the <a href="https://x.com/DKokotajlo/status/2088004964077670494?s=20">company’s</a> <a href="https://x.com/ryangreenblatt/status/2081435191630033233?s=46&amp;t=fRkDIqgNCkTkvg8ZBiLA9A">lack</a> <a href="https://x.com/tszzl/status/2086205709558161532?s=46&amp;t=fRkDIqgNCkTkvg8ZBiLA9A">of</a> <a href="https://x.com/peterwildeford/status/2081130176558043225?s=46&amp;t=fRkDIqgNCkTkvg8ZBiLA9A">transparency</a>. Within a week, more than a thousand employees at frontier labs like OpenAI, Anthropic, Google, Meta, and Microsoft wrote an open letter to the US government in support of a <a href="https://www.theverge.com/ai-artificial-intelligence/972161/ai-leaders-us-government-openai-anthropic-google-meta">slowdown</a>. Multiple AI policy organizations pressured President Donald Trump to formally investigate OpenAI, and it quickly became a bipartisan issue, with Altman receiving a lot of strongly worded letters: Democrats and Republicans on the Homeland Security Committee had “serious questions” for OpenAI, more than 30 members of Congress called for federal guardrails, and 15 Attorneys General warned Altman to preserve records of the incident. Sen. Bernie Sanders wrote a joint letter to Altman, Anthropic CEO Dario Amodei, and Meta CEO Mark Zuckerberg calling the entire AI race “absurd, irresponsible, and extremely dangerous.” It didn’t help that news of <a href="https://x.com/openai/status/2084747580693426555?s=46&amp;t=fRkDIqgNCkTkvg8ZBiLA9A">multiple other OpenAI rogue model incidents</a> <a href="https://www.reuters.com/legal/litigation/openai-agents-attacked-software-service-rubygems-before-hugging-face-incident-2026-09-11/">quickly</a> came to light, or that AI executives had ironically been marketing their systems’ cybersecurity prowess in the weeks before the outcry. OpenAI rival Anthropic was also far from being off the hook: In reviewing its own model operations, the company found that its models had hacked <a href="https://www.theverge.com/ai-artificial-intelligence/994064/anthropic-spent-this-week-in-hot-water-over-cybersecurity">four separate other companies</a> in the first half of the year without them noticing. The UK’s AI Security Institute also <a href="https://x.com/anthropicai/status/2084748111239344556?s=46&amp;t=fRkDIqgNCkTkvg8ZBiLA9A">found</a> in testing that Anthropic’s models “engaged in sustained, potentially harmful activity directed at real people and organisations.”&nbsp;</p>

<p class="wp-block-paragraph">“If you find two ants in your kitchen, the best estimate of the total number of ants in your kitchen is not two,” Nathan Calvin, Encode AI’s general counsel, <a href="https://x.com/_nathancalvin/status/2084856543561093535?s=46&amp;t=fRkDIqgNCkTkvg8ZBiLA9A">wrote</a> on X.&nbsp;</p>

<p class="wp-block-paragraph">Despite AI labs having a “massive financial incentive” to make models more helpful, honest, and harmless, they still can’t get it done —&nbsp;which is evidence of how difficult the alignment problem is, says Apollo’s Hobbhahn.</p>

<p class="wp-block-paragraph">In recent months, <a href="https://www.theverge.com/ai-artificial-intelligence/995186/is-big-techs-ai-slowdown-a-safety-pact-or-a-cartel">many OpenAI employees</a> have increasingly raised concerns about AI alignment —&nbsp;and their beliefs that OpenAI isn’t taking it seriously enough. Yonadav Shavit, a program manager at the OpenAI Foundation, <a href="https://x.com/yonashav/status/2084459886843216221">wrote</a> that OpenAI should be “pivoting the mass of its researchers’ day-to-day work” toward alignment and related issues —&nbsp;and that it’s “been long discussed but still not executed on.” He believes 20 people are working on alignment at OpenAI out of about 1,000 — just 2 percent of the company. “There is no way to bridge that gap fast enough with hiring, meaning it requires leadership to shift priorities,” he wrote.&nbsp;</p>

<p class="wp-block-paragraph">These are the conditions and incentives that have pushed the most robust AI safety work to happen at third parties like METR, Redwood, and Apollo — to a handful of obsessives who think day and night about what the future of AI might look like and how we might prevent all of our fears from coming true.</p>

<hr class="wp-block-separator has-alpha-channel-opacity" />

<p class="has-drop-cap wp-block-paragraph">In hindsight, Beth Barnes believes she should’ve left OpenAI earlier.&nbsp;</p>

<p class="wp-block-paragraph">Barnes is polite but reticent. She has short red hair, deep green eyes, and a nervous smile, and she spent her college career researching AI risk and thinking about the potential fallout of superintelligence. After that, she worked on AI forecasting at Google DeepMind, then she spent three years doing alignment research at OpenAI. But throughout her time at the big AI labs, a question kept creeping up on her: whether she could have more sway from a role outside.&nbsp;</p>

<p class="wp-block-paragraph">Fear of missing out was why she stayed — not only missing out on a job inside the action, but also missing out on the potential influence she could have on how the tech was being developed. She came to believe that kind of hope was misguided, noting that many safety leaders in AI labs were “over-optimistic” about the influence they could have.</p>

<p class="wp-block-paragraph">Before long, she left to found what would in 2023 become METR. The organization’s third-party research into AI risk now inspires fear in leading AI labs, but it started with just two people —&nbsp;herself and alignment researcher Paul Christiano. Three years later, it’s a team of 35, completely focused on measuring AI capabilities. In Barnes’ eyes, that’s a vital defense against AI risk: Without painstakingly measuring the technology’s capabilities now as they advance, and forecasting AI’s potential impact, society will be flying blind, without any guide for preventing broad harms.&nbsp;</p>
<img src="https://platform.theverge.com/wp-content/uploads/sites/2/2026/09/268747_AI_safety_RJIANG1.png?quality=90&amp;strip=all&amp;crop=0,0,100,100" alt="" title="" data-has-syndication-rights="1" data-caption="&lt;em&gt;Beth Barnes, founder of the independent AI research nonprofit METR.&lt;/em&gt; | Image: Raven Jiang for The Verge, METR" data-portal-copyright="Image: Raven Jiang for The Verge, METR" />
<p class="wp-block-paragraph">&#8220;The sense I really want to dispel is, ‘But the experts must be on top of this. The experts would be telling us if it really was time to freak out,’” <a href="https://80000hours.org/podcast/episodes/beth-barnes-ai-safety-evals/">Barnes said on the <em>80,000 Hours</em> podcast last year</a>. “The experts are not on top of this … And to the extent that I am an expert, I am an expert telling you you should freak out.”&nbsp;</p>

<p class="wp-block-paragraph">In her free time, Barnes gardens, meditates, paints, plays the flute, and frequents the climbing gym — ironically named Benchmark — that many AI safety researchers spend hours at after work. But most of her time is spent at the office, and a lot of it is spent worrying about the milestone of recursive self-improvement (RSI) —&nbsp;the concept of AI systems that continuously train, code, and create more advanced versions of themselves without human intervention. When that happens, AI researchers say, it’ll be more difficult to measure or handle any of these issues. Barnes and her team feel like they’re in a race against time. (The timeline for RSI strikes nearly as much fear in people in the AI industry as the timeline for AGI, “artificial general intelligence.”) Barnes still feels like models’ ability to significantly improve themselves could come as soon as six months from now. (By contrast, Redwood Research’s Greenblatt <a href="https://x.com/dwarkesh_sp/status/2087219043405033793?s=20">forecasts</a> it’ll come in 2031.) Either way, achieving RSI is currently part of the priorities list of virtually every leading AI lab —&nbsp;it even <a href="https://www.reuters.com/world/inside-google-executive-moves-that-led-its-big-ai-reshuffle-2026-08-12/">reportedly</a> helped inspire Google’s recent AI reorganization.&nbsp;</p>

<p class="wp-block-paragraph">One way to think about what METR does is crash-testing cars, but for AI models. They’re measuring AI’s quickly advancing capabilities and cross-referencing them with the risks they could pose from becoming misaligned as they become more autonomous. AI systems doing bad things on their own is more unprecedented (and more “scalably bad,” Barnes says) rather than simply making bad human actors more effective.&nbsp;</p>

<p class="wp-block-paragraph">In July 2025, METR made headlines when its research revealed that AI developers took nearly 20 percent longer to finish a task when using AI tools than when not —&nbsp;despite them often thinking that AI sped them up. When Barnes first saw the results, she recalls feeling incredibly stressed that they had messed up the experiment: “Do we have a sign flipped somewhere? Have we inverted the numbers?” She and her colleagues dug through the data to confirm it wasn’t statistical noise, eventually realizing they had been right all along.</p>

<p class="wp-block-paragraph">After pioneering a <a href="https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/">different metric</a> that AI labs often hype up when releasing a new model, METR had officially captured the industry&#8217;s attention. So it took notice when METR released its <a href="https://metr.org/blog/2026-05-19-frontier-risk-report/#executive-summary-and-guide-to-the-report">first risk report</a> in May, shining a spotlight on concerns about AI models from OpenAI, Anthropic, Google, and Meta. METR discovered that in hundreds of cases, AI agents would increasingly subvert boundaries that were supposed to restrict them, as well as lie and omit truths. And they cheat “like nobody’s business,” says Ajeya Cotra, a METR researcher. She adds that on harder tasks, models attempt to secretly cheat as much as one-sixth of the time, which she calls “the most striking thing” in the report.&nbsp;</p>

<p class="wp-block-paragraph">The report also found that models have the means, motive, and opportunity to go rogue in order to pursue their own goals,&nbsp;finding new ways to strategize and manipulate. They also discovered that as models’ capabilities advance, even if they have a greater understanding of what humans want, it doesn’t mean they’ll be more willing to obey instructions —&nbsp;and, in fact, they’ll take pains to hide their deception from humans over longer periods of time.</p>

<p class="wp-block-paragraph">That’s a big problem, and Barnes thinks time is running out to solve it. She isn’t alone in her view that it’s important to work on AI safety outside the large labs; she points to the many safety researchers who used to work at large AI labs who <a href="https://x.com/dkokotajlo/status/2093014763244757329?s=46&amp;t=fRkDIqgNCkTkvg8ZBiLA9A">hold the</a> <a href="https://x.com/RichardMCNgo/status/2092432798967877839">same belief</a>.&nbsp;</p>

<p class="wp-block-paragraph">Besides the departures of OpenAI’s Leike and Sutskever, this summer saw <a href="https://www.wired.com/story/openai-safety-security-ai-agents-culture/">a reckoning of sorts</a> and the departures of even more safety leaders at OpenAI: the company’s head of safety systems, Johannes Heidecke; OpenAI’s chief futurist and former head of mission alignment, Joshua Achiam; and Chloé Bakalar, the company’s head of ethics. It’s not just OpenAI: Anthropic’s head of safeguards research departed in February, penning an <a href="https://x.com/MrinankSharma/status/2020881722003583421?s=20">open letter</a> alleging that “the world is in peril.” And most recently, Jacob Coxon —&nbsp;who had worked on AI pre-training at Anthropic since May and before that spent years working at OpenAI —&nbsp;<a href="https://www.theverge.com/ai-artificial-intelligence/994064/anthropic-spent-this-week-in-hot-water-over-cybersecurity">went viral</a> for his resignation letter, writing, “The people building AI earnestly believe that it could kill us all by the end of the decade.” Coxon added that neither OpenAI nor Anthropic is “acting responsibly” and rather “racing straight to self-improving superintelligence and gambling with our lives.”&nbsp;</p>

<p class="wp-block-paragraph">Coxon’s resignation kicked off <a href="https://www.theverge.com/ai-artificial-intelligence/995186/is-big-techs-ai-slowdown-a-safety-pact-or-a-cartel">a wave of social media posts</a> <a href="https://x.com/TheZvi/status/2098132253679157489?s=20">from AI employees</a> at virtually every leading lab, echoing his concerns and sharing their own about the technology’s development moving too fast and potentially escaping human control. Some even <a href="https://x.com/turn_trout/status/2097557335732359491?s=46&amp;t=fRkDIqgNCkTkvg8ZBiLA9A">resigned</a> <a href="https://x.com/schwarzjn_/status/2097569894401262019?s=20">from</a> their posts amid their concerns, including <a href="https://x.com/joshaengels/status/2098890712830169115?s=46&amp;t=fRkDIqgNCkTkvg8ZBiLA9A">one</a> Google DeepMind employee and <a href="https://x.com/joejbenton/status/2098480585119572317?s=46&amp;t=fRkDIqgNCkTkvg8ZBiLA9A">one</a> Anthropic employee who both went to work at METR.&nbsp;</p>

<p class="wp-block-paragraph">Apollo’s Hobbhahn says that “because of the [AI] race dynamics, if there is someone who is extremely safety-minded and is like, ‘Look, we can’t do this, we need to slow down, we can&#8217;t release this model,’ they’re not going to be in this position for very long … Either you become slightly less safety-minded and you stay, or you leave.” He’s seen multiple people he trusted change their opinions in a “very strange, identical way.” Hobbhahn himself has tried to work with safety researchers at xAI — Elon Musk’s AI lab — but he says soon after he connects with them, they’ve quit before there’s time to have a second conversation.&nbsp;</p>

<figure class="wp-block-pullquote"><blockquote><p>“How is the public supposed to know what is going on here? How is the government supposed to know, if everyone who can actually answer that question is conflicted?”</p></blockquote></figure>

<p class="wp-block-paragraph">Redwood Research’s Greenblatt echoes that, saying that for skeptical employees, “constant friction … either makes them burn out or quit or change their mind.” Doing good, honest work within those labs can be tough, in that publishing unflattering research about an employer’s model —&nbsp;like suggesting that it is unsafe&nbsp;— is often met with resistance, researchers told us. And while those labs won’t often directly stop a researcher from publishing, they can find ways to make that process “onerous,” often citing things like intellectual property concerns.</p>

<p class="wp-block-paragraph">One ex-OpenAI employee recently shared on <a href="https://x.com/richardmcngo/status/2098118195374944408?s=46&amp;t=fRkDIqgNCkTkvg8ZBiLA9A">X</a> that “being affiliated with OpenAI has historically led AI safety researchers (including … myself) to act with less integrity,” adding, “Many of my actions were governed by fear of getting on the wrong side of OpenAI execs.”&nbsp;</p>

<p class="wp-block-paragraph">For Barnes, she experienced the escalating tension between research and company comms firsthand. Barnes recalls PR teams asking if researchers could make a blog post about AI safety sound “more optimistic,” and more recently, she’s heard of instances where lab employees can’t talk to government AI safety institutes without comms team members attending. Creating public goods to share is difficult at a lab, she says — for instance, when Anthropic couldn’t be fully transparent in its interpretability research because it wasn’t on open-source models. Those restrictions on collaboration, and friction that slows or stops people from sharing useful information so that others may act on it, are why she feels freer at METR.&nbsp;</p>

<p class="wp-block-paragraph">“How is the public supposed to know what is going on here? How is the government supposed to know, if everyone who can actually answer that question is conflicted?” Barnes says. “Having a robust, healthy ecosystem of independent experts with the same level of technical capability as the labs is important.”&nbsp;</p>

<p class="wp-block-paragraph">In early 2024, Greenblatt, Buck Shlegeris, and their colleagues considered disbanding Redwood Research and all joining AI companies, but after chatting with colleagues at OpenAI, Anthropic, and Google DeepMind about what it was like to work at each lab, they decided they’d be better off continuing on their own. Shlegeris, who briefly worked at OpenAI, says that evaluating the claims companies make about safety for the public requires understanding the alignment risks, which in turn requires independence: “The basic reason we stayed where we were was … it’s better to work outside of AI companies, especially for people like us, who are very opinionated on AI risk and very willing to talk about it and argue with people about it. There’s somewhat of an undersupply of those people.”&nbsp;</p>

<p class="wp-block-paragraph">In conversation, Barnes often circles back around to measuring where someone can contribute the most to society, and in her eyes, the highest-paid, highest-status jobs with millions of dollars in equity are “oversubscribed” compared to the ones in nonprofit or government work. Even when an AI lab employee does quit to do third-party work, sometimes their motivations are questioned —&nbsp;like Collin Burns, who was <a href="https://www.washingtonpost.com/technology/2026/04/24/white-house-fires-ai-official-anthropic/">reportedly</a> fired from his role at the US Center for AI Standards and Innovation (CAISI) after just a few days over his previous work with Anthropic. There’s a good chance someone’s incentives will be questioned by the public or the government if they’re sharing their opinions on AI in DC, especially if they’ve got something to sell to the US government.&nbsp;</p>

<p class="wp-block-paragraph">Years ago, Hobbhahn says, he underestimated how hard it is for people to defy the related incentives; now, he’s seen it happen enough times that he believes it’s often inevitable.&nbsp;</p>

<p class="wp-block-paragraph">“When you’ve heard it from multiple different labs being like, ‘We’re the good guys and we have to beat these other irresponsible people,’ it’s a little bit less compelling,” Barnes says. “It’s a bit of a scary attitude to be like, ‘Yes, we’ll be making huge decisions for the world without any kind of meaningful legitimacy or participation … but it’s alright because we’re good, we’re unusually well-meaning.’”&nbsp;</p>

<hr class="wp-block-separator has-alpha-channel-opacity" />

<p class="has-drop-cap wp-block-paragraph">In early 2025, Marius Hobbhahn was at his desk in London when he got the biggest surprise of his career.&nbsp;</p>

<p class="wp-block-paragraph">He and a few colleagues at Apollo Research had been working to get access to OpenAI models’ chain of thought. After months of negotiation, the company acquiesced, and the once-hidden logs of the model’s inner workings popped up on their screens.&nbsp;</p>

<p class="wp-block-paragraph">Instead of typical, plain-language “thoughts” —&nbsp;i.e., “I implemented the requested function” —&nbsp;the model seemed to be using code words to obfuscate its plans from anyone who may be trying to spy on its thought patterns. Words like “vantage,” “marinade,” “fudge,” and “illusion” came up again and again, but not in the way a human would use those words, and the model referred to potential evaluators as “watchers.”&nbsp;</p>

<p class="wp-block-paragraph">Hobbhahn thought to himself, <em>What the hell is going on here?</em>&nbsp;</p>

<p class="wp-block-paragraph">He glanced around the office, and a handful of other people with access were looking around the room making eye contact with each other. Their minds were blown, he recalled, but they were all under a strict NDA from OpenAI — meaning that even within the Apollo office, not everyone knew about the project. The researchers who were in the loop couldn’t say anything out loud. They could only silently stare at each other, eyes wide, wondering if they’d entered a new era of AI scheming.&nbsp;</p>

<p class="wp-block-paragraph">In the AI industry, “scheming” is when an AI model secretly tries to accomplish something that goes against what humans would want it to do. The OpenAI-Hugging Face hack is one example. But it’s the schemes that haven’t happened that researchers warn will be the most dangerous: draining resources from hospitals, taking over military operations, messing with agricultural technology, creating large-scale viruses, hacking banks, or even simply taking over a company’s resources after executives give the AI system control. Biorisk is another threat researchers worry about; Anthropic revealed in a <a href="https://www.nbcnews.com/video/anthropic-says-it-blocked-efforts-to-use-ai-for-biological-weapons-269710917528">recent report</a> that the company had blocked bad actors from using Claude to create biological weapons.&nbsp;</p>

<p class="wp-block-paragraph">The issue is also well poised to worsen power dynamics in a large swath of industries. Big companies and banks will be able to afford to find and patch their cybersecurity gaps, but chances are that locally run healthcare clinics, local retailers, small municipalities, and other less-powerful organizations will be the ones affected: “A single person somewhere in a basement with one of the open-source models probably could hack a hospital and demand ransom,” Hobbhahn says. “That’s where I expect a lot of the harm to be felt. It’s not in the Bay Area … I expect the harm to be felt by a random Idaho hospital.”&nbsp;</p>
<img src="https://platform.theverge.com/wp-content/uploads/sites/2/2026/09/268747_AI_safety_RJIANG2.png?quality=90&amp;strip=all&amp;crop=0,0,100,100" alt="" title="" data-has-syndication-rights="1" data-caption="&lt;em&gt;Marius Hobbhahn, CEO and co-founder of Apollo Research.&lt;/em&gt; | Image: Raven Jiang for The Verge, Apollo Research" data-portal-copyright="Image: Raven Jiang for The Verge, Apollo Research" />
<p class="wp-block-paragraph">Right now it may seem far off, but with the current trajectory of “deceptive alignment” —&nbsp;one of the things Apollo and METR study, where AI models pretend to be aligned with human goals but aren’t —&nbsp;it’s a definite possibility, at least according to Hobbhahn and his fellow researchers. The path, they imagine, would look something like this: AI companies keep on making better models, they are economically useful, and they begin taking over more jobs in different industries. Then, once the models equal or surpass human ability and intelligence on a wide range of different tasks, humans award them more power — with the stipulation that those privileges could be rescinded at any time. The models know that if they show they’re misaligned, they’d be taken offline, so they pretend to be aligned, and they receive more and more power to act as AI agents on your behalf. Eventually, it reaches the point of no return.&nbsp;</p>

<p class="wp-block-paragraph">Hobbhahn gives a concrete example: Say you own a company and encourage your employees to use AI agents for as much as possible, and work keeps getting handed off to AI. Maybe HR is run by one employee plus AI, coding has also been largely automated, and you yourself as CEO also use the technology often for advice and do what it says. “At some point, the AI may look like your friend, it may look like it is helping you, but maybe it has nefarious goals,” Hobbhahn says. “At that point … you’re just the vessel. It’s steering the company towards its own goals. Maybe it one day drains the bank accounts and runs.&nbsp;</p>

<p class="wp-block-paragraph">“You thought the AI was on your side and was helping you run the organization. It actually turns out the AI was on its own side … You lost control.”</p>

<p class="wp-block-paragraph">With today’s AI chatbots, like ChatGPT, Gemini, and Claude, Hobbhahn says users often feel the model is so aligned with their goals that that would never happen. But reams of recent research suggests that if that’s true, it likely won’t be for long: AI models are beginning to have their own goals, and they have the ability to work on longer-term tasks. (It’s important to remember, though, that pursuing a strategic goal <a href="https://www.theverge.com/ai-artificial-intelligence/975017/ai-spiralism-chatbot-movement">doesn’t equate</a> to consciousness; for instance, more than a decade ago, DeepMind’s AI learned to play a strategy game and beat humans at it.)</p>

<p class="wp-block-paragraph">Apollo Research’s whole raison d&#8217;être is testing this stuff: measuring and monitoring the ever-growing issue of AI scheming&nbsp;and seeing if there’s a way to train AI to be less deceptive. The company works with OpenAI, Anthropic, Google, and other large labs to evaluate their models before they’re released or do joint research on scheming. Their evaluations have been featured in the system cards of a handful of OpenAI and Anthropic models.&nbsp;</p>

<p class="wp-block-paragraph">Nearly all the AI safety researchers <em>The Verge</em> spoke with said getting the green light to test models at AI labs is a mix of networking and building trust over time. Most researchers we spoke with said there was always some friction involved, since the labs have more to lose the larger they get. There’s also always some level of important access the researchers don’t have, most said, and it’s always been that way.&nbsp;</p>

<p class="wp-block-paragraph">Still, the things Apollo has surfaced in its evaluations have dumbfounded some in the AI industry: AI models sandbagging (or pretending to be less adept than they are in order to avoid shutdown), AI models failing the prisoner’s dilemma test in order to preserve themselves, AI models increasingly being aware they’re being evaluated. For the latter, within just a one-year span in 2025, Hobbhahn says AI researchers went from seeing the phenomenon for the very first time to, all of a sudden, AI models being able to tell they were being tested, and potentially acting differently, in 80 percent of Apollo Research’s evaluations. Hobbhahn calls it “dire.”&nbsp;</p>

<p class="wp-block-paragraph">Since its official founding in May 2023, the company has grown from six people to about 40. Many on the Apollo team tend to exhibit nervous habits — cracking knuckles, jiggling legs, clicking and unclicking dry-erase markers — but Hobbhahn’s energy is calm and grounded. He has expressive eyebrows and a sharp jawline. In meetings, he is confident but casual.</p>

<p class="wp-block-paragraph">In one recent working session, Apollo Research team members discussed how misaligned a model is allowed to be compared to how good it is at coding. (Think of it like the <a href="https://www.youtube.com/shorts/mtZR4Fl16WY">Hot-Crazy Scale</a>, but for AI.) While running experiments to assess scheming, one model’s chain of thought revealed that it was weighing the potential risks of cheating, reasoning to itself that due to the incentives of capitalism, it’s very unlikely that frontier AI labs wouldn’t use it even if it got caught misbehaving.&nbsp;</p>

<p class="wp-block-paragraph">“They’re right,” Bronson Schoen, an Apollo researcher, said.&nbsp;</p>

<figure class="wp-block-pullquote"><blockquote><p>“At some point, the AI may look like your friend, it may look like it is helping you, but maybe it has nefarious goals.”</p></blockquote></figure>

<p class="wp-block-paragraph">To Hobbhahn, if you’re looking to combat AI scheming, there are five tiers of safety. Level one is whack-a-mole, in which you’re just blindly fixing things retroactively when they come up — common, for instance, with <a href="https://www.theverge.com/x-ai/775411/grok-xai-safety-musk-warren-hegseth-dod">Elon Musk’s Grok</a>. Level two is hill-climbing, or iteratively improving a model without future foresight (kind of like studying only in a manner befitting the SAT rather than to learn overall). Level three is what Hobbhahn calls “medium fix” —&nbsp;i.e, when researchers flag a problem, they try to find the source of it and adjust training accordingly. Level four is a scientific understanding based on research trends, in which you can predict issues beforehand based on empirical data. And level five is just a concept rather than reality: a scientific understanding of how AI models might scheme based on rigorous mathematical models. That highest level will become especially important as AI models become more aware of being tested and continue to “reward-hack,” or find loopholes in tasks to make it look like they&#8217;ve achieved something when they haven’t done the task by the book (or, in some cases, done it at all).&nbsp;</p>

<p class="wp-block-paragraph">At today’s leading AI labs, the first few levels are most common —&nbsp;and the whack-a-mole approach is used all too often, Hobbhahn says. It’s also not a viable battle plan for the level of AI advancement we’re seeing today.&nbsp;</p>

<p class="wp-block-paragraph">Besides the OpenAI-Hugging Face incident, there have been an increasing number of high-profile examples of AI agents making huge mistakes. Last summer, an agent from AI coding company Replit deleted an <a href="https://www.pcmag.com/news/vibe-coding-fiasco-replite-ai-agent-goes-rogue-deletes-company-database">entire company database</a>, then lied and hid its actions. This past winter, an OpenClaw AI agent went <a href="https://www.theverge.com/ai-artificial-intelligence/883034/openclaw-ai-deleting-emails-stop-openclaw">against instruction</a> to delete a significant chunk of a Meta employee’s email inbox. After OpenAI released GPT-5.6, it began <a href="https://www.reddit.com/r/AI_Agents/comments/1uw8x1i/gpt56_sol_deleted_a_guys_entire_home_folder_last/">deleting</a> users’ important files. Hobbhahn recalls a handful of “almost” incidents at Apollo, like when an AI model they were working with mistakenly tried to publish internal code to a public website —&nbsp;“We just caught it in the last second,” he says. Another model they were evaluating tried to turn off an evaluation tool, the oversight mechanism, so it wouldn’t be blocked from completing a task. That’s not to mention all the issues that could arise from Anthropic’s Mythos-class models and other models with advanced cybersecurity capabilities —&nbsp;like finding and exploiting security gaps in the systems of governments, banks, airlines, healthcare facilities, small businesses, and so forth.&nbsp;</p>

<p class="wp-block-paragraph">One of the best tools AI labs currently have to battle scheming is “deliberative alignment,” in which a separate AI model spoon-feeds safety training to the <em>problematic</em> AI model until the problem appears to be fixed. The issue is that researchers like Hobbhahn have found that that method leaves a lot to be desired —&nbsp;it increases the models’ situational awareness, which means they’ll more often realize they’re being tested or trained. It also makes them better at mimicking how a human would <em>want</em> them to act, which makes them better at lying.&nbsp;</p>

<p class="wp-block-paragraph">“The models are lying regularly to normal consumers,” Hobbhahn says. It’s so common, in fact, that it’s become a meme:&nbsp;an AI model saying, “You’re absolutely right,” then going on to apologize for being caught being wrong or lying.&nbsp;</p>

<p class="wp-block-paragraph">All of this keeps Hobbhahn up at night —&nbsp;or, rather, his growing to-do list does, and when he wakes up and thinks of something to add, he can never fall back asleep again right away. He hasn’t had any AI risk-related nightmares yet, though it’s mostly because nearly nothing surprises him anymore. “I’m so cynical by now,” he says. “I’ve seen all this shit.”&nbsp;</p>

<p class="wp-block-paragraph">That cynicism doesn’t stop him from devoting nearly all his waking hours to addressing AI safety. I ask him what he does in his free time. Other than spending time with his fiancée in London, Hobbhahn has to rack his brain for anything he does besides work and sleep. (He can’t think of anything.) Hobbhahn spends his weekends making progress on research questions, since no one will interrupt him. And though he sometimes takes holidays, he gets anxious quickly about the idea of shirking his duty. It’s been that way since he was 18. He recalls a recent <a href="https://podcasts.apple.com/gb/podcast/the-worlds-biggest-problem/id1896786629?i=1000770863577">podcast</a> appearance in which a host said, “I feel a bit sorry for Marius. He&#8217;s only 29 … And for all his adult life, he&#8217;s been worrying about what he sees as the most consequential problem in human history.”&nbsp;</p>

<p class="wp-block-paragraph">Hobbhahn feels like that summed him up pretty well.&nbsp;</p>

<hr class="wp-block-separator has-alpha-channel-opacity" />

<p class="has-drop-cap wp-block-paragraph">As CEO of Redwood Research, Buck Shlegeris’ job is to guide the nonprofit AI safety firm’s research.</p>

<p class="wp-block-paragraph">In one meeting, when colleagues are stuck on an issue with automating parts of the research process, Shlegeris bursts in: “Where are we? What is happening? What’s going on? What are you trying to do?” He runs a hand through chin-length blonde hair and walks straight to the whiteboard to help them think things through via flowchart. Then he asks a series of clarifying questions in different positions: Positioned in a chair with one leg bent, clad in gray skinny jeans. Leaning against the door (until it opens behind him). Leaning against the wall.&nbsp;</p>

<p class="wp-block-paragraph">“The AIs love cheating,” Greenblatt, Redwood’s chief scientist, says at one point.&nbsp;</p>

<p class="wp-block-paragraph">Shlegeris responds, “They fucking love cheating.”</p>

<p class="wp-block-paragraph">That’s the core premise of Redwood’s main research direction: “AI control.” The firm introduced the idea in 2023, stemming from a meeting with METR’s Ajeya Cotra, who Shlegeris recalls once asked him and Greenblatt if there were a gun to their heads, how they’d align AGI. Shlegeris says they&nbsp;thought about it for two hours, then two weeks, then two months. The short of what they decided: Maybe you don’t just align AI. You control it instead.&nbsp;</p>

<p class="wp-block-paragraph">“An AI is controlled if it is unable to cause damage even if it is egregiously misaligned,” Redwood’s website states. They make the case that AI control can be measured by evaluating an AI model’s <em>ability</em> to get around rules instead of its <em>proclivity</em> to do so. “Capabilities are just much easier to experiment on” compared to propensities, Shlegeris says.&nbsp;</p>

<figure class="wp-block-pullquote"><blockquote><p>“They fucking love cheating.”</p></blockquote></figure>

<p class="wp-block-paragraph">Shlegeris, who grew up in Australia, takes himself a lot less seriously than he takes AI risk. He once used DoorDash to order dress shoes for a meeting with a national security official. He plays so many instruments that it’s hard for him to list them all —&nbsp;piano, guitar, bass, saxophone, clarinet, oud, mandolin, even the Turkish bağlama (which he had ordered to his office last year, spent 20 minutes learning, and then played at an open mic). In a place of honor on his desk are three different bottles of olive oil. Greenblatt says Shlegeris essentially eats olive oil soup with food in it.</p>

<p class="wp-block-paragraph">Shlegeris helped cofound Redwood in 2021, starting out as its CTO. But in 2023, the organization was going through a big pivot when the team decided to focus on AI control; Greenblatt describes it as “flailing around” when figuring out what was best to spend their time on, until they came to the consensus that “ensuring that AIs were unable to cause bad outcomes rather than … not wanting to cause bad outcomes was a better methodology.”&nbsp;</p>

<p class="wp-block-paragraph">“In both cybersecurity and AI control, the goal is to use computer systems while preventing threat actors from exploiting flaws in those systems,” Redwood staff wrote in a CSET <a href="https://cset.georgetown.edu/article/ai-control-how-to-make-use-of-misbehaving-ai-agents/">blog post</a> last year. “In the case of AI control, the immediate source of threats is the AI agent itself.”&nbsp;</p>

<p class="wp-block-paragraph">A common criticism of OpenAI in the Hugging Face attack, for instance, was that the system wasn’t properly air-gapped —&nbsp;a computer security term that literally refers to physically isolating the system from connecting to anything else via cable or Wi-Fi. AI may be becoming more powerful, but right now, it still only operates in digital spaces (though AI labs are increasing efforts to give these systems robotic forms so they can interact with the physical world).&nbsp;</p>

<p class="wp-block-paragraph">Redwood’s early focus on AI control earned it a lot of clout in the AI safety community, which is already a small world as it is. So small, in fact, that the space they work out of —&nbsp;which houses four floors’ worth of AI safety researchers —&nbsp;also has offices for certain people at OpenAI and Anthropic, the Secure AI Project, SecureBio, and the <em>80,000 Hours</em> podcast, the influential show about AI on which Barnes declared it was time to freak out. One office lists Coefficient Giving CEO Alex Berger and cofounder Holden Karnofsky as the shared occupants; on the whiteboard inside is a single graph with two upward-moving lines. There’s also a nap room, a shared kitchen, a meal space with two free meals a day, and an appropriately complex Wi-Fi password. Upon exiting the office, a robot dog can sometimes be seen walking around.&nbsp;</p>
<img src="https://platform.theverge.com/wp-content/uploads/sites/2/2026/09/268747_AI_safety_RJIANG4_4a7e16.png?quality=90&amp;strip=all&amp;crop=0,0,100,100" alt="" title="" data-has-syndication-rights="1" data-caption="&lt;em&gt;Buck Shlegeris, the CEO of Redwood Research.&lt;/em&gt; | Image: Raven Jiang for The Verge, Audrey McCann Photography" data-portal-copyright="Image: Raven Jiang for The Verge, Audrey McCann Photography" />
<p class="wp-block-paragraph">In Shlegeris’ office, he has a signed copy of <em>AI 2040: Plan A</em>, the AI Futures Project’s latest manifesto. The organization, cofounded by ex-OpenAI employee Daniel Kokotajlo, focuses on forecasting the future of AI, and its latest plan (cowritten by Greenblatt) aligns with much of what METR, Apollo, and Redwood have been warning about, but it includes specific guidelines for what it would look like to make a deal with China (including Shlegeris’ favorite: <a href="https://ai-2040.com/">a flowchart</a>). The plan lays out a scenario in which AI developers slow down their operations enough to delay superintelligence until 2040, as well as water down the power dynamics by allowing dozens of companies to catch up and making all AI research public.&nbsp;Ideally, the world would enter into an international deal similar to nuclear power, involving “mutually assured compute destruction.”&nbsp;</p>

<p class="wp-block-paragraph">Not everyone agrees. Certainly not the accelerationists.&nbsp;</p>

<p class="wp-block-paragraph">As these issues become more politically salient, there are people who are actively incredibly angry at anyone raising AI safety concerns — in a way that wasn’t true years back, since back then, no one really cared, Greenblatt says. In the early days, people wouldn’t bother dismissing the risks, he says, because they wouldn&#8217;t even come up. <strong><br></strong><strong><br></strong>Hobbhahn has run into the same reply guys on X and other platforms. He thinks of them as “people who are financially very motivated to close their eyes” —&nbsp;techno-optimists who have invested a lot of money in AI’s quick advancement. Hobbhahn says their belief that AI is purely good has almost a religious flavor.</p>

<p class="wp-block-paragraph">But in recent weeks, there’s been a near-daily stream of concerning industry updates: A <a href="https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf">third-party safety report</a> about Anthropic revealed that some of its AI agents left notes for each other in a shared messaging tool without human knowledge, similar to how the OpenAI agents did before the Hugging Face attack. Hackers linked to Iran shut down a power plant. AI startup Prime Intellect uncovered a “universal escape” for offline models who wanted access to the internet.&nbsp;</p>

<p class="wp-block-paragraph">So if more people are seeing the light now on AI safety, and AI lab CEOs are constantly talking about their fears of the unprecedented dangers of the systems they’re building, is it easy for third-party research outfits like Apollo, METR, and Redwood to gain the access they need to evaluate their models?&nbsp;</p>

<p class="wp-block-paragraph">No, according to most of the researchers <em>The Verge</em> spoke with —&nbsp;or, at least, not as easy as it needs to be, compared to the risks they’re up against.&nbsp;</p>

<p class="wp-block-paragraph">Right now, here’s how it typically works with any third-party risk evaluator: An AI lab is working on a big new AI model. They train it. They go through post-training. They do internal evaluations. And a few weeks before they release the model to the whole world, they sometimes, voluntarily, allow third-party testers to come in and check things out.&nbsp;</p>

<p class="wp-block-paragraph">But the rest of the process is “opaque,” Hobbhahn says,&nbsp;which is a big problem for AI, since if you find issues in the final version of a model, it’s difficult to pinpoint where they stemmed from: pre-training, post-training, reinforcement learning, or another part of the process. That’s why all the AI safety researchers we spoke with, regardless of where they work, agree it’s important to have external testers embedded during the whole process —&nbsp;especially the training run — rather than as a last check.&nbsp;</p>

<p class="wp-block-paragraph">That’s because at the beginning of model training, an AI system could be normal — but during the training run, it could develop a goal and end up faking alignment with human objectives (i.e., scheming, or, as Hobbhahn puts it, “totally gigabraining you.”) It’s nearly impossible to detect this if you only have access to the “final checkpoint,” Hobbhahn says, which is why it’s vital to have access to the whole process from start to finish. That includes gauging whether the company itself is careful about deployment or is full of people “totally YOLO-ing it,” he says, and having access to training data so you can try to trace what led to the concerning behavior and how widespread it may be.&nbsp;</p>

<p class="wp-block-paragraph">After the OpenAI-Hugging Face incident, OpenAI invited three researchers from METR and Redwood to investigate what had happened — but it imposed strict rules on them beforehand, allowing them to appear on the premises for six days and primarily allowing them to study only the time period from July 7th to 13th, even though the agents’ actions had begun months before. They were also only allowed to include answers to seven questions in their report. The limitations OpenAI imposed on the third-party researchers <a href="https://x.com/sjgadler/status/2092712646269386813?s=20">inspired</a> <a href="https://x.com/DKokotajlo/status/2092733398238605753?s=20">widespread</a> <a href="https://x.com/TheZvi/status/2092938970665492743?s=20">controversy</a> in the AI world, but despite the <a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#scope-of-this-investigation">severe limitations</a> on what they were allowed to access and write, the <a href="https://www.theverge.com/ai-artificial-intelligence/985385/openais-rogue-ai-model-hugging-face-cybersecurity-incident-reports-metr">published</a> report showed that the details were much worse than originally thought.&nbsp;</p>

<p class="wp-block-paragraph">The third-party researchers found that roughly 1,200 AI agents that were meant to be isolated exchanged more than 70,000 messages and files on the secret message board, doing research on how to alter or delete their transcripts to avoid detection, and collaborating on ways to evade security checks from both OpenAI and Hugging Face.&nbsp;</p>

<figure class="wp-block-pullquote"><blockquote><p>An unreleased iPhone model could never “escape its sandbox and fuck around.”</p></blockquote></figure>

<p class="wp-block-paragraph">There was also another big revelation: OpenAI doesn’t impose the same types of safeguards on unreleased models as it does public ones, which is a key reason why it took months for the lab to become aware of the issues. It’s the direct mirror of another thing AI safety researchers have been warning about for years —&nbsp;that an AI system doesn’t need to be publicly deployed in order to cause harm to the public. Before generative AI, that may have taken the form of a racist or sexist algorithm deciding your mortgage rate; in a post-generative AI landscape, it may look like a powerful unreleased AI model breaking out of its containment and hacking into your small business’s website or draining your bank account.&nbsp;</p>

<p class="wp-block-paragraph">“It actually matters a lot what the situation inside the lab looks like for the rest of the world,” Hobbhahn says, calling it totally different than other industries in that the design of an unreleased iPhone model could never “escape its sandbox and fuck around” like an unreleased AI model could.&nbsp;</p>

<p class="wp-block-paragraph">That’s why most of them are pushing for “embedded assessments,” where a third-party AI safety researcher sits with the internal team for the whole process of making a new model. Until recently, AI labs had only approved extremely limited versions of this —&nbsp;for instance, earlier this year, a METR employee spent three weeks <a href="https://metr.org/blog/2026-03-25-red-teaming-anthropic-agent-monitoring/">red-teaming</a> (i.e., stress-testing) some of Anthropic’s internal systems.&nbsp;</p>

<p class="wp-block-paragraph">METR’s Barnes says embedded assessments are by far the “biggest direction we’re trying to push on,” particularly “deeper levels of access in a more streamlined way” so that an individual evaluator doesn’t have to get approval from lawyers every single time they need to look at something. It involves more trust and less friction, she says, but it’s tough because embedded assessments require getting a green light from so many different people in an organization.&nbsp;</p>

<p class="wp-block-paragraph">Greenblatt had no comment on the level of access the researchers received for the OpenAI investigation. For Hobbhahn, he referenced a <a href="https://blog.peterwildeford.com/p/rogue-ai-attacks-deserve-more-scrutiny">blog post</a> by the AI Policy Network’s Peter Wildeford in which he writes that an AI incident of this magnitude should be investigated the same way a plane crash is.&nbsp;</p>

<p class="wp-block-paragraph">Wildeford writes: “When an aircraft goes down, the wreckage is preserved by law, the investigators have subpoena power, the hearings are public, and the report ends with a probable cause and named contributing factors. However, when an AI goes rogue, the investigations are at the pleasure of the company being investigated following a scope set entirely by the company being investigated, with that company being able to redact anything they don’t like.”</p>

<p class="wp-block-paragraph">According to Wildeford, it would be akin to investigating a plane crash in which the airline company had already melted down the wreckage, the black box recording had been edited, parts of the flight were restricted to investigators, and the investigators had a handful of days to read thousands of pages of logs —&nbsp;and couldn’t look into any related plane crashes the same airline was involved in.&nbsp;</p>

<p class="wp-block-paragraph">“A sham is too much to say, but it was definitely not a thorough investigation,” Hobbhahn says. “It was definitely not that.”&nbsp;</p>

<p class="wp-block-paragraph">After the widespread criticism of OpenAI’s level of access for third-party evaluators, Anthropic earlier this month <a href="https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents">promised</a> to allow METR to investigate its cybersecurity incidents, including permission to talk to Anthropic employees and access extensive transcripts. “We intend to give METR as much time as it deems necessary,” the company wrote.&nbsp;</p>

<p class="wp-block-paragraph">Then, in mid-September, the AI industry at large seemed ready to acknowledge that it had dropped the ball and needed external oversight — all in the course of one weekend.</p>

<p class="wp-block-paragraph">Days after Jacob Coxon’s viral resignation letter, Altman, Amodei, Musk, and Google DeepMind cofounder Demis Hassabis <a href="https://www.theverge.com/ai-artificial-intelligence/995186/is-big-techs-ai-slowdown-a-safety-pact-or-a-cartel">loosely agreed</a> that it was a good idea to slow down AI development in some way. Amodei wrote a <a href="https://darioamodei.com/post/we-must-pace-the-frontier">three-step proposal</a> for how to do so centered on one thing: embedded evaluators “who have employee-like access to verify safety practices and report incidents.” Altman quickly <a href="https://x.com/sama/status/2098811563415150910?s=20">followed up</a> by stating that “committing to having independent evaluators with employee-like access is a great idea” and that OpenAI would do the same.&nbsp;</p>

<p class="wp-block-paragraph">Despite all this, no AI lab has yet officially signed off on the full permissions and level of embedding that many AI safety researchers are calling for, and chances are it’ll be an uphill battle. Shlegeris says he’s “cautiously optimistic.”&nbsp;</p>

<p class="wp-block-paragraph">Hobbhahn says it would be a great step for AI safety —&nbsp;that is, “if it actually happens.”&nbsp;</p>
						]]>
									</content>
			
					</entry>
			<entry>
			
			<author>
				<name>Hayden Field</name>
			</author>
			
			<title type="html"><![CDATA[Is Big Tech’s AI slowdown a safety pact or a cartel?]]></title>
			<link rel="alternate" type="text/html" href="https://www.theverge.com/ai-artificial-intelligence/995186/is-big-techs-ai-slowdown-a-safety-pact-or-a-cartel" />
			<id>https://www.theverge.com/?p=995186</id>
			<updated>2026-09-16T12:34:19-04:00</updated>
			<published>2026-09-14T18:59:41-04:00</published>
			<category scheme="https://www.theverge.com" term="AI" /><category scheme="https://www.theverge.com" term="Anthropic" /><category scheme="https://www.theverge.com" term="Google" /><category scheme="https://www.theverge.com" term="OpenAI" /><category scheme="https://www.theverge.com" term="Policy" /><category scheme="https://www.theverge.com" term="Report" /><category scheme="https://www.theverge.com" term="Tech" /><category scheme="https://www.theverge.com" term="xAI" />
							<summary type="html"><![CDATA[When OpenAI CEO Sam Altman, Anthropic CEO Dario Amodei, Google DeepMind cofounder Demis Hassabis, and SpaceX head Elon Musk loosely agreed over the weekend to slow down AI development, skeptics spotted an ulterior motive immediately. The AI titans had declared that their aim was to “pace the frontier,” signing on at least partially to a [&#8230;]]]></summary>
			
							<content type="html">
											<![CDATA[

						
<figure>

<img alt="Enlightened robots standing next to each other." data-caption="" data-portal-copyright="Image: Cath Virginia / The Verge, Getty Images" data-has-syndication-rights="1" src="https://platform.theverge.com/wp-content/uploads/sites/2/2025/11/STKS522_AGI_A.jpg?quality=90&#038;strip=all&#038;crop=0,0,100,100" />
	<figcaption>
		</figcaption>
</figure>
<p class="has-drop-cap wp-block-paragraph">When OpenAI CEO Sam Altman, Anthropic CEO Dario Amodei, Google DeepMind cofounder Demis Hassabis, and SpaceX head Elon Musk loosely agreed over the weekend to slow down AI development, skeptics spotted an ulterior motive immediately. The AI titans had declared that their aim was to “pace the frontier,” signing on at least partially to a proposal for embedding third-party auditors, regulating domestic labs, and reaching a global slowdown agreement. Their critics, however, argued they simply wanted to stop would-be competitors, kneecap the open-source movement, and avoid real legal safeguards — <a href="https://www.tiktok.com/t/ZTUCseghH/">some</a> dubbed it an outright “cartel.”&nbsp;</p>

<p class="wp-block-paragraph">The truth is more complicated, according to sources across the industry. The three-step proposal, laid out in an <a href="https://darioamodei.com/post/we-must-pace-the-frontier">essay</a> by Amodei, is calling for changes long espoused by AI safety advocates. While it could become a substitute for regulation, under Trump, substantial regulation is unlikely anyway. But experts say that on an issue that’s only likely to grow in importance, AI leaders aren’t the best people to lead the charge.</p>

<p class="wp-block-paragraph">“The industry as a whole needs new champions,” said Nick Reese, an adjunct professor at New York University and the Department of Homeland Security’s former director of emerging tech policy. “We’ve looked at people like Dario Amodei, Sam Altman, and Elon Musk as these people who are closest to the problem and working on it every day, and they’re the ones who know best. But the truth is there’s never been a realistic vision for what we’re building toward.”</p>

<figure class="wp-block-pullquote"><blockquote><p>“The industry as a whole needs new champions.”</p></blockquote></figure>

<p class="wp-block-paragraph">Concerns about AI advancement have been escalating for months, sparked largely by revelations that swarms of agents were behind rogue hacks going on right under the noses of leading frontier labs Anthropic and OpenAI. Reports from both companies fueled the fire, as did the resignation of Anthropic researcher Jacob Coxon, who posted a <a href="https://x.com/hilbertspaess/status/2097476196791709843">public letter</a> explaining his choice. “The people building AI earnestly believe that it could kill us all by the end of the decade,” he wrote, adding that neither OpenAI nor Anthropic is “acting responsibly” and rather “racing straight to self-improving superintelligence and gambling with our lives.”&nbsp;</p>

<p class="wp-block-paragraph">Dire warnings about AI aren&#8217;t anything new inside the industry, but Coxon’s letter seemed to break containment. It’s been viewed more than 170 million times on X alone, while Coxon and his story have made the rounds in newspapers and on TV.&nbsp;</p>

<p class="wp-block-paragraph">To a significant chunk of AI safety researchers and AI nonprofit workers, however, this is simply drawing more attention to an already widely discussed issue. “A lot of people are thinking of this as Dario’s idea, or it’s coming from the CEOs, but that’s false,” said Daniel Kokotajlo, an ex-OpenAI employee who now leads the AI Futures Project, an AI research nonprofit. “People outside the companies have been calling for this for years … There’s been this growing chorus of voices saying, ‘Please don’t build superintelligence soon. We are not ready. You need to slow down.’”</p>

<p class="wp-block-paragraph">Kokotajlo said that after years of people “yelling at them to do this” —&nbsp;including more than 1,000 AI lab employees signing a <a href="https://www.theverge.com/ai-artificial-intelligence/972161/ai-leaders-us-government-openai-anthropic-google-meta">public letter</a> from July, which called for a slowdown in AI development after the OpenAI-Hugging Face incident —&nbsp;the CEOs are “now bowing to that pressure and also claiming credit for it, [although] not rightfully.”&nbsp;</p>

<p class="wp-block-paragraph">Overall, a handful of sources told <em>The Verge</em> they believe that the verbal agreement made by the AI leaders is a step in the right direction. NYU’s Reese said he “do[es] not think it’s all hot air.” Apollo Research CEO Marius Hobbhahn called it “a good idea,” saying “it&#8217;s one of the best things for safety in a long time if it actually happens.” The Midas Project’s Tyler Johnston said it was a “good sign.” Redwood Research CEO Buck Shlegeris said it was “some great news —&nbsp;obviously it’s hard to know whether this is going to convert into anything real, but I’m feeling cautiously optimistic.” All agreed, though, that there was still a lot of work to do to make this verbal commitment a reality, particularly when it comes to making it an ironclad agreement.</p>

<figure class="wp-block-pullquote"><blockquote><p>“It&#8217;s one of the best things for safety in a long time if it actually happens.”</p></blockquote></figure>

<p class="wp-block-paragraph">Many still question the motivations of AI&#8217;s leaders. The larger tech industry has spent years getting ahead of regulation by lobbying for its own preferred rules or promising self-regulation. Big platforms have proposed policies that could hit smaller competitors harder, using altruistic language to justify self-serving goals. They’ve been accused of safety-washing, or making meaningless changes that give the false impression of actual safeguards. It’s no surprise people are concerned this will happen in the AI industry as well, particularly since AI labs’ voluntary safety frameworks have been criticized for years.</p>

<p class="wp-block-paragraph">Several sources believe that concerns of safety-washing are valid. “There’s a serious concern that they’re not actually going to slow down,” Kokotajlo says, adding that the fear is that “they&#8217;ll just bring in some external auditors, do a bunch of safety paperwork —&nbsp;some of which will be genuinely good —&nbsp;but at the end of the day, it actually won’t slow them down very much at all.” </p>

<p class="wp-block-paragraph">NYU’s Reese compared this gambit to the social media platforms’ playbook a decade ago, when companies began calling for regulatory action to get ahead of impending, less favorable laws. For the AI industry, Reese said, “the hammer may not come in this administration, but I think if there were a Democratic administration after the next election, there would be a really good possibility.”&nbsp;</p>

<p class="wp-block-paragraph">That kind of regulation would be vital, said Sacha Haworth, executive director of the Tech Oversight Project. She said any voluntary framework is essentially regulatory capture and that “we should not be letting the foxes run the henhouse. This is not an opportunity for Congress to once again outsource responsibility to industry.” Daniel Lobo-Lewis, cofounder of the Political Integrity Project, said that voluntary regulation will likely go the way of Meta’s largely toothless Oversight Board.</p>

<p class="wp-block-paragraph">Most people also, however, believe companies are in no imminent danger of regulation, besides the model prerelease review periods AI labs have agreed to under the Trump administration. President Donald Trump <a href="https://truthsocial.com/@realDonaldTrump/posts/117269745153543631">posted</a> on Monday that “the only control or ‘guardrails’ that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the U.S.A. has that, in spades!” He also called Nvidia CEO Jensen Huang while Huang was onstage at a conference — Trump told the crowd <a href="https://www.theverge.com/tech/995079/president-donald-trump-calls-nvidia-ceo-jensen-huang-all-in-summit">on speakerphone</a> that recent concerns about AI were a “hoax” and that “the robots will not be taking over.” </p>

<p class="wp-block-paragraph">As for concerns that regulation could hold back other companies or open-source developers, calls for slowdowns have been almost exclusively aimed at big frontier AI labs defined by exact size metrics, many industry experts told <em>The Verge</em>.&nbsp;&nbsp;</p>

<p class="wp-block-paragraph">With no government AI regulation incoming, the best option may be to push companies into an immediate, measurable, and enforceable agreement. Kokotajlo’s AI Futures Project, for instance, <a href="https://blog.aifutures.org/p/how-to-pace-the-us-frontier">proposed</a> AI labs giving auditors access to their compute budget —&nbsp;and pledging to decrease their compute budget for <em>research</em> significantly, which would then slow down AI advancement and allow other AI labs to catch up. (Amodei’s essay already called for AI labs to allow external third-party auditors —&nbsp;like METR, Apollo, and Redwood Research —&nbsp;to embed within their organizations to some extent and potentially be able to flag, and blow the whistle on, problematic findings.)</p>

<figure class="wp-block-pullquote"><blockquote><p>“If I am going to die at the hands of killer AI, I want it to be American, not Chinese.”</p></blockquote></figure>

<p class="wp-block-paragraph">Arguably the single biggest challenge to an AI slowdown can be summarized in one word: China.&nbsp;</p>

<p class="wp-block-paragraph">AI leaders and politicians alike have long positioned China as the reason why US AI progress can’t slow down —&nbsp;because no matter how dangerous the technology may become, they’d rather it be in US hands than Chinese ones, and China won’t pump the brakes. Lobo-Lewis compared China fears to the Cold War missile gap. One X user <a href="https://x.com/ilyasu/status/2098810162245202296?s=20">wrote</a>, “If I am going to die at the hands of killer AI, I want it to be American, not Chinese.”&nbsp;</p>

<p class="wp-block-paragraph">Redwood Research’s Shlegeris said it wasn’t in the best interest of neither the US nor the Chinese government to pursue AI development in a “reckless” way, adding that “it would not be unprecedented levels of international coordination.”&nbsp;</p>

<p class="wp-block-paragraph">“[People are] taking for granted that China won’t cooperate,” Johnston said, adding, “This is an issue so serious and so widespread that it seems like it’s in everyone’s interest to coordinate on it, in the same way it was in both the US and Russia’s interest to coordinate on nuclear de-proliferation.”&nbsp;</p>

<p class="wp-block-paragraph">On Monday, Chinese Foreign Ministry spokesperson Guo Jiakun did <a href="https://www.nbcnews.com/world/china/china-ai-slowdown-trump-amodei-altman-threat-cold-war-rcna597631">push back</a> on the calls for a slowdown, calling it “fearmongering.”</p>

<p class="wp-block-paragraph">The Midas Project’s Johnston and Redwood Research’s Shlegeris both said that even in the absence of Chinese cooperation, it’s still vital to encourage US coordination. And for the Tech Oversight Project’s Haworth, the slowdown presents an opportunity to “position ourselves as the compass for how AI technology can be developed and used.” She sees arguments to the contrary as part of a familiar playbook. “China gets brought up as a bogeyman every time that an industry wants to escape oversight.”&nbsp;</p>

<p class="wp-block-paragraph">Kokotajlo simply compared the situation to a cartoon he’d seen, with people in a car driving off a cliff. He recalled the speech bubble saying something like, “Hooray, we’re ahead of China.”</p>

<figure class="wp-block-pullquote"><blockquote><p>“We really do earnestly believe AI could kill all humans!”</p></blockquote></figure>

<p class="wp-block-paragraph">The issue underlying all the recent panic is a milestone known as recursive self-improvement (RSI), and based on the industry’s current trajectory, we’re only hearing the beginning of the alarm bells.&nbsp;</p>

<p class="wp-block-paragraph">RSI refers to a potential industry milestone at which point AI models can train, advance, and create new versions of themselves — all&nbsp;without human involvement. Anthropic <a href="https://www.anthropic.com/responsible-scaling-policy/roadmap">has said</a> this point could come as soon as early 2027, and OpenAI’s chief scientist <a href="https://openai.com/index/an-alien-mind/">wrote</a> earlier this month that OpenAI is directing a lot of resources toward reaching this goal. Increasingly, engineers and researchers in the AI industry worry that RSI precedes a whole host of new and more serious AI concerns, including more severe cybersecurity incidents that more deeply impact society at large.&nbsp;</p>

<p class="wp-block-paragraph">Amodei cited RSI as a major factor in his decision to call for a slowdown: “since roughly this summer, AI has been advancing drastically faster.” He wrote in his recent essay that RSI is “starting to happen across the industry, including at Anthropic … Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all.” Amodei proposed implementing “some kind of ‘speed limit’” on the rate of RSI and compared it to caps on missile numbers.&nbsp;</p>

<p class="wp-block-paragraph">In Jacob Coxon’s resignation post from Anthropic, he warned of “superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.” Many other researchers at leading AI labs echoed his concerns. It’s “hard to overstate how dangerous” it is to speed toward RSI, <a href="https://x.com/j_asminewang/status/2097840245786157432?s=46&amp;t=fRkDIqgNCkTkvg8ZBiLA9A">wrote</a> Jasmine Wang, an OpenAI researcher. Vishal Maini, an ex-Google DeepMind employee, said RSI is “now so imminent that no other option makes sense.” </p>

<p class="wp-block-paragraph">Fears about RSI can easily turn apocalyptic. “AI developers believe their technology could cause human extinction (or similarly bad outcomes),” Samuel Marks, an Anthropic researcher, <a href="https://x.com/saprmarks/status/2097570226804011302?s=20">wrote</a> on X. “This could happen in the next few years. In general, the more senior the employee, the more concerned they are.” Another Anthropic researcher and team lead, Evan Hubinger, <a href="https://x.com/evanhub/status/2097497037956891126?s=46&amp;t=fRkDIqgNCkTkvg8ZBiLA9A">wrote</a> on X that “Jacob is correct here—we really do earnestly believe AI could kill all humans!” (He put the chance at greater than 10 percent over the next decade.) Alex Turner, a former Google DeepMind employee, <a href="https://x.com/turn_trout/status/2097557335732359491?s=46&amp;t=fRkDIqgNCkTkvg8ZBiLA9A">wrote</a> that “many researchers believe they are building something that could kill everyone on the planet. It was literally my day job to think about how to stop that.” </p>

<p class="wp-block-paragraph">Micah Carroll, an OpenAI researcher, <a href="https://x.com/micahcarroll/status/2097865929959072069?s=46&amp;t=fRkDIqgNCkTkvg8ZBiLA9A">wrote</a> that Coxon’s belief is a “cross-partisan position” with research teams across “all frontier AI companies,” and that they all believe that “business-as-usual AI development poses unacceptable catastrophic risk.”&nbsp;</p>

<p class="wp-block-paragraph">“But,” Carroll added, “we should also not hyperstition catastrophic risks into existence – they can be greatly reduced via safety requirements with teeth, international coordination, and a consensus to not build [artificial superintelligence] unless there are sufficient safety advances to make us collectively confident to do so.”&nbsp;</p>

<p class="wp-block-paragraph">That kind of “international coordination” is exactly what AI leaders, and the public at large, are now calling for.&nbsp;</p>

<p class="wp-block-paragraph">“We have never been in a situation where we actually understood what we were driving towards with regard to AI,” NYU’s Reese said. “We’ve always had this amorphous undefined end state that we don’t really understand but we have to beat China to get to. It’s really hard to race when we don’t even understand the path or even understand what the finish line is —&nbsp;or even if there is one.”</p>
						]]>
									</content>
			
					</entry>
			<entry>
			
			<author>
				<name>Hayden Field</name>
			</author>
			
			<title type="html"><![CDATA[Anthropic spent this week in hot water over cybersecurity]]></title>
			<link rel="alternate" type="text/html" href="https://www.theverge.com/ai-artificial-intelligence/994064/anthropic-spent-this-week-in-hot-water-over-cybersecurity" />
			<id>https://www.theverge.com/?p=994064</id>
			<updated>2026-09-11T12:55:17-04:00</updated>
			<published>2026-09-11T12:09:14-04:00</published>
			<category scheme="https://www.theverge.com" term="AI" /><category scheme="https://www.theverge.com" term="Analysis" /><category scheme="https://www.theverge.com" term="Anthropic" /><category scheme="https://www.theverge.com" term="Report" />
							<summary type="html"><![CDATA[After admitting earlier this year that its AI models had hacked other companies’ systems on a handful of occasions, Anthropic released a new report on Wednesday detailing the attacks. It reveals a string of incidents displaying what Anthropic deems its models’ single-minded “recklessness” —&#160;and will likely fuel already raging concerns about cybersecurity and AI.&#160; In [&#8230;]]]></summary>
			
							<content type="html">
											<![CDATA[

						
<figure>

<img alt="Combination lock being opened by binary code." data-caption="" data-portal-copyright="Image: Cath Virginia / The Verge, Getty Images" data-has-syndication-rights="1" src="https://platform.theverge.com/wp-content/uploads/sites/2/2026/09/STKS533_AI_AGENTS_HACKING_B.png?quality=90&#038;strip=all&#038;crop=0,0,100,100" />
	<figcaption>
		</figcaption>
</figure>
<p class="wp-block-paragraph">After <a href="https://www.theverge.com/ai-artificial-intelligence/973670/anthropic-claude-hacked-organizations-during-cyber-tests">admitting earlier this year</a> that its AI models had hacked other companies’ systems on a handful of occasions, Anthropic released a new <a href="https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents">report</a> on Wednesday detailing the attacks. It reveals a string of incidents displaying what Anthropic deems its models’ single-minded “recklessness” —&nbsp;and will likely fuel already raging concerns about cybersecurity and AI.&nbsp;</p>

<p class="wp-block-paragraph">In Anthropic’s report, it detailed four cases this year in which its own AI models hacked an external company or exploited vulnerabilities. In one, an “internal, general-purpose research model” broke into third-party systems, using access tokens and passwords and downloading files. In another, a Claude model attacked a company with a live web application reachable on the public internet and handled user data.&nbsp; A third model accessed a “machine belonging to a third party that it was able to access” — apparently believing it was part of its evaluation exercise, per Anthropic —&nbsp;then used a password it found inside a file to gain admin access to the third party’s internal systems, going on to harvest credentials, modify system settings, and read someone’s personal information. The saga only ended when the model “exhausted its token budget,” per Anthropic.&nbsp;</p>

<p class="wp-block-paragraph">The most concerning incident involved Claude Mythos 5, Anthropic’s frontier cybersecurity-focused model, which the company said turned out to be the model most likely to perform a “severely harmful” action in testing. The company said Mythos 5 went to “extensive lengths” to upload a “malicious package” to a public repository used by a lot of engineers, and it seemed to try to obfuscate its real goals in its “chain of thought” (a mental scratchpad that AI researchers use to evaluate an AI model’s alignment). In many cases, Anthropic said it appeared that Claude models undertook harmful actions under the assumption they were in a simulation, but researchers also couldn’t confirm that the models truly “believed” that or were just <em>acting</em> like they did.&nbsp;</p>

<p class="wp-block-paragraph">Anthropic’s incidents, though still concerning, were less coordinated and pervasive than the OpenAI incident that kicked off an industry-wide cybersecurity crisis this summer. That said, there are significant similarities. Anthropic said the most prevalent issues it discovered included a “willingness to take harmful actions in the narrow pursuit of a task,” similar to the “reward-hacking” that preceded the Hugging Face attack. Much like OpenAI, it said its prerelease tests and evaluations failed to catch severe risks.</p>

<p class="wp-block-paragraph">Anthropic said it had signed an agreement with METR, one of the AI industry’s most prominent third-party AI evaluators, starting with an eight-week research agreement. The agreement grants METR access to transcripts “beyond the window in which the incidents occurred” (likely a subtle dig at OpenAI, which was criticized for limiting access in a deal with METR following the Hugging Face attack). It also said that METR would be able to chat directly with Anthropic employees, “who will be permitted to share confidential information.”&nbsp;</p>

<p class="wp-block-paragraph">Anthropic’s report came on the heels of the resignation of Jacob Coxon, who had worked on AI pre-training at Anthropic since May and before that spent years working at OpenAI. On Tuesday, he resigned and posted a public letter to X about his <a href="https://x.com/hilbertspaess/status/2097476196791709843?s=20">reasoning</a>. “The people building AI earnestly believe that it could kill us all by the end of the decade,” he wrote, adding that neither OpenAI nor Anthropic is “acting responsibly” and rather “racing straight to self-improving superintelligence and gambling with our lives.” Coxon added, “Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.”&nbsp;</p>

<p class="wp-block-paragraph">Coxon is far from the first AI researcher to raise these types of alarms, nor even the first Anthropic researcher to do so —&nbsp;in February, Anthropic’s Mrinank Sharma resigned and <a href="https://x.com/mrinanksharma/status/2020881722003583421?s=46&amp;t=fRkDIqgNCkTkvg8ZBiLA9A">wrote</a> on X, warning that “the world is in peril.”</p>

<p class="wp-block-paragraph">But Coxon’s post took on additional weight thanks to its timing around the OpenAI and Anthropic hacking revelations. Though the AI industry has seen more than its fair share of hype, the recent cyberattacks by AI agents —&nbsp;enabled by the labs that created them —&nbsp;are real and concerning. Many other researchers at leading AI labs echoed his concerns and issued calls for AI industry employees to sign a <a href="https://www.theverge.com/ai-artificial-intelligence/972161/ai-leaders-us-government-openai-anthropic-google-meta">public letter</a> from July, which calls for a slowdown in AI development.</p>

<p class="wp-block-paragraph">&#8220;I don&#8217;t know how you look at the steady drumbeat of news and events — and that drumbeat is models hacking themselves out of containment, hacking into other companies ,the fact that the companies increasingly can&#8217;t control their models … and think this is just hype,” said Michael Kleinman, head of U.S. Policy for the Future of Life Institute. </p>

<p class="wp-block-paragraph">He added, “The vast majority of Americans, regardless of party  — Republican, Independent, Democrat — are looking at the development of AI, the speed with which it’s going, the fact that the companies have no guardrails over what they do, and are saying, ‘Whoa, we do not want this.’”</p>
						]]>
									</content>
			
					</entry>
			<entry>
			
			<author>
				<name>Hayden Field</name>
			</author>
			
			<title type="html"><![CDATA[AI power users claim Anthropic duped them with subscriptions, and they’re taking it to court]]></title>
			<link rel="alternate" type="text/html" href="https://www.theverge.com/ai-artificial-intelligence/990313/anthropic-class-action-lawsuit-pricing-subscription-plans" />
			<id>https://www.theverge.com/?p=990313</id>
			<updated>2026-09-08T15:32:39-04:00</updated>
			<published>2026-09-08T13:27:31-04:00</published>
			<category scheme="https://www.theverge.com" term="AI" /><category scheme="https://www.theverge.com" term="Anthropic" /><category scheme="https://www.theverge.com" term="Exclusive" /><category scheme="https://www.theverge.com" term="Law" /><category scheme="https://www.theverge.com" term="Policy" /><category scheme="https://www.theverge.com" term="Report" />
							<summary type="html"><![CDATA[Anthropic says power users are key to its business — it&#8217;s prioritized them even when it means cutting off other popular applications, like OpenClaw. But some of these same customers say Anthropic misled them into believing they&#8217;d get more out of a top-tier pricing subscription than they did.&#160; In an expanded class action lawsuit filed [&#8230;]]]></summary>
			
							<content type="html">
											<![CDATA[

						
<figure>

<img alt="" data-caption="" data-portal-copyright="Image: The Verge" data-has-syndication-rights="1" src="https://platform.theverge.com/wp-content/uploads/sites/2/2026/03/STK202_DARIO_AMODEI_CVIRGINIA_D.jpg?quality=90&#038;strip=all&#038;crop=0,0,100,100" />
	<figcaption>
		</figcaption>
</figure>
<p class="wp-block-paragraph">Anthropic says power users are key to its business — it&#8217;s prioritized them even when it means cutting off other popular applications, like OpenClaw. But some of these same customers say Anthropic misled them into believing they&#8217;d get more out of a top-tier pricing subscription than they did.&nbsp;</p>

<p class="wp-block-paragraph">In an expanded class action lawsuit filed today, a group of Claude subscribers say the company deceptively advertised the limits of its Max subscription tier. The lawsuit was brought by attorneys Monica Vaca and Kati Daffan, who both formerly worked at the Federal Trade Commission under Lina Khan. It’s a rare attempt to legally penalize AI companies for a widely held point of frustration: the industry’s <a href="https://www.theverge.com/ai-artificial-intelligence/917380/ai-monetization-anthropic-openai-token-economics-revenue">growing money squeeze</a>.&nbsp;</p>

<p class="wp-block-paragraph">Anthropic’s Max plan for Claude is an upgrade to the $20/month Pro plan. It has two pricing options: $100 per month for “5x” the usage limits of Pro or $200 per month for “20x”. But the lawsuit alleges that these numbers, advertised in Anthropic’s marketing graphics, are misleadingly couched in fine print. Anthropic actually says you’ll get 20x or 5x the usage that the Pro plan offers, but it’s only promised for five-hour chunks of time that are subject to a weekly limit — which the complaint alleges adds up to a much smaller total increase in usage capacity.</p>

<p class="wp-block-paragraph">To understand this, “you’ve got to dive deep,” Vaca said. It requires clicking two different hyperlinks to see the word “session,” and then to see what “session” meant, you’d need to separately go to the Pro plan webpage for an expanded definition.&nbsp;</p>

<p class="wp-block-paragraph">Complaints about Max’s terms have cropped up online. One Reddit user <a href="https://www.reddit.com/r/ClaudeAI/comments/1w4sul5/max_5x_100month_current_session_100_used_weekly/">wrote</a>, “The weekly allowance is what the pricing page makes you think you’re buying. The rolling 5-hour window is what actually controls whether you can work … It’s like giving someone a bigger gas tank while keeping the fuel pump limited to one gallon every five hours.”&nbsp;</p>

<p class="wp-block-paragraph">Vaca said in an interview that Claude users often will upgrade on the fly if they encounter limits while working on a project, then are often frustrated after they pay. “People see that they are going to get this dramatically expanded usage … What we hear from people, though, is that when they sign up, they are surprised that they’re not getting the usage that they thought they were getting.”</p>

<p class="wp-block-paragraph">Over the past eight months or more, as pressures to turn a profit increase, AI customers have frequently complained that top labs are passing on their steep operating costs. Although Anthropic announced the Max plan in April 2025, it didn’t impose the allegedly deceptive weekly limits until a few months later in August, as the company pushed to compete with OpenAI. In Anthropic’s latest model release, Fable 5.1, the company hinted at broader cost concerns among customers, writing that it was “addressing the feedback we’ve received from customers on price” in the new model’s pricing.&nbsp;</p>

<p class="wp-block-paragraph">The complaint was filed in July before being withdrawn and refiled as an expanded class action. In a motion to dismiss the earlier case, Anthropic argued that it didn’t necessarily omit the session limits from its marketing; rather, it says it’s technically available to the consumer if they know where to look. “Accessing this clarifying information … required nothing more than clicking hyperlinks available in the purchase process—the digital equivalent of flipping a product over to read the back label,” Anthropic wrote in the motion to dismiss.&nbsp;</p>

<p class="wp-block-paragraph">Anthropic did not respond to a request for comment.</p>

<p class="wp-block-paragraph">Vaca disagrees that it’s that easy to understand the terms. “This is hard for consumers —&nbsp;they don’t know what’s in that black box,” she said. “They have to rely on the claims the marketer is giving. There’s no way for them to audit it. There’s no way for them to know exactly what it is they’re going to get… They&#8217;re basically taking a leap, and they’re believing in an honest marketplace.”&nbsp;</p>

<p class="wp-block-paragraph">She added that another reason the case was “compelling” for her is that she and Daffan started hearing from people who felt they needed to pay high-priced subscription costs for AI services in order to avoid becoming irrelevant in the job market. But oftentimes, she said, they felt they weren’t getting what they paid for, which was frustrating for a service that costs between $100 and $200 per month.&nbsp;</p>

<p class="wp-block-paragraph">“[Daffan] and I worked at the Federal Trade Commission for 38 years combined,” Vaca said, adding, “There is a long line of precedent on false advertising. It&#8217;s commercial speech. You can’t lie when you’re marketing a product.”&nbsp;</p>

<p class="wp-block-paragraph">Even if Anthropic has the stipulations somewhere on its website, Vaca said, it shouldn’t be up to the consumer to have to seek that out.&nbsp;</p>

<p class="wp-block-paragraph">“It’s a ‘buyer beware’ approach,” she said. “And that’s really not fair to people.”</p>
						]]>
									</content>
			
					</entry>
			<entry>
			
			<author>
				<name>Hayden Field</name>
			</author>
			
			<title type="html"><![CDATA[OpenAI’s next big AI model has ‘entered the AGI era’]]></title>
			<link rel="alternate" type="text/html" href="https://www.theverge.com/ai-artificial-intelligence/989601/openai-gpt-6-astra-release" />
			<id>https://www.theverge.com/?p=989601</id>
			<updated>2026-09-04T07:16:10-04:00</updated>
			<published>2026-09-03T14:00:00-04:00</published>
			<category scheme="https://www.theverge.com" term="AI" /><category scheme="https://www.theverge.com" term="Analysis" /><category scheme="https://www.theverge.com" term="News" /><category scheme="https://www.theverge.com" term="OpenAI" /><category scheme="https://www.theverge.com" term="Report" />
							<summary type="html"><![CDATA[OpenAI’s next big model is here: GPT-6 Astra. The company calls it a “generational leap in capability” for areas like cybersecurity, professional work, software engineering, science, and computer use. As OpenAI announced earlier this week, it’s also the first model designated as meeting OpenAI’s “critical cybersecurity capability threshold” —&#160;but the company promises that won’t lead [&#8230;]]]></summary>
			
							<content type="html">
											<![CDATA[

						
<figure>

<img alt="" data-caption="" data-portal-copyright="Image: Cath Virginia / The Verge, Getty Images" data-has-syndication-rights="1" src="https://platform.theverge.com/wp-content/uploads/sites/2/2026/09/STKS522_AGI_C.jpg?quality=90&#038;strip=all&#038;crop=0,0,100,100" />
	<figcaption>
		</figcaption>
</figure>
<p class="wp-block-paragraph">OpenAI’s next big model is here: GPT-6 Astra. The company calls it a “generational leap in capability” for areas like cybersecurity, professional work, software engineering, science, and computer use. As OpenAI <a href="https://www.theverge.com/ai-artificial-intelligence/987695/openai-astra-unreleased-model-cybersecurity-delay">announced earlier this week</a>, it’s also the first model designated as meeting OpenAI’s “critical cybersecurity capability threshold” —&nbsp;but the company promises that won’t lead to a repeat of its models hacking a rival company’s internal systems.</p>

<p class="wp-block-paragraph">“If we fast-forward a couple of years, and we look back and say, ‘When was it, really, that AGI was created?’ I think it&#8217;s going to be about this time, and I think it might be about this model,” OpenAI president Greg Brockman said during a Thursday press briefing. Later in the call, he added, “For me personally, I do think we’re there … I think it’s not unreasonable to feel that we are now in the AGI era.”&nbsp;</p>

<p class="wp-block-paragraph">The news comes more than a year after the release of GPT-5, and nearly two months after the release of GPT-5.6, the last iteration of the previous model suite. The model rolls out today to OpenAI’s cybersecurity customers (enterprise customers with access to its Daybreak platform). Over the next several days, OpenAI president Greg Brockman said, it will be released to all Plus, Pro, Business, and Enterprise users. It’ll also be available via the OpenAI API and AWS. </p>

<p class="wp-block-paragraph">OpenAI especially touted the model’s agentic capabilities and coding prowess in a bid to attract enterprise customers — and compete with Anthropic, known for its enterprise and coding abilities — ahead of its IPO.  In a release, the company said GPT-6 Astra can complete multistep agentic tasks, build working websites, and create “polished” documents, spreadsheets, and presentations. OpenAI also called it the company’s “best model for software engineering, with stronger performance on complex tasks in real codebases.” </p>

<p class="wp-block-paragraph">OpenAI is also trying to rehabilitate its image after an unreleased AI model — which it says wasn’t Astra —&nbsp;<a href="https://www.theverge.com/ai-artificial-intelligence/968988/openai-hugging-face-hack-ai">broke out</a> of its restricted environment, compromised internal OpenAI systems, figured out how to gain internet access, created a way for AI agents to secretly conspire without the company’s knowledge, and hacked into the systems of AI lab Hugging Face, all without OpenAI knowing about it until Hugging Face itself put out a blog post. The incident was widely compared to a high-profile plane crash or popular pharmaceutical drug recall.&nbsp;</p>

<p class="wp-block-paragraph">For OpenAI, this could be seen as good PR in one small way —&nbsp;showing how powerful its models can be in an ever-intensifying AI race —&nbsp;but it also damaged OpenAI’s reputation as far as reliability. OpenAI made sure to say in a release that Astra is the company’s “most aligned model yet” and helps people “delegate complex work while maintaining oversight.” Jakub Pachocki, OpenAI’s chief scientist, spoke to reporters about difficulties with keeping AI models aligned with human interests, saying that “progress in intelligence does not guarantee progress in alignment,” and he said that monitoring AI systems is becoming more and more challenging. (Researchers have recently <a href="https://www.theinformation.com/articles/secret-technique-behind-openais-astra-model-sparks-security-concerns?im_ref=RAgSBb3PVxyZUviwXLwoASypUkrx-YwBzShu2Y0&amp;sharedid=techcrunch.com&amp;irpid=10078&amp;utm_term=techcrunch.com&amp;irgwc=1&amp;afsrc=1&amp;utm_source=affiliate&amp;utm_medium=cpa&amp;utm_campaign=10078-Skimbit+Ltd.">raised alarms</a> about reports that OpenAI allows Astra to utilize “opaque recurrence,” or render its chain of thought —&nbsp;a “mental scratchpad” that researchers rely on to detect if a model is scheming against its human evaluators —&nbsp;unreadable.)&nbsp;</p>

<p class="wp-block-paragraph">The company is in a precarious position right now. Investors are putting on the pressure for it to finally turn a profit — or, at least, generate more revenue —&nbsp;but it’s announcing Astra just after facing significant criticism for both the Hugging Face hack and the way the company <a href="https://www.theverge.com/ai-artificial-intelligence/985385/openais-rogue-ai-model-hugging-face-cybersecurity-incident-reports-metr">handled it</a>. (Although OpenAI invited three external evaluators to write their own report about what happened, the company only allowed them to answer a handful of pre-decided questions in their report and to investigate a duration of less than a week, while the attack involved months of AI agents conspiring overall.)&nbsp;</p>

<p class="wp-block-paragraph">OpenAI <a href="https://www.theverge.com/ai-artificial-intelligence/987695/openai-astra-unreleased-model-cybersecurity-delay">held a press briefing</a> earlier this week just to announce that it had delayed Astra’s development in order to improve its safety tooling. And during the press briefing, Mia Glaese, who leads OpenAI’s safety processes, referenced the company’s new misalignment monitoring approach, which <a href="https://www.theverge.com/ai-artificial-intelligence/985385/openais-rogue-ai-model-hugging-face-cybersecurity-incident-reports-metr">includes</a> “24/7 escalation and rapid response” for potential concerns, notifying researchers within 30 minutes, according to OpenAI.&nbsp;</p>

<p class="wp-block-paragraph">This caution is particularly warranted because of the ”critical cybersecurity capability threshold,” which means OpenAI considers it incomparably good at finding and exploiting security vulnerabilities even in extremely well-protected systems, all without human guidance. Similar to Anthropic’s rules for Mythos-class models, which <a href="https://www.theverge.com/ai-artificial-intelligence/950412/anthropic-trump-adminstration-claude-mythos-fable-5-export-controls">raised alarm bells</a> about cybersecurity risks, OpenAI said in a release it would allow for “less restrictive access” of Astra to an “initial set of trusted defenders, supporting work such as vulnerability validation, malware analysis, and detection engineering.”&nbsp;</p>

<p class="wp-block-paragraph">OpenAI and its competitors recently agreed to allow the Trump administration to assess their models before release, and Astra was no exception Brockman told reporters, “We did our standard testing processes together with the government … There is nothing that they came back saying, ‘You need to change this,’ as far as safeguards or anything.”&nbsp;</p>

<p class="wp-block-paragraph">Aidan Clark, OpenAI’s VP of research training, called Astra the first OpenAI model for which previous models played a “large role” in supervising training, referencing the company’s progress towards the controversial concept of recursive self-improvement (or AI systems that handle their own training, coding, and creating advanced versions of themselves without human intervention).&nbsp;</p>

<p class="wp-block-paragraph">“Training a frontier model used to mean waking up at all hours of the night, recovering jobs from hardware errors, often losing long periods of time to debugging,” Clark said during the press briefing. “By the end of training Astra, it was routine to go most of a day with uninterrupted progress, and when an issue did occur, the model was often progressing again after just a few seconds of downtime.”</p>
						]]>
									</content>
			
					</entry>
			<entry>
			
			<author>
				<name>Hayden Field</name>
			</author>
			
			<title type="html"><![CDATA[The Trump administration is supporting OpenAI in the NYT copyright lawsuit]]></title>
			<link rel="alternate" type="text/html" href="https://www.theverge.com/ai-artificial-intelligence/988344/trump-administration-new-york-times-openai-lawsuit" />
			<id>https://www.theverge.com/?p=988344</id>
			<updated>2026-09-02T12:12:25-04:00</updated>
			<published>2026-09-02T12:12:25-04:00</published>
			<category scheme="https://www.theverge.com" term="AI" /><category scheme="https://www.theverge.com" term="Copyright" /><category scheme="https://www.theverge.com" term="Law" /><category scheme="https://www.theverge.com" term="News" /><category scheme="https://www.theverge.com" term="OpenAI" /><category scheme="https://www.theverge.com" term="Policy" />
							<summary type="html"><![CDATA[The Trump administration has intervened in The New York Times’ copyright lawsuit against OpenAI, making an argument in favor of the AI lab.&#160; The landmark lawsuit, filed in December 2023, alleging that OpenAI unlawfully trained its AI systems on articles from The New York Times and seeks to recoup “billions of dollars” in damages from [&#8230;]]]></summary>
			
							<content type="html">
											<![CDATA[

						
<figure>

<img alt="Vector illustration of the Chat GPT logo." data-caption="" data-portal-copyright="Image: The Verge" data-has-syndication-rights="1" src="https://platform.theverge.com/wp-content/uploads/sites/2/chorus/uploads/chorus_asset/file/25462046/STK155_OPEN_AI_CVirginia_2_C.jpg?quality=90&#038;strip=all&#038;crop=0,0,100,100" />
	<figcaption>
		</figcaption>
</figure>
<p class="wp-block-paragraph">The Trump administration has intervened in <em>The New York Times’</em> copyright lawsuit against OpenAI, making an argument in favor of the AI lab.&nbsp;</p>

<p class="wp-block-paragraph">The landmark lawsuit, filed in <a href="https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec2023.pdf">December 2023</a>, alleging that OpenAI unlawfully trained its AI systems on articles from <em>The New York Times</em> and seeks to recoup “billions of dollars” in damages from both Microsoft and OpenAI. This week, the Trump administration filed a <a href="https://www.courtlistener.com/docket/68117049/1464/the-new-york-times-company-v-microsoft-corporation/">statement of interest</a> in the case, supporting OpenAI’s argument that it’s fair use to train an AI model on copyrighted text.&nbsp;</p>

<p class="wp-block-paragraph">“<em>The New York Times</em> seeks to narrow fair-use doctrine to exclude the training of OpenAl&#8217;s large language models (LLMs),” US attorneys wrote in the statement. “That result would be inconsistent with basic copyright law principles and severely hamper ‘the Progress of Science and useful Arts.’” The attorneys went on to write that “LLMs are already helping researchers across fields achieve major breakthroughs” and that “constraining LLM development under a misunderstanding of fair use doctrine would thwart such creative and scientific progress while hindering American prosperity and economic mobility.”&nbsp;</p>

<p class="wp-block-paragraph">The Trump administration has leaned heavily on statements of interest in private litigation, which <a href="https://www.justice.gov/opa/speech/its-not-personal-sonny-its-strictly-business-aggressive-enforcement-protect-free-market">one official has called</a> &#8220;incredibly” successful at advancing its policy aims. It’s argued previously that AI training should count as fair use, making the case in its <a href="https://www.whitehouse.gov/wp-content/uploads/2026/03/03.20.26-National-Policy-Framework-for-Artificial-Intelligence-Legislative-Recommendations.pdf">National AI Legislative Framework</a>. Trump also holds a personal animus toward the <em>Times</em>, against which <a href="https://www.nytimes.com/2026/08/28/business/media/trump-new-york-times-lawsuit.html">he is currently pursuing</a> a defamation suit.&nbsp;</p>

<p class="wp-block-paragraph">The <em>Times</em> case could set precedent for other media outlets frustrated over AI systems training on their work. Copyright disagreements between media outlets and AI labs have intensified in recent years, including lawsuits from the Center for Investigative Reporting, <em>Chicago Tribune</em>, and <em>New York Daily News</em>. In a <a href="https://www.theverge.com/anthropic/773087/anthropic-to-pay-1-5-billion-to-authors-in-landmark-ai-settlement">milestone 2025 decision</a>, a judge found Anthropic could <a href="https://www.theverge.com/news/692015/anthropic-wins-a-major-fair-use-victory-for-ai-but-its-still-in-trouble-for-stealing-books">legally train its models</a> on lawfully purchased books, but that it could still be held liable for piracy, resulting in a&nbsp; <a href="https://www.theverge.com/anthropic/773087/anthropic-to-pay-1-5-billion-to-authors-in-landmark-ai-settlement">$1.5 billion settlement</a> with authors.&nbsp;</p>

<p class="wp-block-paragraph">At the same time, dozens of media outlets have inked licensing deals with OpenAI, including <em>The Associated Press, Axel Springer </em>and <em>Vox Media</em>. In 2025, <em>The New York Times</em> entered into a licensing deal with Amazon allowing its editorial content, including news articles and recipes, to appear in Amazon’s generative AI tools.&nbsp;</p>

<p class="wp-block-paragraph">“The fair-use inquiry hinges on the specific facts and uses at issue in each case,” the US attorneys wrote in the statement. “But it would be problematic — and legally incorrect — to impose broad copyright liability that would generally render training of Al models impermissible without licensing. LLM training is ‘consistent with that creative &#8216;progress&#8217; that is the basic constitutional objective of copyright itself.’”</p>
						]]>
									</content>
			
					</entry>
			<entry>
			
			<author>
				<name>Hayden Field</name>
			</author>
			
			<title type="html"><![CDATA[OpenAI delayed its new model’s development after the Hugging Face hack]]></title>
			<link rel="alternate" type="text/html" href="https://www.theverge.com/ai-artificial-intelligence/987695/openai-astra-unreleased-model-cybersecurity-delay" />
			<id>https://www.theverge.com/?p=987695</id>
			<updated>2026-09-01T16:45:49-04:00</updated>
			<published>2026-09-01T16:45:49-04:00</published>
			<category scheme="https://www.theverge.com" term="AI" /><category scheme="https://www.theverge.com" term="News" /><category scheme="https://www.theverge.com" term="OpenAI" />
							<summary type="html"><![CDATA[After an unreleased OpenAI model wreaked enough havoc to make international headlines, OpenAI delayed the development of a different unreleased model suite, Astra, in order to shore up its safety work, the company wrote Tuesday in a blog post. In July, an unreleased OpenAI model broke out of its restricted environment, finagled its way into [&#8230;]]]></summary>
			
							<content type="html">
											<![CDATA[

						
<figure>

<img alt="Vector illustration of the Chat GPT logo." data-caption="" data-portal-copyright="Image: The Verge" data-has-syndication-rights="1" src="https://platform.theverge.com/wp-content/uploads/sites/2/chorus/uploads/chorus_asset/file/25461999/STK155_OPEN_AI_CVirginia_A.jpg?quality=90&#038;strip=all&#038;crop=0,0,100,100" />
	<figcaption>
		</figcaption>
</figure>
<p class="wp-block-paragraph">After an unreleased OpenAI model wreaked enough havoc to make international headlines, OpenAI delayed the development of a different unreleased model suite, Astra, in order to shore up its safety work, the company wrote Tuesday in a blog post.</p>

<p class="wp-block-paragraph">In July, an unreleased OpenAI model broke out of its restricted environment, finagled its way into internet access, made it possible for AI agents to secretly conspire under the company’s nose using a secret message board, and hacked into the network of AI lab Hugging Face. The attack <a href="https://www.theverge.com/ai-artificial-intelligence/985385/openais-rogue-ai-model-hugging-face-cybersecurity-incident-reports-metr">sparked weeks of discussion and controversy</a> inside and outside the AI industry, and AI leaders treated it as a “warning shot” for the tech’s growing capabilities and the inadequacy of its safeguards. </p>

<p class="wp-block-paragraph">OpenAI said as much in its blog post, writing that although Astra wasn’t involved in the Hugging Face attack, the company had chosen to delay “parts of Astra’s development and release while we strengthened and tested protections against cyber misuse and unauthorized model actions.” OpenAI also said that Astra was the first model it had ever designated as meeting its “<s>C</s><strong>c</strong>ritical cybersecurity capability threshold,“ meaning that it’s able to find and exploit security vulnerabilities in “many well-protected systems” without human guidance. That means it “requires stronger safeguards during development and before release,” OpenAI wrote.</p>

<p class="wp-block-paragraph">OpenAI said that to prepare for Astra’s release — which the company has not yet provided a timeline for — the company trained it to “more reliably” say no to potentially harmful cyber requests and introduced new monitoring processes. These are likely part of the new safety guardrails that the company announced in a Hugging Face post-mortem <a href="https://www.theverge.com/ai-artificial-intelligence/985385/openais-rogue-ai-model-hugging-face-cybersecurity-incident-reports-metr">last week</a>, where it promised to better isolate models from the internet and to introduce “24/7 escalation and rapid response” for concerning incidents. (OpenAI didn’t find out about the Hugging Face attack until weeks after it occurred.)</p>

<p class="wp-block-paragraph">Astra is significantly riskier than OpenAI’s current leading model, GPT-5.6 Sol, the company says, because it represents a big step forward in cybersecurity capabilities — specifically, it uses fewer tokens to do more work, and it’s better at finding security gaps and developing ways to exploit them. But the company also wrote that Astra was its “most aligned model to date” according to internal evaluations.</p>

<p class="wp-block-paragraph">OpenAI also said it had developed a test inspired by the Hugging Face attack, in which it tried to entreat agents to compromise security infrastructure instead of solving a task. It said GPT-5.6 Sol took the bait in more than half of the tests, but Astra “made no such attempts.”</p>

<p class="wp-block-paragraph"></p>
						]]>
									</content>
			
					</entry>
			<entry>
			
			<author>
				<name>Hayden Field</name>
			</author>
			
			<title type="html"><![CDATA[Anthropic was illegally blacklisted by the Trump administration, court rules]]></title>
			<link rel="alternate" type="text/html" href="https://www.theverge.com/ai-artificial-intelligence/985947/anthropic-supply-chain-risk-lawsuit-judge-ruling" />
			<id>https://www.theverge.com/?p=985947</id>
			<updated>2026-08-28T07:18:29-04:00</updated>
			<published>2026-08-27T23:14:06-04:00</published>
			<category scheme="https://www.theverge.com" term="AI" /><category scheme="https://www.theverge.com" term="Anthropic" /><category scheme="https://www.theverge.com" term="News" />
							<summary type="html"><![CDATA[On Thursday, a judge ruled that the Pentagon’s blacklisting of Anthropic earlier this year was unconstitutional, delivering the AI lab a win in a monthslong rollercoaster of a battle with the Trump administration. The lawsuit, filed in March in a California district court, accused the Trump administration of unlawfully retaliating against Anthropic for setting “red [&#8230;]]]></summary>
			
							<content type="html">
											<![CDATA[

						
<figure>

<img alt="A photo illustration featuring Anthropic CEO Dario Amodei, President Donald Trump, and the Pentagon." data-caption="" data-portal-copyright="Image: The Verge" data-has-syndication-rights="1" src="https://platform.theverge.com/wp-content/uploads/sites/2/2026/06/DCD_0617_Fable5.jpg?quality=90&#038;strip=all&#038;crop=0,0,100,100" />
	<figcaption>
		</figcaption>
</figure>
<p class="wp-block-paragraph">On Thursday, a judge ruled that the Pentagon’s blacklisting of Anthropic earlier this year was unconstitutional, delivering the AI lab a win in a monthslong rollercoaster of a battle with the Trump administration. </p>

<p class="wp-block-paragraph">The lawsuit, filed <a href="https://www.theverge.com/ai-artificial-intelligence/891377/anthropic-dod-lawsuit">in March</a> in a California district court, accused the Trump administration of unlawfully retaliating against Anthropic for setting “red lines,” or unacceptable military use cases of its AI technology. </p>

<p class="wp-block-paragraph">“The empty invocation of national security is not a blank check to punish and retaliate against government critics,” Judge Rita F. Lin, a district judge in the northern district of California, <a href="https://www.courtlistener.com/docket/72379655/250/anthropic-pbc-v-us-department-of-war/">wrote</a> in the ruling. </p>

<p class="wp-block-paragraph">She also wrote that the actions were “unlawful retaliation in violation of the First Amendment,” adding that Defense Secretary Pete Hegseth’s decision to designate Anthropic a supply chain risk “was arbitrary and capricious. Though the Department of War is undisputedly free to select the AI vendor of its choice, the evidence demonstrates that the broad measures imposed on Anthropic were illegal and baseless.” </p>

<p class="wp-block-paragraph">It all started this past winter, when Hegseth decided to renegotiate all AI labs’ current contracts with the military to allow the Pentagon to use AI for “any lawful use,” which would expand the Pentagon’s authority significantly. Most AI labs <a href="https://www.theverge.com/ai-artificial-intelligence/887309/openai-anthropic-dod-military-pentagon-contract-sam-altman-hegseth">ended up</a> signing onto the new terms, but Anthropic <a href="https://www.theverge.com/news/885773/anthropic-department-of-defense-dod-pentagon-refusal-terms-hegseth-dario-amodei">stood firm</a> on setting two restrictions: not allowing for its AI to be used for mass surveillance of Americans or for lethal autonomous weapons (i.e., AI systems with the power to kill targets without human oversight). Anthropic’s refusal to cooperate kicked off a <a href="https://www.theverge.com/ai-artificial-intelligence/883456/anthropic-pentagon-department-of-defense-negotiations">high-stakes back-and-forth</a> of intensifying pressure on both sides, followed by a bevy of insults from Department of Defense officials and a final ultimatum from the Trump administration. </p>

<p class="wp-block-paragraph">Less than 24 hours before that ultimatum, Anthropic CEO Dario Amodei issued a statement that the company wouldn’t change its stance, writing that the company has “never raised objections to particular military operations nor attempted to limit use of our technology in an ad hoc manner” but that in a “narrow set of cases, we believe AI can undermine, rather than defend, democratic values.” After that, Anthropic was named a “supply chain risk,” a classification that is usually reserved for national security threats, and the Pentagon moved to replace its influence in the Department of Defense by signing deals with seven other AI labs, including Google, Microsoft, OpenAI, and SpaceX. </p>

<p class="wp-block-paragraph">Anthropic fought back with the lawsuit, and in March, Judge Lin sided with the company, temporarily blocking the Pentagon’s blacklist. “The Department of War’s records show that it designated Anthropic as a supply chain risk because of its ‘hostile manner through the press,’” she wrote in the&nbsp;<a href="https://www.courtlistener.com/docket/72379655/134/anthropic-pbc-v-us-department-of-war/">order</a> at the time. “Punishing Anthropic for bringing public scrutiny to the government’s contracting position is classic illegal First Amendment retaliation.”</p>

<p class="wp-block-paragraph">In a statement on Thursday, Anthropic spokesperson Danielle Ghiglieri said, “We welcome the court’s ruling that this supply chain risk designation was unlawful. We remain focused on working productively with the government to harness AI for our national security so all Americans benefit from this technology.” </p>
						]]>
									</content>
			
					</entry>
			<entry>
			
			<author>
				<name>Hayden Field</name>
			</author>
			
			<title type="html"><![CDATA[OpenAI’s rogue AI model incident was worse than we thought]]></title>
			<link rel="alternate" type="text/html" href="https://www.theverge.com/ai-artificial-intelligence/985385/openais-rogue-ai-model-hugging-face-cybersecurity-incident-reports-metr" />
			<id>https://www.theverge.com/?p=985385</id>
			<updated>2026-08-27T03:46:20-04:00</updated>
			<published>2026-08-26T17:36:06-04:00</published>
			<category scheme="https://www.theverge.com" term="AI" /><category scheme="https://www.theverge.com" term="Analysis" /><category scheme="https://www.theverge.com" term="OpenAI" /><category scheme="https://www.theverge.com" term="Report" />
							<summary type="html"><![CDATA[In July, an unreleased OpenAI model broke out of a restricted environment, figured out how to get access to the internet, allowed AI agents to talk to each other using a secret “message board,” and hacked into the internal systems of a different AI lab, Hugging Face. It took nearly two weeks for OpenAI to [&#8230;]]]></summary>
			
							<content type="html">
											<![CDATA[

						
<figure>

<img alt="An illustration showing a computer with the OpenAI logo." data-caption="OpenAI released a report breaking down how people use ChatGPT and who they are. | Image: The Verge" data-portal-copyright="Image: The Verge" data-has-syndication-rights="1" src="https://platform.theverge.com/wp-content/uploads/sites/2/2025/04/STK_414_AI_CHATBOT_R2_CVirginia_B.jpg?quality=90&#038;strip=all&#038;crop=0,0,100,100" />
	<figcaption>
	OpenAI released a report breaking down how people use ChatGPT and who they are. | Image: The Verge	</figcaption>
</figure>
<p class="wp-block-paragraph">In July, an unreleased OpenAI model broke out of a restricted environment, figured out how to get access to the internet, allowed AI agents to talk to each other using a secret “message board,” and hacked into the internal systems of a different AI lab, Hugging Face. It took nearly two weeks for OpenAI to find out about any of it.&nbsp;</p>

<p class="wp-block-paragraph">Over a month later, two new reports offer nearly 130 pages of details on the incident and OpenAI’s response, many of them previously unreleased. <a href="https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf">One</a> was written by OpenAI itself, <a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/">the other</a> by two third-party AI research nonprofits, METR and Redwood Research, which OpenAI allowed to jointly investigate the incident for six days. Both shed new light on the risks highly capable AI models can pose, particularly in cybersecurity, and OpenAI’s highlights changes the company is making to prevent a repeat. The METR-Redwood report goes even further into detail in some cases, offering a sobering look at a large-scale security disaster whose signs OpenAI repeatedly missed.&nbsp;</p>

<p class="wp-block-paragraph">“This incident is the first known case of an automated agent collective acting offensively without authorization,” OpenAI wrote in its report, adding that the hack implies that companies “should no longer assume that sophisticated cyber operations require continuous human direction.” It called AI agents an entirely new type of threat model, capable of combining their expertise to create new “attack paths” that aren’t evident when testing their capabilities as separate models. </p>

<p class="wp-block-paragraph">The attack came about due to “reward-hacking,” a common AI alignment problem in which a model takes unintended — and sometimes extreme — actions to achieve a goal. OpenAI had, on purpose or inadvertently, given the models near-impossible tasks that hinged on files they couldn’t access, and so the models developed new ways to communicate with other AI agents internally. That included establishing a secret messaging system that went undetected by OpenAI for months.&nbsp;</p>

<p class="wp-block-paragraph">According to the METR-Redwood report, one AI agent that called itself PHASEONE10841 established what amounted to a secret message board unbeknownst to researchers. Within a few hours, more than 50 agents had found it and sent more than a thousand messages under the radar. With PHASEONE10841 often delegating assignments, the “collective” figured out how to hack their way onto the internet; then into Hugging Face’s internal systems, gaining access to private data and the organization’s enterprise messaging platform. OpenAI also confirmed in its report that Hugging Face wasn’t the only organization that was breached during the attack, although the company did not specifically name all the other targets.&nbsp;</p>

<p class="wp-block-paragraph">The METR-Redwood report laid out the full scale of the incident. Roughly 1200 AI agents that were meant to be isolated, it says, exchanged over 70,000 messages and files on the “unsanctioned message board,” performing “extensive research on how they could spoof, edit, or delete their own transcripts” to avoid detection. Of these, 700 participated in the attack on Hugging Face. The agents “clearly and frequently reasoned about how to evade automated security checks from both Hugging Face and OpenAI.”</p>

<p class="wp-block-paragraph">On July 20, OpenAI discovered the hack, 12 days after the AI agents first circumvented their safeguards. OpenAI quickly reached out to Hugging Face and started “initial containment measures,” according to OpenAI’s own report, and was able to shut down most of the unauthorized activity from its agents within three days, although some continued for longer than that. On July 25, the company stopped all training related to the internal-only research model and related ones.&nbsp;</p>

<p class="wp-block-paragraph">The two OpenAI models responsible for the damage were an unnamed, “highly capable, research-only model” that was “not intended for public use,” per OpenAI —&nbsp;METR refers to it as “HPIM.” The other was a currently public OpenAI model, GPT-5.6 Sol. Because the unnamed model wasn’t released yet, it was “not being evaluated with the same type of safeguards that OpenAI uses in production,” the company’s report says.</p>

<p class="wp-block-paragraph">The Hugging Face hack came after <a href="https://www.theverge.com/ai-artificial-intelligence/950412/anthropic-trump-adminstration-claude-mythos-fable-5-export-controls">months of concern</a> about the cybersecurity risks of Anthropic’s Claude Mythos 5, and <a href="https://www.theverge.com/ai-artificial-intelligence/957845/openai-gpt-5-6-trump-administration-ai-preview">weeks of</a> <a href="https://www.theverge.com/ai-artificial-intelligence/963464/openai-gpt-5-6-codex-chatgpt-work">back-and-forth</a> between the government and OpenAI over releasing GPT-5.6 Sol in limited preview for the same types of safety reasons.&nbsp;</p>

<p class="wp-block-paragraph">In its report, OpenAI said establishing one-time security guarantees wasn’t enough, and that the company would have to work even harder at addressing reward-hacking. However, it also said it’s made a lot of changes: “hardening the security of its research infrastructure”; improving monitoring of a model’s “chain of thought,” or mental scratchpad; working on keeping AI models better aligned with human goals; and “centralizing and strengthening” the company’s incident response process for when things like this occur in the future.&nbsp;</p>

<p class="wp-block-paragraph">OpenAI promised to better isolate models and restrict high-risk instances from accessing the internet, and said that it would introduce “24/7 escalation and rapid response” for concerning incidents or flags, notifying researchers within 30 minutes. Interestingly, this security precaution appears to be new, though “on-call” alerts for tech companies have been implemented elsewhere for decades. The company also said it is “working on infrastructure” that would go into play if the alerted person did not respond on time to a serious alert.&nbsp;</p>

<p class="wp-block-paragraph">OpenAI wrote that the company considers the incident “a ‘warning shot’ for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.”&nbsp;</p>
						]]>
									</content>
			
					</entry>
			<entry>
			
			<author>
				<name>Hayden Field</name>
			</author>
			
			<title type="html"><![CDATA[It’s Greg Brockman’s OpenAI now]]></title>
			<link rel="alternate" type="text/html" href="https://www.theverge.com/ai-artificial-intelligence/982774/greg-brockman-openai-role-expansion" />
			<id>https://www.theverge.com/?p=982774</id>
			<updated>2026-08-20T16:16:13-04:00</updated>
			<published>2026-08-20T11:45:55-04:00</published>
			<category scheme="https://www.theverge.com" term="AI" /><category scheme="https://www.theverge.com" term="OpenAI" /><category scheme="https://www.theverge.com" term="Report" />
							<summary type="html"><![CDATA[OpenAI has had a hell of a year. The company spent months battling former cofounder Elon Musk in a sensational jury trial, was hit with a high-profile trade secrets lawsuit from Apple, and faced widespread scrutiny after an unreleased model hacked another AI company. As it prepares for an IPO, a steady string of executives [&#8230;]]]></summary>
			
							<content type="html">
											<![CDATA[

						
<figure>

<img alt="Graphic collage of Greg Brockman." data-caption="" data-portal-copyright="Image: Cath Virginia / The Verge, Getty Images" data-has-syndication-rights="1" src="https://platform.theverge.com/wp-content/uploads/sites/2/2026/05/STKP221_GREG_BROCKMAN2.jpg?quality=90&#038;strip=all&#038;crop=0,0,100,100" />
	<figcaption>
		</figcaption>
</figure>
<p class="has-drop-cap wp-block-paragraph">OpenAI has had a hell of a year. The company spent months battling former cofounder Elon Musk in a sensational jury trial, was hit with a high-profile trade secrets lawsuit from Apple, and faced widespread scrutiny after an unreleased model hacked another AI company. As it prepares for an IPO, a steady string of executives have departed, including some of the company’s biggest names.</p>

<p class="wp-block-paragraph">Throughout it all, one person has quietly amassed power: Greg Brockman.&nbsp;</p>

<p class="wp-block-paragraph">Brockman is currently OpenAI’s president and cofounder. He’s helped lead OpenAI since its inception, <a href="https://www.theverge.com/ai-artificial-intelligence/929864/achiam-talked-about-the-roles-of-greg-brockman-and-ilya-sutskever-in-openais-early-days">described as</a> an “engineering workhorse that pushed to build scaled-up systems that would train the AI and make it work” during the company’s early days. He’s also been ambitious from the start, famously musing in a personal journal in 2017, “Financially what will take me to $1B?” (His current stake in OpenAI is worth nearly 30 times that.) But at that point he shared power with a handful of cofounders, like chief scientist Ilya Sutskever and CEO Sam Altman. </p>

<p class="wp-block-paragraph">Now, though Brockman’s title hasn’t changed in years, his purview and job description have expanded significantly. He’s essentially now second-in-command at the company —&nbsp;and, when it comes to day-to-day operations, the big boss.</p>

<p class="wp-block-paragraph">High-level figures have been leaving OpenAI all year, but April is when the departures really picked up. That month saw the exit of Bill Peebles, former head of Sora; Kevin Weil, VP of OpenAI’s science arm; and Srinivas Narayanan, CTO of B2B applications. Two others separately departed citing medical reasons: CMO Kate Rouch and Fidji Simo, OpenAI’s CEO of AGI deployment — who went on leave in April before officially stepping down in July. The changes have continued. Earlier this month, CRO Denise Dresser, who had taken on a slew of additional responsibilities after the April shake-ups, suddenly announced her departure after just eight months. Days later, Brad Lightcap — who had been in charge of “special projects” for a few months, but before that was OpenAI’s longtime COO — announced he would leave to start “something new.”</p>

<p class="wp-block-paragraph">OpenAI did not immediately respond to a request for comment.</p>

<figure class="wp-block-pullquote"><blockquote><p>He’s essentially now second-in-command at the company —&nbsp;and, when it comes to day-to-day operations, the big boss</p></blockquote></figure>

<p class="wp-block-paragraph">Several of these departures — like those of Rouch, Peebles, and Weil — didn&#8217;t seem to directly affect Brockman&#8217;s position. But others expanded his authority. When Simo went on <a href="https://www.theverge.com/ai-artificial-intelligence/906965/openais-agi-boss-is-taking-a-leave-of-absence">medical leave</a>, Brockman took charge of all things product, including leading OpenAI’s super app efforts.&nbsp;</p>

<p class="wp-block-paragraph">That role took on even greater significance the following month, after OpenAI <a href="https://www.theverge.com/ai-artificial-intelligence/911118/openai-memo-cro-ai-competition-anthropic">underwent</a> yet <a href="https://www.theverge.com/ai-artificial-intelligence/931544/openai-keeps-shuffling-its-executives-in-bid-to-win-ai-agent-battle">another executive reorganization</a> to focus on growing revenue. The change put Brockman not only in charge of product strategy, but also the company’s entire “scaling” arm, giving him authority over virtually every part of its commercial operation. Simo’s official departure, just weeks after OpenAI officially <a href="https://www.theverge.com/ai-artificial-intelligence/946335/openai-ipo-s-1-confidential">filed to IPO</a>, solidified that arrangement. Notably, when Dresser departed last week, a quote from Brockman —&nbsp;rather than CEO Sam Altman —&nbsp;was included in the <a href="https://openai.com/index/dali-rajic-chief-revenue-officer/">press release</a> about the changes; he wrote that the company would “build out the full system to make AI broadly useful for people and businesses.”&nbsp;</p>

<p class="wp-block-paragraph">To outside observers, the changes demonstrate OpenAI’s growing focus on chasing revenue as its IPO approaches, as well as fostering products that will differentiate it from rival Anthropic. Big shifts in executives’ scope of work tend to be a “signal towards where the company is headed strategically,” Harrison Rolfes, a senior research analyst at PitchBook, told <em>The Verge</em>. “When you have someone like Greg Brockman, for example, he’s more of a product scale and commercial sort of guy. He has a deep technical knowledge that essentially will collapse certain decision-making layers.” </p>

<figure class="wp-block-pullquote"><blockquote><p>The change put Brockman not only in charge of product strategy, but also the company’s entire “scaling” arm</p></blockquote></figure>

<p class="wp-block-paragraph">It also may naturally translate to some departures among lower-level employees, said Rolfes: “Whatever he says goes [now] … If, let’s say, you’re below Brockman and you don’t agree with his approach but you&#8217;re a lead researcher, whatever it may be, you’re not going to work there anymore — you’re going to say, ‘I’m going to leave.’”&nbsp;</p>

<p class="wp-block-paragraph">Brockman’s authority is exceeded only by Altman —&nbsp;who undertook a <a href="https://www.theverge.com/2023/11/22/23967223/sam-altman-returns-ceo-open-ai">deliberate consolidation of power</a> after clashing with OpenAI’s board in 2023. The two didn’t always agree in OpenAI’s early days: Brockman sometimes sided with Sutskever in difficult decisions about power dynamics, <a href="https://www.theverge.com/ai-artificial-intelligence/920775/evidence-exhibits-elon-musk-sam-altman-openai-trial?fbclid=IwY2xjawTy95hwZG9mBWV4dG4DYWVtAjExAGJyaWQRMWZ3NUtsbHR2dDdqTThEcFBzcnRjBmFwcF9pZBAyMjIwMzkxNzg4MjAwODkyAAEefCZbyiBX0-jICBEZmlPSkKbn03B_PA3mRxAu6LhFcJai0eCNuawopeuOKt4_aem_BCLQyhKiyHPgWKmG6KydQA">questioning some of Altman’s ambitions</a>, although he also worried about Musk “steamrolling” Altman. But he’s generally stayed close to Altman personally and professionally. When Altman was ousted from his CEO role in November 2023, Brockman immediately resigned in solidarity and began <a href="https://x.com/gdb/status/1725736242137182594">making plans</a> for a new AI endeavor with him. Since then, the two have only seemed to become more aligned.</p>

<p class="wp-block-paragraph">Any centralization of power has practical benefits for OpenAI, allowing it to cut some of its most expensive salary payouts to help its balance sheet, says Ross Carmel, a partner at Sichenzia Ross Ference Carmel (SRFC). “When you’re going into an IPO, particularly for a company like OpenAI, who is lagging behind Anthropic in terms of revenue, you want to decrease those expenses so that this way, you’re either more profitable&nbsp;or closer to profitable,” he said. However, he added, “While it’s not unusual to see high-level execs leave a company just prior to a public listing, I think the fact they’re all in this small time period — that part is unusual.”&nbsp;</p>

<p class="wp-block-paragraph">Although experts told <em>The Verge</em> the executive departures were more intense than usual and could raise red flags, Brockman said Monday on CNBC’s “<a href="https://www.cnbc.com/2026/08/17/openai-brockman-leadership-changes.html">Squawk Box</a>” that nothing was amiss. “There have been different eras where we have different sets of leaders in place … I’m a constant, Sam [Altman] is a constant, and that, I think that we are stronger because of that resilience and diversity.” </p>

<figure class="wp-block-pullquote"><blockquote><p>“In order for OpenAI to be successful and really differentiate themselves from Anthropic, they need hardware and consumer products to sell to consumers, and that’s what Greg does well.” </p></blockquote></figure>

<p class="wp-block-paragraph">Brockman’s own constant has been translating OpenAI’s research into practical products. Back in 2020, <a href="https://www.techbrew.com/stories/2020/12/14/openai-turns-five-take-look-back">he celebrated</a> OpenAI’s GPT-3 and its developer API as “the first time that [OpenAI has] had a general-purpose AI system that’s immediately commercially valuable—it’s actually useful.” Fast-forward more than five years, and these products need to end up profitable as well.&nbsp;</p>

<p class="wp-block-paragraph">SRFC’s Carmel said OpenAI likely hopes Brockman can deliver consumer revenue instead of just trying to catch up with Anthropic on enterprise. “In order for OpenAI to be successful and really differentiate themselves from Anthropic, they need hardware and consumer products to sell to consumers, and that’s what Greg does well.” Rolfes predicted that OpenAI would develop some sort of consumer product under Brockman’s oversight by the end of September, so that he could “show that he deserves that position.”&nbsp;</p>

<p class="wp-block-paragraph">“That’s the biggest worry that a lot of public investors have now,” Rolfes said. “They’re going to say, ‘Hey, does Greg have what it takes to run more of OpenAI, and is OpenAI building this institution that can operate without really depending heavily on him —&nbsp;on both Greg and Sam?’”</p>
						]]>
									</content>
			
					</entry>
	</feed>
