Advertisement

Over the past five years, Anthropic has repeatedly sounded the alarm about the catastrophic risks posed by cutting-edge artificial intelligence: from enabling mass devastation and eroding social stability to a long list of other severe harms to people and communities. Yet at the same time, the company has emerged as one of the most powerful drivers advancing AI capabilities today. Anthropic now ranks among the world’s top developers and distributors of state-of-the-art AI models, and counts clients including the U.S. military among its partners. Most recently, the company hit a valuation of nearly $1 trillion.

On the surface, Anthropic’s dire public warnings and its aggressive business and technical growth look like a fundamental contradiction. But for many people inside the company, this tension is not a contradiction at all. To understand their perspective, you first have to grasp the two core beliefs that shape how Anthropic operates. The first is that artificial intelligence is the most transformative technology humanity has ever created, and its widespread arrival is unavoidable. The only open question, the company argues, is whether this shift will end in global catastrophe or deliver unprecedented prosperity.

The second core belief, according to multiple former Anthropic employees who spoke to WIRED on condition of anonymity, is that the world will be safer and better off if Anthropic maintains its position at the cutting edge of the global AI race. Two of these sources noted that company leaders and staff regularly frame themselves as the “good guys” of the AI industry: responsible stewards of the technology, unlike their competitors. For Anthropic, accumulating power—whether that means capital, computing infrastructure, top research talent, or political influence—is not a goal in and of itself. Instead, it is a necessary cost to fulfill the company’s stated mission: “to ensure the world safely makes the transition through transformative AI.”

Helen Toner, executive director of Georgetown University’s Center for Security and Emerging Technology and a former OpenAI board member, uses a vivid analogy to explain Anthropic’s worldview. She compares the development of powerful AI to a forest that holds both magical, life-changing treasures and deadly, dangerous monsters. All the nearby villagers are rushing into the woods, drawn by the promise of treasure. In this framing, Anthropic wants to push deeper into the forest than any other group, while pouring massive resources into taming the monsters—meaning capturing AI’s enormous benefits while keeping its catastrophic risks in check.

“What makes Anthropic distinct is their core stance: people are going into the forest no matter what, so we have to get there first,” Toner explained to me. “This is their explicit strategy: build cutting-edge AI so they can be a serious voice at the table, able to speak to what advanced AI systems actually look like, what risks they carry, and push for thoughtful, reasonable safeguards. They’re completely open about this strategy—it’s just unusual enough that many people struggle to wrap their heads around it.”

Anthropic CEO Dario Amodei laid out this approach clearly in a conversation with his co-founders posted to the company’s careers page. “You have to find a way to actually be competitive, to actually lead the industry in some cases, and yet manage to do things safely,” he said. “If you can do that, the gravitational pull you exert is so great.”

Anthropic was launched in 2021 by a group of former OpenAI employees who left the company after losing confidence in OpenAI leadership—particularly CEO Sam Altman—to develop transformative AI safely. That original skepticism still defines the company’s culture today. Two of the former employees I spoke with noted that in internal discussions, Anthropic executives often hold up Altman and OpenAI, and to a lesser degree Meta and Elon Musk’s xAI, as cautionary tales that reinforce Anthropic’s own commitment to responsible development.

In many ways, Anthropic looks just like any other Silicon Valley startup. Countless new tech companies frame themselves as plucky Davids taking on outdated, entrenched Goliaths in the industries they aim to disrupt. Google, Meta (then Facebook), and Apple all launched with idealistic founding principles that later became blurred or abandoned entirely as the companies grew richer, larger, and more powerful.

But former employees say Anthropic stands out for how deeply its team believes in its mission, and how clearly it communicates to staff that technological and commercial power are only tools to reach that end. One former employee explained that in job interviews, Anthropic emphasizes to candidates that it is not a typical company driven purely by market pressures. It is structured as a public benefit corporation, which lets it prioritize the “long-term benefit of humanity” over shareholder profits. Even so, the company frames financial success and building the most powerful AI models as critical to that larger mission—a non-negotiable prerequisite to leading the industry on AI safety.

“None of us wanted to found a company, we just felt like it was our duty,” Sam McCandlish, Anthropic co-founder and chief architect, said in that same careers page conversation. “We have to do this thing. This is the way we’re gonna make things go better with AI.”

Anthropic declined to comment for this story.

The Good Guy Problem

Anthropic describes itself on its website as a “high-trust, low-ego organization” with little internal political infighting, a description former employees say is largely accurate. They note that compared to leaders at other major AI labs, Anthropic staff generally trust Amodei to be honest with them about the company’s technical progress, its engagements with government officials, and its views on global geopolitics.

But ideological diversity is a key pillar of accountability, and it is often missing here. Shazeda Ahmed, a postdoctoral scholar at UCLA who has researched the ideological roots of the AI safety movement, notes that organizations like Anthropic often struggle with a lack of ideological pluralism. Her work finds that the AI safety movement, which has deep roots in communities like effective altruism, suffers from groupthink and a default preference for industry self-governance over external oversight.

“When you surround yourself with people who all share the same core beliefs, your ideas never get challenged,” Ahmed says. “When your main measure of success is how well you live up to your own ideological commitments, you don’t often stop to ask: what if this goes wrong because we aren’t actually the right people to hold this much power? They don’t always examine their own blind spots.”

One former employee I spoke with says there is a robust culture of internal debate at Anthropic, and staff critiques often prompt detailed, lengthy responses from company leadership. But another former employee paints a darker picture: more open criticism stays locked in private group chats, and rarely turns into direct pushback against Amodei’s decisions. They described the company’s regular all-hands meetings with Amodei, which staff call “Dario Vision Quests,” as similar to “going to a sermon to hear a priest.”

One of the largest internal controversies in Anthropic’s history erupted in fall 2024, when the company became the first AI lab to partner with Palantir to deliver AI services to U.S. intelligence and defense agencies. Several former employees I spoke with said questions about the deal were raised internally, but those debates never led to changes to the company’s policy.

At the time, Anthropic employee Evan Hubinger wrote in a post on the online forum LessWrong that the company was “extremely forthright” with staff about the Palantir partnership. While he acknowledged there were some lines that should not be crossed without careful deliberation, he called the deal overall a positive step. “If you take catastrophic risks from AI seriously, the U.S. government is an extremely important actor to engage with, and trying to just block the U.S. government out of using AI is not a viable strategy,” he wrote.

Less than two years later, multiple reports indicate the Pentagon has begun using Anthropic’s Claude AI for tasks including identifying strike targets during the Israel-Iran conflict. When asked in a recent Bloomberg interview whether Anthropic’s models were used in an attack on an Iranian elementary school that killed more than 120 people, Amodei said he did not know, but added that the use would have been approved by the company as long as a human made the final decision to strike. The moment is a stark example of how Anthropic’s vision of responsible AI does not always align with the broader public’s understanding of that term.

Conflicts over Anthropic’s rules for how Claude can and cannot be used have played out in other public contexts too. Earlier this month, Anthropic launched its latest cutting-edge model, Claude Fable 5, with a uniquely untransparent safeguard built in: if researchers tried to use the model for independent frontier AI development, a move that would violate Anthropic’s terms of service, the company would secretly sabotage their work. The move drew immediate backlash from researchers across the AI industry, and Anthropic backed away from the policy within a few days, saying it would make the safeguard visible to users. In a statement at the time, Anthropic said it had failed to strike the right balance, and that its goal was to block misuse by U.S. foreign adversaries.

Power Struggles

Amodei has publicly acknowledged the risk of concentrating too much power over AI development in the hands of just a small number of companies, including his own. “It is somewhat awkward to say this as the CEO of an AI company, but I think the next tier of risk is actually AI companies themselves,” he wrote in an essay earlier this year. But the solutions he proposes—that AI companies “be carefully watched” and potentially make public pledges to “not take certain actions”—do little to fundamentally redistribute that concentrated power.

In longer passages of the essay, Amodei reflects on the enormous scale of his own influence and the responsibility that comes with it. But he largely avoids framing this concentration of power as a personal issue, instead positioning it as a challenge facing all of humanity: “Humanity is about to be handed almost unimaginable power, and it is deeply unclear whether our social, political, and technological systems possess the maturity to wield it,” he writes. He goes on to argue that it is the responsibility of “those closest to the technology to simply tell the truth about the situation humanity is in, which I have always tried to do.”

A common critique of Anthropic’s stance is that the company claims it understands the “truth about the situation humanity is in” better than anyone else. It frames AI as both extraordinarily powerful and ultimately governable, as long as the right people lead its development. But the reality is that no one knows exactly how AI will reshape the world—and some people just get far more say over that future than others.

This piece is an edition of Maxwell Zeff’s Model Behavior newsletter. Read previous newsletters here.

Related Article