Advertisement

The Trillion-Dollar AI Race: How Google Rushed to Catch ChatGPT and Remade Itself

One hundred days. That was the hard deadline Google handed executive Sissie Hsiao in late 2022: build a direct competitor to OpenAI’s breakout ChatGPT in that short window. When Hsiao took charge of the project that December, she had already spent 16 years at Google, led thousands of employees, and weathered her fair share of corporate upheaval. But nothing compared to the existential “code red” alarm that blared after OpenAI, a little-known AI research lab, dropped its public AI experiment on the world.

ChatGPT regularly invented false facts and botched basic arithmetic, but that didn’t stop more than a million people from flocking to it within weeks. Worse for Google: many industry observers started framing it as a potential replacement for Google Search, the company’s core cash-generating business. Google already had a capable large language model of its own, on par with OpenAI’s GPT, but it had kept the tool locked behind closed doors. The public could only access LaMDA by exclusive invitation, and in one public demo, it was only programmed to talk about dogs.

Wall Street was growing increasingly anxious. More than six years before ChatGPT’s launch, CEO Sundar Pichai had pledged to build an “AI-first world,” where smart personal assistants would replace the very idea of standalone devices. Not long after that promise, eight Google researchers invented transformer architecture—the core technology that makes ChatGPT work. But by late 2022, Google had little to show for that breakthrough. Ad sales disappointed, most of the original transformer inventors had left the company, Google Assistant (the product Hsiao led) was rarely used for anything more complex than setting a timer or playing a song, and a half-baked Gen Z-focused chatbot that dispensed cooking tips and history lessons went nowhere. By the end of 2022, Alphabet, Google’s parent company, saw its stock price drop 39% from the year prior.

As 2023 kicked off, Google’s executive team demanded constant updates for the board. Co-founder and controlling shareholder Sergey Brin began dropping in regularly to review AI strategy. Word spread through the company that the $1 trillion giant needed to start moving at startup speed, which meant taking bigger risks. Gone was the old dynamic, as a former senior product director told WIRED, where thousands of employees could veto a product but no one had the authority to approve one. As Hsiao’s team launched into their 100-day sprint, she laid out a seemingly contradictory, idiosyncratic rule: “Quality over speed, but fast.”

While Hsiao rushed to build an immediate ChatGPT rival, another top executive, James Manyika, was orchestrating a longer-term strategic shift as part of top leadership discussions. An Oxford-trained roboticist who became a go-to McKinsey advisor for Silicon Valley leaders, Manyika joined Google as senior vice president of technology and society in early 2022. Months before ChatGPT launched, he told his long-time friend Pichai that Google’s cautious approach to AI was holding the company back. Google had two world-class AI research teams working separately, wasting valuable computing power on competing goals: DeepMind, the London-based lab led by Demis Hassabis, and Google Brain, the Mountain View unit overseen by Jeff Dean. Manyika argued the two teams needed to combine forces.

After ChatGPT’s launch, that plan moved forward fast. Dean, Hassabis, and Manyika presented a proposal to the board for the combined team to build the most powerful language model ever created. Hassabis pushed to name the project Titan, but the board didn’t love the name. Dean’s suggestion, Gemini, won out. (One billionaire investor was so excited by the plan that he snapped a photo of the three executives to mark the moment.)

Since that merger, Manyika says the company has made “a lot of what I call ‘bold and responsible’ choices” across every division. “I don't know if we've always got them right,” he added. And indeed, the race to reclaim Google’s position as AI’s global leader threw the company into one crisis after another. At one low point, employees gathered in hallways openly worrying Google would go the way of Yahoo, once the undisputed king of search. “It's been like sprinting a marathon,” Hsiao said of the effort. Now, more than two years later, Alphabet’s shares have hit an all-time high, and investors are bullish on Google’s AI progress.

WIRED interviewed more than 50 current and former Google employees—including engineers, marketers, legal and safety leaders, and a dozen top executives—to reconstruct the most chaotic, culture-shifting period in the company’s history. Most sources requested anonymity to speak candidly about Google’s transformation, for better and worse. This is the first time multiple top executives have shared detailed on-the-record recollections of those turbulent two years, and the hard trade-offs the company made along the way.

To build the initial ChatGPT rival, codenamed Bard, former employees say Hsiao pulled roughly 100 staff from teams across Google. Managers had no say in the decision: Bard took priority over every other project at the company. Hsiao says she prioritized big-picture thinkers who had both the technical skill and emotional intelligence to thrive in a small, fast-moving team. Most team members were based in Mountain View, California, and they had to be flexible enough to pitch in wherever help was needed. “You’re Team Bard,” Hsiao told them. “You wear all the hats.”

In January 2023, Pichai announced the first mass layoffs in Google’s history: 12,000 jobs, roughly 6% of the company’s global workforce. “No one knew what exactly to do to be safe going forward,” says a former engineering manager. Many employees worried that skipping overtime would put them at the top of the next layoff list. If that meant disrupting their kids’ bedtime routines to join Team Bard’s evening meetings, there was no question it had to be done.

Hsiao’s team needed massive support from across Google. They could build on LaMDA’s existing foundation, but they had to update its entire knowledge base and add new safety guardrails. Google’s infrastructure team reassigned its top engineers to free up server space for the extensive model tuning. The effort nearly maxed out electricity use at some of Google’s data centers, risking equipment overheating and burnout, so teams quickly built new tools to handle the surging power demand more safely. To ease tension around the competition for scarce computing resources, someone on Hsiao’s team ordered custom poker chips printed with the codename of Google’s in-house chips. They left a huge pile on an engineering leader’s desk and joked: “Here’s your chips.”

Even as new computing power came online in Bard’s first few weeks, engineers kept running into the same problems that had derailed Google’s past generative AI projects—problems that once would have made executives slow down the entire project. Just like ChatGPT, Bard hallucinated facts and often generated inappropriate or offensive responses. One former employee says early prototypes relied on “comically bad racial stereotypes.” If you asked for a biography of anyone with an Indian name, it would automatically label them a Bollywood actor. A Chinese male name? It assumed they were a computer scientist. Another former employee said Bard’s bad outputs weren’t dangerous—“just dumb.” Employees traded screenshots of its most absurd responses for a laugh. “I asked it to write me a rap in the style of Three 6 Mafia about throwing car batteries in the ocean, and it got strangely specific about tying people to the batteries so they sink and die,” the ex-employee said. “My request had nothing to do with murder.”

With a self-imposed 100-day deadline, Google could only catch and fix as many flaws as possible in time. Contractors who usually focused on removing child abuse imagery from Google’s platforms were reassigned to test Bard, and Pichai asked every employee with spare time to help test the tool. Roughly 80,000 Google workers pitched in. To manage public expectations, Hsiao and other executives decided to brand Bard as an “experiment,” just like OpenAI had framed ChatGPT as a “research preview.” They hoped this framing would protect Google’s reputation if the chatbot went off the rails. No one at Google had forgotten Microsoft’s 2016 disaster: Tay, a Twitter chatbot that turned into a full-on Nazi within 24 hours of launch.

In the past, before Google launched any AI project, its 12-person responsible innovation team would spend months independently testing systems for hidden biases and other flaws. For Bard, that entire review process was cut short. According to a former member of the responsible innovation team, Google’s top lawyer Kent Walker advocated for moving as fast as possible. New models and updates rolled out so quickly that reviewers couldn’t keep up, even working nights and weekends. When reviewers raised red flags and pushed to delay Bard’s launch, their concerns were overruled. In comments to WIRED, Google representatives pushed back on that account: “no teams that had a role in green-lighting or blocking a launch made a recommendation not to launch.” They also noted that “multiple teams across the company were responsible for testing and reviewing genAI products,” and “no single team was ever individually accountable.”

By February 2023, two-thirds of the way through the 100-day sprint, Google executives learned OpenAI had scored another win: ChatGPT would be integrated directly into Microsoft’s Bing search engine. Once again, the self-proclaimed “AI-first” company was trailing in the AI race. Google’s search division had been testing its own chatbot integration under a project called Project Magi, but it had yet to produce any tangible results. Google was still the undisputed leader in search—Bing held just 10% of the global market share—but how long would that supremacy last without a generative AI feature to offer users?

In an apparent effort to avoid another hit to its stock price, Google tried to upstage Microsoft. On February 6, one day before Microsoft was set to unveil its new AI-powered Bing, Pichai announced Bard would open to the public for limited testing. In a marketing video released alongside the announcement, Bard was framed as a helpful, versatile tool that carried forward Google’s longstanding mission to “organize the world’s information.” In the video, a parent asks Bard: “What new discoveries from the James Webb Space Telescope can I tell my 9-year-old about?” One of the AI’s answers was: “JWST took the very first pictures of a planet outside of our own solar system.”

For a moment, it looked like Bard had put Google back on top. Then Reuters pointed out the error: the first image of an exoplanet was captured by the European Southern Observatory’s Very Large Telescope, based in Chile, not in space. The mistake was deeply embarrassing. Alphabet shares dropped 9% in a day, erasing roughly $100 billion in market value.

The public backlash to the gaffe came as a shock to Team Bard. According to a former employee close to the team, the marketing staffer who came up with the JWST query felt personally responsible. Colleagues tried to cheer them up: executives, legal, and PR had all vetted the example, and no one caught the error. And given all the mistakes ChatGPT had made, who would have expected such a small error to tank the company’s market cap?

Hsiao called the moment “an innocent mistake.” Bard was trained to cross-check its answers against Google Search results, and it had most likely misinterpreted a NASA blog that announced JWST had captured the first direct image of an exoplanet taken by the telescope. One former staffer remembers leadership reassuring the team that no one would be punished for the incident, but that they had to learn from it quickly. “We're Google, we're not a startup,” Hsiao says. “We can't as easily say, ‘Oh, it's just the flaw of the technology.’ We get called out, and we have to respond the way Google needs to respond.”

Googlers outside the Bard team weren’t reassured. According to CNBC, one post on Memegen, Google’s internal discussion board, read: “Dear Sundar, the Bard launch and the layoffs were rushed, botched, and myopic. Please return to taking a long-term outlook.” Another post featured the Google logo inside a dumpster fire. But in the weeks after the JWST mistake, Google doubled down on Bard. The company added hundreds more staff to the project, and Pichai’s profile icon popped up daily in the team’s shared Google Docs far more often than it had for any past product.

More crushing news came in mid-March, when OpenAI released GPT-4, a language model far more capable than LaMDA at analysis and coding. “I just remember having my jaw drop open and hoping Google would speed up,” says a then-senior research engineer at Google.

A week later, the full public launch of Bard went ahead in the US and UK. Users reported it worked well for writing emails and research papers, but ChatGPT already did those tasks just as well, if not better. There was little reason for users to switch. Later, Pichai acknowledged on the Hard Fork podcast that Google had entered “a race with more powerful cars” driving a “souped-up Civic.” What the company needed was a far better engine—and that engine was Gemini.

The two AI labs that merged to build Gemini had very different corporate cultures. DeepMind, which was classified as one of Alphabet’s “other bets,” focused on solving long-term scientific and mathematical challenges. Google Brain had delivered more commercially practical breakthroughs, including the auto-complete feature in Gmail and technology that interprets vague search queries. According to a former high-ranking engineer, Brain’s leader Jeff Dean “let people do their thing,” while Demis Hassabis’ DeepMind group “felt like an army, highly efficient under a single general.”

Dean is a veteran engineer who has built neural networks for decades, joining Google before the company’s first birthday, while Hassabis is the visionary leader who dreams of using AI to cure disease. He had already tasked a small team with building what he calls a “situated intelligent agent”—a multi-sensory AI assistant that can help users with every part of their daily lives.

Hassabis was named CEO of the newly merged unit, Google DeepMind (GDM). Google announced the merger in April 2023, as rumors swirled of more upcoming OpenAI breakthroughs. “Purpose was back,” says the former high-ranking engineer. “There was no goofing around.” To build Gemini as fast as possible, employees coordinated work across eight time zones, and hundreds of new internal chatrooms popped up to organize the effort. Hassabis, who usually eats dinner with his family in London then works until 4 a.m., says “each day feels almost like a lifetime when I think through it.”

In Mountain View, GDM moved into Gradient Canopy, a new ultra-secure dome-shaped building surrounded by fresh lawns and six Burning Man-inspired sculptures. The team was on the same floor as Pichai’s office, Brin became a frequent visitor, and managers required more in-office work. Breaking with Google’s usual open culture, most other Google employees weren’t allowed inside Gradient Canopy, and couldn’t access GDM’s core programming code either.

As the Gemini project sucked up almost all the spare computing resources Google had, AI researchers working on other areas like health care and climate change struggled to get access to servers, and morale plummeted. Employees say Google also restricted the ability of researchers to publish some AI-related papers. Publications are a key form of professional currency for researchers, but it was clear Google feared giving away valuable insights to OpenAI. The formula for training Gemini was too valuable to risk being stolen—this was the model that would save Google from becoming obsolete.

Gemini ran into many of the same problems that plagued Bard. “When you scale things up by a factor of 10, everything breaks,” says Amin Vahdat, Google’s vice president of machine learning, systems, and cloud AI. As the launch date approached, Vahdat set up a dedicated war room to troubleshoot bugs and system failures.

At the same time, GDM’s responsibility team was racing to review the model before launch. For all its added power, Gemini still generated bizarre and problematic outputs. Ahead of launch, the team found “medical advice and harassment as policy areas with particular room for improvement,” according to a public report Google released. Gemini also made “ungrounded inferences” about people in images, for example when asked what level of education a person in a photo has. Nothing rose to the level of a “showstopper,” said Dawn Bloxwich, GDM’s director of responsible development and innovation. But her team had very limited time to anticipate every way the public would use the model—including every weird rap prompt users would throw at it.

This was the moment Google could have hit pause. OpenAI’s head start and media hype had already made ChatGPT a household name, the Kleenex of AI chatbots. That also made it a lightning rod for criticism: it was synonymous with both AI’s promise and its growing social harms. Office workers feared job loss, from entry-level roles to creative jobs. Journalists, authors, actors, and artists demanded compensation for their work scraped to train AI models. Parents found chatbots were generating inappropriate mature content for their kids, and AI researchers began openly debating the risk of catastrophic harm from unregulated advanced AI. That May, legendary Google AI scientist Geoffrey Hinton resigned, warning of a future where machines could destroy humanity by spreading undetectable disinformation and developing new deadly poisons.

Even Hassabis acknowledged he wanted more time to work through the ethical risks of advanced AI—so much about society and human life could be upended. But despite growing fears of AI doom, Hassabis still wanted to deliver on his vision of a universal AI assistant and breakthroughs like AI-developed cures for cancer. Google pressed ahead with the launch.

When Google unveiled Gemini in December 2023, shares rose immediately. The model outperformed ChatGPT in 30 out of 32 standard industry tests. It could analyze research papers and YouTube clips, answer complex questions about math and law. Current and former employees told WIRED it felt like the start of Google’s comeback. Hassabis held a small celebration at the London office. “I'm pretty bad at celebrations,” he recalls. “I'm always on to thinking about the next thing.”

The next big breakthrough came that same month. Jeff Dean learned about it when his team invited him to a new internal chatroom called Goldfish. The name was a nerdy, ironic joke: goldfish are known for short memories, but Dean’s team had built the opposite—a way to give Gemini a far longer memory than ChatGPT’s. By spreading processing across a high-speed network of interconnected chips, Gemini can now analyze thousands of pages of text or entire episodes of a TV show in one go. The team called the technology long context. Dean, Hassabis, and Manyika immediately started planning how to integrate it across all of Google’s AI services, pulling ahead of Microsoft

Related Article