The India Decade

AI's new scaling law shifts the battle from training to thinking

As tech companies run out of high-quality web data to train larger systems, they are finding that letting models spend more time calculating before they speak yields massive leaps in logic.

By The India Decade

Published

OpenAI logo
OpenAI logo · “OpenAI logo 2025 (wordmark)” by OpenAI (PUBLIC DOMAIN) via Wikimedia Commons

The race to build ever-larger artificial intelligence models is hitting a wall, forcing the industry's biggest players to pivot to a new strategy: teaching their systems how to think before they speak.

For years, the recipe for building smarter AI was simple. Engineers gathered more data, plugged in more microchips, and built a larger foundation model. But as high-quality human text on the public internet runs dry, and the financial cost of training these behemoths reaches billions of dollars, the returns of this brute-force approach are flattening.

Instead of focusing solely on the training phase, researchers are now scaling what they call "test-time compute" or inference-time reasoning. By using reinforcement learning to reward models for showing their work, labs are training systems to generate internal chains of thought, allowing them to identify and correct their own mistakes before presenting a final answer to the user.

This shift means AI is moving from instant, autocomplete-style pattern matching to active problem-solving. When faced with a complex prompt, these reasoning models do not immediately output the first word that comes to mind. Instead, they run an internal monologue, testing hypotheses, discarding dead ends, and refining their logic.

OpenAI pioneered this public shift with the release of its o1 and o3 models, designed to pause for several seconds—or even minutes—while they map out a logical path to a prompt. This approach has allowed systems to achieve human-expert performance on complex math, science, and coding benchmarks that previously baffled even the largest standard models.

However, the proprietary hold on this technology did not last long. The competitive dynamics shifted radically when the Chinese research firm DeepSeek released its R1 model. Not only did R1 match the reasoning performance of leading American models on key benchmarks, but its creators claimed to have trained it for a fraction of the cost, open-sourcing the weights so that researchers worldwide could study and replicate the technique.

This development changes how AI performance is evaluated and deployed. Traditional benchmarks tested instant recall. Under the new paradigm, a model might take several minutes to solve a complex software bug or verify a chemical formula, but the accuracy of the result is far higher.

It also opens the door to truly autonomous AI agents. If a model can systematically plan, test its own code, and debug its failures in a closed loop, it can operate independently for hours to solve complex engineering tasks without human intervention.

For the industry, the bottleneck is no longer just finding enough raw data on the internet. The new challenge is securing the computational power to let millions of users run these long-winded, computationally expensive internal thoughts at the same time. While tech giants scramble to secure power grids for future datacentres, the immediate battle is over efficiency: making these thinking models fast enough and cheap enough for everyday business applications.

Found an error in this story? Write to editor@theindiadecade.com. Our corrections policy explains how we put mistakes right.

Key numbers

DeepSeek-R1 claimed training cost
Under $6 million
Source: DeepSeek research paper

In this story

  • OpenAI — Pioneered the public shift to reasoning models with its o1 and o3 systems.
  • DeepSeek — Released the open-source R1 reasoning model, challenging proprietary dominance.

Topics

The day’s reporting, once each morning. No advertising.

More on this story