Chapter 13
The Consolidation of Compute and Foundation Models
The server hall in The Dalles had no windows, only the constant roar of manufactured air. Technicians in anti-static jackets moved with practiced economy between the racks, their headlamps bobbing in the gloom. Each rack held forty-four tensor processing units, Google’s custom answer to the compute problem. The units were dense, identical, humming with the work of matrix multiplications conducted at a scale unimaginable a decade prior. Outside, the Columbia River carved through the high desert, a witness to a different kind of engine room. Here, the power was not hydroelectric but computational, measured in kilowatts per square foot and the relentless depreciation of silicon against an insatiable demand curve. This was not a place of discovery. It was a place of execution.
The pursuit of scale was not merely a technical preference; it was an economic imperative forged in the fire of industrial capitalism. The transformer architecture and its subsequent scaling revealed a truth that resonated with a much older logic of production: the law of increasing returns. In classical manufacturing, scaling production led to lower marginal costs. In foundation model training, scaling compute led to not just linear but super-linear gains in capability, as captured by the power-law scaling laws. This created a dynamic where the entity that could invest first and heaviest could achieve a capability so superior that competitors would be permanently locked out of the market for the most advanced general-purpose intelligence services. The first laboratory to train a model with, say, 100 trillion parameters and achieve a qualitative leap in reasoning would effectively establish a new market reality. This was the core of the arms race: a race not to a finish line of general intelligence, but to a point of market irreversibility where compute-driven consolidation became the permanent state of the ecosystem. The venture capital flowing into OpenAI and Anthropic, the tens of billions committed by Microsoft and Google for internal cloud AI divisions, was not speculative funding for breakthrough ideas. It was a strategic outlay for industrial capacity, akin to the vast investments in steel mills or transcontinental railroads in the nineteenth century. The goal was to build the intellectual and physical infrastructure of a new monopoly, where the product—the foundational model—would be so expensive to replicate that its dominance would be assumed.
This economic logic reshaped the geography of innovation. Silicon Valley, the historical crucible of agile software startups, found itself cultivating entities that resembled state-owned industrial combines. The ideal AI CEO was no longer a visionary tinkerer but a logistics savant capable of managing global supply chains for semiconductors, negotiating power purchase agreements for atomic-scale energy consumption, and navigating the geopolitics of data sovereignty. The boardroom discussions shifted from user experience and product-market fit to methane emission offsets, underwater cable routes, and the placement of server halls along with geothermal vents. A new kind of institutional actor was consolidated: the Hyperscaler-Laboratory hybrid. Google, Amazon Web Services, and Microsoft Azure were not just selling cloud compute; they were consuming vast swathes of their own capacity to train proprietary models, creating a feedback loop where their platform dominance funded their model dominance, which in turn locked customers into their platform for API access. This vertical integration erased the distinction between infrastructure provider and application developer, consolidating the entire stack from silicon to syntax.
The human capital dimension of this consolidation was equally profound and historically specific. The field attracted a new legion of practitioners: not the polymath computer scientists of the early AI era, but specialized engineers in distributed systems, compiler optimization, and hardware-software co-design. The celebrated researchers became those who could write efficient CUDA kernels or design fault-tolerant training loop orchestration, not those who could conceive a new model of reasoning. A cultural schism emerged within computer science. Theorists and systematists found themselves fetching the equivalent of water for the massive growth engines of the foundation models. The “scaling hypothesis,” often presented as a scientific insight, was in reality an engineering manifesto that effectively declared the profession of classical AI research, as practiced for decades, economically obsolete. The graduate students at top universities learned not to question the transformer’s supremacy, but to optimize its hyper-parameters, to fine-tune it with novel adapters, or to apply it to new data modalities—silently accepting the underlying paradigm as immutable ground truth. This was not a failure of imagination, but a rational adaptation to a resource landscape that had been utterly transformed. The consolidation of compute was, in this sense, also a consolidation of intellectual ambition, channeling a generation of talent into a single, hyper-optimized stream.
The empirical discovery of scaling laws was the moment the theoretical became the tactical. Publications from Google (e.g., the 2020 paper “Scaling Laws for Neural Language Models”) and OpenAI (“Scaling Laws for Autoregressive Generative Modeling”) provided more than research results; they issued a roadmap for capital allocation. These laws demonstrated that key metrics of a model’s capability, like its loss on a prediction task or its performance on a question-answering benchmark, would follow smooth, predictable curves as one increased the number of model parameters (N), the size of the dataset (D), and the amount of compute used for training (C). Crucially, the relationships were power laws, meaning gains continued unabated over many orders of magnitude. This transformed AI research and development from a speculative gamble into a calculable, if extraordinarily expensive, construction project. A tech giant could now draft a “capability forecast”: by predicting hardware efficiency gains, data procurement costs, and the price of capital, they could chart a five-year plan to build a model of specified capability. The race was on to build the cleanest, most efficient pipeline to follow these prophetic curves.
This procedural foreknowledge had a homogenizing effect on research. The diverse evolutionary pressures of algorithmic innovation were replaced by a single selective force: compatibility with the scaling playbook. Architectures were evaluated not on elegance, theoretical promise, or interpretability, but on how efficiently they could scale. Could they leverage parallelism? Did they have a stable underlying structure amenable to incremental expansion? The transformer, with its uniform layers and matrix-centric operations, answered yes to these engineering questions overwhelmingly well. Consequently, variations like Switch Transformers or Mixture-of-Experts models emerged not as conceptual departures, but as optimizations within the scaling paradigm, designed to make the scaling curve shallower (i.e., achieve the same performance with less compute) or steeper more favorably. The foundational algorithmic concept was locked in; the race became one of marginal efficiency gains along a pre-ordained trajectory. This created a powerful conservatism at the heart of the ostensible frontier of innovation. The new, once they appeared, were typically techniques for managing the consequences of scale (like reinforcement learning from human feedback for alignment) rather than for displacing the scaled entity itself.
The data ingestion itself became a colossal, industrial operation. The era of carefully curated, small-scale datasets like ImageNet ended. Now, the “web-scale” dataset was a superset of the entire Common Crawl, hundreds of terabytes of text scraped from websites, forums, books, and code repositories. This required an entirely new sub-industry of data curation: deduplication algorithms to remove repeated passages, classifiers to filter out toxic or copyrighted content, and systems to balance domain representation. This process was as critical and capital-intensive as building the compute infrastructure. The data pipeline became a strategic asset. Meta’s scroll through vast social media archives, Google’s index of the web, and the carefully licensed books and articles compiled by academic or commercial consortia became the proprietary oil wells of the AI economy. The consolidation was therefore double-layered: consolidation of the physical means of computation and consolidation of the refined informational feedstock that computation processed. An entity without access to both was relegated to the sidelines, using the curated outputs of others or working on narrow, simulated problems detached from the web-sourced foundation.
The result was a fundamental shift in the locus of value creation in computer science. The value was no longer in a clever algorithm that could solve a specific problem with minimal resources, but in the ownership of the capital-intensive infrastructure and the licensed data corpus required to train a general-purpose model. This reoriented the entire field away from efficiency and toward abundance. The new courtesies of the field were not about proving a theorem about convergence rates, but about announcing a doubling of training compute, or a 2x reduction in inference cost, or the inclusion of a novel data modality like molecular protein structures or astronomical sensor logs. Each announcement further cemented the path of consolidation, as it demonstrated the futility of competing without committing to an infrastructure arms race of one’s own. The academic paper, once the currency of idea exchange, became in many domains a trailing indicator, a formalization of techniques already in production at the labs, with the most consequential details—the exact data mixtures, the proprietary fine-tuning recipes, the hyper-parameter search grids—kept as unpublished lore, part of the closed industrial practice of the compute oligopoly.
The physical infrastructure that enabled this consolidation became a defining stock character in the narrative. The modern data center, a windowless cathedral to computation, emerged in locations dictated by a calculus of energy, climate, and law. Places like Council Bluffs, Iowa; Luleå, Sweden; or Quincy, Washington, were chosen for their cool climates (reducing cooling costs), access to abundant and preferably cheap hydroelectric or wind power, and favorable tax regimes. These facilities were monuments to planned obsolescence on a grand scale. The lifespan of the specialized AI accelerators within them was measured in years, not decades. Each new chip generation from NVIDIA or Google offered a leap in performance-per-watt, forcing a continuous, costly reinvestment simply to keep pace. This “Red Queen’s race” meant that to train a frontier model was to commit to a depreciating asset of breathtaking scale. The server hall, therefore, was not just a tool but a terminus—a final destination for capital, where abstract financial investment was physically transformed into the fleeting, energetic act of training a model, after which the hardware began its swift slide into economic obsolescence, to be replaced and the cycle repeated.
This infrastructure demanded a new social contract with the host communities and states. The data center’s voracious appetite for power tied the fate of leading-edge AI development to local energy grids and, by extension, to national energy policy. In turn, the promise of high-tech jobs and transformative investment allowed these corporations to negotiate unprecedented concessions. They became integrated into planning processes for national energy infrastructure, influencing the development of new power generation projects—sometimes with clean energy mandates, sometimes without—further weaving their strategic needs into the physical fabric of governance. The consolidation of AI was thus not a clean, digital phenomenon divorced from messy reality; it was a physical occupation of territory, a drawing of new maps of resource power that belonged as much to the nineteenth-century legacy of resource extraction and heavy industry as it did to the twenty-first-century imagination of sleek, dematerialized technology.
The feedback loop completed itself inside these temples of compute. The immense cost of each training run created an institutional aversion to risk and a preference for incremental, verifiable progress. A six-month training run costing over a hundred million dollars could not be jeopardized by a wild experimentation with a radically different model architecture. Consequently, algorithmic research became heavily simulation-based and scaled-down. Teams might test a new idea on a small model (e.g., a few billion parameters) to see if it worked, and if it did, the only question became: “Does it still work predictably at 1000x scale?” This focus on scalability as the primary virtue of an idea became the filter through which all innovation had to pass. It created a profound conservatism, favoring extensions of the known transformer template over true paradigm shifts. The institution of the modern AI lab, therefore, had morphed from a place of exploration into a titan’s workshop, where the scale of the operation dictated the scope of the permissible thought.
Beyond the handful of frontier laboratories lived the vast majority of the AI research community: academics, independent researchers, and scientists at smaller companies. For them, the consolidation of compute struck like a gravitational force, bending their orbit and limiting their trajectory. The most powerful research tool in AI—massive, pre-trained foundation models—was not a tool they owned or even fully understood. It was a service, accessed via APIs, operated as a black box by distant corporations. This created a dynamic of dependence. Researchers could study the fine-tuning of these models (“prompt engineering”), their societal impacts, their biases, or their applications in specific domains like medicine or law. What they could not do easily was interrogate the model’s internal representations at scale, experiment with fundamental architectural changes, or train a competitive counterpart from scratch. The frontiers of the field moved behind a paywall, measured in dollars-per-million-tokens of API access rather than in peer-reviewed papers.
This dependency fostered a new kind of scholarly practice, often called “API science.” Research proposals began to formulate questions around the capabilities and limitations of specific, proprietary model versions (e.g., “GPT-4-turbo” or “Claude 2.1”). A finding might be valid only for a particular snapshot of a model’s training data and alignment tuning, creating a reproducibility crisis of a new kind. The underlying platform was a shifting sand, updated without notice by the owning company. This undermined the traditional scientific goals of building cumulative, stable knowledge. Furthermore, it colonized the intellectual space. The most pressing questions became those relevant to the models that existed: how to make them safe, how to make them useful, how to interpret their outputs. Questions about alternative, more efficient, or more interpretable architectures became niche, theoretical, or, worse, “irrelevant” to the dominant paradigm. The consolidation had thus achieved a remarkable feat: it had made the study of AI largely the study of its own industrial outputs.
The institutions meant to foster open inquiry began to adapt in troubling ways. University computer science departments found themselves at a disadvantage in recruiting top faculty, who could command orders-of-magnitude higher salaries and have access to unmatched computational resources at the private labs. To remain competitive, some departments sought partnerships with the tech giants, receiving access to models or compute resources in exchange for focusing on pre-approved areas of research or for providing human talent through internship pipelines. While these collaborations accelerated certain applied work, they risked turning academic departments into extensions of corporate R&D, subtly orienting the fundamental research agenda toward the strategic interests of a handful of companies. The consolidation of compute was, in this way, also a consolidation of intellectual authority. The questions deemed important were increasingly those that could be asked and answered within the walled gardens of the hyperscalers.
This state of affairs represented a profound historical rupture. Previous waves of technological transformation, from personal computers to the internet, were characterized by a feedback loop between decentralized, entrepreneurial tinkering and large-scale industrial production. The inventor in the garage could sometimes leapfrog the corporation. In the era of consolidated foundation models, the garage was irrelevant. The scale of entry was not a million dollars of venture capital, but tens of billions tied to access to semiconductor fabrication pipelines and global, primary energy contracts. The consolidation was not gradual; it was a phase transition, a crystallization of the AI field around a new, immovable center of gravity defined by capital and physical infrastructure. The server hall in The Dalles was not just a place where models were trained; it was a physical manifestation of the new order’s logic—a place where the future was being quietly and inexorably written in the silent, accumulated weight of its own enabling constraints.
This chapter’s argument about institutional consolidation is the direct consequence of the scaling laws detailed in Chapter 12. The predictable mechanics of massive scale did not just guide research; they dictated a new industrial reality. The logic of increasing returns, once quantified, became a blueprint for monopoly. The field’s long debate between scale and cleverness was no longer a philosophical question. It was a financial one, settled by the balance sheets of the hyperscalers. The victory of the old ideas—backpropagation and neural networks—was complete, but their triumph had produced a world where the power to explore those ideas was concentrated in a few hands. The next chapter examines the final, physical dimension of this concentration: the consolidation of compute and capital itself, the drawing of new maps where power flows not from algorithms but from energy contracts and semiconductor supply chains.