Chapter 1
The Exiled Architects of the Connectionist Revival
The mathematics did not announce itself with triumph. It arrived quietly, in paper form, embedded in an algebraic correction. The year was 1986, and in a two-volume collection of conference proceedings, a chapter titled “Learning Internal Representations by Error Propagation” presented a method for training networks with more than one layer of artificial neurons. The authors were David Rumelhart, Geoffrey Hinton, and Ronald Williams. The algorithm was backpropagation. Its publication did not trigger a rush of funding or a wave of industrial adoption. It triggered a quiet, persistent insurgency.
To grasp the deep learning revolution that would, decades later, produce systems of consummate power—GPT-4, ChatGPT—one must first look not at the sprawling data centers of the 2020s but at the cramped offices and ill-equipped universities of the late 1980s. The foundational mechanisms of deep learning did not emerge from well-funded corporate labs or government mega-projects. They emerged from the margins, and for the next twenty-five years, they would largely stay there. Their eventual triumph was not the inevitable arc of a superior technology. It was the achievement of a scattered, stubborn minority who refused to abandon an intellectually elegant but computationally starved idea. This is the story of the exiled architects: Geoffrey Hinton in Toronto, Yann LeCun in New York, Yoshua Bengio in Montréal. Their biographical persistence—their will to keep designing, publishing, and training miniature versions of grand ideas in a climate of institutional indifference—was the first, crucial precondition for everything that followed.
Before the exile, there had been a brief, intoxicating spring. The 1980s saw what historians of artificial intelligence call the connectionist revival. After decades dominated by symbolic AI—the paradigm of hand-coded rules, logical inference trees, and expert systems—neural networks, inspired loosely by the brain’s architecture, offered a different path. Instead of programming intelligence, you would let a network learn from data. The key practical problem was credit assignment: how could you adjust the weights on connections deep within a network to reduce errors at the output? Backpropagation, building on earlier work by Paul Werbos and others, provided a computationally tractable answer. It used the chain rule of calculus to propagate error gradients backward through the layers, allowing the system to learn.
For a brief moment, the idea gathered real energy. Conferences buzzed. Defense agencies, intrigued by the prospect of machines that could learn to recognize speech or terrain, allocated modest grants. Commercial ventures, like the startup founded by Terry Sejnowski and Charles Rosenberg to build a neural network chip called NETtalk, made the pages of technical magazines. The atmosphere was one of earnest, slightly naive hope. A simple machine, trained on phonemes, could read aloud from the Wall Street Journal. It felt like a foundation.
But the foundation was built on sand, and the tide of history was turning against it. The symbolic AI establishment, centered at elite institutions like Stanford and MIT, viewed connectionism with skepticism and, often, open contempt. Marvin Minsky, the most formidable voice in the field, along with Seymour Papert, had published Perceptrons in 1969, a book whose mathematical critique of single-layer networks was so devastating it had effectively buried the entire subfield for a decade. Now, the rise of backpropagation and multilayer perceptrons was a direct challenge to that narrative. The counter-offensive was intellectual and material.
Intellectually, the symbolic camp argued that neural networks were a form of “just” pattern matching, lacking the logical rigor and explanatory power of their own systems. Networks learned statistical correlations, but they did not understand. They could not reason in human-comprehensible terms. An expert system for medical diagnosis could explain its chain of logic; a neural network was a black box that gave answers. In the 1980s and 1990s, when computer memory was expensive and processing was slow, the symbolic approach offered a more efficient use of scarce resources. You could encode knowledge explicitly; you didn’t need to waste precious cycles trying to rediscover what was already known.
Materially, the institutional and funding landscape froze out the connectionists. Government grants, particularly from the newly confident field of expert systems, flowed to the paradigms that worked—paradigms that produced demonstrable, if narrow, results. The Defense Advanced Research Projects Agency (DARPA), which had flirted with neural networks, pulled back sharply in the early 1990s, influenced by scathing outside reports that questioned their potential. The subsequent “AI winter” was broad, but its chill fell unevenly. The establishments for symbolic AI, while disappointed, had tenured professors, established departments, and graduate students. The connectionists were becoming homeless.
This is where the story turns from a technological narrative into a human one. The survival of the multilayer perceptron and backpropagation as living, researched ideas depended almost entirely on the decisions of a handful of individuals. These were not men of supreme political skill or institutional power. They were researchers with an uncommon, almost pigheaded commitment to a line of inquiry that the rest of the world had declared futile. Their laboratories became arks, their graduate students a fellowship of the faithful. The exile forged them.
Geoffrey Hinton became the movement’s reluctant patriarch. The great-great-grandson of the logician George Boole, Hinton had come to AI through a dissatisfaction with its symbolic strain. He saw connectionism not as mere engineering but as a path to understanding the mind. After positions in the United States and a stint at the Canadian Institute for Advanced Research (CIFAR), he landed at the University of Toronto in 1987. The timing was catastrophic. He arrived just as the winter deepened. His lab, initially funded by CIFAR’s unique, long-term funding model, became a sanctuary. In the years that followed, securing conventional grants from Canada’s Natural Sciences and Engineering Research Council (NSERC) was a constant struggle. Proposals that mentioned “neural networks” often went straight to the reject pile.
Hinton’s response was not to abandon the core ideas, but to double down on them, pursuing a strategy of intellectual refinement and exotic cultivation. If neural networks were beaten by support vector machines and other kernel methods in head-to-head benchmarks, the problem, Hinton reasoned, was not the concept but the practice. We needed better ways to initialize weights, better architectures, better learning rules. He fostered a culture of intense, rigorous theorizing. His lab became a greenhouse for novel architectural ideas: the Boltzmann machine, a stochastic neural network that could learn internal representations; later, the idea of using neural networks for dimensionality reduction. These were not proof-of-concept demos for the world; they were deep intellectual explorations. The lab produced seminal papers, but their impact was felt only in the small, global community of true believers. The physical space was often cluttered, the funding precarious. The currency was ideas, and the students who could generate them. This fostered a specific kind of resilience: a deep familiarity with the mechanism’s theoretical limits and a contempt for the shallow benchmarks that unfairly condemned it.
Yann LeCun’s exile was more explicitly architectural, and more directly tied to the physical world. If Hinton was the theorist, LeCun was the engineer possessed by a single, powerful insight: that a network’s structure should mirror the structure of its problem. Working first at AT&T Bell Labs, then later at the NEC Research Institute, and eventually back in academia at NYU, LeCun immersed himself in the problem of visual recognition. He understood that a densely connected multilayer perceptron, where every neuron in one layer connects to every neuron in the next, was a brute-force solution to vision. It was computationally expensive and ignored the powerful spatial hierarchies present in images: a cat’s ear is a local feature that can appear anywhere in the frame.
His solution was the convolutional neural network (CNN). By imposing a grid-like topology on the network, making local connections, forcing neurons in a given region to share weights, and using pooling operations to gradually build spatial invariance, LeCun created an architecture that was specifically designed to process visual data. It was a magnificent piece of reverse-engineering from first principles. In the early 1990s, his team built a system that could read ZIP codes on handwritten envelopes for the US Postal Service, and a check-reading system used by banks. It worked, commercially. Yet it was seen, again, as a clever but limited specialized tool. The broader AI community saw it as a niche application of neural networks, not a general-purpose breakthrough. LeCun’s work at Bell Labs was eventually defunded as the institution’s priorities shifted. His move to NYU was a retreat to academia, but a strategic one, positioning his CNN work at the heart of curriculum and research. The architecture he designed would lie in wait for twenty years, until the data and compute arrived to unlock its potential.
Yoshua Bengio, based at the Université de Montréal, represented the third pillar: the rigorous theoretician who sought to place neural networks on the firmest possible mathematical footing. While others might be content to show that something worked, Bengio wanted to prove why it worked, and to formalize the conditions under which it could be expected to succeed. His early work, often in collaboration with Hinton, focused on the properties of recurrent neural networks for sequences and the challenges of learning long-term dependencies—the infamous vanishing gradient problem. This wasn’t flirtation with a fad; this was a committed, long-term effort to solve the hard theoretical problems that plagued the field.
Bengio’s laboratory in Montréal, like Hinton’s in Toronto, became a legendary incubator. It was a place where the core tenets of connectionism were debated with philosophical intensity. There was a shared sense of mission, of being the custodians of a suppressed truth. The intellectual culture was insular by necessity. When mainstream conferences like NeurIPS (then NIPS) became hostile reception grounds for neural network papers, the community organized its own workshops. CIFAR, under Hinton’s early leadership, ran a neural computation program that effectively created a parallel academic society. It funded travel, hosted small, closed-door workshops, and allowed the core group to maintain a critical mass. The isolation bred a kind of tribal cohesion. You couldn’t just dabble in this work; to continue, you had to believe.
This belief was, for most of the period, irrational by conventional metrics. The empirical results were mixed. On standard machine learning benchmarks—such as those for character recognition or simple classification tasks—neural networks were often outperformed by support vector machines (SVMs), random forests, or other methods that were easier to train, less finicky, and came with stronger statistical learning theory guarantees. The SVM camp, which became dominant from the late 1990s onward, could point to powerful optimization theories and elegant dual formulations. Neural networks seemed like a brute-force, heuristic approach from an earlier era. The hardware was simply not there to test them at scale. A network with a million parameters was considered large; training it could take weeks on the CPUs of the day.
Thus, the architectural ingenuity of the exiled architects was not born of abundance, but of scarcity. They were forced to be profound precisely because they could not afford to be merely large. The convolutional architecture was a profound piece of algorithmic efficiency. The breakthroughs in understanding gradient flow (Bengio’s work on vanishing gradients, Sepp Hochreiter and Jürgen Schmidhuber’s Long Short-Term Memory for recurrent nets) were attempts to solve fundamental learning problems with smarter design, not more compute. The development of ReLU (Rectified Linear Unit) activation functions in the early 2000s, which helped mitigate the vanishing gradient problem and allowed for faster training, was another piece of architectural insight, one that would prove vital later. In this period, the “bitter lesson” as articulated by Richard Sutton decades later—that methods that leverage computation tend to win in the end—was not yet evident. Computation was anemic. The only way forward was cleverness.
This period culminated in a crisis that perfectly illustrates the structural tension within the field. On one side, the symbolic and statistical establishments looked down on neural networks as theoretically impoverished and empirically unimpressive. On the other side, the neural network tribe, armed with powerful architectural and mathematical ideas, saw a universe of potential stymied by the material world’s failure to catch up. The ideas were ready. The world was not.
The first major signal that the world might be changing came not from inside the field, but from the broader economy. In the late 1990s and early 2000s, the internet was exploding. This created, almost as a by-product, two things that would later serve as the tinder and kindling for deep learning: vast repositories of digital data, and a massive market for computer graphics, which spurred the development of highly parallel processors for rendering visual scenes—the graphics processing unit (GPU).
In 2006, Hinton published a paper that, though technical, was intended as a shot across the bow. He introduced the concept of “deep belief nets” and a layer-wise pre-training procedure that made it feasible to train networks with many layers. The paper was important because it showed that deep networks could be made to work again; it was a technical reactivation of the core idea. But its impact was still muted. The world did not yet have the means to take the next step.
That step would be provided by an unlikely coalition. Fei-Fei Li, a computer vision scientist at Stanford, had a vision for a dataset that would dwarf all others in scale and diversity: ImageNet. Her goal was to organize a massive subset of the world’s images into a hierarchy of concepts and use it to train the next generation of computer vision systems. This was not a project to revive neural networks; it was a project to impose order on the emerging digital visual world. The resulting dataset, completed in 2009, contained over 14 million labeled images across 20, 000 categories. It was an asset of unprecedented scale, a massive fuel depot for machine learning.
The engine to consume this fuel was provided by the gaming industry. NVIDIA, in its relentless pursuit of faster and more realistic graphics for video games, had developed a massively parallel processor architecture: the GPU. Its key attribute was not clock speed, but throughput—the number of operations it could perform simultaneously. By the mid-2000s, a handful of researchers, like Andrew Ng at Stanford, were beginning to understand that GPUs could be repurposed for the parallel matrix multiplications that sat at the heart of neural network training. In 2009, a landmark paper by Rajat Raina, Anand Madhavan, and Andrew Ng demonstrated that training a deep belief network on a GPU could be over 70 times faster than training it on a standard CPU. This was an inflection point, though it was not yet widely recognized as such.
The stage was now set. The theories were refined in exile. The architectures for vision (CNNs) and sequence processing (RNNs/LSTMs) were waiting. The math for training deep networks was understood. The massive, curated dataset (ImageNet) was available. And the hardware (GPUs), designed for alien planets and shoot-em-ups, had been found to be strangely, miraculously compatible with the mathematical demands of deep learning.
The exiled architects did not create ImageNet, nor did they design the GPU. But they had kept the mechanism alive, refining its theoretical and architectural foundations through the long winter. They had maintained a global network of believers, training the next generation (like Alex Krizhevsky and Ilya Sutskever in Hinton’s Toronto lab). When the fuel and the engine arrived, they were the only ones with a blueprint for a machine that could use them.
The counter-argument emerges here with force: was this not, in fact, a triumph of architectural ingenuity? The symbolic AI paradigm also had access to more data and more compute (expert systems could be run on powerful mainframes). It failed to scale. The kernel methods that dominated the 2000s were also helped by more data and faster computers. They, too, plateaued. It was the specific architecture of the convolutional neural network, with its weight-sharing and pooling, that allowed it to leverage scale efficiently. It was the specific fix of ReLU activations and improved initialization that allowed gradients to flow through deep layers. Without those pieces of architectural insight, simply throwing more data and compute at a dense, randomly initialized multilayer perceptron would have yielded dismal, overfitted results.
This is the core of the structural tension that defines this period. The exiled architects were not passive victims of history, simply waiting for the hardware to catch up. They were actively engineering the theoretical and architectural tools that would allow their ideas to exploit scale once it became available. Their isolation forced them to focus on the essence of the mechanism, to solve the hard problems within it. The “wasteland” of the AI winter was, in retrospect, a period of fertile, focused R&D.
Thus, the eventual victory of deep learning was not a sudden miracle. It was the convergence of two independent trajectories that had been evolving in parallel. One was the trajectory of scale—driven by internet data, consumer electronics, and video games—a trajectory entirely external to and largely ignorant of the other. The second was the trajectory of neural network architecture and theory, driven by the exiled architects in their university labs. Neither trajectory alone was sufficient. Scale without architectural insight produced only brute-force, inefficient computation. Architectural insight without scale produced only elegant, unproven theories.
The end of the exile was not a surrender on either side. It was an accidental partnership. The GPUs were purpose-built for massively parallel floating-point math. The data was labeled not by AI researchers but by thousands of human workers on Amazon Mechanical Turk and in Fei-Fei Li’s army of student annotators. The connectionist researchers, meanwhile, had spent decades understanding how to structure the mathematical operations that the GPU was now ready to accelerate. They understood how to build a network that could learn hierarchical features from raw pixels. They were the architects of the cathedral; the world had finally delivered the quarry of stone and the legions of laborers.
This biographical persistence, this refusal to let the mechanism die, set the foundation for everything. The textbooks and lecture notes from Hinton, LeCun, and Bengio courses during the 1990s and 2000s trained the cadre of engineers who would later staff the AI labs of Google, Facebook, and DeepMind. The technical blogs and open-source code from their labs provided the templates. The very idea that neural networks were a legitimate, if marginal, area of research survived because they continued to publish in it.
The stage was now set for the collision. The mechanisms, exiled and refined for a quarter-century, were about to meet the material world that had ignored them. The story was about to shift from the quiet persistence of architects to the noisy, capital-fueled revolution of engineers and entrepreneurs. The data was labeled. The GPUs were humming. The first true test was about to be administered, not in a patent office or a medical diagnosis tool, but in the loud, public arena of image recognition. And the world of computer science, which had spent decades dismissing these exiled architects, was about to have its eyes forced open. The mechanism origins had led to this. The next chapter would begin with the fuel and the silicon—how datasets were assembled and chips were repurposed—setting the stage for the shock to come.
The genealogy of the algorithm itself deserves closer scrutiny, for the chain of rediscovery reveals how marginal ideas circulate through the academy like contraband. The mathematical core of backpropagation—the application of the chain rule to compute gradients of a composite function—was not born in 1986. Its intellectual DNA can be traced to Seppo Linnainmaa’s 1970 master’s thesis at the University of Helsinki, which described a method of reverse-mode automatic differentiation for nested composite functions. Linnainmaa’s work was not concerned with neural networks at all; it was a contribution to numerical computing, a way to efficiently compute derivatives for scientific algorithms. The connection to learning in multilayer networks was made, with considerable foresight, by Paul Werbos, an applied mathematician working at Harvard in the mid-1970s. Werbos’s doctoral dissertation, completed in 1974 under the supervision of several committee members skeptical of its relevance to psychology, explicitly proposed using backpropagation as a learning algorithm for neural networks. He even prefigured the idea that such networks could model cognitive processes. The thesis was published in a relatively obscure venue and did not penetrate the computer science mainstream. Werbos continued to advocate for the approach for years, publishing a book in 1994 titled The Roots of Backpropagation that sought to reclaim the intellectual history. But in the competitive ecology of the 1970s and 1980s, where would-be discoverers jostled for credit, the fact that the method existed in a usable mathematical form years before Rumelhart, Hinton, and Williams’s celebrated 1986 paper meant that the “discovery” was as much about community acceptance as about novelty. The 1986 paper succeeded in catalyzing the revival not because it contained a radically new equation, but because it arrived at a moment when enough people were ready to listen, and the parallel distributed processing framework gave the algorithm a compelling cognitive gloss. This layering of rediscovery hints at a deeper pattern: the mechanisms of deep learning were repeatedly found and repeatedly lost, not for lack of mathematical insight, but for lack of institutional traction.
The shadow cast by Marvin Minsky and Seymour Papert’s Perceptrons (1969) elongated far beyond its immediate mathematical content. The book’s most famous result demonstrated that a single-layer perceptron could not compute the XOR function—a simple logical operation that requires, in effect, the ability to carve a non-linear decision boundary. The proof was elegant and, in its limited scope, correct. But the way it was framed, and the way it was received, amplified its meaning far beyond the theorem. Minsky and Papert’s concluding chapters suggested, with considerable rhetorical force, that the limitations of single-layer perceptrons were indicative of fundamental limitations in the neural network approach as a whole. The implication—that scaling up to multilayer networks would not rescue the paradigm—was widely absorbed as an established fact, even though Perceptrons did not contain a rigorous mathematical proof that multilayer perceptrons were hopeless. In the American academic ecosystem of the early 1970s, where Minsky’s authority at MIT was near absolute and where the newly created field of AI was struggling for legitimacy, the book functioned as a political instrument. It effectively redirected the funding of an entire generation. Graduate students were steered away from neural networks. Journal editors became skeptical of submissions that referenced connectionist models. The word “perceptron” became a mark of obsolescence.
The depth of this chill can be measured by what did not happen in the 1970s. There were researchers who understood, even then, that multilayer networks with appropriate nonlinearities could overcome the XOR limitation. Notably, Alexey Ivakhnenko in the Soviet Union had developed a form of multilayer nonlinear modeling through his work on the Group Method of Data Handling (GMDH), a layer-by-layer approach to growing polynomial networks, published through the 1960s and 1970s. Kunihiko Fukushima in Japan designed the neocognitron, a hierarchical, multi-layered neural architecture for pattern recognition, in 1979—an architecture that clearly anticipated the convolutional neural network. But these works existed in academic cultures that were, for different reasons, disconnected from the American mainstream: Ivakhnenko in the Soviet scientific establishment, which operated under entirely different institutional pressures; Fukushima in the Japanese research communities, where neural network work would continue but with limited influence on American funding agencies. The result was a peculiar form of intellectual xenophobia: the United States, having declared the field dead, was largely ignorant that useful work continued elsewhere. The mechanism was allowed to atrophy not because it had been definitively disproven, but because the most powerful gatekeepers in the most powerful academic system had pronounced it moribund.
The institutional dynamics of the AI winter of the late 1980s and 1990s were not a simple story of universal deprivation. The winter fell unevenly, and its cold was most bitter for those whose ideas were least aligned with the prevailing paradigm. Symbolic AI researchers saw their budgets shrink but did not lose their departments or their graduate student pipelines. The LISP machine companies—Symbolics, LMI, Texas Instruments’ Explorer line—had collapsed, taking with them a pillar of the AI industrial ecosystem. The Japanese Fifth Generation Computer project, which had inspired a wave of funding in the early 1980s with its ambitious promise of logic-based machine reasoning, was being quietly acknowledged as a failure by the late 1980s. Expert systems, which had been the brightest commercial star of early-1980s AI, proved brittle: they worked within their narrow domains but could not gracefully handle ambiguity, novelty, or the ambiguity of natural language. DARPA’s strategic assessment under the leadership of officials who were, by the early 1990s, increasingly skeptical of the field’s return on investment, led to a pronounced shift toward information technology projects—database management, logistics, communication networks—where results could be more easily quantified. Programs like the Strategic Computing Initiative, which had bankrolled a significant swathe of AI research in the mid-1980s, lost momentum and funding authority.
Against this backdrop, the National Science Foundation’s review panels became increasingly hostile to neural network proposals. Former program officers and reviewers from this period recall, in later interviews and historical accounts, a pattern of dismissive reviews. Proposals that mentioned “neural networks” or “connectionism” were sometimes sent back with notes suggesting the applicants redirect their efforts toward “real” machine learning—which meant statistical methods, support vector machines, or Bayesian approaches—or toward the analysis of biological neurons, which was the domain of neuroscience, not computer science. The very ambiguity of neural networks—they lived between biology and mathematics, between engineering and cognitive science—meant they did not fit neatly into departmental structures. They were too mathematical for the biologists, too biological for the mathematicians, and too empirically suspect for the engineers. This interstitial status intensified their vulnerability. When budgets tightened, the interdisciplinary niche was the first to lose its grip.
David Rumelhart, the third member of the 1986 backpropagation paper’s authorship, bears mentioning precisely because his trajectory illustrates the personal costs that even partial engagement with the connectionist camp could exact. Rumelhart was a cognitive psychologist by training, based at Stanford and later at UC San Diego. He was the intellectual architect of the parallel distributed processing (PDP) research group, which produced the two-volume Parallel Distributed Processing: Explorations in the Microstructure of Cognition in 1986—a landmark work that went far beyond the technical report on backpropagation to articulate a comprehensive vision of cognition as emergent from distributed, learned representations. Rumelhart was not a computer scientist by temperament; he was interested in the mind, and he saw in neural networks a vehicle for a new cognitive science. But by the late 1990s, he was stricken with Pick’s disease, a form of frontotemporal dementia that progressively robbed him of language and memory. He could no longer participate in the field he had helped to launch. There is a painful irony here: the man who had championed the idea that memory and learning were properties of a network of connections was, himself, losing his own connections. Rumelhart died in 2011, just two years before the full flowering of the deep learning revolution. His co-authors carried the torch; his illness became a silent emblem of the period’s fragility, a reminder that the custody of ideas depended on the physical and cognitive health of a small number of mortal individuals.
The role of industry laboratories in sustaining connectionist research during this period was paradoxical—simultaneously more enabling and more precarious than academic positions. Bell Labs, the storied research arm of AT&T, occupied a unique position in the American scientific landscape. Its generous funding model, sustained by the regulated monopoly revenues of the telephone company, supported fundamental research across physics, mathematics, and computer science with an openness that rivaled the best universities. It was in this environment that Claude Shannon had developed information theory, that the transistor was invented, and that the programming language C and the operating system Unix were born. When Yann LeCun joined Bell Labs’ Adaptive Systems Research Department, he entered a culture that valued practical elegance and that rewarded demonstrations of working systems. The handwritten digit recognition system, deployed in the early 1990s to process checks for NCR Corporation and other financial institutions, was a remarkable achievement in this context: a neural network trained end-to-end that outperformed all competing approaches on a commercially valuable task. The system was custom-designed hardware as well as software—a multi-chip board that could execute the forward pass of a convolutional network in real time.
But the lesson of Bell Labs for the connectionist cause was double-edged. The corporate laboratory’s purpose was not the preservation of intellectual traditions; it was the creation of value for its parent organization. AT&T’s divestiture and the subsequent restructuring of the telecommunications industry in the 1990s and early 2000s subjected Bell Labs to repeated cycles of strategic re-evaluation. The emphasis shifted from blue-sky research to nearer-term technology development. LeCun’s group, along with other fundamental research efforts, was progressively marginalized. The labs were relocated, reorganized, and eventually spun off during the 2005 Lucent-Alcatel merger. By the time the deep learning revolution arrived, the specific institutional home that had nurtured CNNs had essentially ceased to exist in its original form. The intellectual property, the institutional memory, and the physical infrastructure were scattered. That LeCun’s ideas survived such dispersal testified not to the stability of their institutional home, but to the fact that the ideas had been encoded—literally, in published papers—and that LeCun himself had the tenacity to carry them forward into a new academic setting at NYU’s Courant Institute. The lesson was stark: in the cold, no institution was safe. Only the researcher, with the mechanism inscribed in their mind and their publication record, was portable.
The culture of Bell Labs, for all its pressures, also gave LeCun something that pure academic isolation could not have provided: a direct, unsentimental encounter with the practical constraints of real-world systems. Check-reading is not a glamorous problem. But it imposed hard constraints—low latency, high accuracy, tolerance of messy inputs—that shaped the architecture in ways that benchmark competitions on clean datasets could not. A system that could reliably read the scrawled “12345” on the bottom of a check in a bank’s high-speed reader had been forged in contact with reality. This engineering discipline—the need to make things that actually worked, reliably, in production—would later prove immensely valuable. When the field entered its industrial phase in the 2010s, the architects who had spent time in industry possessed a practical fluency that pure theorists often lacked. The exile, in this sense, was not monolithic; it came in different forms. The academic exile bred theoretical depth and a defensive, insular culture. The industrial exile bred pragmatic resilience and an understanding of engineering constraints. Both were necessary.
A more granular look at the daily economics of connectionist research reveals how much the survival of the field depended on the peculiarities of non-American funding ecosystems. In Canada, the natural granting agencies operated on a different rhythm and held different values. The Natural Sciences and Engineering Research Council (NSERC) is funded through parliamentary appropriation and tends toward smaller, more stable grants distributed across a broad base. In the 1990s, a typical NSERC discovery grant for a computer science researcher might run between $30, 000 and $70, 000 CAD per year, a figure that was modest even by the standards of the time but that could cover a graduate student’s stipend and a small amount of computing. The critical difference was not the amount but the evaluation criteria. NSERC’s peer review committees, while not immune to trends, were less susceptible to the herd behavior that characterized big-grant agencies like DARPA or the NSF’s most competitive programs. A researcher who continued to publish solid work in any well-defined subfield of computer science could expect to maintain funding, even if that subfield was currently unfashionable. This stability was a lifeline. It meant that Hinton and, later, Bengio could maintain small but continuous streams of graduate students and postdoctoral fellows, preserving the institutional knowledge of how to train, debug, and theorize about neural networks.
CIFAR, the Canadian Institute for Advanced Research, provided a different, complementary form of support. Founded in 1982, CIFAR funds networks of scholars—called programs—across multiple institutions, providing flexible, multi-year support that is not tied to specific deliverables. Its Neural Computation and Learning program, initiated in 2004 (with roots in earlier, related efforts in the 1990s), brought together Hinton, Bengio, and later LeCun, along with dozens of other researchers, into a coordinated intellectual community. The program held regular meetings, often in secluded locations, where participants presented early-stage work, debated ideas, and formed collaborations that crossed institutional lines. The funding was modest in absolute terms—CIFAR’s annual budget for all its programs has historically been in the tens of millions of Canadian dollars, a tiny fraction of the money flowing through American federal agencies—but its impact was disproportionate. CIFAR funds went directly to people, not to equipment or overhead. They paid for plane tickets, for sabbaticals, for the time to think. In a field starved by the results-driven logic of project-based grants, this kind of support was transformative. It created a protected space in which the architects could focus on long-term, difficult questions without the constant pressure to produce short-term, publishable results.
Elsewhere in the diaspora, the picture was more varied and more precarious. In continental Europe, neural network research was sustained by a patchwork of national funding bodies, each with its own priorities and norms. France, which would later emerge as a major center of deep learning research in part through Bengio’s networks and the attractions of Montréal for French-speaking scientists, had a complex relationship with the field. The French academic establishment, centered on the grandes écoles and the CNRS (Center National de la Recherche Scientifique), tended to favor mathematical rigor and was skeptical of the empirically-driven, heuristic-heavy style of American neural network research. This meant that French researchers who did engage with connectionism often did so through the lens of more formally accepted mathematical frameworks—statistical physics, optimization theory, dynamical systems—producing a body of work that was sometimes more theoretically sophisticated than its American counterpart but that existed in a parallel academic culture with limited cross-fertilization. The United Kingdom had a few active centers, notably at the University of Edinburgh and Aston University, but the overall environment was similarly fragmented and underfunded. In Japan, where Fukushima’s neocognitron had established a precedent, neural network research continued at several universities and at the Advanced Telecommunications Research Institute (ATR) in Kyoto, but the Japanese AI community was itself moving in other directions, increasingly focused on robotics and human-computer interaction.
The conferences that served as the connective tissue of this dispersed community merit attention as institutional artifacts in their own right. The Neural Information Processing Systems (NeurIPS) conference, founded in 1987, occupied a peculiar position. It was created by a group of neural network enthusiasts, many of whom came from the computational neuroscience side, in the hope of bridging the gap between biology and machine learning. In its early years, NeurIPS was small, intimate, and committed to the connectionist cause. Papers on recurrent networks, backpropagation, and representation learning could be presented without the risk of being drowned out by the symbolic AI establishment that dominated larger conferences like AAAI (the Association for the Advancement of Artificial Intelligence) and IJCAI (the International Joint Conference on Artificial Intelligence). By the mid-1990s, however, NeurIPS began to attract researchers from the statistical machine learning community, who were interested in kernel methods, Gaussian processes, and related approaches. This influx was both a sign of health—people were paying attention—and a source of cultural tension. Neural network papers competed in sessions alongside statistical methods papers, and the latter often seemed to have the advantage in terms of theoretical polish and the clean guarantees that reviewers valued. The result was a conference that served as the neural network community’s most important venue but that was also the site of recurring and sometimes bitter debates about the rightful place of neural networks within the broader field of machine learning.
The smaller, more focused workshops were equally critical. The Connectionist Models Summer School, organized periodically by Hinton and others, provided a pedagogical incubator for the next generation. These were intensive, immersive events where students from around the world gathered to learn the fundamentals of neural network theory and practice, often from the architects themselves. The atmosphere was informal, even bohemian, with a near-religious intensity. Participants recall long evenings of heated discussion, whiteboard sessions that lasted into the small hours, and a pervasive sense of being part of something that the wider world did not understand or appreciate. Many of the engineers and researchers who would later play key roles in the deep learning revolution—including Ilya Sutskever, who would co-found OpenAI after a stint at Google Brain, and Alex Krizhevsky, whose AlexNet would ignite the 2012 breakthrough—passed through these summer schools and similar venues. The workshops functioned as secret seminaries, where the ideas that the mainstream considered dead were treated as living gospel.
In the 2000s, the landscape of machine learning was dominated by an alternative: kernel methods, and in particular the support vector machine (SVM). Introduced by Vladimir Vapnik and Corinna Cortes at AT&T Bell Labs in the mid-1990s, the SVM offered a mathematically rigorous, convex-optimization-based approach to classification that was, in many respects, the antithesis of the neural network philosophy. Where neural networks were trained by gradient descent on a non-convex loss surface, producing different results depending on random initialization, the SVM found a unique global optimum. Where neural networks required careful tuning of architecture and hyperparameters, the SVM had a single key knob—the choice of kernel function and a regularization parameter. Moreover, Vapnik’s statistical learning theory provided bounds on generalization error, giving practitioners a theoretical guarantee that neural networks, with their tendency to overfit, conspicuously lacked. The SVM community produced a series of striking empirical victories: on handwritten digit recognition (the same domain where LeCun’s CNNs had previously ruled), on text classification, on bioinformatics benchmarks. They also produced a cohort of brilliant mathematicians—Bernhard Schölkopf, Jason Weston, and others—who could argue with formidable theoretical sophistication for the superiority of their approach.
For the neural network community, the rise of SVMs was not merely another competing method; it was an existential challenge. In head-to-head comparisons on standardized benchmarks, neural networks frequently lost, or won only by small margins after extensive architectural tuning. This was deeply demoralizing. It suggested that the connectionist faith—that learning distributed representations from data was the right approach to intelligence—might be wrong, or at least that it was not the most efficient path to practical performance. Some neural network researchers defected to the SVM camp. Others grew disillusioned and left the field entirely. The few who persisted did so with the understanding that the benchmark paradigm itself was flawed: standardized benchmarks, by their nature, favored methods that could be quickly tuned to small, clean datasets, and penalized methods whose true strength lay in scalability and in learning rich, hierarchical representations from large, messy data. This insight—that neural networks were poor sprinters but potentially great marathon runners—was a conviction, not yet a demonstrated fact. Holding it required a kind of faith that, by conventional metrics, bordered on delusion.
The mathematical discourse between the two camps was frequently laced with a tone that went beyond academic disagreement. At NeurIPS workshops and in the pages of machine learning journals, the arguments could be sharp and personal. Proponents of kernel methods accused neural network researchers of being theoretically unsophisticated, relying on computational brute force rather than elegant mathematics. Neural network proponents responded that SVMs were fundamentally limited by the need to hand-design a kernel function, which amounted to encoding prior knowledge about the problem in a way that was no different, in spirit, from the hand-engineered features of symbolic AI. The SVM operated in a high-dimensional feature space induced by the kernel; a deep neural network, by contrast, learned its own feature space from scratch.