Skip to main content

A Short History of AI: Rules to Generative Models

Seventy years in four eras — symbolic AI, statistical machine learning, deep learning and generative AI — and the ideas that waited decades for the hardware.

Verdeshell Team · 10 min read · Last reviewed

The history of AI is mostly a history of where knowledge comes from. First people wrote the rules; then machines learned them from data; then deep networks learned their own features; now models learn enough to produce new content.

A short history of artificial intelligence, from symbolic AI to generative AIA timeline in four eras. Symbolic AI: the term artificial intelligence is coined in a 1955 proposal for the 1956 Dartmouth workshop; rule-based and expert systems in the 1970s and 80s, whose commercial bubble burst in the late 1980s — the clearest "AI winter"; IBM Deep Blue beats the world chess champion in 1997 using search rather than learning. Statistical machine learning through the 1990s and 2000s. Deep learning from AlexNet in 2012 and the Transformer architecture in 2017. Generative AI from ChatGPT in 2022, and reasoning models and agents from 2024.Symbolic AIStatistical MLDeep learningGenerative AI1956Dartmouthterm coined in 19551970s–80sRules and expertsthen a late-80s bust1997Deep Blue winssearch, not learning1990s–2000sLearning from datastatistical ML2012AlexNetdeep learning takes off2017The Transformerbuilt on attention2022ChatGPTGenAI goes mainstream2024–Reasoning, agentsmodels that plan and actSpacing is by era, not to scale
Four eras, each defined by how the system gets its knowledge: written by people, learned from data, learned in layers, and learned well enough to generate.

Why the history is worth ten minutes

Most of today’s vocabulary — machine learning, neural networks, even “artificial intelligence” itself — is decades old. Knowing which idea came from where makes it much easier to tell a genuine shift from a rebranding, and explains why the field has been through booms and busts before.

The thread to follow is where the system’s knowledge comes from. Every era below answers that question differently.

1950–1970s · Symbolic AI: people write the rules

In 1950 Alan Turing published “Computing Machinery and Intelligence”, replacing the question “can machines think?” with a testable game of imitation. The term artificial intelligence arrived five years later, in a 1955 proposal by John McCarthy, Marvin Minsky, Nathaniel Rochester and Claude Shannon for a summer research project at Dartmouth College, held in 1956.

The dominant approach for the next three decades was symbolic: represent knowledge as symbols and rules, and reason by manipulating them. It produced programs like ELIZA (1966), which imitated a psychotherapist by matching keywords in what you typed — convincing enough to unsettle its own creator, with no understanding at all.

A second idea ran in parallel. In 1958 Frank Rosenblatt demonstrated the perceptron, a simple artificial neuron that learned to sort inputs into two classes by correcting its own mistakes. In 1969 Minsky and Papert’s book “Perceptrons” set out the mathematical limits of single-layer networks, and interest in neural networks cooled for more than a decade.

The 1980s · Expert systems, and the bust

Rule-based AI found commercial success in expert systems: large collections of if-then rules written with specialists. The best-known, R1 — later XCON — configured computer orders for Digital Equipment Corporation from the early 1980s.

Funding had already been cut sharply in the UK after the critical 1973 Lighthill Report, and that period is often called the first “AI winter”, though historians dispute how far it went beyond Britain. The clearer winter came at the end of the 1980s, when the expert-system boom collapsed: rule bases were expensive to build, brittle outside their narrow domain and hard to maintain, and large government programmes were wound down.

The quieter event of the decade mattered more in hindsight. In 1986 David Rumelhart, Geoffrey Hinton and Ronald Williams showed that backpropagation could train networks with multiple layers to learn useful internal representations — the technique almost every modern model still depends on. It was not practical at scale for another twenty-five years.

1990s–2000s · Statistical machine learning: learn it from data

The field shifted from writing rules to learning them from examples, using statistical methods that worked well on the data and hardware available. This is the era that produced the spam filter, the recommendation engine and the fraud score — machine learning that quietly works rather than talks.

In 1997 IBM’s Deep Blue beat the reigning world chess champion, Garry Kasparov, in a match under standard tournament conditions. It is a useful counter-example: Deep Blue won by searching enormous numbers of positions with hand-tuned evaluation, not by learning. Impressive AI, very little machine learning.

2012–2017 · Deep learning: networks learn their own features

Two things changed: large labelled datasets and cheap parallel computing. The ImageNet dataset (2009) supplied the first; graphics processors supplied the second. In 2012 AlexNet — a deep convolutional network from Alex Krizhevsky, Ilya Sutskever and Hinton, trained on GPUs — won the ImageNet recognition challenge by a wide margin, and neural networks were back.

Progress followed quickly: word2vec (2013) represented words as vectors whose distances carry meaning; generative adversarial networks (2014) learned to generate new images by pitting two networks against each other; and in 2016 DeepMind’s AlphaGo, combining deep networks with search, beat Lee Sedol at Go, a game long considered out of reach.

In 2017 a Google team published “Attention Is All You Need”, introducing the Transformer: an architecture built entirely on attention, where every part of the input can weigh every other part directly. It trained efficiently in parallel, and it scaled. Nearly every large language model since is built on it.

2018 onwards · Generative AI: learned well enough to produce

Transformers pre-trained on very large amounts of text turned out to learn a great deal about language. BERT (2018) was built to understand text; OpenAI’s GPT series was built to generate it. GPT-3 (2020) showed that a large enough model could pick up a new task from a few examples written into the prompt, without any retraining.

The step that made these models usable by anyone was training them to follow instructions. OpenAI’s InstructGPT work (2022) fine-tuned a model on human-written demonstrations and on human rankings of its answers — reinforcement learning from human feedback. ChatGPT, launched as a research preview on 30 November 2022, put that kind of model in front of the public.

Recognition caught up with the researchers: the 2018 Turing Award went to Yoshua Bengio, Geoffrey Hinton and Yann LeCun for deep learning, and the 2024 Nobel Prize in Physics to John Hopfield and Hinton for foundational work on neural networks, with half of that year’s Chemistry prize going to Demis Hassabis and John Jumper for AI-based protein structure prediction.

2024 onwards · Reasoning models and agents

Two shifts define the current period. The first is reasoning models, trained to work through a problem step by step before answering — OpenAI’s o1 (September 2024) was the first widely available example, and DeepSeek released R1 with open weights in January 2025.

The second is agents: models given tools and a goal rather than a single question, which choose their own next actions. Open standards for connecting models to tools (the Model Context Protocol, November 2024) and to each other (Agent2Agent, April 2025) are what that shift is being built on. Our building track covers how those systems are structured.

What the history teaches

The ideas are older than the hype. Neural networks, learning from feedback and generating text were all proposed decades before they worked; what changed each time was data, computing power and engineering. That is a reason for caution about any claim that a technique is entirely new.

Booms end when promises outrun what systems reliably do. The expert-system bust was not caused by the technology failing to work — it worked, narrowly — but by it being sold as more general than it was. The same test applies now: ask what a system does reliably, not what it can be shown doing once.

And “AI” has always meant several things at once. A 1960s rule engine, a 2000s spam filter and a 2020s language model all fit the label. Our piece on AI, machine learning and generative AI sorts out which is which.

Next in the pathAI, Machine Learning and Generative AI: What the Words Mean

Want this built properly?

We design and run these systems for clients. Tell us the problem and we will tell you whether an agent is the right shape for it.