A Short History of AI: Rules to Generative Models
Seventy years in four eras — symbolic AI, statistical machine learning, deep learning and generative AI — and the ideas that waited decades for the hardware.
BeginnerVerdeshell Team · 10 min read · Last reviewed
The history of AI is mostly a history of where knowledge comes from. First people wrote the rules; then machines learned them from data; then deep networks learned their own features; now models learn enough to produce new content.
Key takeaways
- The term “artificial intelligence” was coined in a 1955 proposal for the 1956 Dartmouth workshop.
- AI has moved through four eras: symbolic rules, statistical machine learning, deep learning and generative AI.
- Most core ideas — neural networks, learning from feedback — are decades old; data and computing power are what changed.
- The expert-system bust of the late 1980s came from promises outrunning what the systems reliably did.
Why the history is worth ten minutes
Most of today’s vocabulary — machine learning, neural networks, even “artificial intelligence” itself — is decades old. Knowing which idea came from where makes it much easier to tell a genuine shift from a rebranding, and explains why the field has been through booms and busts before.
The thread to follow is where the system’s knowledge comes from. Every era below answers that question differently.
1950–1970s · Symbolic AI: people write the rules
In 1950 Alan Turing published “Computing Machinery and Intelligence”, replacing the question “can machines think?” with a testable game of imitation. The term artificial intelligence arrived five years later, in a 1955 proposal by John McCarthy, Marvin Minsky, Nathaniel Rochester and Claude Shannon for a summer research project at Dartmouth College, held in 1956.
The dominant approach for the next three decades was symbolic: represent knowledge as symbols and rules, and reason by manipulating them. It produced programs like ELIZA (1966), which imitated a psychotherapist by matching keywords in what you typed — convincing enough to unsettle its own creator, with no understanding at all.
A second idea ran in parallel. In 1958 Frank Rosenblatt demonstrated the perceptron, a simple artificial neuron that learned to sort inputs into two classes by correcting its own mistakes. In 1969 Minsky and Papert’s book “Perceptrons” set out the mathematical limits of single-layer networks, and interest in neural networks cooled for more than a decade.
The 1980s · Expert systems, and the bust
Rule-based AI found commercial success in expert systems: large collections of if-then rules written with specialists. The best-known, R1 — later XCON — configured computer orders for Digital Equipment Corporation from the early 1980s.
Funding had already been cut sharply in the UK after the critical 1973 Lighthill Report, and that period is often called the first “AI winter”, though historians dispute how far it went beyond Britain. The clearer winter came at the end of the 1980s, when the expert-system boom collapsed: rule bases were expensive to build, brittle outside their narrow domain and hard to maintain, and large government programmes were wound down.
The quieter event of the decade mattered more in hindsight. In 1986 David Rumelhart, Geoffrey Hinton and Ronald Williams showed that backpropagation could train networks with multiple layers to learn useful internal representations — the technique almost every modern model still depends on. It was not practical at scale for another twenty-five years.
1990s–2000s · Statistical machine learning: learn it from data
The field shifted from writing rules to learning them from examples, using statistical methods that worked well on the data and hardware available. This is the era that produced the spam filter, the recommendation engine and the fraud score — machine learning that quietly works rather than talks.
In 1997 IBM’s Deep Blue beat the reigning world chess champion, Garry Kasparov, in a match under standard tournament conditions. It is a useful counter-example: Deep Blue won by searching enormous numbers of positions with hand-tuned evaluation, not by learning. Impressive AI, very little machine learning.
2012–2017 · Deep learning: networks learn their own features
Two things changed: large labelled datasets and cheap parallel computing. The ImageNet dataset (2009) supplied the first; graphics processors supplied the second. In 2012 AlexNet — a deep convolutional network from Alex Krizhevsky, Ilya Sutskever and Hinton, trained on GPUs — won the ImageNet recognition challenge by a wide margin, and neural networks were back.
Progress followed quickly: word2vec (2013) represented words as vectors whose distances carry meaning; generative adversarial networks (2014) learned to generate new images by pitting two networks against each other; and in 2016 DeepMind’s AlphaGo, combining deep networks with search, beat Lee Sedol at Go, a game long considered out of reach.
In 2017 a Google team published “Attention Is All You Need”, introducing the Transformer: an architecture built entirely on attention, where every part of the input can weigh every other part directly. It trained efficiently in parallel, and it scaled. Nearly every large language model since is built on it.
2018 onwards · Generative AI: learned well enough to produce
Transformers pre-trained on very large amounts of text turned out to learn a great deal about language. BERT (2018) was built to understand text; OpenAI’s GPT series was built to generate it. GPT-3 (2020) showed that a large enough model could pick up a new task from a few examples written into the prompt, without any retraining.
The step that made these models usable by anyone was training them to follow instructions. OpenAI’s InstructGPT work (2022) fine-tuned a model on human-written demonstrations and on human rankings of its answers — reinforcement learning from human feedback. ChatGPT, launched as a research preview on 30 November 2022, put that kind of model in front of the public.
Recognition caught up with the researchers: the 2018 Turing Award went to Yoshua Bengio, Geoffrey Hinton and Yann LeCun for deep learning, and the 2024 Nobel Prize in Physics to John Hopfield and Hinton for foundational work on neural networks, with half of that year’s Chemistry prize going to Demis Hassabis and John Jumper for AI-based protein structure prediction.
2024 onwards · Reasoning models and agents
Two shifts define the current period. The first is reasoning models, trained to work through a problem step by step before answering — OpenAI’s o1 (September 2024) was the first widely available example, and DeepSeek released R1 with open weights in January 2025.
The second is agents: models given tools and a goal rather than a single question, which choose their own next actions. Open standards for connecting models to tools (the Model Context Protocol, November 2024) and to each other (Agent2Agent, April 2025) are what that shift is being built on. Our AI agents topic covers how those systems are structured.
What the history teaches
The ideas are older than the hype. Neural networks, learning from feedback and generating text were all proposed decades before they worked; what changed each time was data, computing power and engineering. That is a reason for caution about any claim that a technique is entirely new.
Booms end when promises outrun what systems reliably do. The expert-system bust was not caused by the technology failing to work — it worked, narrowly — but by it being sold as more general than it was. The same test applies now: ask what a system does reliably, not what it can be shown doing once.
And “AI” has always meant several things at once. A 1960s rule engine, a 2000s spam filter and a 2020s language model all fit the label. Our piece on AI, machine learning and generative AI sorts out which is which.
More in AI Basics
Want this built properly?
We design and build AI systems for clients. Tell us the problem and we will tell you honestly whether AI — and which kind — is the right fit for it.