# 2024 Nobel Prize in Physics: Memory as a landscape: the physics behind machine learning

> The 2024 Physics prize honours the physics that underpins machine learning. John Hopfield showed that a simple network of connected nodes can store a memory as the low point of an energy landscape and recall it from a noisy or partial clue. Geoffrey Hinton extended that idea into the Boltzmann machine, a network that learns the hidden patterns in data on its own and helped launch today's deep learning.

- Laureates: John J. Hopfield (share 1/2, Princeton University, Princeton, NJ, USA); Geoffrey Hinton (share 1/2, University of Toronto, Toronto, Canada)
- Official citation: “for foundational discoveries and inventions that enable machine learning with artificial neural networks”
- Course page: https://nobelexplained.com/physics/2024-machine-learning
- Last updated: 2026-06-21

## Explained simply (ELI5)

Imagine a hilly landscape with a few deep valleys. Roll a ball anywhere on it and the ball runs downhill until it settles at the bottom of the nearest valley.

John Hopfield showed that a network of tiny connected switches can work the same way. Each thing you want it to remember, like a picture, becomes its own valley. If you then hand the network a smudged or half-missing version of that picture, it acts like the rolling ball. It keeps adjusting its switches to move downhill until it reaches the nearest valley, which is the clean memory. That is how it fills in and completes the clue.

> **Remembering is rolling downhill** A memory is stored as the bottom of a valley. Give the network a noisy hint and it slides down to the closest valley, recalling the full pattern. No searching through a list is needed.

Geoffrey Hinton took this further. He built a network that studies many examples and quietly learns the patterns hiding inside them, so it can even make new examples of its own. These physics ideas about energy and chance are the seeds of the machine learning we use every day.

**Video transcript (35 seconds):** John Hopfield grew up without a TV. Both his parents were physicists. Years later, he pictured memory as a hilly landscape. A ball dropped on a slope rolls down into the nearest valley. His computer network kept each picture in its own valley. Shown a smudged picture, it rolled downhill to the clean one. Geoffrey Hinton built on that idea. His machine learned from lots of examples and could even make up new ones. The two scientists shared the 2024 Nobel Prize in Physics.

## Explained for undergraduates

A **Hopfield network** is a single layer of nodes where every node connects to every other node. Each node holds a value of on or off, and each connection has a weight. Crucially the weights are symmetric: the link from node i to node j is as strong as the link from j to i.

Hopfield borrowed an idea from the physics of magnets. He gave the whole network a single number, its **energy**, computed from every node value and every connection strength. Stored patterns are arranged to sit at low-energy points. To recall a memory, the network visits its nodes one at a time and flips any node whose flip would lower the total energy. Step by step the energy falls and the state settles into the nearest low point, the stored pattern closest to the starting clue.

> **Hebbian weights carve the valleys** Patterns are written into the weights with a Hebbian rule: nodes that should be on together get a positive connection, nodes that disagree get a negative one. Each stored pattern then becomes a local minimum of the energy, a valley the network can roll into. Add too many patterns and the valleys start to merge, which sets a limit on how much one network can remember.

Hinton asked a different question: instead of storing fixed patterns, could a network learn the structure of data by itself? His **Boltzmann machine** adds **hidden nodes** that are not tied to the input. Using the Boltzmann distribution from statistical physics, where low-energy states are the most likely, it adjusts its weights until the patterns it generates on its own match the patterns it was shown. It can then classify images or create new examples of the kind it learned.

- **1982**: Hopfield publishes the associative-memory network, using just 30 nodes and fewer than 500 connections.
- **1985**: Hinton and Terrence Sejnowski introduce the Boltzmann machine, a network that learns from examples.
- **1986**: Hinton and colleagues popularise backpropagation, training deeper networks efficiently.
- **2012**: Hinton's group wins the ImageNet contest with a deep network, sparking the modern boom.
- **2024**: Hopfield and Hinton share the Nobel Prize in Physics.

**Video transcript (62 seconds):** Around 1965, a teenage Geoffrey Hinton heard a school friend say memories might spread across the brain like a hologram. He has thought about the brain ever since. In 1982, John Hopfield borrowed from magnets, where each atom's spin nudges its neighbours. His network gives each state an energy, with stored memories at the lowest points. Hinton and Terrence Sejnowski added hidden nodes and Ludwig Boltzmann's equation. Their 1985 Boltzmann machine learned from examples and generated new patterns. In the 1990s, many researchers lost interest in neural networks, but Hinton kept going. In 2006, his team stacked Boltzmann machines to pretrain deep networks. Hopfield's first network had 30 nodes. Today's language models hold over a trillion parameters. Both shared the 2024 Nobel Prize in Physics.

## Explained for experts

The Hopfield network is a recurrent system of N binary units si taking values -1 or +1, with symmetric weights wij = wji and no self-connection. Its state is scored by an energy function taken straight from the Ising model of interacting spins: **E = -1/2 Σ wij si sj**. Each asynchronous update sets a unit to the sign of its local field, a step that can never raise E, so the network slides down a fixed energy surface and halts at a local minimum.

> **Hebbian storage and the 0.138N limit** To store P patterns, the weights are set by the Hebbian outer-product rule w_ij = (1/N) Σ ξ_i ξ_j summed over the patterns. Each pattern becomes an attractor, a basin in the energy landscape. Because spurious minima and pattern overlap accumulate, a standard Hopfield network reliably stores only about 0.138N patterns before recall breaks down, a capacity later lifted enormously by dense associative memories.

Hopfield's contribution reframed memory as a dynamical-systems problem: content-addressable recall becomes gradient descent on E, and robustness to noise is the size of each attractor's basin. The same machinery links neural computation to spin glasses, which is why the prize sits in physics rather than computer science.

Hinton's **Boltzmann machine** turns the energy idea into a generative model. Units are split into visible v and hidden h, and a configuration has probability p(v,h) = exp(-E(v,h)) / Z, the Boltzmann distribution, with Z the partition function. Training maximises the likelihood of the data; the gradient is a difference between two correlations, one measured while the data is clamped on the visible units and one measured while the model runs freely. This contrast nudges the weights until the model's own samples resemble the training set.

> **Restricted Boltzmann machines and deep learning** A general Boltzmann machine trains slowly. Restricting it so hidden units connect only to visible units, with no links inside a layer, gives the restricted Boltzmann machine, whose conditional independence makes learning fast. Stacking these layer by layer, then fine-tuning, was an early recipe for training deep networks and helped ignite the deep learning era.

> “Hopfield likened searching the network for a saved state to rolling a ball through a landscape of peaks and valleys, with friction that slows its movement.”
> (The Royal Swedish Academy of Sciences, 2024)

**Video transcript (66 seconds):** Before dawn, in a cheap California hotel, Geoffrey Hinton took a call from Stockholm. He first suspected a spoof. Several Swedish accents convinced him. The shared prize goes back to John Hopfield in 1982. His network pairs binary nodes and symmetric Hebbian weights with an energy borrowed from magnetism, which can only fall as nodes update. Memories are local minima, and corrupted inputs settle into the nearest one. Hinton and Terrence Sejnowski made it stochastic. A state's probability falls exponentially with its energy, and hidden nodes let the Boltzmann machine learn distributions. But each learning step needed slow equilibrium simulations. The restricted version links visible nodes only to hidden ones, and contrastive divergence trains it fast. In 2006, Hinton and colleagues stacked them to pretrain deep networks before backpropagation fine-tuning. In a randomised screening study, machine learning clearly improved breast cancer detection on mammograms.

## Think first

**Q:** A network is shown a blurry, half-erased photo of a face it has seen before. With no database lookup, it cleans the image up and returns the original. How can a web of simple on/off nodes do that?

**A:** **It treats memory as a downhill roll.** Each stored pattern sits at the bottom of a valley in an energy landscape. The blurry photo starts partway up a slope, and the network keeps flipping nodes to lower its energy, so the state slides down into the nearest valley. The valley it lands in is the stored pattern most similar to the clue, which is the cleaned-up face.

**Q:** You never tell the network what a cat is. You just show it thousands of pictures. Later it can sketch a brand new cat-like image on its own. What did it actually learn?

**A:** **The statistics of the data, not a list of rules.** A Boltzmann machine has hidden nodes that are not pinned to the picture. By adjusting its connection strengths until the patterns it tends to produce match the patterns it was shown, it captures the recurring features of cats. Because those features now live in its weights, it can generate fresh examples that share them.

## A trillion parameters grew from fewer than 500

Hopfield's original 1982 network had 30 nodes and fewer than 500 connections to adjust, simple enough to run on the computers of the day. The large language models built on the same basic idea now juggle more than a trillion parameters, a jump of more than a billionfold in barely four decades.

## Key terms

- **Hopfield network**: A recurrent network where every node connects to every other with symmetric weights. It stores patterns as low-energy states and recalls them by settling into the nearest one.
- **Energy function**: A single number for the whole network, E = -1/2 times the sum of w_ij s_i s_j over all node pairs, borrowed from the physics of magnetic spins. Stored patterns sit at its lowest points.
- **Associative memory**: A memory addressed by content rather than by location. Given part of a pattern, it returns the complete stored pattern that best matches.
- **Hebbian learning**: A rule for setting connection weights so that nodes which should be active together are linked positively. It carves each stored pattern into a valley of the energy landscape.
- **Boltzmann machine**: A network with hidden units that learns the statistical structure of data using the Boltzmann distribution, and can generate new examples like those it was trained on.
- **Hidden units**: Nodes in a Boltzmann machine that are not fixed to the input. They let the network represent features that are not directly given in the data.

## Check yourself

1. In a Hopfield network, what does a stored memory correspond to?
   - A peak in the energy landscape
   - **A valley, a local minimum, in the energy landscape** (correct)
   - A single special node
   - The strongest connection
   - Why: Each stored pattern is written into the weights so that it sits at a local minimum of the network's energy. Recall is the network rolling downhill into the nearest such valley.

2. You feed a Hopfield network a noisy, partial version of a stored pattern. What happens?
   - It raises its energy until the noise grows
   - **It flips nodes to lower its energy until it settles in the closest stored pattern** (correct)
   - It deletes the pattern
   - Nothing, it needs an exact match
   - Why: The update rule only ever lowers the total energy, so the state slides down to the nearest minimum, which is the stored pattern most similar to the noisy clue. That is how the network completes and cleans up the input.

3. What did Hinton's Boltzmann machine add beyond the Hopfield network?
   - It removed all the weights
   - **It used hidden units and statistical physics to learn patterns in data and generate new examples** (correct)
   - It made the network smaller
   - It only worked on text
   - Why: The Boltzmann machine introduces hidden units and uses the Boltzmann distribution, training the weights until the network's own samples match the data. This lets it discover features on its own and generate new examples, not just store fixed memories.

## The laureates

### John J. Hopfield

Born in Chicago in 1933, Hopfield was a physicist who moved into biology at Caltech. In 1982 he showed that a network of simple connected nodes can act as an associative memory, storing patterns as low points of an energy landscape and recalling them from noisy or partial clues.

### Geoffrey Hinton

Born in London in 1947, Hinton is a computer scientist long based at the University of Toronto. He built the Boltzmann machine on top of Hopfield's idea, using statistical physics to let a network learn the structure of data by itself, work that helped start the modern growth of machine learning.

## Sources

- [The Nobel Prize in Physics 2024, popular science background (NobelPrize.org)](https://www.nobelprize.org/prizes/physics/2024/popular-information/)
- [The Nobel Prize in Physics 2024, press release (NobelPrize.org)](https://www.nobelprize.org/prizes/physics/2024/press-release/)
- [Hopfield network (Wikipedia)](https://en.wikipedia.org/wiki/Hopfield_network)
- [Boltzmann machine (Wikipedia)](https://en.wikipedia.org/wiki/Boltzmann_machine)

---

From Nobel Explained (https://nobelexplained.com), a free, independent course on every Nobel Prize since 2020. Not affiliated with the Nobel Foundation. Facts are pinned from the official Nobel Prize API.
