Tech●●●●●Difficulty 4 of 5

Why does a processor keep a tiny memory right next to itself?

From 1986 to 2000, processors got 55 percent faster every year. Memory improved by only 10. Something had to give.

▶ Start the story

Moore's law kept packing more transistors onto chips, so processors got faster almost every year. Main memory did not keep up. From 1986 to 2000, CPU speed improved at about 55 percent a year, while the response time of memory sitting off the chip improved at only about 10 percent. That growing gap is called the memory wall, and it means a processor can spend much of its time simply waiting for memory to answer.

CPU speed vs. memory speed, annual growth (1986-2000)

% per year

Bar chart: CPU speed vs. memory speed, annual growth (1986-2000). (% per year)
Annual improvement
CPU speed55 % per year
Off-chip memory speed10 % per year
From 1986 to 2000, CPU speed improved about 55% a year while memory response time improved only about 10% a year, opening the memory wall.

The fix is a cache: a small, fast memory placed right next to the processor core, holding copies of data the processor has used recently. Checking the cache first means the processor often avoids a trip to main memory, which can be tens to hundreds of times slower to reach. A modern chip usually layers several of these caches: L1, closest to the core and fastest; L2, a bit farther away and slower but bigger; and often L3, shared across multiple cores.

Caches work because of a pattern called locality of reference: a processor tends to use the same memory locations again soon (temporal locality), and it tends to use memory locations near each other (spatial locality). Guess right about what's about to be needed, and the cache already has it ready.

It isn't a free upgrade. Cache memory is built from SRAM, which needs four or six transistors to hold a single bit; the DRAM of main memory needs just one transistor and a capacitor. Back in the 386 era, main memory could have latencies up to 120 nanoseconds, while SRAM cache ran at roughly 10 to 25 nanoseconds. That speed costs chip space, which is exactly why caches stay small while memory stays big.

Quiz me

0/3

  1. 1.Why does a processor use a small, fast cache instead of just making all memory fast?
  2. 2.A program reads a long list of numbers one after another. Which pattern lets the cache help most?
  3. 3.Why is L1 cache smaller than L2 or L3 cache?

Recap

Caches work because of locality of reference: processors tend to reuse the same memory locations soon (temporal locality) and ones nearby (spatial locality), so a small fast cache can usually guess right about what's needed next.

Surprising fact · From 1986 to 2000, CPU speed improved about 55 percent a year while off-chip memory improved only about 10 percent a year.

Sources (5)

No source, no claim. Every fact in this lesson (14 claims) cites at least one of these.

  1. [1]CPU cache · Wikipedia
  2. [2]Locality of reference · Wikipedia
  3. [3]Random-access memory · Wikipedia
  4. [4]Memory hierarchy · Wikipedia
  5. [5]Dynamic random-access memory · Wikipedia
More lessons in 💻 Tech (3) See all tech lessons →

One more light on your map.

Get one lesson like this every day, about the things you love. Free, in two or five minutes.

Get the share card for this lesson ↗