MCS 510 · Advanced Computer Architecture

From the ISA to the energy bill

How processors are organized, how today's designs grew out of sixty years of trade-offs, and how to reason quantitatively about speed and efficiency. Each part ends with tools to try and papers to read.

Part A

How a computer is organized

Layers of abstraction

A computing system is a stack of contracts. Each layer hides the one below it, and architecture lives at the boundary where software stops and hardware begins. The instruction set architecture (ISA) is that boundary: the registers, instructions, data types, addressing modes and memory model a program can rely on. The microarchitecture is one particular way of implementing it. x86-64 has been implemented by dozens of very different microarchitectures; the same binary runs on all of them.

Application / algorithmWhat the problem needs: data structures, access patterns, parallelism
Language / compilerTurns source into instructions; scheduling, vectorization, register allocation
Operating systemVirtual memory, privilege levels, scheduling across cores
ISAThe hardware/software contract: x86-64, Arm64, RISC-V
MicroarchitecturePipeline depth, caches, predictors, issue width, number of cores
Logic / circuitsGates, flip-flops, SRAM cells, clock distribution
DevicesTransistors (FinFET, gate-all-around), interconnect, packaging

The classic organization is the von Neumann model: a single memory holds both instructions and data, and a processor fetches, decodes and executes one instruction at a time. Its weakness, the shared path between processor and memory, is still called the von Neumann bottleneck. A Harvard organization separates instruction and data paths; modern cores are effectively a hybrid, with split L1 instruction and data caches backed by a unified memory.

ISAs differ in philosophy. CISC designs (x86) offer variable-length instructions that can operate on memory directly. RISC designs (Arm, RISC-V, MIPS) use fixed-length, load/store instructions that are simple to decode and pipeline. Modern x86 cores decode CISC instructions into RISC-like micro-operations internally, so the distinction now matters most for decoder complexity, code density and licensing, rather than raw speed.

Datapath and pipelining

The datapath is the hardware that moves and transforms data (register file, ALU, memory ports); the control unit tells it what to do each cycle. Pipelining splits instruction execution into stages so several instructions are in flight at once, like an assembly line. The textbook five-stage RISC pipeline is Fetch, Decode, Execute, Memory access and Write-back.

Instructionc1c2c3c4c5c6c7c8
lw x1, 0(x2)IFIDEXMEMWB
add x3, x1, x4IFIDstallEXMEMWB
sub x5, x6, x7IFstallIDEXMEMWB

A load-use hazard: add needs x1 before the load has read memory, so even with forwarding the pipeline inserts one bubble.

Pipelines are limited by hazards. Structural hazards occur when two instructions need the same hardware. Data hazards occur when an instruction needs a result that is not ready yet; forwarding (bypassing) solves most of them. Control hazards occur at branches, because the next instruction to fetch is unknown until the branch resolves. Deeper pipelines allow faster clocks but make every misprediction and stall more expensive, which is one reason clock frequencies stopped climbing in the mid-2000s.

Memory hierarchy

Processors got faster much more quickly than DRAM, a gap known as the memory wall. Caches close the gap by exploiting temporal locality (recently used data is reused) and spatial locality (nearby data is used soon). Each level is larger, slower and cheaper per byte than the one above it.

Register
~0.3 ns
L1 cache
~1 ns
L2 cache
~4 ns
L3 cache
~12 ns
DRAM
~80 ns
NVMe SSD
~20 µs
Disk seek
~5 ms

Typical orders of magnitude for a current desktop or server part; bar length is logarithmic. Exact values vary by product.

Cache design comes down to a few parameters: capacity, block (line) size, associativity, replacement policy, and write policy (write-through or write-back). Misses fall into the "three Cs": compulsory, capacity and conflict, with a fourth, coherence, on multicore chips. Virtual memory adds a translation layer: page tables map virtual to physical addresses, and the TLB caches those translations so most accesses skip the page-table walk.

Instruction-level parallelism

A single core finds parallelism inside one instruction stream. Superscalar cores issue several instructions per cycle. Out-of-order execution lets an instruction run as soon as its operands are ready, using register renaming to remove false dependences and a reorder buffer to retire results in program order. The core idea traces back to Tomasulo's algorithm for the IBM System/360 Model 91.

Branch prediction and speculative execution keep the pipeline full by guessing the outcome of branches and running ahead. Modern predictors are right well over 95% of the time on typical code. Speculation has a cost beyond wasted work: Spectre and Meltdown showed that squashed instructions can still leave measurable traces in caches, turning a performance feature into a security vulnerability.

Parallel and specialized architectures

Flynn classMeaningWhere you see it today
SISDSingle instruction, single dataA conceptual scalar core
SIMDOne instruction applied to many data elementsAVX-512, Arm SVE, RISC-V V; GPU warps (SIMT)
MISDMany instructions on one data streamRare; some fault-tolerant systems
MIMDIndependent instruction streams on independent dataMulticore CPUs, clusters, supercomputers

Multicore chips share memory, so they need cache coherence (snooping protocols such as MESI, or directories at larger scale) and a defined memory consistency model that says which orderings of loads and stores other cores may observe. x86 provides fairly strong ordering (TSO); Arm and RISC-V are weaker and rely on explicit fences.

GPUs trade single-thread speed for massive throughput: thousands of lightweight threads hide memory latency by switching among themselves. Domain-specific accelerators such as Google's TPU go further, building systolic arrays of multiply-accumulate units with low-precision arithmetic and software-managed memory. Heterogeneous systems now combine big and little CPU cores, GPUs, neural engines and media blocks on one die or across chiplets connected in one package.

Part B

How we got here

Timeline

  1. 1945
    EDVAC report

    Von Neumann's draft describes the stored-program computer, with instructions held in the same memory as data.

  2. 1964
    IBM System/360

    The first family of machines sharing one ISA across a wide price range, establishing the separation of architecture from implementation and binary compatibility as a business strategy.

  3. 1965
    Moore's observation

    Transistor counts per chip double on a regular cadence, later settled at roughly every two years.

  4. 1967
    Tomasulo's algorithm

    Dynamic scheduling on the System/360 Model 91 introduces ideas still at the heart of out-of-order cores.

  5. 1971
    Intel 4004

    A complete CPU on one chip; the microprocessor era begins.

  6. 1974
    Dennard scaling

    Shrinking transistors keeps power density constant, so each generation gets faster and denser for free.

  7. 1976
    Cray-1

    Vector registers and pipelined functional units define supercomputing for a decade.

  8. 1980s
    RISC

    Berkeley RISC and Stanford MIPS show that simple, pipelinable instructions plus good compilers beat complex ISAs. Arm and SPARC follow.

  9. 1995
    Pentium Pro

    x86 instructions are translated into internal micro-ops and run out of order, blending CISC compatibility with RISC execution.

  10. 2004
    The power wall

    Dennard scaling breaks down as leakage grows. Clock rates plateau and vendors pivot to multicore designs.

  11. 2007
    CUDA

    General-purpose GPU programming goes mainstream, setting up the deep-learning hardware boom.

  12. 2010
    RISC-V

    An open, royalty-free ISA begins at Berkeley and becomes a base for research, embedded and custom silicon.

  13. 2015
    Google TPU

    A domain-specific accelerator for neural-network inference enters production data centers.

  14. 2018
    Spectre and Meltdown

    Speculative execution leaks data across security boundaries; security becomes a first-class architectural constraint.

  15. 2019
    Chiplets go mainstream

    AMD's Zen 2 splits CPUs into smaller dies in one package, improving yield and letting each die use the best process node.

  16. 2020
    Apple M1

    A wide Arm core, unified memory and integrated accelerators show the efficiency of tightly integrated SoCs in laptops.

  17. 2022
    Exascale

    Frontier passes 1018 floating-point operations per second on HPL, built mainly from GPUs.

What history left us

Compatibility is sticky. The System/360 decision to fix an ISA and vary the implementation is why x86 and Arm binaries outlive the chips they were compiled for, and why new ISAs succeed mainly in new markets.

RISC won inside CISC. The RISC argument for simple, pipelinable operations was absorbed by x86 through micro-op translation. Today's debate is less RISC versus CISC and more open versus licensed ISAs.

The free lunch ended. With Dennard scaling gone, performance has to come from parallelism and specialization rather than clock speed. That shift explains multicore chips, GPUs in data centers, and dark silicon: parts of a chip that cannot all be powered at once.

Specialization is the new scaling. Hennessy and Patterson call this a new golden age: domain-specific architectures, domain-specific languages, open hardware and agile chip design, all driven by the end of general-purpose scaling.

Data movement dominates. Moving data now costs far more time and energy than computing on it, pushing designs toward larger caches, stacked high-bandwidth memory, chiplets and processing near or in memory.

Part C

Performance and efficiency

The iron law of processor performance

CPU time = Instruction count × CPI × Clock cycle time

Each term belongs to different layers of the stack. The algorithm, compiler and ISA set the instruction count. The microarchitecture and memory system set cycles per instruction (CPI). The circuit technology and pipeline depth set the clock. Improving one term often worsens another, which is why a single metric like GHz or MIPS misleads.

Worked example. A program executes 2.0×109 instructions at CPI 1.25 on a 3 GHz core: 2.0×109 × 1.25 / 3×109 = 0.833 s. A compiler change cuts instructions by 10% but raises CPI to 1.40: 1.8×109 × 1.40 / 3×109 = 0.840 s. Fewer instructions, slower program.

AMAT = Hit time + Miss rate × Miss penalty

Average memory access time applies the same idea to caches. With a 1 ns L1, 5% L1 miss rate, a 4 ns L2 that misses 20% of the time, and 80 ns DRAM: AMAT = 1 + 0.05 × (4 + 0.2 × 80) = 2.0 ns.

Amdahl and Gustafson

Speedup = 1 / ((1 − p) + p / n)

Amdahl's law bounds the speedup from improving a fraction p of the work by a factor n. The serial part sets a hard ceiling of 1/(1 − p) no matter how many cores you add. Gustafson's law gives the complementary view: as machines grow, people solve bigger problems, so the parallel fraction grows with them. Hill and Marty applied Amdahl's law to multicore chip design and showed why asymmetric big-plus-little cores can beat both extremes.

6.40× speedup
Ceiling with infinite cores: 10.0×. Efficiency 40%.
Speedup versus cores (log scale) for the current p; the dashed line is the Amdahl ceiling.

Roofline

Attainable FLOP/s = min(Peak FLOP/s, Memory bandwidth × Arithmetic intensity)

Arithmetic intensity is floating-point operations per byte moved from memory. Kernels left of the ridge point are memory-bound and gain nothing from more compute; kernels to the right are compute-bound. The model tells you which optimization to try: improve data reuse (blocking, fusion) to move right, or vectorize and parallelize to climb toward the roof.

200 GFLOP/s attainable
Memory-bound. Ridge point at 10 FLOP/byte.
Log-log roofline; the dot is your kernel.

Power and energy

Pdynamic ≈ α · C · V² · f

Dynamic power scales with activity factor, switched capacitance, the square of supply voltage, and frequency. Lowering voltage and frequency together (DVFS) cuts power roughly with the cube of frequency, which is why many slower cores can beat one fast core on throughput per watt. Static (leakage) power flows even when transistors are idle and is what ended Dennard scaling.

Operation (45 nm)EnergyRelative to 32-bit add
8-bit integer add0.03 pJ0.3×
32-bit integer add0.1 pJ1×
32-bit float multiply3.7 pJ37×
32-bit read, 8 KB SRAM5 pJ50×
32-bit read, DRAM640 pJ6,400×

Figures from Horowitz (ISSCC 2014). Newer nodes lower every value, but the ratio between compute and off-chip access remains in the hundreds to thousands.

Measuring honestly

Use execution time of real workloads as the ground truth. Standard suites such as SPEC CPU, MLPerf and HPL make results comparable; summarize normalized ratios with the geometric mean. Report efficiency alongside speed (performance per watt, per dollar, per mm²), and for services, report tail latency (p99) as well as averages. Be wary of peak numbers, MIPS, clock rate alone, and benchmarks tuned for one machine.

Go further

Resources

Interactive tools and simulators

Peer-reviewed reading

  • Hennessy, J. L., & Patterson, D. A. (2019). A new golden age for computer architecture. Communications of the ACM, 62(2), 48–60.Turing Award lecture; the best single overview of where the field is heading.doi.org/10.1145/3282307
  • von Neumann, J. (1993). First draft of a report on the EDVAC. IEEE Annals of the History of Computing, 15(4), 27–75.The 1945 stored-program design, reprinted with commentary.doi.org/10.1109/85.238389
  • Moore, G. E. (1998). Cramming more components onto integrated circuits. Proceedings of the IEEE, 86(1), 82–85.Reprint of the 1965 paper behind Moore's law.doi.org/10.1109/JPROC.1998.658762
  • Dennard, R. H., et al. (1974). Design of ion-implanted MOSFET's with very small physical dimensions. IEEE Journal of Solid-State Circuits, 9(5), 256–268.The scaling rules whose breakdown ended rising clock speeds.doi.org/10.1109/JSSC.1974.1050511
  • Flynn, M. J. (1972). Some computer organizations and their effectiveness. IEEE Transactions on Computers, C-21(9), 948–960.Origin of the SISD/SIMD/MISD/MIMD taxonomy.doi.org/10.1109/TC.1972.5009071
  • Tomasulo, R. M. (1967). An efficient algorithm for exploiting multiple arithmetic units. IBM Journal of Research and Development, 11(1), 25–33.The foundation of dynamic scheduling and register renaming.doi.org/10.1147/rd.111.0025
  • Patterson, D. A., & Ditzel, D. R. (1980). The case for the reduced instruction set computer. ACM SIGARCH Computer Architecture News, 8(6), 25–33.The short paper that launched the RISC movement.doi.org/10.1145/641914.641917
  • Blem, E., Menon, J., & Sankaralingam, K. (2013). Power struggles: Revisiting the RISC vs. CISC debate on contemporary ARM and x86 architectures. Proc. IEEE HPCA, 1–12.Measured evidence that ISA matters less than microarchitecture.doi.org/10.1109/HPCA.2013.6522302
  • Wulf, W. A., & McKee, S. A. (1995). Hitting the memory wall: Implications of the obvious. ACM SIGARCH Computer Architecture News, 23(1), 20–24.Coined the memory wall.doi.org/10.1145/216585.216588
  • Amdahl, G. M. (1967). Validity of the single processor approach to achieving large scale computing capabilities. Proc. AFIPS Spring Joint Computer Conference, 483–485.The original statement of Amdahl's law.doi.org/10.1145/1465482.1465560
  • Gustafson, J. L. (1988). Reevaluating Amdahl's law. Communications of the ACM, 31(5), 532–533.Scaled speedup when problem size grows with the machine.doi.org/10.1145/42411.42415
  • Hill, M. D., & Marty, M. R. (2008). Amdahl's law in the multicore era. IEEE Computer, 41(7), 33–38.Symmetric, asymmetric and dynamic multicore trade-offs.doi.org/10.1109/MC.2008.209
  • Williams, S., Waterman, A., & Patterson, D. (2009). Roofline: An insightful visual performance model for multicore architectures. Communications of the ACM, 52(4), 65–76.The roofline model used above.doi.org/10.1145/1498765.1498785
  • Asanović, K., et al. (2009). A view of the parallel computing landscape. Communications of the ACM, 52(10), 56–67.Berkeley's analysis of the shift to parallelism after the power wall.doi.org/10.1145/1562764.1562783
  • Esmaeilzadeh, H., et al. (2011). Dark silicon and the end of multicore scaling. Proc. ISCA, 365–376.Why power limits keep much of a chip idle.doi.org/10.1145/2000064.2000108
  • Horowitz, M. (2014). Computing's energy problem (and what we can do about it). Proc. IEEE ISSCC, 10–14.Source of the energy-per-operation table.doi.org/10.1109/ISSCC.2014.6757323
  • Jouppi, N. P., et al. (2017). In-datacenter performance analysis of a tensor processing unit. Proc. ISCA, 1–12.The case study for domain-specific accelerators.doi.org/10.1145/3079856.3080246
  • Dally, W. J., Turakhia, Y., & Han, S. (2020). Domain-specific hardware accelerators. Communications of the ACM, 63(7), 48–57.Where accelerator gains really come from.doi.org/10.1145/3361682
  • Leiserson, C. E., et al. (2020). There's plenty of room at the Top: What will drive computer performance after Moore's law? Science, 368(6495), eaam9744.Performance from software, algorithms and hardware architecture.doi.org/10.1126/science.aam9744
  • Thompson, N. C., & Spanuth, S. (2021). The decline of computers as a general purpose technology. Communications of the ACM, 64(3), 64–72.The economics behind specialization.doi.org/10.1145/3430936
  • Kocher, P., et al. (2019). Spectre attacks: Exploiting speculative execution. Proc. IEEE Symposium on Security and Privacy, 1–19.How speculation became an attack surface.doi.org/10.1109/SP.2019.00002
  • Kim, Y., et al. (2014). Flipping bits in memory without accessing them: An experimental study of DRAM disturbance errors. Proc. ISCA, 361–372.The Rowhammer paper; reliability and security at the memory level.doi.org/10.1109/ISCA.2014.6853210