How a computer is organized
Layers of abstraction
A computing system is a stack of contracts. Each layer hides the one below it, and architecture lives at the boundary where software stops and hardware begins. The instruction set architecture (ISA) is that boundary: the registers, instructions, data types, addressing modes and memory model a program can rely on. The microarchitecture is one particular way of implementing it. x86-64 has been implemented by dozens of very different microarchitectures; the same binary runs on all of them.
The classic organization is the von Neumann model: a single memory holds both instructions and data, and a processor fetches, decodes and executes one instruction at a time. Its weakness, the shared path between processor and memory, is still called the von Neumann bottleneck. A Harvard organization separates instruction and data paths; modern cores are effectively a hybrid, with split L1 instruction and data caches backed by a unified memory.
ISAs differ in philosophy. CISC designs (x86) offer variable-length instructions that can operate on memory directly. RISC designs (Arm, RISC-V, MIPS) use fixed-length, load/store instructions that are simple to decode and pipeline. Modern x86 cores decode CISC instructions into RISC-like micro-operations internally, so the distinction now matters most for decoder complexity, code density and licensing, rather than raw speed.
Datapath and pipelining
The datapath is the hardware that moves and transforms data (register file, ALU, memory ports); the control unit tells it what to do each cycle. Pipelining splits instruction execution into stages so several instructions are in flight at once, like an assembly line. The textbook five-stage RISC pipeline is Fetch, Decode, Execute, Memory access and Write-back.
| Instruction | c1 | c2 | c3 | c4 | c5 | c6 | c7 | c8 |
|---|---|---|---|---|---|---|---|---|
| lw x1, 0(x2) | IF | ID | EX | MEM | WB | |||
| add x3, x1, x4 | IF | ID | stall | EX | MEM | WB | ||
| sub x5, x6, x7 | IF | stall | ID | EX | MEM | WB |
A load-use hazard: add needs x1 before the load has read memory, so even with forwarding the pipeline inserts one bubble.
Pipelines are limited by hazards. Structural hazards occur when two instructions need the same hardware. Data hazards occur when an instruction needs a result that is not ready yet; forwarding (bypassing) solves most of them. Control hazards occur at branches, because the next instruction to fetch is unknown until the branch resolves. Deeper pipelines allow faster clocks but make every misprediction and stall more expensive, which is one reason clock frequencies stopped climbing in the mid-2000s.
Memory hierarchy
Processors got faster much more quickly than DRAM, a gap known as the memory wall. Caches close the gap by exploiting temporal locality (recently used data is reused) and spatial locality (nearby data is used soon). Each level is larger, slower and cheaper per byte than the one above it.
Typical orders of magnitude for a current desktop or server part; bar length is logarithmic. Exact values vary by product.
Cache design comes down to a few parameters: capacity, block (line) size, associativity, replacement policy, and write policy (write-through or write-back). Misses fall into the "three Cs": compulsory, capacity and conflict, with a fourth, coherence, on multicore chips. Virtual memory adds a translation layer: page tables map virtual to physical addresses, and the TLB caches those translations so most accesses skip the page-table walk.
Instruction-level parallelism
A single core finds parallelism inside one instruction stream. Superscalar cores issue several instructions per cycle. Out-of-order execution lets an instruction run as soon as its operands are ready, using register renaming to remove false dependences and a reorder buffer to retire results in program order. The core idea traces back to Tomasulo's algorithm for the IBM System/360 Model 91.
Branch prediction and speculative execution keep the pipeline full by guessing the outcome of branches and running ahead. Modern predictors are right well over 95% of the time on typical code. Speculation has a cost beyond wasted work: Spectre and Meltdown showed that squashed instructions can still leave measurable traces in caches, turning a performance feature into a security vulnerability.
Parallel and specialized architectures
| Flynn class | Meaning | Where you see it today |
|---|---|---|
| SISD | Single instruction, single data | A conceptual scalar core |
| SIMD | One instruction applied to many data elements | AVX-512, Arm SVE, RISC-V V; GPU warps (SIMT) |
| MISD | Many instructions on one data stream | Rare; some fault-tolerant systems |
| MIMD | Independent instruction streams on independent data | Multicore CPUs, clusters, supercomputers |
Multicore chips share memory, so they need cache coherence (snooping protocols such as MESI, or directories at larger scale) and a defined memory consistency model that says which orderings of loads and stores other cores may observe. x86 provides fairly strong ordering (TSO); Arm and RISC-V are weaker and rely on explicit fences.
GPUs trade single-thread speed for massive throughput: thousands of lightweight threads hide memory latency by switching among themselves. Domain-specific accelerators such as Google's TPU go further, building systolic arrays of multiply-accumulate units with low-precision arithmetic and software-managed memory. Heterogeneous systems now combine big and little CPU cores, GPUs, neural engines and media blocks on one die or across chiplets connected in one package.
How we got here
Timeline
- 1945EDVAC report
Von Neumann's draft describes the stored-program computer, with instructions held in the same memory as data.
- 1964IBM System/360
The first family of machines sharing one ISA across a wide price range, establishing the separation of architecture from implementation and binary compatibility as a business strategy.
- 1965Moore's observation
Transistor counts per chip double on a regular cadence, later settled at roughly every two years.
- 1967Tomasulo's algorithm
Dynamic scheduling on the System/360 Model 91 introduces ideas still at the heart of out-of-order cores.
- 1971Intel 4004
A complete CPU on one chip; the microprocessor era begins.
- 1974Dennard scaling
Shrinking transistors keeps power density constant, so each generation gets faster and denser for free.
- 1976Cray-1
Vector registers and pipelined functional units define supercomputing for a decade.
- 1980sRISC
Berkeley RISC and Stanford MIPS show that simple, pipelinable instructions plus good compilers beat complex ISAs. Arm and SPARC follow.
- 1995Pentium Pro
x86 instructions are translated into internal micro-ops and run out of order, blending CISC compatibility with RISC execution.
- 2004The power wall
Dennard scaling breaks down as leakage grows. Clock rates plateau and vendors pivot to multicore designs.
- 2007CUDA
General-purpose GPU programming goes mainstream, setting up the deep-learning hardware boom.
- 2010RISC-V
An open, royalty-free ISA begins at Berkeley and becomes a base for research, embedded and custom silicon.
- 2015Google TPU
A domain-specific accelerator for neural-network inference enters production data centers.
- 2018Spectre and Meltdown
Speculative execution leaks data across security boundaries; security becomes a first-class architectural constraint.
- 2019Chiplets go mainstream
AMD's Zen 2 splits CPUs into smaller dies in one package, improving yield and letting each die use the best process node.
- 2020Apple M1
A wide Arm core, unified memory and integrated accelerators show the efficiency of tightly integrated SoCs in laptops.
- 2022Exascale
Frontier passes 1018 floating-point operations per second on HPL, built mainly from GPUs.
What history left us
Compatibility is sticky. The System/360 decision to fix an ISA and vary the implementation is why x86 and Arm binaries outlive the chips they were compiled for, and why new ISAs succeed mainly in new markets.
RISC won inside CISC. The RISC argument for simple, pipelinable operations was absorbed by x86 through micro-op translation. Today's debate is less RISC versus CISC and more open versus licensed ISAs.
The free lunch ended. With Dennard scaling gone, performance has to come from parallelism and specialization rather than clock speed. That shift explains multicore chips, GPUs in data centers, and dark silicon: parts of a chip that cannot all be powered at once.
Specialization is the new scaling. Hennessy and Patterson call this a new golden age: domain-specific architectures, domain-specific languages, open hardware and agile chip design, all driven by the end of general-purpose scaling.
Data movement dominates. Moving data now costs far more time and energy than computing on it, pushing designs toward larger caches, stacked high-bandwidth memory, chiplets and processing near or in memory.
Performance and efficiency
The iron law of processor performance
Each term belongs to different layers of the stack. The algorithm, compiler and ISA set the instruction count. The microarchitecture and memory system set cycles per instruction (CPI). The circuit technology and pipeline depth set the clock. Improving one term often worsens another, which is why a single metric like GHz or MIPS misleads.
Worked example. A program executes 2.0×109 instructions at CPI 1.25 on a 3 GHz core: 2.0×109 × 1.25 / 3×109 = 0.833 s. A compiler change cuts instructions by 10% but raises CPI to 1.40: 1.8×109 × 1.40 / 3×109 = 0.840 s. Fewer instructions, slower program.
Average memory access time applies the same idea to caches. With a 1 ns L1, 5% L1 miss rate, a 4 ns L2 that misses 20% of the time, and 80 ns DRAM: AMAT = 1 + 0.05 × (4 + 0.2 × 80) = 2.0 ns.
Amdahl and Gustafson
Amdahl's law bounds the speedup from improving a fraction p of the work by a factor n. The serial part sets a hard ceiling of 1/(1 − p) no matter how many cores you add. Gustafson's law gives the complementary view: as machines grow, people solve bigger problems, so the parallel fraction grows with them. Hill and Marty applied Amdahl's law to multicore chip design and showed why asymmetric big-plus-little cores can beat both extremes.
Roofline
Arithmetic intensity is floating-point operations per byte moved from memory. Kernels left of the ridge point are memory-bound and gain nothing from more compute; kernels to the right are compute-bound. The model tells you which optimization to try: improve data reuse (blocking, fusion) to move right, or vectorize and parallelize to climb toward the roof.
Power and energy
Dynamic power scales with activity factor, switched capacitance, the square of supply voltage, and frequency. Lowering voltage and frequency together (DVFS) cuts power roughly with the cube of frequency, which is why many slower cores can beat one fast core on throughput per watt. Static (leakage) power flows even when transistors are idle and is what ended Dennard scaling.
| Operation (45 nm) | Energy | Relative to 32-bit add |
|---|---|---|
| 8-bit integer add | 0.03 pJ | 0.3× |
| 32-bit integer add | 0.1 pJ | 1× |
| 32-bit float multiply | 3.7 pJ | 37× |
| 32-bit read, 8 KB SRAM | 5 pJ | 50× |
| 32-bit read, DRAM | 640 pJ | 6,400× |
Figures from Horowitz (ISSCC 2014). Newer nodes lower every value, but the ratio between compute and off-chip access remains in the hundreds to thousands.
Measuring honestly
Use execution time of real workloads as the ground truth. Standard suites such as SPEC CPU, MLPerf and HPL make results comparable; summarize normalized ratios with the geometric mean. Report efficiency alongside speed (performance per watt, per dollar, per mm²), and for services, report tail latency (p99) as well as averages. Be wary of peak numbers, MIPS, clock rate alone, and benchmarks tuned for one machine.
Resources
Interactive tools and simulators
Peer-reviewed reading
- Hennessy, J. L., & Patterson, D. A. (2019). A new golden age for computer architecture. Communications of the ACM, 62(2), 48–60.Turing Award lecture; the best single overview of where the field is heading.doi.org/10.1145/3282307
- von Neumann, J. (1993). First draft of a report on the EDVAC. IEEE Annals of the History of Computing, 15(4), 27–75.The 1945 stored-program design, reprinted with commentary.doi.org/10.1109/85.238389
- Moore, G. E. (1998). Cramming more components onto integrated circuits. Proceedings of the IEEE, 86(1), 82–85.Reprint of the 1965 paper behind Moore's law.doi.org/10.1109/JPROC.1998.658762
- Dennard, R. H., et al. (1974). Design of ion-implanted MOSFET's with very small physical dimensions. IEEE Journal of Solid-State Circuits, 9(5), 256–268.The scaling rules whose breakdown ended rising clock speeds.doi.org/10.1109/JSSC.1974.1050511
- Flynn, M. J. (1972). Some computer organizations and their effectiveness. IEEE Transactions on Computers, C-21(9), 948–960.Origin of the SISD/SIMD/MISD/MIMD taxonomy.doi.org/10.1109/TC.1972.5009071
- Tomasulo, R. M. (1967). An efficient algorithm for exploiting multiple arithmetic units. IBM Journal of Research and Development, 11(1), 25–33.The foundation of dynamic scheduling and register renaming.doi.org/10.1147/rd.111.0025
- Patterson, D. A., & Ditzel, D. R. (1980). The case for the reduced instruction set computer. ACM SIGARCH Computer Architecture News, 8(6), 25–33.The short paper that launched the RISC movement.doi.org/10.1145/641914.641917
- Blem, E., Menon, J., & Sankaralingam, K. (2013). Power struggles: Revisiting the RISC vs. CISC debate on contemporary ARM and x86 architectures. Proc. IEEE HPCA, 1–12.Measured evidence that ISA matters less than microarchitecture.doi.org/10.1109/HPCA.2013.6522302
- Wulf, W. A., & McKee, S. A. (1995). Hitting the memory wall: Implications of the obvious. ACM SIGARCH Computer Architecture News, 23(1), 20–24.Coined the memory wall.doi.org/10.1145/216585.216588
- Amdahl, G. M. (1967). Validity of the single processor approach to achieving large scale computing capabilities. Proc. AFIPS Spring Joint Computer Conference, 483–485.The original statement of Amdahl's law.doi.org/10.1145/1465482.1465560
- Gustafson, J. L. (1988). Reevaluating Amdahl's law. Communications of the ACM, 31(5), 532–533.Scaled speedup when problem size grows with the machine.doi.org/10.1145/42411.42415
- Hill, M. D., & Marty, M. R. (2008). Amdahl's law in the multicore era. IEEE Computer, 41(7), 33–38.Symmetric, asymmetric and dynamic multicore trade-offs.doi.org/10.1109/MC.2008.209
- Williams, S., Waterman, A., & Patterson, D. (2009). Roofline: An insightful visual performance model for multicore architectures. Communications of the ACM, 52(4), 65–76.The roofline model used above.doi.org/10.1145/1498765.1498785
- Asanović, K., et al. (2009). A view of the parallel computing landscape. Communications of the ACM, 52(10), 56–67.Berkeley's analysis of the shift to parallelism after the power wall.doi.org/10.1145/1562764.1562783
- Esmaeilzadeh, H., et al. (2011). Dark silicon and the end of multicore scaling. Proc. ISCA, 365–376.Why power limits keep much of a chip idle.doi.org/10.1145/2000064.2000108
- Horowitz, M. (2014). Computing's energy problem (and what we can do about it). Proc. IEEE ISSCC, 10–14.Source of the energy-per-operation table.doi.org/10.1109/ISSCC.2014.6757323
- Jouppi, N. P., et al. (2017). In-datacenter performance analysis of a tensor processing unit. Proc. ISCA, 1–12.The case study for domain-specific accelerators.doi.org/10.1145/3079856.3080246
- Dally, W. J., Turakhia, Y., & Han, S. (2020). Domain-specific hardware accelerators. Communications of the ACM, 63(7), 48–57.Where accelerator gains really come from.doi.org/10.1145/3361682
- Leiserson, C. E., et al. (2020). There's plenty of room at the Top: What will drive computer performance after Moore's law? Science, 368(6495), eaam9744.Performance from software, algorithms and hardware architecture.doi.org/10.1126/science.aam9744
- Thompson, N. C., & Spanuth, S. (2021). The decline of computers as a general purpose technology. Communications of the ACM, 64(3), 64–72.The economics behind specialization.doi.org/10.1145/3430936
- Kocher, P., et al. (2019). Spectre attacks: Exploiting speculative execution. Proc. IEEE Symposium on Security and Privacy, 1–19.How speculation became an attack surface.doi.org/10.1109/SP.2019.00002
- Kim, Y., et al. (2014). Flipping bits in memory without accessing them: An experimental study of DRAM disturbance errors. Proc. ISCA, 361–372.The Rowhammer paper; reliability and security at the memory level.doi.org/10.1109/ISCA.2014.6853210