LECTURE 02

The Architecture of Computing Systems

Duration: 120 min of teaching Level: undergraduate, year III - recommended after Lecture 01 Course: Embedded Systems PDF: download the notes RO versiunea română

This lecture opens the "black box" inside an embedded system: what an instruction set (ISA) is, how the addressing models evolved from three-address machines to stack-based ones, what separates a CISC architecture from a RISC one - and, perhaps the most practically useful of all, how the real performance of a processor is computed and interpreted, beyond the clock-frequency figure printed on the box.

1The subject and structure of the lecture6 min

In the last lecture we treated embedded systems "from the outside" - what they are, where they appear, what distinguishes them from an ordinary computer. This lecture goes one level down: what happens, concretely, inside the processing unit when it executes an instruction - and why the choice of architecture matters so much for a system with limited resources.

Recap of the previous lecture
  • An embedded system is optimized for a dedicated function, not for flexibility
  • Designing it means a compromise between performance, cost, power and reliability
  • Most microcontrollers in embedded systems use RISC architectures (ARM Cortex-M, RISC-V, AVR) - a statement we shall at last be able to explain technically today

Learning outcomes

  • To explain what an ISA is and why it separates the software from the hardware implementation
  • To recognize and compare the four models of instruction addressing
  • To explain the fundamental difference between CISC and RISC architectures
  • To write down and apply the fundamental equation of processor performance
  • To compute CPI, IPC and the execution time for a given program
  • To explain why the clock frequency, taken in isolation, says nothing about real performance

2Instruction Set Architecture12 min

Instruction Set Architecture (ISA)
The interface between software and hardware: the set of available instructions, the registers accessible to the programmer, the data types, the addressing modes, the organization of memory and the mechanisms of interrupts and exceptions that any hardware implementation must observe.

The essential idea: the ISA describes what the processor can do, not how it does it internally. The compiler and the assembly-language programmer work exclusively at the level of the ISA, with no need to know what the circuits that execute each instruction look like.

Programmer - develops the application Application and compiler C / C++ / Rust → machine code Instruction Set Architecture (ISA) instructions · registers · addressing modes · memory · interrupts Microarchitecture pipeline · cache · ALU · control Hardware - CPU, memory, buses
Fig. 1 - The ISA as the interface between software and hardware. The programmer and the compiler "see" only the ISA; everything below it (the microarchitecture, the hardware proper) may vary freely between generations or manufacturers, as long as the ISA stays identical.
Analogy The ISA is like a driving contract: the accelerator pedal, the brake, the steering wheel and the indicators work the same way, whoever made the car. What happens under the bonnet - a petrol engine, an electric one, with 4 or 8 cylinders - may differ completely, without the driver noticing any change in the way they drive. In the same way, two processors may implement the same ISA with totally different microarchitectures and still execute the same program correctly.
Example - x86-64: one ISA, two families of processors

Intel Core and AMD Ryzen processors both implement the x86-64 architecture, but with completely different internal microarchitectures (pipeline, cache, branch prediction). Both execute exactly the same code compiled for x86-64 - this is the power of separating the ISA from the implementation.

For embedded systems, the choice of ISA directly influences performance, energy consumption, the size of the executable code and the availability of software tools. Most modern microcontrollers use RISC architectures (ARM Cortex-M, RISC-V, AVR), because of their energy efficiency and the simplicity of the instruction set - a thread we shall take up again in the sections on CISC and RISC, below.

3Three-address machines10 min

The most intuitive instruction format specifies explicitly both source operands and the location in which the result will be stored - the three-address model, frequently used as an intermediate representation by compilers, precisely because it reflects directly the mathematical expressions of high-level languages.

General form
ADD DEST, OP1, OP2 → DEST = OP1 + OP2

Consider the expression A = B + C × D − E + F + A. With a three-address machine:

three_address.asm
MUL T, C, D
ADD T, T, B
SUB T, T, E
ADD T, T, F
ADD A, T, A

Every instruction is easy to follow - you can see directly what is added to what, and where the result ends up. The price paid: every instruction has to encode three addresses, so the instructions are longer and the executable code takes up more memory - a real problem for an embedded system with limited Flash.

4Two-address machines8 min

The two-address model reduces the explicit fields: one of the operands is used at the same time as the source and as the destination of the result.

General form
ADD OP1, OP2 → OP1 = OP1 + OP2
two_address.asm
LOAD T, C
MUL  T, D
ADD  T, B
SUB  T, E
ADD  T, F
ADD  A, T

The instructions are shorter than in the three-address model, but the code is slightly harder to follow, because the initial value of the first operand is overwritten.

Example - x86 uses exactly this model

The instruction ADD EAX, EBX adds the value in EBX to EAX, and the result stays in EAX - the register EAX is at the same time the source and the destination. You can now recognize why x86 syntax looks the way it does.

5One-address machines8 min

The one-address model goes one step further: it introduces a special register, called the accumulator (AC), used implicitly by almost every arithmetic and logic operation. The instruction specifies explicitly a single operand; the other operand and the destination of the result are always the accumulator.

General form
ADD OP → AC = AC + OP
one_address.asm
LOAD  C
MUL   D
ADD   B
SUB   E
ADD   F
ADD   A
STORE A

The instructions are shorter still, and hardware decoding is simpler - but the price is a larger number of memory accesses (every LOAD/STORE moves data to and from the accumulator) and reduced flexibility for complex computations.

A historical note

The first generations of microprocessors - the Intel 8080, the MOS Technology 6502 - made extensive use of accumulator-type registers. Modern architectures mainly use general registers, but the concept remains useful for understanding why architectures evolved the way they did.

6Zero-address machines and a comparison of the models12 min

The most compact model: zero addresses, based on a stack data structure (LIFO - Last In, First Out). Arithmetic instructions no longer contain any explicit operand - they automatically use the first two elements at the top of the stack.

General form
ADD → TOS = TOS₋₁ + TOS (TOS = Top Of Stack)
Stack 7 5 TOS = top ADD POP, POP, ADD, PUSH Stack 12 result: 5+7
Fig. 2 - A zero-address machine: ADD implicitly takes the two values off the top of the stack, adds them, and puts the result back on top. No operand appears explicitly in the instruction.
Evaluating the expression (A+B)*C on a stack machine: put the instructions in order
zero_address.asm
PUSH B
PUSH C
PUSH D
MUL
ADD
PUSH E
SUB
PUSH F
ADD
PUSH A
ADD
POP  A
Example - the Java virtual machine

The JVM (Java Virtual Machine) uses exactly this model: bytecode instructions such as iadd, imul, isub operate implicitly on the top of the execution stack. The choice contributes directly to the portability of Java applications.

Comparing all four models

Characteristic3 addresses2 addresses1 address0 addresses
Explicit operands3210
Implicit registernonoaccumulatorstack
Instruction lengthlargemediumsmallvery small
Hardware complexityhighmediumlowlow
Ease of programmingvery goodgoodmediumlow
Match each addressing model with its real example
To remember At present, most commercial processors use architectures based on general registers (a variant close to the two/three-address model, but with a large set of registers instead of a single accumulator) - the best practical compromise between performance and flexibility. The historical models remain important because they explain why modern architectures look the way they do.

7The CISC architecture10 min

CISC (Complex Instruction Set Computer) architectures appeared at a time when memory was expensive and compilers had limited possibilities of optimization. The objective: to reduce the number of instructions an algorithm needs, by introducing complex instructions able to carry out several operations "in one go".

CISC processors include, as a rule, hundreds of instructions, numerous addressing modes and instructions of variable length - some can simultaneously access memory, do an arithmetic operation and update registers. The result: more compact programs, but a processor far more complex to implement.

Microcode

Internally, a CISC processor breaks each complex instruction down into simpler micro-operations, through a mechanism called microcode. The programmer "sees" a complex instruction; the hardware, behind the scenes, executes a sequence of elementary steps - exactly the way in which new instructions can be added without changing the basic circuits of the processor.

The Intel x86 family is the best-known example of a CISC architecture. Interestingly, even modern x86 processors internally translate complex instructions into micro-operations similar to those of RISC architectures - a convergence we discuss below.

8The RISC architecture10 min

The RISC concept (Reduced Instruction Set Computer) came out of research at Berkeley and Stanford in the 1980s, which showed that most programs in fact use only a fraction of the instructions available in a CISC architecture. The conclusion: simplify the instruction set and optimize the execution of those that remain.

RISC processors use instructions of fixed length, a large number of general registers and a load/store architecture: arithmetic operations are carried out exclusively on registers, and memory is accessed only through dedicated load and store instructions.

Why this matters for an embedded system Fixed-length instructions and a reduced set radically simplify decoding - fewer transistors devoted to control, more room (and energy) left for execution proper. The simplicity also allows an efficient pipeline: several instructions "in flight" at once, each in a different stage of execution. The combination explains why ARM Cortex-M, AVR and RISC-V dominate modern microcontrollers: good performance, low energy consumption.

The typical flow of execution in a classic RISC pipeline has five stages: Fetch (fetching the instruction), Decode (decoding and reading the registers), Execute (the operation in the ALU), Memory Access (accessing memory, if needed) and Write Back (writing the result into a register) - stages that can overlap, different instructions being, at every moment, in different stages of the pipeline.

9RISC vs. CISC: comparison and convergence8 min

Although for a long time regarded as competing approaches, modern implementations borrow elements from both worlds: x86 processors have a CISC ISA, but internally execute RISC-like micro-operations; modern ARM processors include branch prediction and speculative execution - techniques historically associated with high-performance processors, whatever the family.

CharacteristicRISCCISC
Complexity of the instructionslowhigh
Length of the instructionsfixedvariable
Number of registerslargesmall/medium
Pipelinevery efficientmore difficult
Memory accessload/storedirectly in the instructions
Size of the codelargersmaller
ExamplesARM, AVR, RISC-VIntel x86, AMD64

The comparison must be seen through the lens of engineering trade-offs, not as a "better / worse" verdict - which is why servers and desktops (where software compatibility matters enormously) carry on using x86, while embedded systems (where power and cost dominate) have migrated almost universally to RISC.

10The performance of a processor: CPI, IPC, execution time7 min

The performance of an embedded system does not come down to "how fast" it executes instructions - a braking controller needs determinism, not maximum speed. All the same, in order to discuss performance rigorously we need a precise formula, not intuitions.

The fundamental equation of performance
Execution Time = Instruction Count × CPI × Clock Cycle
where Instruction Count (IC) = the total number of instructions executed, CPI (Cycles Per Instruction) = the average number of cycles per instruction, and Clock Cycle = 1 / the clock frequency.

The central observation: execution time can be reduced along three independent paths - fewer instructions (a better algorithm/compiler), a smaller CPI (a more efficient architecture) or a higher frequency (at an energy and thermal cost). None of them, taken separately, tells the whole story.

The calculator for the performance equation

11CPI, IPC and the performance equation9 min

The same workload, three architectures: where the time goes
CPI - Cycles Per Instruction CPI = Clock Cycles / Instruction Count

The smaller it is, the more efficient the architecture. Instructions do not all cost the same: a simple addition may take 1 cycle, a memory access or a conditional branch, several.
IPC - Instructions Per Cycle IPC = Instruction Count / Clock Cycles = 1 / CPI

A classic scalar processor has IPC ≤ 1. Modern superscalar processors, with multiple execution units, can complete several instructions per cycle - IPC > 1.
A worked exercise - a higher frequency does not always mean faster

Two processors execute the same program of 1,000,000 instructions. Processor A runs at 1 GHz with CPI = 2. Processor B runs at 800 MHz with CPI = 1. Which is faster?

See the solution

T_A = (1,000,000 × 2) / 10⁹ = 2 ms

T_B = (1,000,000 × 1) / (800 × 10⁶) = 1.25 ms

Although processor B has the lower frequency, it is the faster one - the reduced CPI amply compensates for the difference in frequency. The practical conclusion: clock frequency, in isolation, is not a sufficient indicator of real performance.

In practice, real performance depends on three categories of factor acting at the same time: the architecture of the processor (frequency, pipeline, cache, number of cores), the efficiency of the code executed (algorithm, data structures, locality of access) and the compiler and its optimizations (register allocation, elimination of redundant code, instruction selection). None of them alone guarantees performance - all three have to be aligned.

The performance equation of the processor

12Frequent mistakes5 min

  • "A higher frequency always means a faster processor." False - see the worked exercise above. CPI matters just as much, sometimes more, than frequency. Always compare real execution time, not merely the frequency printed on the datasheet.
  • "RISC means fewer instructions in the program." as a rule it is the other way round: a RISC program needs more instructions (each of them simpler) than the CISC equivalent, but each executes in fewer cycles, and the sum may nevertheless be faster. Separate "the number of instructions" from "the total execution time" - the performance equation links them both, they are not the same thing.
  • "CISC and RISC are completely separate categories today." They are no longer cleanly separate - x86 (CISC as an ISA) translates internally into RISC-like micro-operations, and ARM (RISC) has added advanced performance mechanisms historically associated with CISC. Treat RISC/CISC as a spectrum of design philosophies, not as a fixed binary label.

13Summary and glossary5 min

The ISA is the stable contract between software and hardware - what allows the internal implementation to evolve freely without breaking the compatibility of existing programs. The addressing models (three, two, one address, zero addresses) represent different points on the flexibility-versus-compactness axis. RISC and CISC are design philosophies that today partly converge. And the performance of a processor is measured rigorously through the equation Execution Time = IC × CPI × Clock Cycle - not through frequency alone.

ISA
Instruction Set Architecture - the interface between software and hardware.
Microarchitecture
the concrete internal implementation of an ISA (pipeline, cache).
Accumulator
the implicit register used by the instructions of one-address machines.
Load/Store
a model in which memory is accessed only through dedicated instructions.
CISC
Complex Instruction Set Computer - complex instructions, a rich ISA.
RISC
Reduced Instruction Set Computer - simple instructions, of fixed length.
CPI
Cycles Per Instruction - the average cycles needed per instruction.
IPC
Instructions Per Cycle - the reciprocal of CPI.
Pipeline
the overlapping of the execution stages of several instructions.

14Self-check questions8 min

  1. What does the ISA represent and what is its role in the relationship between hardware and software?
  2. What are the advantages and disadvantages of a three-address machine, and of a one-address machine?
  3. How does a stack-based machine (zero addresses) work? Give a real example.
  4. What are the main characteristics of CISC and of RISC architectures?
  5. Why are RISC architectures preferred in embedded systems?
  6. Define CPI and IPC and state the relationship between them.
  7. Write down the performance equation of the processor and explain every term.
  8. Why is clock frequency, taken in isolation, not sufficient for assessing performance?
  9. Compute the execution time for a program of 2,000,000 instructions, CPI = 2, at 500 MHz.

15Where to go next2 min

The next lecture continues directly: now that we know what an instruction looks like and how the performance of a processor is measured, we move on to the memory hierarchy and to the particularities of the ARM architectures - the most used today in microcontrollers and SoCs for embedded systems.

The frequency-versus-performance concept discussed here returns, applied concretely, in Laboratory 01 (lectures 4-6), where you measure on a real Raspberry Pi 5 how execution time and power consumption change when you vary the frequency of the processor.