This lecture answers two questions with direct practical consequences: how the memory of an embedded system is organized (von Neumann, Harvard, modified Harvard) and what exactly makes ARM processors - by far the most widespread in today's microcontrollers - look and behave the way they do: registers, pipeline, operating modes, interrupts and the two instruction sets, A32 and Thumb/Thumb-2.
1The subject and structure of the lecture6 min
So far we have discussed what an embedded system is and how the performance of a processor is measured "at an abstract level" (ISA, CPI, IPC). Today we go one floor further down: what the memory of a microcontroller looks like concretely and how the processor that accesses it is organized - with the emphasis on the ARM family, present in the great majority of modern microcontrollers.
- The ISA separates the software visible to the programmer from the internal hardware implementation
- RISC (simple instructions, fixed length, load/store) dominates embedded systems because of its energy efficiency and simplicity of implementation
- Real performance is computed, not guessed from the frequency: Execution Time = IC × CPI × Clock Cycle
Learning outcomes
- To explain the difference between the von Neumann, Harvard and modified Harvard architectures
- To identify the ARM profile (A, R or M) suited to a given application
- To locate the role of the registers R0-R15 in a simple piece of ARM/Cortex-M code
- To explain how a three-stage pipeline works and why a branch "flushes" it
- To read the flags N, Z, C, V from a status register and anticipate a conditional branch
- To explain why Thumb/Thumb-2 exists and what trade-off it resolves
2The von Neumann architecture10 min
The main advantage is flexibility: the memory space can be divided dynamically between the program and the data, without the size of the two regions being fixed by the hardware. A single address space simplifies the programming model - code can be treated as stored information, just like any other data.
In order to compute the sum of two values from memory, the processor must: (1) fetch the instruction, (2) read the two operands, (3) write the result. If all of these transfers pass through the same memory interface, they compete for the same resource - even if the addition itself takes a single ALU cycle, the total time may be dominated by the memory accesses.
3The Harvard architecture10 min
The separation allows the processor to fetch the next instruction at the same time as the current instruction reads or writes an operand - access to the program no longer competes with access to the data. The property is valuable above all in pipelined processors and in applications that repetitively process streams of data (DSPs).
The physical separation also allows each memory to be adapted to its role: the program memory can use a non-volatile technology (Flash), with a word wide enough for a complete instruction, and the data memory can use fast SRAM, optimized for frequent reads and writes - exactly the Flash+RAM combination of any modern microcontroller.
A microcontroller has 256 KB of Flash for the program and 64 KB of SRAM for variables. Even if the application uses only half the Flash, the space left over cannot be transferred automatically to the SRAM in order to extend the stack - the two memories have separate address spaces. That is the price paid for simultaneous access.
Because the program and the data occupy different address spaces, the processor cannot ordinarily access the program memory with the same instructions used for data - constants, lookup tables and firmware updating need dedicated mechanisms.
4Modified Harvard and a comparison of the models11 min
Strict separation improves throughput, but reduces the flexibility of the software; a completely unified memory is easy to program, but reintroduces competition for access. The modified Harvard architecture aims to combine the advantages of both: there is no single universal implementation, but the most common variant in modern processors keeps a unified memory space at the program level, with separate caches for instructions (I-cache) and data (D-cache) close to the core.
When the program writes an area of memory that is then to be executed as code (for example, when loading new firmware), the data written may remain temporarily only in the D-cache, while the I-cache still holds an old copy. The system software must explicitly synchronize and invalidate the caches before jumping into the new code - a classic source of bugs that are hard to reproduce in systems that do OTA updates.
| Characteristic | von Neumann | Harvard | Modified Harvard |
|---|---|---|---|
| Organization of the memory | common | distinct | separation only at certain levels |
| Concurrent access | limited | simultaneous | simultaneous through I-cache/D-cache |
| Address spaces | unified | separate | unified, with locally separate paths |
| Flexibility | high | lower | high, with more complex hardware |
| Representative uses | the classic conceptual model | DSPs, strict microcontrollers | most modern processors |
5The ARM family of processors10 min
ARM is by far the most widespread family of processor architectures in the world - from microcontrollers with minimal resources up to telephones, automotive systems and servers. Its success is tied to the RISC principles, energy efficiency and adaptability to very different requirements. The name started from Acorn RISC Machine (in the 1980s), then Advanced RISC Machines; today "Arm" is used.
ARM (the company) does not manufacture processors - it licenses architectures and processing cores to manufacturers of integrated circuits. That is why two different microcontrollers may use the same Cortex-M core, yet have completely different peripherals and memory organization - the core is only one piece of the puzzle.
The ARM family is organized into three architectural profiles, each optimized for a distinct class of systems, not merely "more/less powerful":
| Profile | Main field | Memory management | Operating system | Examples |
|---|---|---|---|---|
| Cortex-A (Application) | systems with a complex OS | MMU, virtual memory, advanced cache | Linux, Android, Windows | telephones, routers, servers |
| Cortex-R (Real-time) | critical, deterministic control | MPU, predictable TCM memories | an RTOS or dedicated | automotive, storage, industrial |
| Cortex-M (Microcontroller) | embedded microcontrollers | optional MPU, no classic MMU | bare-metal or an RTOS | IoT, sensors, controllers |
In this lecture we concentrate on the elements of the classic AArch32 model (useful for understanding the evolution of the architecture) and on the particularities of Cortex-M - the most relevant profile for the embedded systems you build in the laboratory.
6The registers of the ARM architecture11 min
Registers are the fastest storage space available to the processor - they hold the operands, the memory addresses and the control information of the execution. Efficient use of the registers reduces the number of accesses to main memory and directly increases performance.
The classic AArch32 model: R0-R15
| Register | Role |
|---|---|
| R0-R3 | function parameters and return values |
| R4-R11 | local variables, preserved across function calls |
| R12 | a temporary register (used by the compiler / linker) |
| R13 (SP) | Stack Pointer - the top of the current stack |
| R14 (LR) | Link Register - the return address from a function call |
| R15 (PC) | Program Counter - the address of the current/next instruction |
BL function ; saves the return address in LR, jumps to "function" ; ... the program continues ... function: ADD R0, R0, #1 BX LR ; returns using the address in LR
If the function in turn calls another function, it must save LR on the stack first, otherwise the original return address is lost:
function:
PUSH {R4, LR}
MOV R4, R0
BL other_function
ADD R0, R0, R4
POP {R4, PC} ; loading directly into PC performs the returnCortex-M particularities
Cortex-M keeps R0-R15 with similar roles, but introduces two distinct stack pointers: MSP (Main Stack Pointer), used after reset and in Handler Mode, and PSP (Process Stack Pointer), used by the application or by the tasks of an RTOS. The separation isolates the stack of the kernel and of the exception routines from the stacks used by the application - a useful protection when an application error must not corrupt the system's exception handling.
7The Program Counter and pipelined execution10 min
The PC register holds the address of the next instruction. In a processor without a pipeline, its value would correspond directly to the instruction being executed. In a pipelined processor - every modern ARM processor - several instructions are simultaneously in different stages, and the exact meaning of the value read from PC depends on the organization of the architecture.
This observation explains a practical rule of embedded programming: code with many unpredictable conditional branches may be slower than "the number of instructions" would suggest - the pipeline pays a penalty at every mispredicted branch.
8The status registers CPSR and xPSR10 min
Besides the general registers, ARM processors keep the status information - flags resulting from arithmetic operations, the privilege level, the state of the control mechanisms - in dedicated registers. The structure differs between the classic AArch32 model and Cortex-M.
CPSR (Current Program Status Register) - classic AArch32
| Flag | Meaning |
|---|---|
| N | the result has its sign bit set (negative) |
| Z | the result of the operation is zero |
| C | the operation produced a carry or a borrow |
| V | the signed operation produced an arithmetic overflow |
CMP R0, #0 ; a "phantom" subtraction: updates N,Z,C,V without storing the result BEQ value_zero ; branch taken only if the Z flag is set
If R0 holds the value zero, the Z flag is set and the conditional branch is taken - exactly the
mechanism behind any if compiled for ARM.
xPSR - Cortex-M
Cortex-M gathers conceptually the same information in the xPSR register, made up of three components: APSR (the flags N, Z, C, V, with a role identical to those in CPSR), IPSR (the number of the active exception/interrupt) and EPSR (the state of the execution). The essential difference: Cortex-M does not use the "mode" field of CPSR, nor the IRQ/FIQ/Supervisor modes of classic AArch32 - it has its own, simpler model, discussed in the following section.
9Operating modes and interrupt handling10 min
Operating modes
The classic AArch32 model defines several processor modes: User (unprivileged) and the privileged modes System, Supervisor, IRQ, FIQ, Abort, Undefined. The switch to an exception mode is automatic, on the occurrence of the corresponding event. FIQ is designed for a fast response and has additional banked registers - it reduces the need to save general registers on entering the routine.
Cortex-M simplifies this radically: only two modes, Thread Mode (the execution of the application) and Handler Mode (the handling of exceptions, always privileged). Thread Mode may be privileged or not, depending on the CONTROL register - an organization suited to microcontrollers and to the real-time operating systems (RTOS) that run on them.
Interrupt handling
Interrupts allow the processor to react to events without continuously interrogating the state of every peripheral (polling). Instead of the application constantly checking whether a timer has expired, the peripheral "asks for the attention" of the processor only when the relevant event occurs - a direct saving of cycles and of energy.
Handling an interrupt follows, broadly, four steps: (1) temporarily suspending the current program, (2) saving the context needed to resume it, (3) executing the handling routine (ISR), (4) restoring the context and resuming the interrupted program. Cortex-M automates the first two steps and the last one in hardware, through the exception entry/exit mechanism, which significantly simplifies writing an interrupt routine compared with the classic AArch32 model.
10Conditional execution, the Barrel Shifter, Thumb and AMBA10 min
We close this overview of the ARM architecture with three mechanisms that explain the density and the efficiency of the code generated: conditional execution, the Barrel Shifter, and the Thumb/Thumb-2 instruction sets.
Conditional execution
In the A32 encoding, many instructions include a condition field, attached as a suffix to the mnemonic, allowing an instruction to be executed only if the status flags satisfy a condition - reducing the number of branches needed for short decisions.
CMP R0, #0 MOVEQ R1, #1 ; executed only if R0 == 0 MOVNE R1, #0 ; executed only if R0 != 0
The T32 set (Thumb-2) is more restrictive: it mainly uses direct conditional branches, and some
architectures allow the IT (If-Then) block for making a short sequence of instructions
conditional.
The Barrel Shifter
A hardware block that shifts or rotates an operand before it enters the ALU - combining a shift with an arithmetic or logic operation, in a single instruction.
ADD R0, R1, R2, LSL #2 ; R0 = R1 + 4 × R2
Directly useful for computing the address of an element of an array whose elements occupy 4 bytes -
exactly the pattern the compiler generates for an array[i] access.
| Operation | Full name | Behaviour |
|---|---|---|
| LSL | Logical Shift Left | logical left shift - equivalent to × 2ⁿ |
| LSR | Logical Shift Right | logical right shift, introduces zeros |
| ASR | Arithmetic Shift Right | preserves the sign bit - approximate division by 2ⁿ |
| ROR | Rotate Right | rotation - the bits shifted out re-enter at the other end |
Thumb and Thumb-2: code density
A32 instructions, although uniform, always occupy 32 bits - programs can become large, which costs Flash memory, energy in transferring them and bandwidth. The original Thumb set solved this with 16-bit instructions (a compact subset of A32 functionality, with some restrictions on registers and constants). Thumb-2 combines 16- and 32-bit encodings in the same set (T32): frequent operations stay compact, complex ones can use the wide encoding when it is needed.
All Cortex-M processors execute T32 - but not identically: Cortex-M0/M0+ implement a restricted subset, optimized for minimum area and power; Cortex-M3/M4 and more capable cores include a richer Thumb-2 set and optional extensions (DSP instructions on Cortex-M4, for instance). The practical conclusion: code compiled for a Cortex-M4 does not run unchanged on a Cortex-M0 - the compiler must be configured explicitly for the target core.
AMBA interconnect (briefly)
A modern SoC connects the processor, the memory, the DMA and the peripherals through a family of protocols standardized by ARM, called AMBA: AXI for high-throughput interconnects (processor, memory, accelerators), AHB/AHB-Lite for a system bus with good performance and lower complexity (frequent in Cortex-M microcontrollers), and APB for simple peripherals (GPIO, UART, timers), where simplicity and low power matter more than throughput.
11Frequent mistakes5 min
- "Harvard always means two physically separate, completely isolated memories." In modern processors, "Harvard" most often refers to modified Harvard: a unified main memory, but separate paths (I-cache/D-cache) close to the core. Check whether the discussion refers to the architecture visible to the programmer or to the internal microarchitectural organization - they often differ.
- "CPSR and xPSR are the same thing, just under a different name." Conceptually they partly overlap (both hold N,Z,C,V), but Cortex-M does not have the IRQ/FIQ/Supervisor modes of classic AArch32 - the exception model is structurally different. When you read ARM documentation, check first whether it refers to the classic A/R profile or to Cortex-M - many confusions start here.
- "Fewer Thumb instructions always means smaller code." One 32-bit A32 instruction may replace two 16-bit Thumb instructions - the real advantage depends on the application, the compiler and the distribution of the instructions used, not merely on the nominal length of the encoding. Do not compare code density from the length of an isolated instruction - measure the size of the final compiled binary.
12Summary and glossary5 min
The organization of the memory (von Neumann, Harvard, modified Harvard) and the architecture of the processor (registers, pipeline, modes, interrupts, instruction sets) together determine how fast, predictable and energy-efficient an embedded system can be. The ARM family dominates microcontrollers today precisely because it offers, through the A/R/M profiles, solutions adapted to every combination of requirements - from telephones running Linux to battery-powered sensors with no operating system.
13Self-check questions10 min
- What is the fundamental difference between the von Neumann and the Harvard architectures?
- What is the modified Harvard architecture and why is it the most frequently met today?
- What are the three ARM architectural profiles and what is each optimized for?
- What is the role of the registers SP, LR and PC in the AArch32 model?
- What are MSP and PSP in Cortex-M and why are there two stacks?
- How does a three-stage pipeline work and what happens at a conditional branch?
- What flags do the CPSR/xPSR registers hold and what does each of them mean?
- What do the Thumb and Thumb-2 instruction sets solve?
- What role does the Barrel Shifter play in an ARM instruction?
- Name the three AMBA protocols and the field of use of each.
14Where to go next2 min
With the architecture of the processor and the organization of the memory clarified, the next lecture moves on to a constraint that returns constantly in embedded systems: energy consumption. We shall link today's concepts directly (operating modes, interrupts, pipeline) to the concrete techniques by which a microcontroller reduces its consumption - sleep, wake-on-interrupt, frequency scaling.
Many of the concepts of this lecture (registers, pipeline, interrupts) become tangible only in the laboratory - follow them concretely in the laboratories of the course as you work directly on ARM Cortex-A (Raspberry Pi 5) and Cortex-M/Xtensa (ESP32) hardware.