LECTURE 03

Memory Architecture and ARM Processors

Duration: 120 min of teaching Level: undergraduate, year III - recommended after Lecture 02 Course: Embedded Systems PDF: download the notes RO versiunea română

This lecture answers two questions with direct practical consequences: how the memory of an embedded system is organized (von Neumann, Harvard, modified Harvard) and what exactly makes ARM processors - by far the most widespread in today's microcontrollers - look and behave the way they do: registers, pipeline, operating modes, interrupts and the two instruction sets, A32 and Thumb/Thumb-2.

1The subject and structure of the lecture6 min

So far we have discussed what an embedded system is and how the performance of a processor is measured "at an abstract level" (ISA, CPI, IPC). Today we go one floor further down: what the memory of a microcontroller looks like concretely and how the processor that accesses it is organized - with the emphasis on the ARM family, present in the great majority of modern microcontrollers.

Recap of the previous lecture
  • The ISA separates the software visible to the programmer from the internal hardware implementation
  • RISC (simple instructions, fixed length, load/store) dominates embedded systems because of its energy efficiency and simplicity of implementation
  • Real performance is computed, not guessed from the frequency: Execution Time = IC × CPI × Clock Cycle

Learning outcomes

  • To explain the difference between the von Neumann, Harvard and modified Harvard architectures
  • To identify the ARM profile (A, R or M) suited to a given application
  • To locate the role of the registers R0-R15 in a simple piece of ARM/Cortex-M code
  • To explain how a three-stage pipeline works and why a branch "flushes" it
  • To read the flags N, Z, C, V from a status register and anticipate a conditional branch
  • To explain why Thumb/Thumb-2 exists and what trade-off it resolves

2The von Neumann architecture10 min

The von Neumann architecture
A model in which the instructions of the program and the data are kept in the same memory and accessed through a common addressing mechanism. Its origin: the 1945 report of John von Neumann, for the EDVAC project - the idea of the "stored program".
CPU control · ALU · registers Unified memory instructions + data, one space a single bus (instructions / data compete) CPU control · ALU · registers Program memory Data memory
Fig. 1 - Von Neumann (top): a common bus, so instructions and data compete for the same access. Harvard (bottom, anticipating the following section): separate paths, simultaneous access possible.

The main advantage is flexibility: the memory space can be divided dynamically between the program and the data, without the size of the two regions being fixed by the hardware. A single address space simplifies the programming model - code can be treated as stored information, just like any other data.

The von Neumann bottleneck Using the same path for instructions and data means that the processor cannot fetch a new instruction and, at the same time, read or write an operand through the same interface. In the modern sense, the "von Neumann bottleneck" refers not merely to a single physical bus, but to the ever-growing gap between the speed of the processor and the speed at which memory can deliver data to it. Caches, prefetching and burst transfers mitigate the problem without eliminating it entirely.
Example

In order to compute the sum of two values from memory, the processor must: (1) fetch the instruction, (2) read the two operands, (3) write the result. If all of these transfers pass through the same memory interface, they compete for the same resource - even if the addition itself takes a single ALU cycle, the total time may be dominated by the memory accesses.

3The Harvard architecture10 min

The Harvard architecture
A model in which the instructions and the data are kept in distinct memories and transferred along separate paths - possibly with different address spaces, word widths and technologies for each.

The separation allows the processor to fetch the next instruction at the same time as the current instruction reads or writes an operand - access to the program no longer competes with access to the data. The property is valuable above all in pipelined processors and in applications that repetitively process streams of data (DSPs).

The physical separation also allows each memory to be adapted to its role: the program memory can use a non-volatile technology (Flash), with a word wide enough for a complete instruction, and the data memory can use fast SRAM, optimized for frequent reads and writes - exactly the Flash+RAM combination of any modern microcontroller.

An example - and its limitation

A microcontroller has 256 KB of Flash for the program and 64 KB of SRAM for variables. Even if the application uses only half the Flash, the space left over cannot be transferred automatically to the SRAM in order to extend the stack - the two memories have separate address spaces. That is the price paid for simultaneous access.

Because the program and the data occupy different address spaces, the processor cannot ordinarily access the program memory with the same instructions used for data - constants, lookup tables and firmware updating need dedicated mechanisms.

4Modified Harvard and a comparison of the models11 min

Strict separation improves throughput, but reduces the flexibility of the software; a completely unified memory is easy to program, but reintroduces competition for access. The modified Harvard architecture aims to combine the advantages of both: there is no single universal implementation, but the most common variant in modern processors keeps a unified memory space at the program level, with separate caches for instructions (I-cache) and data (D-cache) close to the core.

Mind the I-cache / D-cache coherence

When the program writes an area of memory that is then to be executed as code (for example, when loading new firmware), the data written may remain temporarily only in the D-cache, while the I-cache still holds an old copy. The system software must explicitly synchronize and invalidate the caches before jumping into the new code - a classic source of bugs that are hard to reproduce in systems that do OTA updates.

Characteristicvon NeumannHarvardModified Harvard
Organization of the memorycommondistinctseparation only at certain levels
Concurrent accesslimitedsimultaneoussimultaneous through I-cache/D-cache
Address spacesunifiedseparateunified, with locally separate paths
Flexibilityhighlowerhigh, with more complex hardware
Representative usesthe classic conceptual modelDSPs, strict microcontrollersmost modern processors
To remember Labelling a processor "von Neumann" or "Harvard" can hide important details: a processor may present a single address space to the software, but internally use separate caches and multiple buses. For embedded systems, the choice influences response time, energy consumption, the complexity of the software and the possibility of updating the firmware - exactly the threads we discussed in Lecture 01.

5The ARM family of processors10 min

ARM is by far the most widespread family of processor architectures in the world - from microcontrollers with minimal resources up to telephones, automotive systems and servers. Its success is tied to the RISC principles, energy efficiency and adaptability to very different requirements. The name started from Acorn RISC Machine (in the 1980s), then Advanced RISC Machines; today "Arm" is used.

The licensing model

ARM (the company) does not manufacture processors - it licenses architectures and processing cores to manufacturers of integrated circuits. That is why two different microcontrollers may use the same Cortex-M core, yet have completely different peripherals and memory organization - the core is only one piece of the puzzle.

The ARM family is organized into three architectural profiles, each optimized for a distinct class of systems, not merely "more/less powerful":

ProfileMain fieldMemory managementOperating systemExamples
Cortex-A (Application)systems with a complex OSMMU, virtual memory, advanced cacheLinux, Android, Windowstelephones, routers, servers
Cortex-R (Real-time)critical, deterministic controlMPU, predictable TCM memoriesan RTOS or dedicatedautomotive, storage, industrial
Cortex-M (Microcontroller)embedded microcontrollersoptional MPU, no classic MMUbare-metal or an RTOSIoT, sensors, controllers
Match the ARM profile with the application it suits best

In this lecture we concentrate on the elements of the classic AArch32 model (useful for understanding the evolution of the architecture) and on the particularities of Cortex-M - the most relevant profile for the embedded systems you build in the laboratory.

6The registers of the ARM architecture11 min

The memory hierarchy: energy cost per access

Registers are the fastest storage space available to the processor - they hold the operands, the memory addresses and the control information of the execution. Efficient use of the registers reduces the number of accesses to main memory and directly increases performance.

The classic AArch32 model: R0-R15

RegisterRole
R0-R3function parameters and return values
R4-R11local variables, preserved across function calls
R12a temporary register (used by the compiler / linker)
R13 (SP)Stack Pointer - the top of the current stack
R14 (LR)Link Register - the return address from a function call
R15 (PC)Program Counter - the address of the current/next instruction
Example - calling and returning from a function
arm_call.s
BL   function     ; saves the return address in LR, jumps to "function"
; ... the program continues ...

function:
  ADD R0, R0, #1
  BX  LR          ; returns using the address in LR

If the function in turn calls another function, it must save LR on the stack first, otherwise the original return address is lost:

arm_nested_call.s
function:
  PUSH {R4, LR}
  MOV  R4, R0
  BL   other_function
  ADD  R0, R0, R4
  POP  {R4, PC}    ; loading directly into PC performs the return

Cortex-M particularities

Cortex-M keeps R0-R15 with similar roles, but introduces two distinct stack pointers: MSP (Main Stack Pointer), used after reset and in Handler Mode, and PSP (Process Stack Pointer), used by the application or by the tasks of an RTOS. The separation isolates the stack of the kernel and of the exception routines from the stacks used by the application - a useful protection when an application error must not corrupt the system's exception handling.

7The Program Counter and pipelined execution10 min

The PC register holds the address of the next instruction. In a processor without a pipeline, its value would correspond directly to the instruction being executed. In a pipelined processor - every modern ARM processor - several instructions are simultaneously in different stages, and the exact meaning of the value read from PC depends on the organization of the architecture.

Instr. 1 Fetch Decode Execute Instr. 2 Fetch Decode Execute Instr. 3 Fetch Decode Execute Instr. 4 Fetch Decode Execute
Fig. 2 - A three-stage pipeline: once "filled", one instruction is completed every cycle, although each individual instruction traverses the whole pipeline.
The speed-up given by the pipeline and the cost of hazards
The price of a branch When the processor executes a conditional branch, the instructions already in Fetch and Decode (on the "wrong" path) have to be discarded, and the pipeline has to be "refilled" from the new address - a direct performance penalty. High-performance processors use branch prediction to reduce that cost; simple microcontrollers, with a short pipeline, accept it in exchange for a simpler and more predictable implementation.

This observation explains a practical rule of embedded programming: code with many unpredictable conditional branches may be slower than "the number of instructions" would suggest - the pipeline pays a penalty at every mispredicted branch.

8The status registers CPSR and xPSR10 min

Besides the general registers, ARM processors keep the status information - flags resulting from arithmetic operations, the privilege level, the state of the control mechanisms - in dedicated registers. The structure differs between the classic AArch32 model and Cortex-M.

CPSR (Current Program Status Register) - classic AArch32

FlagMeaning
Nthe result has its sign bit set (negative)
Zthe result of the operation is zero
Cthe operation produced a carry or a borrow
Vthe signed operation produced an arithmetic overflow
Example - updating and reading the flags
cpsr_flags.s
CMP R0, #0      ; a "phantom" subtraction: updates N,Z,C,V without storing the result
BEQ value_zero  ; branch taken only if the Z flag is set

If R0 holds the value zero, the Z flag is set and the conditional branch is taken - exactly the mechanism behind any if compiled for ARM.

xPSR - Cortex-M

Cortex-M gathers conceptually the same information in the xPSR register, made up of three components: APSR (the flags N, Z, C, V, with a role identical to those in CPSR), IPSR (the number of the active exception/interrupt) and EPSR (the state of the execution). The essential difference: Cortex-M does not use the "mode" field of CPSR, nor the IRQ/FIQ/Supervisor modes of classic AArch32 - it has its own, simpler model, discussed in the following section.

Why the difference matters A program written for Cortex-M must not assume the existence of the IRQ/FIQ modes or of the classic CPSR register, even though they appear frequently in material about "the ARM architecture" in general. Always check whether the documentation refers to the classic A/R profile or to Cortex-M.

9Operating modes and interrupt handling10 min

Operating modes

The classic AArch32 model defines several processor modes: User (unprivileged) and the privileged modes System, Supervisor, IRQ, FIQ, Abort, Undefined. The switch to an exception mode is automatic, on the occurrence of the corresponding event. FIQ is designed for a fast response and has additional banked registers - it reduces the need to save general registers on entering the routine.

Cortex-M simplifies this radically: only two modes, Thread Mode (the execution of the application) and Handler Mode (the handling of exceptions, always privileged). Thread Mode may be privileged or not, depending on the CONTROL register - an organization suited to microcontrollers and to the real-time operating systems (RTOS) that run on them.

Interrupt handling

Interrupts allow the processor to react to events without continuously interrogating the state of every peripheral (polling). Instead of the application constantly checking whether a timer has expired, the peripheral "asks for the attention" of the processor only when the relevant event occurs - a direct saving of cycles and of energy.

Handling an interrupt follows, broadly, four steps: (1) temporarily suspending the current program, (2) saving the context needed to resume it, (3) executing the handling routine (ISR), (4) restoring the context and resuming the interrupted program. Cortex-M automates the first two steps and the last one in hardware, through the exception entry/exit mechanism, which significantly simplifies writing an interrupt routine compared with the classic AArch32 model.

Analogy Polling is like checking every 5 seconds whether an email has arrived. An interrupt is like a notification: you carry on with what you were doing, and are told exactly when it matters. The difference in efficiency is enormous, above all on a battery-powered system that can "fall asleep" between interrupts.
Handling an exception on ARM Cortex-M: put the steps in order

10Conditional execution, the Barrel Shifter, Thumb and AMBA10 min

We close this overview of the ARM architecture with three mechanisms that explain the density and the efficiency of the code generated: conditional execution, the Barrel Shifter, and the Thumb/Thumb-2 instruction sets.

Conditional execution

In the A32 encoding, many instructions include a condition field, attached as a suffix to the mnemonic, allowing an instruction to be executed only if the status flags satisfy a condition - reducing the number of branches needed for short decisions.

cond_exec.s
CMP   R0, #0
MOVEQ R1, #1    ; executed only if R0 == 0
MOVNE R1, #0    ; executed only if R0 != 0

The T32 set (Thumb-2) is more restrictive: it mainly uses direct conditional branches, and some architectures allow the IT (If-Then) block for making a short sequence of instructions conditional.

The Barrel Shifter

A hardware block that shifts or rotates an operand before it enters the ALU - combining a shift with an arithmetic or logic operation, in a single instruction.

Example - indexing an array in a single instruction
barrel_shifter.s
ADD R0, R1, R2, LSL #2   ; R0 = R1 + 4 × R2

Directly useful for computing the address of an element of an array whose elements occupy 4 bytes - exactly the pattern the compiler generates for an array[i] access.

OperationFull nameBehaviour
LSLLogical Shift Leftlogical left shift - equivalent to × 2ⁿ
LSRLogical Shift Rightlogical right shift, introduces zeros
ASRArithmetic Shift Rightpreserves the sign bit - approximate division by 2ⁿ
RORRotate Rightrotation - the bits shifted out re-enter at the other end

Thumb and Thumb-2: code density

A32 instructions, although uniform, always occupy 32 bits - programs can become large, which costs Flash memory, energy in transferring them and bandwidth. The original Thumb set solved this with 16-bit instructions (a compact subset of A32 functionality, with some restrictions on registers and constants). Thumb-2 combines 16- and 32-bit encodings in the same set (T32): frequent operations stay compact, complex ones can use the wide encoding when it is needed.

All Cortex-M processors execute T32 - but not identically: Cortex-M0/M0+ implement a restricted subset, optimized for minimum area and power; Cortex-M3/M4 and more capable cores include a richer Thumb-2 set and optional extensions (DSP instructions on Cortex-M4, for instance). The practical conclusion: code compiled for a Cortex-M4 does not run unchanged on a Cortex-M0 - the compiler must be configured explicitly for the target core.

AMBA interconnect (briefly)

A modern SoC connects the processor, the memory, the DMA and the peripherals through a family of protocols standardized by ARM, called AMBA: AXI for high-throughput interconnects (processor, memory, accelerators), AHB/AHB-Lite for a system bus with good performance and lower complexity (frequent in Cortex-M microcontrollers), and APB for simple peripherals (GPIO, UART, timers), where simplicity and low power matter more than throughput.

11Frequent mistakes5 min

  • "Harvard always means two physically separate, completely isolated memories." In modern processors, "Harvard" most often refers to modified Harvard: a unified main memory, but separate paths (I-cache/D-cache) close to the core. Check whether the discussion refers to the architecture visible to the programmer or to the internal microarchitectural organization - they often differ.
  • "CPSR and xPSR are the same thing, just under a different name." Conceptually they partly overlap (both hold N,Z,C,V), but Cortex-M does not have the IRQ/FIQ/Supervisor modes of classic AArch32 - the exception model is structurally different. When you read ARM documentation, check first whether it refers to the classic A/R profile or to Cortex-M - many confusions start here.
  • "Fewer Thumb instructions always means smaller code." One 32-bit A32 instruction may replace two 16-bit Thumb instructions - the real advantage depends on the application, the compiler and the distribution of the instructions used, not merely on the nominal length of the encoding. Do not compare code density from the length of an isolated instruction - measure the size of the final compiled binary.

12Summary and glossary5 min

The organization of the memory (von Neumann, Harvard, modified Harvard) and the architecture of the processor (registers, pipeline, modes, interrupts, instruction sets) together determine how fast, predictable and energy-efficient an embedded system can be. The ARM family dominates microcontrollers today precisely because it offers, through the A/R/M profiles, solutions adapted to every combination of requirements - from telephones running Linux to battery-powered sensors with no operating system.

Von Neumann bottleneck
the throughput limitation of a common program/data bus.
I-cache / D-cache
separate caches for instructions and for data respectively.
Cortex-A/R/M
the ARM profiles for applications, real time, microcontrollers.
SP, LR, PC
Stack Pointer, Link Register, Program Counter (R13-R15).
MSP / PSP
Main/Process Stack Pointer - the two stacks of Cortex-M.
CPSR / xPSR
status registers - classic AArch32 and Cortex-M respectively.
Barrel Shifter
a hardware block for shifting/rotating combined with the ALU.
Thumb / Thumb-2
compact instruction sets, of 16 and 16+32 bits.
AMBA (AXI/AHB/APB)
a family of on-chip interconnect protocols.

13Self-check questions10 min

  1. What is the fundamental difference between the von Neumann and the Harvard architectures?
  2. What is the modified Harvard architecture and why is it the most frequently met today?
  3. What are the three ARM architectural profiles and what is each optimized for?
  4. What is the role of the registers SP, LR and PC in the AArch32 model?
  5. What are MSP and PSP in Cortex-M and why are there two stacks?
  6. How does a three-stage pipeline work and what happens at a conditional branch?
  7. What flags do the CPSR/xPSR registers hold and what does each of them mean?
  8. What do the Thumb and Thumb-2 instruction sets solve?
  9. What role does the Barrel Shifter play in an ARM instruction?
  10. Name the three AMBA protocols and the field of use of each.

14Where to go next2 min

With the architecture of the processor and the organization of the memory clarified, the next lecture moves on to a constraint that returns constantly in embedded systems: energy consumption. We shall link today's concepts directly (operating modes, interrupts, pipeline) to the concrete techniques by which a microcontroller reduces its consumption - sleep, wake-on-interrupt, frequency scaling.

Many of the concepts of this lecture (registers, pipeline, interrupts) become tangible only in the laboratory - follow them concretely in the laboratories of the course as you work directly on ARM Cortex-A (Raspberry Pi 5) and Cortex-M/Xtensa (ESP32) hardware.