Last lecture showed how an RTOS schedules tasks in time; this one shows what happens when those tasks need to cooperate - to share data without corrupting each other. We start from the race condition and the classic synchronization mechanisms, go through the priority-inversion trap and the two protocols that solve it (Priority Inheritance and Priority Ceiling), then move up a level: how the kernel itself is organized (monolithic vs. microkernel), how memory isolates tasks from each other, what multicore means in real time, and how an RTOS is actually chosen for a concrete project.
1The subject and structure of the lecture6 min
A correctly scheduled task that corrupts another task's data is not a working system - it is a system that fails unpredictably. Correct synchronization of access to shared resources is just as essential as the temporal scheduling studied in Lecture 07, and the two problems are linked more tightly than they first appear: a poorly chosen synchronization mechanism can destroy the schedulability guarantees demonstrated earlier.
- A task moves through the Running, Ready, Blocked, Suspended states
- RMS offers fixed priorities, with the sufficient Liu-Layland test; EDF offers dynamic priorities, with the U ≤ 1 test
- The blocking time of a high-priority task, caused by others with lower priority, must be included in the temporal analysis
Learning outcomes
- Identify a race condition and explain why it occurs in an apparently simple operation
- Choose between a mutex, a semaphore and an event group for a given scenario
- Explain priority inversion and apply Priority Inheritance or Priority Ceiling
- Compare the monolithic architecture with the microkernel from the perspective of fault isolation
- Describe the role of the MPU/MMU and the difference between AMP and SMP on multicore platforms
- List the practical criteria for choosing an RTOS in a real project
2Race conditions and critical sections8 min
Even an operation that looks atomic in code can be, at processor level, a sequence of three separate steps:
uint32_t event_count = 0;
void TaskA() { ++event_count; }
void TaskB() { ++event_count; }
event_count++ is not necessarily a single instruction: it decomposes into
reading the value from memory, incrementing it and writing the result back.
If TaskA and TaskB execute this sequence nearly simultaneously, the following can happen:
τA reads 10 → τB reads 10 → τA writes 11 → τB writes 11
Although both tasks incremented the variable, the final value is 11, not 12 - the second write overwrites the result of the first, without knowing the first ever happened. τA's update has vanished completely, silently, with no error signaled.
N_tasks_in_critical_section(R) ≤ 1A critical section must be as short as possible: no blocking operations, no voluntary delays, no computation that could be moved outside it.
Variant A holds a mutex during filtering, writing to Flash and sending over the network. Variant B copies the data under the mutex's protection, then processes it outside the lock. Which one is better, and why?
See the solution
// Variant A - excessively long critical section
lock(data_mutex);
read_shared_data();
perform_complex_filtering();
write_external_flash();
send_network_packet();
unlock(data_mutex);
// Variant B - fast copy, processing outside the mutex
SensorData local_copy{};
lock(data_mutex);
local_copy = shared_sensor_data;
unlock(data_mutex);
perform_complex_filtering(local_copy);
write_external_flash(local_copy);
send_network_packet(local_copy);
Variant B is correct: only the actual copy needs protection, the remaining operations can run
without the mutex held. The maximum duration of the critical section directly bounds the blocking
time of other tasks requesting the same resource: B_i ≥ C_critical,max - a long
critical section in Variant A would propagate large delays to any task blocked on that mutex.
Temporarily disabling interrupts can protect very short sequences between a task and an ISR
(L_IRQ,max ≥ T_interrupts_disabled,max), but does not work on multicore -
masking on one core does not stop another core from accessing the same memory. The
volatile qualifier only affects compiler optimizations; it does not provide
mutual exclusion and does not turn a read-modify-write sequence into an atomic
operation.
3Mutexes and semaphores10 min
Three synchronization primitives cover most of the needs of an embedded system: the mutex, the semaphore and the event group.
| Primitive | Model | Typical use |
|---|---|---|
| Mutex | has an owner; only whoever acquires it can release it | protecting a shared resource (mutual exclusion) |
| Semaphore | counter 0 ≤ S ≤ S_max, no owner | signaling between tasks or between an ISR and a task |
| Event group | set of bits, each marking a condition | waiting on a combination of several conditions |
// Mutex - protecting a shared resource, with a timeout
if (lock_mutex(configuration_mutex, mutex_timeout)) {
shared_configuration = new_configuration;
unlock_mutex(configuration_mutex);
} else {
handle_lock_timeout();
}
An ADC finishes a conversion and raises an interrupt; the processing task waits for that event without spinning the CPU in active waiting.
// in the ISR
void ADC_IRQHandler() {
read_conversion_result();
give_semaphore_from_isr(adc_ready_semaphore);
}
// in the task
void ProcessingTask() {
while (true) {
take_semaphore(adc_ready_semaphore, portMAX_DELAY);
process_adc_sample();
}
}
The binary semaphore (S ∈ {0,1}) fits here: the ISR "gives" the semaphore, and the task "takes" it - blocking efficiently, without polling, until the event arrives.
As in Lecture 07: almost any wait on a mutex or semaphore should have a timeout - the awaited resource might never arrive (a lost message, a faulty peripheral), and a task blocked indefinitely on a resource that never comes is, from the system's perspective, equivalent to a dead task.
4Event groups and message queues6 min
Event groups let a task wait on a logical combination of conditions, not just a single one:
E = e₀ ∨ e₁ ∨ ... ∨ eₘ₋₁ - each bit is set independently by a
different source (a task or an ISR), and a task can wait on any AND/OR combination of these
bits.Message queues carry data, not just signaling - one task "puts" an element, another "gets" one, blocking (with a timeout) when the queue is full or empty. One aspect easy to overlook: ownership of the data transmitted through a queue must be defined explicitly - once a buffer has been sent, the sending task should no longer modify it, otherwise a race condition reappears, just moved from a global variable to an object passed through a channel.
- "I used a message queue, so I'm protected from race conditions." Only if data ownership is also strictly respected - if the sender keeps writing to the buffer it already sent after placing it in the queue, the problem returns in a different shape. Transfer ownership completely: after sending, the sender never touches that buffer again.
5Priority inversion8 min
The three primitives above solve the correctness of data access - but they introduce a new, purely temporal problem when tasks with different priorities share the same resource.
Three tasks: τH (high priority), τM (medium priority), τL (low priority). τL holds a mutex M. The problematic scenario:
- τL acquires mutex M and enters the critical section.
- τH becomes Ready, preempts τL, but then also requests mutex M → τH blocks, waiting for τL to release it.
- τM becomes Ready. Since τL has lower priority than τM, τM preempts τL.
- τM runs for as long as it needs, without using mutex M at all - but while τM runs, τL cannot finish the critical section, so it cannot release M, so τH remains blocked.
The solution is not "avoid mutexes" - it is to explicitly bound how long the blocking can last, through dedicated priority-control protocols, the subject of the next two sections.
6Priority Inheritance Protocol8 min
Priority Inheritance Protocol (PIP) reduces priority inversion by temporarily raising the priority of the task holding the resource.
p_effective_L = max(p_base_L, p_H). After
releasing the resource: p_effective_L → p_base_L.Applied to the earlier scenario:
- τL holds the mutex.
- τH requests the mutex and blocks.
- τL inherits τH's priority.
- τM can no longer preempt τL, because τL's effective priority is now higher than τM's.
- τL quickly finishes the critical section and releases the mutex.
- τH is unblocked and runs.
- τL returns to its base priority.
τH's blocking is now bounded by the duration of τL's critical section, not by the interference of an indeterminate number of intermediate tasks.
| Advantage | Limitation |
|---|---|
| Eliminates interference from intermediate-priority tasks | Does not automatically prevent deadlock |
| Built into many RTOS mutexes, without changing the programming model | Can allow chained blocking |
| Keeps the usual acquire/release model for the resource | Only works if the correct mutex is used, not a binary semaphore |
The task holding the mutex must not be suspended, voluntarily delayed or blocked on a slow operation inside the critical section - if τL, even with an inherited priority, performs a slow, uncontrolled operation in the critical section (for instance a blocking Flash write), the blocking remains long regardless of the protocol. PIP bounds who can delay τH, not how long the critical section itself takes - that discipline remains the designer's responsibility.
Also, in many RTOSes, Priority Inheritance is implemented for mutexes, not for binary semaphores - replacing a mutex with a semaphore, in code that looks equivalent, can completely remove the protection against priority inversion, without any warning.
7Priority Ceiling and design rules10 min
Protocols of the Priority Ceiling type assign each resource a value called the priority ceiling.
Π(Rₖ) = max(pᵢ), over all tasks τᵢ that can use
Rₖ - that is, the highest priority among the tasks that can request the resource.Several variants exist. In the Immediate Ceiling Priority Protocol (the most commonly used variant in practice), the task that acquires a resource is immediately given the priority of that resource's ceiling:
p_effective_i = max(p_base_i, Π(Rₖ))Resource R is used by τH, τM and τL; its ceiling is Π(R) = pH (the highest priority among its
users). When τL acquires R: p_effective_L = max(pL, Π(R)) = pH - τL runs immediately
at the highest priority possible for that resource, even before any higher-priority task actually
requests it. τM cannot preempt the critical section, and the temporal analysis becomes simpler
than under PIP, because blocking no longer depends on when the concurrent request
arrives.
Bᵢ ≤ max(Cⱼ,critical), over the lower-priority tasks τⱼ that touch a resource
relevant to τᵢ.| Protocol | Main advantage | Main limitation |
|---|---|---|
| No protocol | Minimal implementation | Potentially unbounded priority inversion |
| Priority Inheritance | Eliminates interference from intermediate priorities | Does not automatically prevent deadlock and chained blocking |
| Priority Ceiling (original) | Prevents deadlock and bounds blocking | Requires managing the system's ceiling |
| Immediate Ceiling | Predictable behavior, efficient implementation | Requires correctly setting each resource's ceiling |
The choice between Priority Inheritance and Priority Ceiling depends on the services offered by the RTOS, whether all users of a resource can be known in advance, the requirements of formal analysis, and the system's criticality level - Priority Ceiling is better suited to static systems analyzed thoroughly beforehand; it can be difficult to apply in highly dynamic architectures, where not all possible users of a resource are known in advance.
8Kernel architectures: monolithic vs. microkernel10 min
The kernel's architecture establishes how responsibilities are distributed between privileged code and application components - a decision that affects service latency, the cost of communication between components, fault isolation, the surface of privileged code, and the possibility of certification.
The single-address-space (monolithic) architecture
The kernel, the drivers and the application access the same memory, usually linked into a
single executable image: Application + services + kernel + drivers → firmware image.
A service call is typically direct: C_service ≈ C_call + C_kernel, with no address
space switch and no message copying.
Advantages: low latency, small overhead, efficient memory use, simple integration on microcontrollers without an MMU. Main limitation: error propagation - an invalid pointer in a driver can corrupt kernel memory or another task's stack, because there is no hardware boundary between them.
Microkernel and separated services
A microkernel keeps only the essential mechanisms in privileged mode: scheduling, thread
management, inter-process communication (IPC), address spaces, interrupts, and access control to
resources. Drivers, file systems and network stacks run in separate components, reached through
message exchange: client → IPC → server, with an extra cost compared to a direct
call:
C_microkernel = C_call + C_IPC + C_schedule + C_service + C_return
- larger than the direct monolithic call, but with an important structural benefit.The benefit: if an unprivileged driver fails, the kernel can prevent it from accessing other components' memory, and the system can stop the driver, restart the service, or fall back to a degraded mode - without compromising the rest of the system.
These properties depend on the correctness of the kernel, the configuration of permissions,
the services running on top of it, the IPC protocol and error handling. The relevant principle is
least privilege: Permissions(Cᵢ) = minimum necessary - a misconfigured
microkernel can grant a service excessive access to memory or peripherals, effectively negating
the architecture's advantage.
| Criterion | Monolithic / single address space | Microkernel |
|---|---|---|
| Service latency | low | higher (IPC cost) |
| Fault isolation | weak, no dedicated hardware | good, through IPC boundaries |
| Best fit | small microcontrollers, strict latency requirements | mixed-criticality systems, where isolation matters |
| Complexity | low | higher (design + communication) |
9Memory protection: MPU and MMU8 min
The software isolation described above needs hardware support. Two central components are the MPU (Memory Protection Unit) and the MMU (Memory Management Unit).
P(Rₖ) = (rₖ, wₖ, xₖ, qₖ), where rₖ allows reading,
wₖ writing, xₖ execution, and qₖ is the required privilege level.The MMU additionally adds translation of virtual addresses to physical addresses:
A_physical = M(A_virtual) - it enables separate address spaces and paging, at a
higher hardware and software cost than a plain MPU.
In a system with an MPU, an unprivileged task may have access to its own stack, its own data, an explicitly shared buffer and certain peripherals - but not to another task's memory or the kernel's. An access attempt outside the allowed regions raises an exception, and the fault handler can identify the offending task, the address accessed and the type of operation.
10Multicore and mixed-criticality systems10 min
Modern embedded processors increasingly contain multiple cores. Two main organization models: AMP (Asymmetric Multiprocessing) and SMP (Symmetric Multiprocessing).
| Model | Principle | Example organization |
|---|---|---|
| AMP | Core0 → RTOS0, Core1 → RTOS1 - each core runs a separate software instance | one core for control, another for communications, another for non-critical functions |
| SMP | RTOS → {Core0, Core1, ..., Coreₘ₋₁} - a single kernel instance schedules tasks across all cores | a single Ready queue, a single global scheduler |
AMP offers a clearer architectural separation, with explicit communication between cores through shared memory, hardware mailboxes or inter-core interrupts. SMP allows flexible use of the cores, but introduces new complexities: true parallel execution, task migration between cores, cross-core locks, contention on the shared cache and additional difficulties in WCET analysis.
U ≤ m -
tasks can have affinity restricted to a subset of cores
(Aᵢ ⊆ {0,1,...,m-1}), serial sections that cannot be parallelized, shared resources
that introduce cross-core blocking, and hardware interference (cache, buses) that is hard to bound
analytically. Schedulability analysis for multicore is significantly more complex than the
single-processor analysis studied in Lecture 07.Mixed-criticality systems
In a mixed-criticality system, functions with different levels of importance (for example critical control and a graphical interface; a safety function and connectivity) share the same hardware platform. Separation must be achieved both spatially and temporally:
Mᵢ ∩ Mⱼ = ∅ for the private regions of components i and
j. Temporal: Cᵢ ≤ Qᵢ in every interval Pᵢ, where Qᵢ is the budget allocated to the
component, and Pᵢ is the budget's replenishment period.A mixed-criticality system can switch into a high-criticality mode when a critical function's execution time exceeds the normal estimate - in that mode, non-critical activities can be reduced or suspended to preserve the essential functions, much like a form of controlled degradation under overload.
11Modern RTOSes: a comparative look10 min
The embedded operating-system ecosystem spans a wide spectrum: compact kernels for microcontrollers, POSIX-like systems, real-time-preemptible Linux platforms, and microkernels with formally verified isolation.
The comparison must be made at the level of a specific version and configuration - compiler options, drivers, the BSP and the application's configuration can significantly change a product's real maximum latency, memory footprint, certifiability and security.
| Platform | Dominant orientation | What to evaluate |
|---|---|---|
| FreeRTOS | compact kernel, direct integration on small microcontrollers | the additional services the product needs (networking, filesystem) are not included |
| Zephyr | integrated, configurable ecosystem (Devicetree, Kconfig), for connected devices | the complexity of configuration and updating |
| Apache NuttX | POSIX-like interfaces (threads, sockets, filesystem) | the cost of the services enabled |
| Linux PREEMPT_RT | general-purpose functionality and low latency, on processors with an MMU | the full hardware-software configuration must be measured, not assumed |
| Commercial microkernel (VxWorks, QNX) | isolation, professional support, certification packages | licensing cost and vendor dependence |
| seL4 | capability-based microkernel, formally verified for certain configurations | building the services and integrating the application remain the team's responsibility |
FreeRTOS provides tasks, priorities, preemptive or cooperative scheduling, queues, semaphores,
mutexes and event groups - exactly the primitives studied in this lecture - in a static image
linked directly with the application. Zephyr separates the hardware description (Devicetree) from
the application logic, letting the same application run on different platforms. NuttX brings
embedded development closer to the classic POSIX model (pthread_create, file
descriptors, sockets). Linux with PREEMPT_RT shrinks the non-preemptible regions of the kernel, but
remains suitable only for platforms with a capable processor, an MMU and enough memory - and, as
everywhere in this field, using PREEMPT_RT alone does not prove that a deadline is met; the
platform must be measured in the product's actual configuration.
Linux ↔ RTOS. There is no need to pick a single "winner" for the whole system - the
choice is made per component, based on its temporal requirements.12Criteria for choosing an RTOS8 min
There is no universal, optimal RTOS. Selection must start from the system's requirements, not from a technology's popularity, and must answer a single central question:
"Which platform lets us demonstrate the product's requirements, at an acceptable cost and risk over its entire lifetime?" - not "which platform is the most popular" or "which one has the most features".
The criteria must be weighted differently depending on context: for a simple device, memory and cost may dominate the decision; for a critical system, isolation, certification and temporal predictability usually take priority.
| Dimension | What must be evaluated |
|---|---|
| Functional and temporal requirements | tasks, IPC, networking, maximum interrupt latency, jitter, scheduling policy - each marked mandatory/recommended/optional |
| Resources | M_RAM = M_kernel + ΣM_stack,i + M_heap + M_buffer + M_driver, measured with the real drivers and protocols, not just the bare kernel |
| Safety, security, isolation | applicable standard, certification documentation, toolchain qualification, threat model, secure boot |
| Ecosystem and total cost | C_total = C_license + C_integration + C_development + C_validation + C_maintenance + C_migration |
A safety certificate is relevant only if the version used is the one it covers, the configuration is permitted, and the processor and compiler are among the supported ones - a certification package for a different RTOS version does not transfer automatically.
The recommended process goes through six stages: defining measurable requirements, an eliminatory filter of candidates that fail the mandatory requirements, a weighted evaluation of the remaining ones, building a prototype on the real target platform, actual measurements (latency, jitter, RAM/Flash, consumption, behavior under overload), and finally a documented decision - with the eliminated candidates, the assumptions used, and the conditions that would call for a later re-evaluation.
13Frequent mistakes5 min
- "I protected the variable with volatile, so it's thread-safe." False - volatile only affects compiler optimizations, it does not provide mutual exclusion and does not turn a read-modify-write sequence into an atomic operation. Use a mutex or a dedicated atomic operation for real concurrent access.
- "A binary semaphore can always replace a mutex." No - semaphores have no concept of ownership, and Priority Inheritance, in many RTOSes, is implemented only for mutexes. Replacing a mutex with a semaphore can silently remove the protection against priority inversion. Use a mutex to protect resources, a semaphore for signaling.
- "Priority Inheritance fully solves priority inversion, no matter how long the critical section takes." No - PIP bounds who can delay the high-priority task, but it does not shorten the critical section itself; a poorly designed critical section (slow operations, blocking I/O) remains a problem regardless of the protocol. Keep critical sections short, regardless of which priority-control protocol is used.
- "A microkernel alone guarantees a safer system than a monolithic one." Not automatically - a microkernel reduces the privileged code and isolates components, but the actual safety depends on correctly configuring permissions and applying the principle of least privilege. Evaluate the concrete configuration, not just the generic architectural classification.
14Summary and glossary5 min
Correct synchronization of concurrent tasks requires mutual exclusion for shared resources (mutex) and dedicated signaling mechanisms (semaphore, event group, message queue) - the wrong choice between them can silently reintroduce race conditions or remove protections the application unknowingly relies on. When tasks with different priorities share a resource, priority inversion can occur; Priority Inheritance and Priority Ceiling are the two classic protocols that bound it, with different trade-offs between simplicity and predictability. At the kernel level, the monolithic architecture offers minimal latency but weak isolation, the microkernel offers better isolation at a communication cost; the MPU/MMU provide the hardware support isolation needs, and multicore systems (AMP/SMP) together with mixed-criticality systems significantly complicate temporal analysis. Choosing a concrete RTOS must be guided by the product's verifiable requirements, not by a platform's popularity.
15Self-check questions6 min
- Why isn't the operation
counter++necessarily atomic, at processor level? - What is the fundamental difference between a mutex and a semaphore?
- Describe the minimal priority-inversion scenario, with three tasks.
- How does Priority Inheritance reduce the blocking time of a high-priority task?
- What does Priority Ceiling additionally guarantee compared to Priority Inheritance?
- What is the main difference between a monolithic kernel and a microkernel?
- What can, and can't, memory protection through an MPU detect?
- What is the difference between AMP and SMP on a multicore processor?
16Where to go next2 min
The next lecture moves from synchronization and RTOS structure to a different direction: reconfigurable systems and hardware acceleration - why and when it is worth moving a computation from software to dedicated hardware (FPGA, accelerator), and what trade-offs that decision brings.
The mutexes, semaphores, event groups and priority-control protocols discussed here become real code, verifiable on hardware, in Laboratory 03, alongside the concepts from Lecture 07.