The last lecture explained where the consumption of an embedded system comes from and how it is measured. This lecture moves on to the engineering part: how that consumption is reduced methodically - from treating optimization as a constrained problem, to concrete techniques (reducing switching activity, advanced DVFS, parallelism with Amdahl's Law, pipelines, VLIW accelerators) and to the dynamic power management that decides, automatically, when and how to apply them.
1The subject and structure of the lecture6 min
Lecture 04 answered the question "how much does the system consume and why". Today we answer the next, harder question: "how do I reduce that consumption without breaking the other requirements of the system?" - because energy optimization isolated from performance, latency and reliability is never the right answer.
- CMOS dynamic power ≈ αC_effV²f - reducing the voltage brings quadratic savings
- DVFS, clock gating and power gating are the basic hardware techniques
- Autonomy depends on the average power, not on the peak power
Learning outcomes
- To formulate energy optimization as a constrained minimization problem
- To distinguish local from global optimization, at the level of the system
- To choose between the "race to idle" and "pace to idle" strategies for a given scenario
- To apply Amdahl's Law in order to estimate the limit of speed-up through parallelism
- To explain why a hardware accelerator or a VLIW architecture can reduce energy
- To compute the break-even time for a transition into a low-power mode
2Energy optimization as a constrained problem10 min
The objective of energy optimization is not always the lowest instantaneous consumption. A system may run at low power, but for a very long time - in which case the total energy may be greater; in other situations, executing a piece of work quickly and then returning immediately to rest is more efficient. Optimization has to be assessed over the whole operating scenario, not over an isolated snapshot.
minimize E(x), where x = the adjustable parameters (frequency, voltage, algorithm, number
of active cores, the power-management policy), subject to the constraints:T_execution(x) ≤ deadline · P_max(x) ≤ P_allowed ·
T_junction(x) ≤ T_max · Q_service(x) ≥ Q_minAn operating point that minimizes the energy but does not meet the deadline is not a valid solution - just as reducing the sampling rate of a sensor is not acceptable if it leads to critical events being missed. Real energy optimization always plays out inside these constraints, not instead of them.
The levels at which one can optimize
| Level | What can be influenced |
|---|---|
| Circuit and technology | the supply voltage, the switched capacitances, the leakage currents |
| Digital logic | the number of transitions, useless activity, switching caused by hazards |
| Microarchitecture | pipeline, cache, accelerators, parallelism, clock/power gating |
| Operating system / runtime | performance states, rest modes, task scheduling |
| Algorithm and application | the number of operations, memory transfers, volume of data, activation rate |
3Local optimization and global optimization10 min
Reducing the consumption of a single component does not guarantee a reduction of the energy of the whole system. Compressing the data may increase the energy consumed by the processor, but it reduces the energy needed for communication by more; a faster algorithm may require an external memory with high consumption, becoming, at the level of the system, less efficient than the "slow" variant.
E_total = E_processor + E_memory + E_communication + E_sensors +
E_actuators + E_conversionAn optimization has to be assessed by the variation of the whole expression, not of a single term.
In a system, communication represents 60% of the total energy. A compression technique reduces the communication energy by 40%, but increases the energy of the processor by an amount equal to 5% of the initial energy of the system. Is the technique worth applying?
See the solution
The saving in communication: 0.60 × 0.40 = 0.24 (24% of the initial total energy).
The increase at the processor: 0.05 (5%).
The net saving: 0.24 - 0.05 = 0.19 → the total energy falls by 19%, even though
the energy of the processor, taken in isolation, has grown. A "locally" negative optimization of one
subsystem may nevertheless be right at the level of the system.
The fundamental trade-offs of energy optimization extend beyond energy: performance (throughput, execution time), latency (meeting deadlines), precision and quality, reliability (thermal margins), cost/area, and the complexity of the software needed to implement the optimization. There is no universally optimal technique - the right solution depends on the real profile of the application.
4Energy per unit of work: an addition to the metrics of Lecture 046 min
Let us briefly recall the metrics relevant to comparing implementation variants - a subject treated at length in Lecture 04, but with one important addition here: energy per unit of work.
E_work = ∫ P(t) dt, integrated only over the duration of the piece of
work being analysed (processing one sample, one control iteration, one measurement cycle) - it allows
implementations with different execution times to be compared, unlike simple average power.The practical rule stays the same as for any efficiency metric: using instantaneous power or peak current alone, without context, can lead to wrong conclusions - a system with a high peak power may nevertheless have the best total energy, if it finishes quickly and goes back to rest.
5Race to idle and pace to idle10 min
For a piece of work with a deadline, two opposite strategies are possible.
The minimum frequency needed
If a piece of work needs N cycles and has a relative deadline D, the ideal minimum frequency is
f_min = N / D. In practice a safety margin M_f > 1 is needed, for interrupts, the variation
of the execution time, unpredictable memory accesses and the latency of changing the frequency:
f_selected = M_f × N_max / D.
Reducing the frequency is not always "free": it may increase latency, it may violate deadlines, it may increase the share of static energy (the useful activity falls, but the leakage currents carry on), and it may require recalibrating the peripherals derived from the processor clock (baud rate, timers, buses) - details easily missed in a first implementation.
6DVS, DFS and DVFS: the cost of transitions10 min
Three distinct mechanisms often hide under the label "DVFS": DFS (Dynamic Frequency Scaling - changes only the frequency), DVS (Dynamic Voltage Scaling - changes only the voltage), and DVFS proper, which coordinates both.
Every transition between operating points has a cost of its own: the time needed to stabilize the voltage, recalibrating the clock, the energy consumed during the transition and the complexity of the control algorithm that decides when to change the operating point. For very short pieces of work, the cost of the transition may exceed the saving - exactly why a good power-management system assesses transitions over the medium term, not at every millisecond.
7When clock gating and power gating are worth it6 min
Lecture 04 introduced clock gating and power gating as hardware techniques. The practical question that remains: when exactly are they worth applying? The answer depends on a threshold computation, identical in principle to the one we formalize completely in the section on dynamic power management (DPM) below: a transition into a lower-consumption mode pays off only if the time spent there exceeds the break-even time of the transition - the energy saved by switching off must exceed the energy spent on switching off and back on.
Clock gating has a much lower break-even threshold than power gating (it does not lose the state, so coming back is almost instantaneous) - suited to pauses of the order of microseconds. Power gating has a higher threshold (it requires re-powering and, often, reinitialization) - worth it only for pauses long enough to cover that cost.
8Parallelism, energy efficiency and Amdahl's Law13 min
Parallelism can reduce energy, but not automatically - the mechanism by which it helps is often counter-intuitive: it is not executing "several things at once" that saves energy, but the possibility of reducing the voltage while keeping the same throughput.
S(p) = 1 / ((1-q) + q/p), and the limit for p → ∞ is
S_max = 1 / (1-q).If 10% of a piece of work is necessarily serial (q = 0.9), the maximum speed-up possible, whatever
number of cores you add, is S_max = 1/0.1 = 10×. Adding hundreds of cores beyond that point
cannot eliminate the serial part - the parallel efficiency η_p = S(p)/p falls steadily. Check
with the widget above: at q = 0.9 and p = 8, S(p) is far from 8×.
The energy model of parallel execution
The total dynamic power of p identical cores: P_dynamic,p = p·α·C_eff·V_p²·f_p. Parallelism
is energetically advantageous only if E_p < E_1 - merely reducing the execution time
does not guarantee that this condition is met.
One core carries out a piece of work at V₁=1.2 V, f₁=1 GHz. The same work, distributed ideally over two cores at V₂=0.9 V, f₂=500 MHz (keeping the same total throughput). How do the total dynamic powers compare?
See the solution
P₂/P₁ = 2 × (0.9/1.2)² × (0.5/1) = 2 × 0.5625 × 0.5 = 0.5625
The total dynamic power falls to approximately 56.3% of the initial value - although we now have two active cores. The saving comes exclusively from the reduction in voltage allowed by the lower frequency per core; if the voltage had stayed constant, two cores at half the frequency would have consumed about as much as a single core at full frequency.
The costs of parallelism
Parallel execution introduces additional resources - threads, buffers, semaphores, barriers, message
queues, coherence mechanisms - with an energy cost of their own
(E_additional = E_synchronization + E_communication + E_coherence + E_control). For very short
pieces of work, that cost may exceed the saving obtained by reducing the execution time. Moreover, if the
load is not distributed evenly, the cores that finish earlier wait idle -
T_p = max(T₁,...,T_p) + T_synchronization - and carry on consuming static power while they
wait.
9Pipeline and energy efficiency8 min
A deeper pipeline can increase throughput (more instructions completed per second), but it introduces energy costs of its own: every additional stage requires pipeline registers (which switch every cycle, even if they do no visible "useful work"), and a hazard or a mispredicted branch flushes a deeper pipeline at a greater cost (more instructions "on the wrong path" have to be discarded - exactly the mechanism discussed in Lecture 03).
| Effect of a deeper pipeline | Consequence |
|---|---|
| A shorter critical path per stage | a higher maximum frequency becomes possible |
| More pipeline registers | additional energy at every cycle, even with no hazards |
| A larger penalty on a mispredicted branch | more "wasted" instructions have to be discarded |
| A voltage reduction made possible at a higher frequency | partly compensates for the cost of the extra registers |
10VLIW architectures and specialized accelerators11 min
In a superscalar processor, the hardware dynamically detects which instructions can run in parallel. A VLIW architecture (Very Long Instruction Word) moves that decision to the compiler: a single instruction contains several operations, destined for different functional units, scheduled statically, before execution.
ADD R1, R2, R3 | MUL R4, R5, R6 | LOAD R7, [R8] ; three independent operations, issued simultaneously
The energy advantage: the dynamic scheduling logic (expensive, present in any superscalar processor) is
replaced by the compiler - less control logic, hence less energy spent merely on "deciding what to execute".
The limit: the efficiency depends directly on the ability of the compiler to find enough independent
operations; if the parallelism available is low, the unused slots are filled with NOPs, and the utilization
U_VLIW = useful slots / total slots falls - larger code, lower energy efficiency.
Specialized accelerators
An accelerator implements a function directly in hardware (DSP, cryptography, DMA, image processing, neural networks). The sources of its energy efficiency: it eliminates the repeated fetching/decoding of instructions, it uses data paths dedicated to the operation, it reuses the data locally (reducing the costly transfers to external memory) and it can run massively in parallel at a reduced frequency.
In many modern applications, the energy of transferring data may exceed the energy of the
arithmetic operation itself. That is why an efficient architecture maximizes local reuse:
R_reuse = operations / external memory accesses - the more times a transferred value is used
before being "thrown away", the more energy-efficient the accelerator. This is exactly the technique behind
accelerators for convolution and filtering.
E_HW = E_configuration + E_transfer + E_execution + E_reactivation.
Acceleration is energetically advantageous only if E_HW < E_SW (the equivalent energy run
on the general-purpose processor).A cryptographic algorithm consumes E_CPU = 20 µJ per block on the processor. On the accelerator: E_configuration = 2 µJ, E_transfer = 1 µJ, E_computation = 3 µJ. Is it worth accelerating?
See the solution
E_acc = 2 + 1 + 3 = 6 µJ
Saving = (20 - 6) / 20 × 100% = 70%
Yes, decidedly - but notice that, for a very small volume of data (a single block, occasionally), the configuration energy (2 out of 6 µJ, a third of the total) could become dominant. Accelerators are efficient for repeated volumes of data, not for isolated, rare operations.
11Dynamic Power Management (DPM)11 min
DPM (Dynamic Power Management) controls the energy state of the resources according to the activity of the system - unlike DVFS, which adapts the performance of an active resource, DPM decides between states: active, idle, sleep, deep sleep, off.
T_BE = E_tr / (P_A - P_S). The resting state is advantageous only if
T_idle > T_BE - and if the wake-up latency respects the maximum limit accepted by
the application.Reactive vs. predictive policies
T_idle ≥ T_timeout. Simple, cheap to implement, with behaviour that is easy to
verify. Limitations: energy wasted while waiting for the timeout, a slow reaction to long periods of
inactivity, and choosing the threshold is a difficult compromise.In real systems, power management often coordinates several components at once (the processor, the radio, the sensors, the peripherals) - each with its own states, thresholds and wake-up sources - and the policy has to balance the saving of energy against meeting the quality of service required by the application (maximum acceptable latency, minimum sampling rate).
12Frequent mistakes5 min
- "More cores always mean faster and more efficient execution." Amdahl's Law shows the limit clearly: the serial part of the program sets a maximum speed-up, whatever number of cores you add. Beyond that point, the extra cores merely consume additional static power with no benefit. Always compute S_max = 1/(1-q) before investing in further parallelization.
- "A hardware accelerator is always more efficient than software on the general-purpose processor." False for small volumes of data - the configuration and transfer energy may dominate the computation energy proper entirely. Check the condition E_HW < E_SW for the real data volume of the application, not for an ideal case with a large volume.
- "Entering the deepest low-power mode available is always the best choice." Only if T_idle > T_BE. For short pauses, the transition costs more than it saves, and the long wake-up latency may violate response requirements. Compute the break-even time of every available state before designing the power-management policy.
13Summary and glossary5 min
Energy optimization is a constrained minimization problem, not a race to the lowest instantaneous consumption - it must be assessed over the whole system and the whole operating scenario, at levels that range from the circuit up to the algorithm. Race to idle and pace to idle are the two extreme strategies for work with a deadline. Parallelism helps energetically above all through the voltage reduction made possible by the lower frequency per core, limited however by Amdahl's Law. VLIW and specialized accelerators reduce energy by moving complexity out of the control hardware into the compiler or into a dedicated data path. And dynamic power management formalizes, through the break-even time, exactly the question any embedded engineer asks constantly: is this transition worth it?
14Self-check questions6 min
- Formulate energy optimization as a constrained minimization problem.
- Give an example in which a "locally" negative optimization nevertheless improves the system.
- When is the "race to idle" strategy advantageous over "pace to idle"?
- What does Amdahl's Law say about the limit of speed-up through parallelism?
- Why can parallelism reduce energy even if the total number of operations stays unchanged?
- What is the energy-efficiency condition of a hardware accelerator?
- What is the break-even time and why does it matter for power management?
15Where to go next2 min
The next lecture moves from the hardware/architectural techniques to the software level: how an embedded programmer writes code that actively exploits these mechanisms - from explicit calls to power-management APIs, to structuring algorithms so as to maximize the time spent at rest.
Amdahl's Law and the DVFS model discussed today become concrete in Laboratory 01, where you measure directly, on a multi-core Raspberry Pi 5, how close (or far) the real speed-up gets to the theoretical limit computed here.