LECTURE 05

Techniques for Optimizing Energy Consumption

Duration: 119 min of teaching Level: undergraduate, year III - recommended after Lecture 04 Course: Embedded Systems Associated laboratory: Laboratory 01 PDF: download the notes RO versiunea română

The last lecture explained where the consumption of an embedded system comes from and how it is measured. This lecture moves on to the engineering part: how that consumption is reduced methodically - from treating optimization as a constrained problem, to concrete techniques (reducing switching activity, advanced DVFS, parallelism with Amdahl's Law, pipelines, VLIW accelerators) and to the dynamic power management that decides, automatically, when and how to apply them.

1The subject and structure of the lecture6 min

Lecture 04 answered the question "how much does the system consume and why". Today we answer the next, harder question: "how do I reduce that consumption without breaking the other requirements of the system?" - because energy optimization isolated from performance, latency and reliability is never the right answer.

Recap of Lecture 04
  • CMOS dynamic power ≈ αC_effV²f - reducing the voltage brings quadratic savings
  • DVFS, clock gating and power gating are the basic hardware techniques
  • Autonomy depends on the average power, not on the peak power

Learning outcomes

  • To formulate energy optimization as a constrained minimization problem
  • To distinguish local from global optimization, at the level of the system
  • To choose between the "race to idle" and "pace to idle" strategies for a given scenario
  • To apply Amdahl's Law in order to estimate the limit of speed-up through parallelism
  • To explain why a hardware accelerator or a VLIW architecture can reduce energy
  • To compute the break-even time for a transition into a low-power mode

2Energy optimization as a constrained problem10 min

The objective of energy optimization is not always the lowest instantaneous consumption. A system may run at low power, but for a very long time - in which case the total energy may be greater; in other situations, executing a piece of work quickly and then returning immediately to rest is more efficient. Optimization has to be assessed over the whole operating scenario, not over an isolated snapshot.

Energy optimization as a constrained problem
minimize E(x), where x = the adjustable parameters (frequency, voltage, algorithm, number of active cores, the power-management policy), subject to the constraints:
T_execution(x) ≤ deadline · P_max(x) ≤ P_allowed · T_junction(x) ≤ T_max · Q_service(x) ≥ Q_min

An operating point that minimizes the energy but does not meet the deadline is not a valid solution - just as reducing the sampling rate of a sensor is not acceptable if it leads to critical events being missed. Real energy optimization always plays out inside these constraints, not instead of them.

The levels at which one can optimize

LevelWhat can be influenced
Circuit and technologythe supply voltage, the switched capacitances, the leakage currents
Digital logicthe number of transitions, useless activity, switching caused by hazards
Microarchitecturepipeline, cache, accelerators, parallelism, clock/power gating
Operating system / runtimeperformance states, rest modes, task scheduling
Algorithm and applicationthe number of operations, memory transfers, volume of data, activation rate
To remember Optimizations at the higher levels (algorithm, application) often have the greatest impact, because they change the total volume of work demanded of the hardware - but they must be supported by hardware mechanisms actually able to switch off or adapt the resources. A more efficient algorithm does not help if the hardware cannot enter a low-power mode while it has nothing to do.

3Local optimization and global optimization10 min

Reducing the consumption of a single component does not guarantee a reduction of the energy of the whole system. Compressing the data may increase the energy consumed by the processor, but it reduces the energy needed for communication by more; a faster algorithm may require an external memory with high consumption, becoming, at the level of the system, less efficient than the "slow" variant.

The total energy of the system
E_total = E_processor + E_memory + E_communication + E_sensors + E_actuators + E_conversion

An optimization has to be assessed by the variation of the whole expression, not of a single term.

A worked exercise

In a system, communication represents 60% of the total energy. A compression technique reduces the communication energy by 40%, but increases the energy of the processor by an amount equal to 5% of the initial energy of the system. Is the technique worth applying?

See the solution

The saving in communication: 0.60 × 0.40 = 0.24 (24% of the initial total energy).

The increase at the processor: 0.05 (5%).

The net saving: 0.24 - 0.05 = 0.19 → the total energy falls by 19%, even though the energy of the processor, taken in isolation, has grown. A "locally" negative optimization of one subsystem may nevertheless be right at the level of the system.

The fundamental trade-offs of energy optimization extend beyond energy: performance (throughput, execution time), latency (meeting deadlines), precision and quality, reliability (thermal margins), cost/area, and the complexity of the software needed to implement the optimization. There is no universally optimal technique - the right solution depends on the real profile of the application.

4Energy per unit of work: an addition to the metrics of Lecture 046 min

Let us briefly recall the metrics relevant to comparing implementation variants - a subject treated at length in Lecture 04, but with one important addition here: energy per unit of work.

Energy per unit of work
E_work = ∫ P(t) dt, integrated only over the duration of the piece of work being analysed (processing one sample, one control iteration, one measurement cycle) - it allows implementations with different execution times to be compared, unlike simple average power.

The practical rule stays the same as for any efficiency metric: using instantaneous power or peak current alone, without context, can lead to wrong conclusions - a system with a high peak power may nevertheless have the best total energy, if it finishes quickly and goes back to rest.

5Race to idle and pace to idle10 min

For a piece of work with a deadline, two opposite strategies are possible.

Race to idle The processor runs at a high frequency, finishes the work quickly, then enters deep rest. Advantageous if: the power at rest is very low, the transition into rest is fast, running fast does not demand an excessive increase in voltage, and the other components can be switched off immediately afterwards.
Pace to idle The processor runs at a lower frequency, using a larger part of the interval available. Advantageous if: the voltage can be reduced significantly, the resting power is not much lower than the active one, there are thermal or peak-power constraints, or the transitions between states are costly.
To remember The right choice is made on the basis of the energy measured over the whole cycle, not by intuition. The two strategies are the two extremes of the same axis (the working frequency/voltage), and real systems often choose an intermediate point, validated empirically.

The minimum frequency needed

If a piece of work needs N cycles and has a relative deadline D, the ideal minimum frequency is f_min = N / D. In practice a safety margin M_f > 1 is needed, for interrupts, the variation of the execution time, unpredictable memory accesses and the latency of changing the frequency: f_selected = M_f × N_max / D.

Side effects of reducing the frequency

Reducing the frequency is not always "free": it may increase latency, it may violate deadlines, it may increase the share of static energy (the useful activity falls, but the leakage currents carry on), and it may require recalibrating the peripherals derived from the processor clock (baud rate, timers, buses) - details easily missed in a first implementation.

6DVS, DFS and DVFS: the cost of transitions10 min

Three distinct mechanisms often hide under the label "DVFS": DFS (Dynamic Frequency Scaling - changes only the frequency), DVS (Dynamic Voltage Scaling - changes only the voltage), and DVFS proper, which coordinates both.

Why DFS is not used on its own Reducing the frequency alone lowers the dynamic power, but - as we computed in Lecture 04 - does not necessarily produce an important reduction of the energy, because the execution time grows proportionally. Reducing the voltage is far more energy-efficient (a quadratic dependence), but it limits the maximum frequency the circuit can sustain. That is why real systems use validated pairs of voltage and frequency - operating points established by the manufacturer, not arbitrary combinations.
Recap: the DVFS calculator (from Lecture 04)

Every transition between operating points has a cost of its own: the time needed to stabilize the voltage, recalibrating the clock, the energy consumed during the transition and the complexity of the control algorithm that decides when to change the operating point. For very short pieces of work, the cost of the transition may exceed the saving - exactly why a good power-management system assesses transitions over the medium term, not at every millisecond.

7When clock gating and power gating are worth it6 min

Lecture 04 introduced clock gating and power gating as hardware techniques. The practical question that remains: when exactly are they worth applying? The answer depends on a threshold computation, identical in principle to the one we formalize completely in the section on dynamic power management (DPM) below: a transition into a lower-consumption mode pays off only if the time spent there exceeds the break-even time of the transition - the energy saved by switching off must exceed the energy spent on switching off and back on.

Clock gating has a much lower break-even threshold than power gating (it does not lose the state, so coming back is almost instantaneous) - suited to pauses of the order of microseconds. Power gating has a higher threshold (it requires re-powering and, often, reinitialization) - worth it only for pauses long enough to cover that cost.

8Parallelism, energy efficiency and Amdahl's Law13 min

Amdahl's Law: the speed-up plateaus, however many cores you add

Parallelism can reduce energy, but not automatically - the mechanism by which it helps is often counter-intuitive: it is not executing "several things at once" that saves energy, but the possibility of reducing the voltage while keeping the same throughput.

Amdahl's Law
If q = the parallelizable fraction of a program and (1-q) the strictly serial fraction, the speed-up on p processors is S(p) = 1 / ((1-q) + q/p), and the limit for p → ∞ is S_max = 1 / (1-q).
Explore Amdahl's Law
Example - why extra cores have diminishing returns

If 10% of a piece of work is necessarily serial (q = 0.9), the maximum speed-up possible, whatever number of cores you add, is S_max = 1/0.1 = 10×. Adding hundreds of cores beyond that point cannot eliminate the serial part - the parallel efficiency η_p = S(p)/p falls steadily. Check with the widget above: at q = 0.9 and p = 8, S(p) is far from 8×.

The energy model of parallel execution

The total dynamic power of p identical cores: P_dynamic,p = p·α·C_eff·V_p²·f_p. Parallelism is energetically advantageous only if E_p < E_1 - merely reducing the execution time does not guarantee that this condition is met.

A worked exercise - reducing the voltage through parallelism

One core carries out a piece of work at V₁=1.2 V, f₁=1 GHz. The same work, distributed ideally over two cores at V₂=0.9 V, f₂=500 MHz (keeping the same total throughput). How do the total dynamic powers compare?

See the solution

P₂/P₁ = 2 × (0.9/1.2)² × (0.5/1) = 2 × 0.5625 × 0.5 = 0.5625

The total dynamic power falls to approximately 56.3% of the initial value - although we now have two active cores. The saving comes exclusively from the reduction in voltage allowed by the lower frequency per core; if the voltage had stayed constant, two cores at half the frequency would have consumed about as much as a single core at full frequency.

The costs of parallelism

Parallel execution introduces additional resources - threads, buffers, semaphores, barriers, message queues, coherence mechanisms - with an energy cost of their own (E_additional = E_synchronization + E_communication + E_coherence + E_control). For very short pieces of work, that cost may exceed the saving obtained by reducing the execution time. Moreover, if the load is not distributed evenly, the cores that finish earlier wait idle - T_p = max(T₁,...,T_p) + T_synchronization - and carry on consuming static power while they wait.

Amdahl's Law: the limit of speed-up through parallelism

9Pipeline and energy efficiency8 min

A deeper pipeline can increase throughput (more instructions completed per second), but it introduces energy costs of its own: every additional stage requires pipeline registers (which switch every cycle, even if they do no visible "useful work"), and a hazard or a mispredicted branch flushes a deeper pipeline at a greater cost (more instructions "on the wrong path" have to be discarded - exactly the mechanism discussed in Lecture 03).

Effect of a deeper pipelineConsequence
A shorter critical path per stagea higher maximum frequency becomes possible
More pipeline registersadditional energy at every cycle, even with no hazards
A larger penalty on a mispredicted branchmore "wasted" instructions have to be discarded
A voltage reduction made possible at a higher frequencypartly compensates for the cost of the extra registers
To remember Choosing the depth of the pipeline is, like almost every other decision in this lecture, a compromise: a deeper pipeline helps if the application has predictable code, with few unpredictable branches; a shallower pipeline is often preferred in embedded microcontrollers precisely for simplicity, low consumption and smaller penalties on code with frequent branches - typical of control code, not of intensive computation.

10VLIW architectures and specialized accelerators11 min

In a superscalar processor, the hardware dynamically detects which instructions can run in parallel. A VLIW architecture (Very Long Instruction Word) moves that decision to the compiler: a single instruction contains several operations, destined for different functional units, scheduled statically, before execution.

vliw_conceptual.asm
ADD R1, R2, R3  |  MUL R4, R5, R6  |  LOAD R7, [R8]
; three independent operations, issued simultaneously

The energy advantage: the dynamic scheduling logic (expensive, present in any superscalar processor) is replaced by the compiler - less control logic, hence less energy spent merely on "deciding what to execute". The limit: the efficiency depends directly on the ability of the compiler to find enough independent operations; if the parallelism available is low, the unused slots are filled with NOPs, and the utilization U_VLIW = useful slots / total slots falls - larger code, lower energy efficiency.

Specialized accelerators

An accelerator implements a function directly in hardware (DSP, cryptography, DMA, image processing, neural networks). The sources of its energy efficiency: it eliminates the repeated fetching/decoding of instructions, it uses data paths dedicated to the operation, it reuses the data locally (reducing the costly transfers to external memory) and it can run massively in parallel at a reduced frequency.

The energy of moving data

In many modern applications, the energy of transferring data may exceed the energy of the arithmetic operation itself. That is why an efficient architecture maximizes local reuse: R_reuse = operations / external memory accesses - the more times a transferred value is used before being "thrown away", the more energy-efficient the accelerator. This is exactly the technique behind accelerators for convolution and filtering.

The efficiency condition of an accelerator
E_HW = E_configuration + E_transfer + E_execution + E_reactivation. Acceleration is energetically advantageous only if E_HW < E_SW (the equivalent energy run on the general-purpose processor).
A worked exercise

A cryptographic algorithm consumes E_CPU = 20 µJ per block on the processor. On the accelerator: E_configuration = 2 µJ, E_transfer = 1 µJ, E_computation = 3 µJ. Is it worth accelerating?

See the solution

E_acc = 2 + 1 + 3 = 6 µJ

Saving = (20 - 6) / 20 × 100% = 70%

Yes, decidedly - but notice that, for a very small volume of data (a single block, occasionally), the configuration energy (2 out of 6 µJ, a third of the total) could become dominant. Accelerators are efficient for repeated volumes of data, not for isolated, rare operations.

11Dynamic Power Management (DPM)11 min

DPM (Dynamic Power Management) controls the energy state of the resources according to the activity of the system - unlike DVFS, which adapts the performance of an active resource, DPM decides between states: active, idle, sleep, deep sleep, off.

The state-based model
To each state S_i are associated a power P_i, a wake-up latency, the context preserved and the reactivation sources. Moving between states S_i and S_j has an energy E_ij and a duration T_ij of its own - power management means, in essence, choosing correctly the moment of the transition.
The break-even time
For a transition with a total energy (entry+exit) E_tr, between an active state (P_A) and a resting one (P_S): T_BE = E_tr / (P_A - P_S). The resting state is advantageous only if T_idle > T_BE - and if the wake-up latency respects the maximum limit accepted by the application.
To remember A policy that changes states very frequently may consume more than a simpler policy - every transition has a fixed cost, however short the pause was. This is exactly why an over-aggressive timeout (entering sleep quickly) may be worse than staying in idle.
The decision of a DPM policy when an idle period appears: put the steps in order

Reactive vs. predictive policies

A reactive policy (timeout) Enters sleep if T_idle ≥ T_timeout. Simple, cheap to implement, with behaviour that is easy to verify. Limitations: energy wasted while waiting for the timeout, a slow reaction to long periods of inactivity, and choosing the threshold is a difficult compromise.
A predictive policy Estimates the future duration of the inactivity from history, periodicity, the state of the application or statistical/machine-learning models. It can react faster and more precisely, but it adds computational complexity and the risk of a wrong prediction (entering sleep prematurely, with a latency penalty).

In real systems, power management often coordinates several components at once (the processor, the radio, the sensors, the peripherals) - each with its own states, thresholds and wake-up sources - and the policy has to balance the saving of energy against meeting the quality of service required by the application (maximum acceptable latency, minimum sampling rate).

12Frequent mistakes5 min

  • "More cores always mean faster and more efficient execution." Amdahl's Law shows the limit clearly: the serial part of the program sets a maximum speed-up, whatever number of cores you add. Beyond that point, the extra cores merely consume additional static power with no benefit. Always compute S_max = 1/(1-q) before investing in further parallelization.
  • "A hardware accelerator is always more efficient than software on the general-purpose processor." False for small volumes of data - the configuration and transfer energy may dominate the computation energy proper entirely. Check the condition E_HW < E_SW for the real data volume of the application, not for an ideal case with a large volume.
  • "Entering the deepest low-power mode available is always the best choice." Only if T_idle > T_BE. For short pauses, the transition costs more than it saves, and the long wake-up latency may violate response requirements. Compute the break-even time of every available state before designing the power-management policy.

13Summary and glossary5 min

Energy optimization is a constrained minimization problem, not a race to the lowest instantaneous consumption - it must be assessed over the whole system and the whole operating scenario, at levels that range from the circuit up to the algorithm. Race to idle and pace to idle are the two extreme strategies for work with a deadline. Parallelism helps energetically above all through the voltage reduction made possible by the lower frequency per core, limited however by Amdahl's Law. VLIW and specialized accelerators reduce energy by moving complexity out of the control hardware into the compiler or into a dedicated data path. And dynamic power management formalizes, through the break-even time, exactly the question any embedded engineer asks constantly: is this transition worth it?

Race to idle / Pace to idle
opposite strategies: fast-then-rest vs. slow-and-steady.
DVS / DFS / DVFS
scaling voltage only, frequency only, or both in a coordinated way.
Amdahl's Law
S(p) = 1/((1-q)+q/p) - the limit of speed-up through parallelism.
VLIW
Very Long Instruction Word - static scheduling, by the compiler.
DPM
Dynamic Power Management - automatic choice of the energy state.
Break-even time
the minimum resting time that justifies a transition.

14Self-check questions6 min

  1. Formulate energy optimization as a constrained minimization problem.
  2. Give an example in which a "locally" negative optimization nevertheless improves the system.
  3. When is the "race to idle" strategy advantageous over "pace to idle"?
  4. What does Amdahl's Law say about the limit of speed-up through parallelism?
  5. Why can parallelism reduce energy even if the total number of operations stays unchanged?
  6. What is the energy-efficiency condition of a hardware accelerator?
  7. What is the break-even time and why does it matter for power management?

15Where to go next2 min

The next lecture moves from the hardware/architectural techniques to the software level: how an embedded programmer writes code that actively exploits these mechanisms - from explicit calls to power-management APIs, to structuring algorithms so as to maximize the time spent at rest.

Amdahl's Law and the DVFS model discussed today become concrete in Laboratory 01, where you measure directly, on a multi-core Raspberry Pi 5, how close (or far) the real speed-up gets to the theoretical limit computed here.