LECTURE 06

Software Power Management and Energy Optimization at the Software Level

Duration: 118 min of teaching Level: undergraduate, year III - recommended after Lecture 05 Course: Embedded Systems Associated laboratory: Laboratory 01 PDF: download the notes RO versiunea română

Lectures 04 and 05 treated energy from the point of view of the hardware. This lecture goes straight to the code you write: how the choice of algorithm, the access to memory and the structure of the loops influence the measurable consumption of a program, how a software energy-management loop is built, and how a program can anticipate future load in order to choose proactively the right DVFS operating point.

1The subject and structure of the lecture6 min

There is a frequent and mistaken assumption: that saving energy is mainly a hardware problem, and that the software "merely runs" on what it is given. In reality, the same requirements can be implemented with differences in consumption of the order of tens of per cent, purely through the way the code is written - with no change of hardware at all.

Recap of the previous lectures
  • CMOS dynamic power depends quadratically on voltage - DVFS exploits exactly that (Lectures 04-05)
  • Amdahl's Law limits the benefit of parallelism (Lecture 05)
  • The break-even time decides whether a transition into a low-power mode pays off (Lecture 05)

Learning outcomes

  • To recognize algorithm, code and memory optimizations with a real energy impact
  • To design a software loop for monitoring and controlling energy
  • To decide when DMA is more energy-efficient than a CPU-controlled transfer
  • To choose between reactive, predictive and hybrid power-management policies
  • To apply and compare simple methods of load prediction for software DVFS

2The role of software in energy consumption6 min

Attention turns implicitly, and often, towards hardware when we discuss consumption: the processor, the memories, the converters, the sensors. But the software decides how long those components stay active, at what frequency, in what order and with what data - so it decides the greatest part of the real energy profile of the system.

To remember The hardware sets the limits of what is possible (the best power per operation, the power modes available); the software decides how close the real system gets to those limits. An excellent microcontroller, badly programmed, can consume more than a modest one that is well programmed.

3Energy optimization at the level of algorithm and code9 min

The most effective way of reducing software energy is eliminating the work that should not be done at all - an optimization at the level of the algorithm can avoid millions of instructions, while a local micro-optimization sometimes saves only a few cycles.

The energy of an algorithm
E_algorithm = N_op·E_op + N_mem·E_mem + E_control + E_I/O. The asymptotic complexity (the O notation) describes how the number of operations grows, but does not include the different cost of the operations, the memory accesses or the optimizations of the compiler - a theoretically "better" algorithm may be more energy-costly in practice.
A classic example - linear versus binary search

For a sorted array with N=1,000,000 elements: a linear search needs, in the worst case, N=1,000,000 comparisons; a binary search needs ⌈log₂N⌉ ≈ 20 comparisons - a difference of orders of magnitude. But: a binary search requires sorted data and random access; if the array has to be sorted for a single search, the cost of sorting may exceed the saving. The right choice depends on the number of searches, the frequency with which the data change and the memory available - not merely on the asymptotic complexity in isolation.

Incremental computation instead of recomputation

A moving average computed "directly" sums all N elements of the window at every new sample. The incremental variant updates only the difference:

incremental_average.cpp
// S[k] = S[k-1] + x[k] - x[k-N]
running_sum += new_value;
running_sum -= oldest_value;
const float average = running_sum / static_cast<float>(window_size);

The number of operations falls radically - but a circular buffer with the previous values has to be kept, so the energy saved on computation has to be weighed against the energy of the extra buffer accesses. The pattern "recognize the redundant work and eliminate it" returns constantly in software energy optimization.

4Optimizations at code level8 min

TechniqueThe ideaWatch out for...
Moving invariant computation out of the loopdo not recompute a constant value at every iterationan optimizing compiler does this automatically, if it can prove the invariance
A data type with an explicit size (uint8_t)saves memory and bandwidththe processor may compute on 32-bit registers anyway - it guarantees no speed
A shift instead of a division by 2ⁿa cheaper instructionthe compiler usually makes the transformation automatically; for signed values, the result may differ
Short-circuit evaluation (&&, ||)put the cheap, frequently false test firstdo not change the order if the functions have side effects
Loop unrollingfewer control comparisons/brancheslarger code may increase Flash accesses - the opposite of the intended effect
Dynamic memory allocation

Dynamic allocation and deallocation introduce management time, fragmentation, variable temporal behaviour and additional memory accesses. In real-time systems with limited memory, static buffers, object pools or allocation done once, at initialization, are often preferred - more predictability, less management cost.

A worked exercise - how much does shortening the active time save?

A piece of work initially runs for 20 ms at P_active = 50 mW. After the code is optimized, it takes 12 ms. What is the energy saving, from shortening the active time alone?

See the solution

E₁ = 50 mW × 20 ms = 1 mJ

E₂ = 50 mW × 12 ms = 0.6 mJ

Saving = (1 - 0.6)/1 × 100% = 40%

If the system then goes to rest, the 8 ms freed produce an additional, indirect saving. But careful: if shortening the time required raising the voltage (in order to run faster), the total energy may grow, not fall - always assess the whole operating point, not merely the time.

Energy optimization of code: in what order the interventions are made

5Optimizing memory access and data transfers13 min

In modern embedded systems, the energy of moving data may be comparable with, or even greater than, the energy of the arithmetic operations carried out on it - optimizing memory is as much an energy problem as it is a performance one.

The energy of a transfer
E_transfer = N_accesses·E_access + N_bits·E_bit. It is reduced by lowering the number of accesses, the number of bits transferred, or the distance between the memory used and the computing unit.

Memory is organized hierarchically: registers (the fastest, of minimal capacity) → cache/TCM → internal SRAM → internal Flash → external RAM/Flash → external storage/network (the slowest, of maximum capacity, with the highest energy per access). The key principle is locality: temporal locality (a value used recently will probably be used again soon) and spatial locality (after a location has been accessed, its neighbours are likely to be accessed) - the properties on which the whole efficiency of cache memories rests.

Example - sequential versus indirect access
sequential_access.cpp
for (std::size_t i = 0; i < size; ++i) sum += data[i];          // sequential - efficient
for (std::size_t i = 0; i < size; ++i) sum += data[index[i]];   // indirect - more cache misses

Sequential access makes efficient use of the cache lines and of burst transfers; indirect access may jump to distant locations, generating additional cache misses, plus the cost of reading the array of indices. Similarly, for matrices in C/C++ (stored by rows), traversing for(row) for(column) is far more efficient than reversing the loops.

Block processing (blocking/tiling) keeps a region of data in local memory and reuses it before loading the next block - the technique behind matrix multiplication, image processing and efficient convolutions. The reuse is measured by R_reuse = N_operations / N_external_transfers - a large value means that every value loaded is used many times before being "thrown away".

DMA versus a processor-controlled transfer

The efficiency condition of DMA
E_DMA = E_configuration + E_transfer + E_completion versus E_CPU = E_instructions + E_transfer + E_waiting. DMA is advantageous if E_DMA < E_CPU - for very short transfers, configuring the DMA may be more costly than a simple copy; for large blocks or periodic transfers, DMA wins almost always, because the processor can enter rest while the transfer takes place.

Even the data structures matter: the fields of a structure may be aligned by the compiler with empty gaps (padding), increasing the memory occupied. Reordering the fields from the largest to the smallest type often reduces that padding, with no functional change to the code at all.

6Power-aware software and the energy control loop11 min

An energy-efficient application reduces at the same time the duration of the activity and the number of activations of the resources.

The energy of one operating cycle
E_cycle = Σ Pᵢ·Tᵢ + Σ E_transition,j - the energy can be reduced by shortening the high-power states, selecting states with lower power, reducing the number of transitions, grouping activities and eliminating useless activations.
Grouping activities (batching)

If a sensor has to be read several times within a short interval, it may be more efficient to leave it active until the whole group of measurements is finished, rather than switching it on and off for each individual sample. If, on the other hand, the measurements are separated by long intervals, keeping the sensor active between them wastes energy - the decision depends on the duration of the inactivity, the settling time and the start-up energy, exactly the break-even-time logic discussed in Lecture 05.

The software loop for monitoring and controlling energy

A typical "power-aware" piece of software works as a control loop with five steps: (1) monitoring the state of the system (CPU load, battery level, temperature, events), (2) estimating the performance required, (3) choosing the energy configuration (frequency, voltage, state), (4) applying the configuration, (5) checking the result and the deadlines.

Constrained optimization, at the software level
minimize E_total, subject to T_response ≤ T_max, N_missed_deadlines = 0, Q_service ≥ Q_min. The policy does not pursue the minimization of energy exclusively - a control system has to respond in time, and a communication device has to respect the windows of the protocol.

The quality of service (Q_service) may mean the sampling rate, the precision of the measurements, the throughput of the communication or the frame rate, depending on the application - defining it correctly is the first step of any well-designed energy control loop.

7Software models for estimating consumption6 min

In order to take good decisions, the energy control loop needs an estimate of the consumption - direct measurement, in real time, with precision instruments, is not always possible on the target device. There are two complementary approaches: models based on instructions/blocks of code (each type of instruction or functional block is associated with an estimated energy, from a calibrated table) and models based on profiling and states (energy associated directly with the execution states observed: active, memory access, peripheral active, sleep).

To remember Any software model for estimating energy must be calibrated and validated through real measurements on the target platform - a generic model, not confirmed empirically, may lead to wrong control decisions precisely in the limit situations where it matters most (an almost empty battery, a tight deadline).

8Software management of the energy states8 min

The management of energy states discussed at the hardware level in Lecture 05 (DPM) has a direct counterpart at the software level: the operating system or the firmware has to decide explicitly when it changes the state of a resource, what context it saves before the transition, and how it restores the context on waking.

The typical wake-up sources (a timer, an external pin, a watchdog, a communication peripheral) must be configured explicitly by the software before entering the low-power mode - a forgotten wake-up source means a system that never wakes again, or one that stays permanently active "to be safe", cancelling any energy benefit.

9Break-even time, revisited at the software level6 min

Race-to-idle against pace-to-idle, on the same deadline

The break-even time (T_BE = E_tr / (P_A - P_S), introduced in Lecture 05) is, at the software level, the explicit criterion that a power-management routine evaluates before every possible transition. The practical difference: at the software level, T_BE has to be estimated online, from measurements or from recent history, not merely read statically from a datasheet - the real behaviour may vary with the temperature, the state of the battery or the version of the firmware.

For systems with several successive power states (Idle → Sleep → Deep Sleep), each threshold has its own T_BE, and the software policy has to choose the deepest state whose T_BE is exceeded by the estimated inactivity - not automatically the deepest one available.

10Software policies: reactive, predictive and hybrid8 min

Lecture 05 introduced the reactive (timeout) and predictive policies at the conceptual level. At the software level, hybrid variants appear as well, with additional stabilizing mechanisms.

PolicyPrincipleTrade-off
Reactive (timeout)enters rest after T_idle ≥ T_timeoutsimple, but energy is wasted waiting for the timeout
Predictiveestimates future inactivity from history/patternsmore precise, but more complex and exposed to prediction errors
Hybrid with hysteresisdifferent thresholds for entering and leaving restavoids rapid oscillation between states ("chattering")
Why hysteresis matters

Without hysteresis, a system sitting right at the timeout threshold can oscillate rapidly between active and resting if the load varies slightly around the threshold - every transition costs energy (see T_BE), so repeated oscillation can consume more than simply staying active. Practical policies explicitly limit the rate of transitions, even at the price of a slightly slower reaction.

The break-even time: is it worth entering rest?

11Load prediction and software-controlled DVFS13 min

For software-controlled DVFS, the system has to anticipate the future load in order to choose the right frequency proactively, not merely reactively, once the delay has already appeared. The load of an interval is defined as W_n = T_active,n / T_interval (for example, W=0.7 means the processor was busy 70% of the interval).

MethodFormulaCharacter
PASTŴ_{n+1} = W_na fast reaction, sensitive to sudden variations
FLATŴ_{n+1} = W_fixedstable, but does not adapt to the real load
LONG_SHORTŴ_{n+1} = α·W_short + (1-α)·W_longa reaction/trend compromise, adjustable through α
AGED_AVERAGEŴ_{n+1} = β·W_n + (1-β)·Ŵ_nan exponential average, filters the variations
CYCLEŴ_{n+1} = W_{n+1-k}good for periodic loads with a known period k
PEAKŴ_{n+1} = max(W_{n-k+1},...,W_n)avoids underestimating, but tends towards higher frequencies
A worked exercise

The recent loads: W_{n-2}=40%, W_{n-1}=55%, W_n=70%. What is the prediction for the next interval, with the PAST, AGED_AVERAGE (β=0.3, previous prediction 50%) and PEAK (a window of 3) methods?

See the solution

PAST: Ŵ_{n+1} = 70% (the last value observed)

AGED_AVERAGE: Ŵ_{n+1} = 0.3 × 70 + 0.7 × 50 = 56%

PEAK: Ŵ_{n+1} = max(40, 55, 70) = 70%

The three methods produce visibly different results (56% vs. 70%) for the same sequence of observations - choosing the right method depends on how "noisy" or periodic the real load of the application is. There is no universally superior method.

Once a load prediction has been chosen, selecting the DVFS operating point has to add a safety margin (similar to the M_f of Lecture 05), in order to cover the prediction error without missing deadlines - and the predictor itself has to be evaluated continuously through: the energy consumed, the prediction error, the number of DVFS transitions triggered and the number of deadlines missed. A "precise" predictor that triggers very frequent DVFS transitions may be, all told, more costly than a simpler and more stable one.

12Power management in modern embedded systems6 min

In modern embedded systems, power management is no longer a single isolated loop, but a coordination between several levels: the hardware offers the states and the transitions possible, the operating system (or the firmware, on bare-metal) schedules the tasks taking energy into account - not merely priority - and the application reports quality-of-service constraints to the levels below.

"Energy-aware" task scheduling can group compatible tasks so as to maximize the common windows of rest, instead of treating them independently. Robust systems add telemetry as well (periodic reporting of the real energy state) and modes of degraded operation - a device with an almost empty battery can explicitly reduce its sampling rate or disable non-essential functions, instead of shutting down abruptly and unpredictably.

13Frequent mistakes5 min

  • "An algorithm with a better asymptotic complexity is always more energy-efficient." False - the O notation ignores the different cost of the operations, the memory accesses and the optimizations of the compiler. A binary search may lose to a linear one if it requires costly prior sorting for a single search. Always assess the complete scenario (the number of real operations, their frequency, the cost of initialization), not merely the complexity formula.
  • "Aggressive manual optimization of the code is always beneficial." Modern compilers already do dead-code elimination, constant propagation and vectorization automatically - manual optimization can sometimes prevent those transformations. Always measure the real effect (time and energy), do not assume it - premature, unverified optimization can do more harm than good.
  • "The most aggressive policy for entering rest is always the best." Without hysteresis or a limit on the rate of transitions, an over-aggressive policy can oscillate rapidly between states, paying the transition cost many times with no net benefit. Always add hysteresis or a minimum time limit between consecutive transitions.

14Summary and glossary5 min

The software does not merely run on the hardware available - it decides, through every choice of algorithm, data structure and memory-access pattern, how close the real system gets to the theoretical energy limits of the hardware. An explicit loop for monitoring and controlling energy, with clear quality-of-service constraints, turns energy management from an implicit concern into an active part of the software architecture. And for proactive DVFS, the choice of the load-prediction method (PAST, AGED_AVERAGE, CYCLE, PEAK and so on) has a measurable impact, which depends on the real pattern of the application's load.

Temporal/spatial locality
the basis of the efficient working of cache memories.
Batching
grouping the activations of a resource in order to reduce transitions.
Power-aware software
code that actively monitors and adjusts the energy configuration.
Hysteresis
different thresholds for entering/leaving a state, prevents oscillation.
Load (workload) W_n
the fraction of the interval in which the processor was active.
PAST / AGED_AVERAGE / CYCLE / PEAK
simple methods of predicting future load.

15Self-check questions6 min

  1. Why does asymptotic complexity not completely describe the energy of an algorithm?
  2. What is batching (grouping activities) and when is it advantageous?
  3. Describe the five steps of a software energy control loop.
  4. When is DMA more energy-efficient than a transfer controlled directly by the processor?
  5. What is hysteresis in a power-management policy and why does it prevent "chattering"?
  6. Compare the PAST and AGED_AVERAGE methods of load prediction.
  7. Why may a very "reactive" predictor (such as PAST) not be the best choice?

16Where to go next2 min

Lectures 04-06 have formed a coherent block about energy: where it comes from, how it is reduced through hardware, and how it is reduced through software. The next lecture changes the subject completely: models of execution and scheduling in real-time embedded systems - how we guarantee that a critical task always meets its deadline, whatever else is running in parallel on the same processor.

The techniques of this lecture (batching, compact data structures, DMA) apply directly in Laboratory 01 and Laboratory 02, where you can measure their real effect on consumption.