LECTURE 09

Reconfigurable Systems and Hardware Acceleration for Embedded Applications

Duration: 120 min of teaching Level: undergraduate, year III - recommended after Lecture 08 Course: Embedded Systems Associated laboratory: Laboratory 05 PDF: download the notes RO versiunea română

So far, everything discussed ran on a processor executing instructions one at a time. This lecture opens a different direction: what happens when we move part of the computation directly into hardware, using an FPGA. We start from the motivation for this decision (profiling, not guessing), go through the internal architecture of an FPGA - LUTs, DSP blocks, configuration memory - through the hardware design flow, up to the central practical question: how is an application split between the processor and reconfigurable logic, and why accelerating a single function rarely accelerates the application by the same amount.

1The subject and structure of the lecture6 min

A fast processor inevitably runs into a limit: raising the clock frequency raises the power consumed, and lowering the voltage is limited by the correct operation of the transistors. The dynamic power of a digital circuit can be approximated by P_dynamic = αC_LV²f, where α is the activity factor, C_L the switched capacitance, V the voltage, and f the frequency - doubling the frequency roughly doubles the power consumed. For this reason, modern architectures no longer chase only "faster", but "compute structures specialized for the task at hand".

Recap from Lectures 04-06
  • Dynamic power depends on V², which is why DVFS reduces energy aggressively
  • Hardware-accelerating an operation directly reduces the time spent at high frequency/voltage for that operation
Today we shift the focus from "how does the code run more efficiently on the CPU" to "what do we gain if it does not run on the CPU at all".

Learning outcomes

  • Explain where the FPGA sits between the general-purpose processor and a dedicated ASIC circuit
  • Identify a candidate function for acceleration based on profiling, not intuition
  • Describe the role of LUTs, DSP blocks and configuration memory in an FPGA
  • Follow the FPGA design flow and interpret a negative timing slack
  • Calculate the efficiency threshold of an accelerator, including the transfer cost
  • Apply Amdahl's Law to the partial acceleration of an application

2Reconfigurable computing and spatial parallelism8 min

Reconfigurable computing occupies an intermediate position between software execution on a general-purpose processor and implementing a function in a dedicated circuit: the hardware functionality can be changed after the device is manufactured, by loading a new configuration. The main platform for this type of computing is the FPGA (Field-Programmable Gate Array) - programmable logic blocks, registers, memories and interconnects, configured to form a circuit specific to the application.

PlatformFlexibilityHardware efficiency
General-purpose processorvery high (you change the program)limited by the general architecture
FPGAhigh (you change the bitstream)high, custom parallelism
ASICfixed (set at manufacturing)very high
The essential difference from programming A program describes a sequence of instructions executed by the same compute units, reused in turn. A hardware description defines the structure of a circuit - where y = a·b + c·d can use, simultaneously, two parallel multipliers followed by an adder, instead of two successive multiplications on the same arithmetic unit.

The central advantage is spatial parallelism: for N independent channels, each with processing time T_C, a sequential implementation has T_sequential ≈ N·T_C, while N simultaneous hardware units can reach T_parallel ≈ T_C, plus the latencies of synchronization and collecting the results.

Example - eight independent channels

Eight channels, each needing 10 µs: sequentially, 8 × 10 µs = 80 µs. With eight parallel units, the time stays close to 10 µs - but the reduction is not automatic: it requires replicating resources, enough memory ports, adequate bandwidth and the absence of data dependencies.

A second, equally important form of parallelism is the pipeline: the algorithm is split into stages, each processing a different element simultaneously. Once the pipeline is filled, a result can be produced every initiation interval (II - the number of cycles between two successive inputs), even though the total latency remains several cycles - the latency/throughput distinction is revisited in detail later in this lecture.

3The motivation for hardware acceleration: profiling, not guessing8 min

Hardware acceleration should not be applied to the whole application. Almost always, only a small part of the code consumes a significant proportion of the total time - these portions are called critical regions, hotspots or dominant functions. The correct process starts with profiling the application, not with picking an accelerator in advance.

A function's contribution to the total time
pᵢ = Tᵢ / T_total
Worked exercise

An application needs 100 ms; image filtering consumes 70 ms. What proportion does the filtering represent, and is it worth accelerating?

See the solution

p_filter = 70/100 = 0.7

70% of the total time is an excellent candidate - accelerating the filter can significantly improve the application. By contrast, accelerating a function that consumes only 2% of the time will have a small effect, no matter how efficient the accelerator built for it is.

A function is a good candidate for hardware if it contributes heavily to the total time, runs frequently, operates on regular data, can be parallelized, and the data transfer does not cancel the gain. This last point is easy to underestimate: the actual time of an accelerator is T_accelerated = T_send + T_setup + T_execution + T_receive - for very short operations, the cost of CPU-FPGA communication can exceed the time saved by hardware execution, a topic we revisit in detail in the partitioning section.

The application's total time has several components

T_application = T_control + T_compute + T_memory + T_I/O. Accelerating the T_compute component can have a small effect if the application is dominated by memory or input-output transfers - the entire data path must be analyzed, not just the mathematical function in isolation.

4Processor, FPGA and ASIC: comparison and cost threshold10 min

The same function, four platforms: energy per operation

Three main categories of compute platforms: the general-purpose processor, the FPGA and the ASIC (Application-Specific Integrated Circuit). They are not mutually exclusive - a product can use a processor, an FPGA and several integrated ASIC blocks at the same time.

CharacteristicProcessorFPGAASIC
Flexibilityvery highhighvery low
Custom parallelismlimited by the architecturehighvery high
Energy efficiencymoderatehigh for suitable tasksvery high
Initial costlowlow or moderatevery high
Unit cost (high volume)depends on the platformrelatively highlow
Updatethrough softwarethrough software and bitstreamgenerally impossible

The ASIC offers maximum performance and minimum consumption for the exact function it was designed for, but at a very high non-recurring cost (NRE) and with no possibility of modification after manufacturing. The FPGA uses extra programmable resources (LUTs, routing, configuration memory) to stay flexible, which generally leads to lower area, speed and power efficiency than an equivalent ASIC.

Total cost and the break-even threshold
C_total = C_NRE + N·C_unit. The volume threshold at which ASIC and FPGA reach the same total cost: N* = (C_NRE,ASIC - C_NRE,FPGA) / (C_FPGA - C_ASIC)
Below N*, the FPGA wins; above N*, the ASIC wins The ASIC has a large NRE and a small unit cost; the FPGA has a small NRE and a larger unit cost. Below the N* threshold, the FPGA's total cost stays lower (the low NRE dominates); above N*, the ASIC's lower unit cost offsets the large initial investment. A common industry strategy: prototype on FPGA → validation → ASIC at high volume - but the migration is not automatic, the hardware description, memories and DSP blocks specific to the FPGA must be adapted to ASIC technology.
Choose the right platform for each scenario

5The internal architecture of an FPGA10 min

A modern FPGA is an array of configurable hardware resources, connected through a programmable network. The configuration establishes both the functions implemented in the logic blocks and the signal paths between them. The main resource categories: configurable logic blocks (CLB), registers, programmable interconnects, input-output blocks, internal memories (Block RAM), specialized arithmetic blocks (DSP) and clock-distribution resources.

The configurable logic block is the basic unit - it typically includes one or more LUTs, flip-flop registers, multiplexers, fast carry-chain logic for efficient adders, and connections to the routing network.

Routing delay can dominate logic delay T_path = T_logic + T_routing, and the maximum frequency is limited by the slowest synchronous path: f_max ≤ 1/(T_clk→q + T_logic + T_routing + T_setup). In many implementations, the routing delay is comparable to, or even larger than, the logic delay - two functionally equivalent descriptions can lead to different maximum frequencies after place and route, a result often surprising to someone coming from the software world, where "equivalent" code is usually just as fast.

Configuration memory and the bitstream

The FPGA's configuration is described by a binary file called the bitstream, which establishes the contents of the LUTs, the state of the routing switches, the I/O block configuration and the DSP/memory resource modes. In SRAM-based devices, the configuration is volatile - the bitstream must be reloaded on every power-up, from external memory or through a processor.

Configuration loading time
T_config ≈ N_bitstream / R_config
Worked exercise

A 32 Mbit bitstream is loaded at a rate of 100 Mbit/s. How long does the configuration take?

See the solution

T_config ≈ 32/100 = 0.32 s

In practice, initialization, verification and internal-resource startup times are added - the real time is longer than this simplified calculation. Some FPGAs support partial reconfiguration: a region of the device can be modified without fully stopping the logic in the other regions, useful for dynamically swapping an accelerator or updating a function without stopping the whole system - at the cost of greater design and verification complexity.

6LUTs and sequential logic8 min

The fundamental logic unit of an FPGA is the LUT (Look-Up Table) - conceptually, a small memory in which the inputs form the address, and the output is the value stored at that address.

The LUT principle
F(x₃,x₂,x₁,x₀) = LUT[x₃x₂x₁x₀] - any boolean function with at most k inputs can be implemented directly in a LUT with k inputs.
Example - the XOR function on two inputs

The truth table of F = A ⊕ B: (0,0)→0, (0,1)→1, (1,0)→1, (1,1)→0. These four values, stored at addresses 00, 01, 10, 11 of a 2-input LUT, implement the function completely - with no "real" logic gate at all, just a memory read.

If a function has more inputs than the LUT size, the synthesis tool decomposes it into several connected LUTs - the decomposition can increase the number of logic levels, the delay and the consumption, which is why the number of inputs of a LUT influences both the flexibility and the efficiency of the architecture.

Combinational vs. sequential logic

A combinational description produces an output that depends only on the current inputs: y(t) = F(x₀(t), x₁(t), ...). A sequential description includes the previous state: q[n+1] = F(q[n], x[n]) - registers attached to LUTs enable the implementation of state machines, counters and pipeline registers.

-- synchronous register with reset and enable, in VHDL
process(clk)
begin
  if rising_edge(clk) then
    if reset = '1' then
      q <= (others => '0');
    elsif enable = '1' then
      q <= d;
    end if;
  end if;
end process;
Pipelining reduces the length of combinational paths Without pipeline registers, three operations with delays T₁, T₂, T₃ chain together: T_comb = T₁ + T₂ + T₃. With registers between stages, the clock period is limited by the slowest individual stage: T_clk ≥ max(T₁,T₂,T₃) + T_registers - the latency expressed in cycles increases, but the throughput can increase significantly, exactly the trade-off studied in detail below, in the performance section.

7Specialized hardware resources: BRAM, DSP, clocks8 min

Implementing every function purely through LUTs and registers would consume excessive resources. Modern FPGAs include dedicated blocks, more efficient for frequently encountered operations.

Block RAM and distributed memory

Distributed memory uses LUTs as storage elements - suitable for small tables and shift registers. Block RAM (BRAM) is a dedicated physical memory, configurable as single-port or dual-port, with variable depth and word width.

The capacity of a memory
M = N_words · W_word. For 1024 words of 16 bits: M = 1024 × 16 = 16 384 bits.

A dual-port BRAM allows two accesses in the same cycle - but not an unlimited number of simultaneous users; more parallel reads require replicating the memory, splitting it into banks, or arbitrating access.

DSP blocks - Multiply-Accumulate

The central operation of a DSP block
P = A·B + C (Multiply-Accumulate) - appears in FIR filters, transforms, convolutions and neural networks.

A FIR filter with N coefficients: y[n] = Σ h[k]x[n-k], k=0..N-1. A fully parallel implementation can use N multipliers simultaneously; a serialized implementation reuses a smaller number of DSP blocks, but needs more cycles - a trade-off expressed conceptually as resources · time ≈ amount of computation. Numeric precision directly influences resource usage: an 8-bit multiplier consumes far fewer resources than a 32-bit one.

Clocks and crossing clock domains

Clock-management blocks (PLL/MMCM) can multiply the frequency, divide it, shift the phase, or generate several independent clock domains. Signals transferred between different clock domains need Clock Domain Crossing (CDC) mechanisms - a two-register synchronizer for a single control bit, an asynchronous FIFO, or Gray code for data buses.

Connecting asynchronous clock domains directly is dangerous

Without a CDC mechanism, a bus connected directly between two independent clock domains can produce metastability and incoherent data - a risk easy to overlook for someone used only to software programming, where this concept has no direct equivalent.

8The FPGA design flow8 min

Developing an FPGA application differs fundamentally from software development: compiling a program produces instructions for an existing processor, while implementing a hardware description produces the configuration of a new digital circuit.

  1. Architecture description - in VHDL/Verilog (RTL level) or through HLS (High-Level Synthesis, starting from a C/C++ function).
  2. Functional simulation - the module is stimulated through a testbench, and the outputs are compared with the expected results: y_HDL(x) = y_reference(x), for a set of scenarios that must include boundary values, reset and overflow situations, not just the nominal case.
  3. Logic synthesis - the description is transformed into a network of LUTs, registers, DSPs and connections.
  4. Placement and routing - every logic element gets a physical position, and routing selects the paths between resources.
  5. Timing analysis - checks whether the circuit meets the desired frequency.
  6. Bitstream generation and hardware validation.
HLS does not automatically guarantee parallelism A simple C++ function (for example a loop that adds two vectors element by element) described for HLS does not automatically become parallel hardware - the tool decides whether the loop is executed sequentially, partially unrolled, fully unrolled, or turned into a pipeline, based on the synthesis directives. The code must be treated as a description of a possible architecture, not just as an ordinary program.
Timing margin (slack)
T_arrival = T_clk→q + T_logic + T_routing; T_required = T_clk - T_setup - T_uncertainty; Slack = T_required - T_arrival. The requirement is met if Slack ≥ 0.
Negative slack - correct in simulation, wrong on hardware

A negative slack indicates a timing violation: the circuit may produce correct results in functional simulation (which ignores physical delays), but may fail on hardware at the requested frequency. Fixing it involves adding extra pipeline registers, reducing the logic levels between registers, or lowering the frequency - but incorrect timing constraints are dangerous in the opposite direction too: if a real path is not analyzed at all, the tool may incorrectly report that the implementation meets the requirements.

The utilization of a resource is U_R = N_R,used / N_R,available - a utilization close to 100% can lead to routing difficulties and a lower maximum frequency; simply fitting within the device's capacity is not always enough for an efficient implementation.

The FPGA design flow: put the stages in order

9Heterogeneous CPU-FPGA systems8 min

A heterogeneous architecture combines a general-purpose processor with reconfigurable logic: the processor runs the control software, and the FPGA implements accelerators and interfaces with strict timing requirements. Communication is split into a control path (configuring the accelerator, writing parameters, reading status) and a data path (the large volumes being processed).

// conceptual control of a memory-mapped accelerator
accelerator->source_address = input_buffer_physical_address;
accelerator->destination_address = output_buffer_physical_address;
accelerator->data_length = number_of_elements;
accelerator->control = start_command;

while ((accelerator->status & completed_flag) == 0U) {
    wait_for_interrupt_or_poll();
}

DMA - avoiding element-by-element copying

Transferring every element through CPU instructions can consume more time than the actual processing. For large volumes, DMA (Direct Memory Access) is used: the processor configures the transfer, but does not copy every word.

The time of a DMA transfer
T_transfer = T_init + D/B_effective, where B_effective is smaller than the theoretical bandwidth, due to arbitration and memory latency.
Double buffering overlaps the transfer with the processing With two buffers (one being processed, the other being transferred, swapping roles every iteration), the duration of an iteration tends toward T_iteration ≈ max(T_transfer, T_accelerator), instead of the sum of the two - a gain that matters especially when the transfer and the computation have comparable durations.

Shared memory and cache coherence

If the processor uses a cache, a modified value may not be written immediately to the memory visible to the accelerator. Before a CPU→FPGA transfer, a flush(cache) may be needed; after the accelerator writes the result, an invalidate(cache) may be needed before the CPU reads it. Synchronization can use polling (simple, but consumes processor time) or interrupts (lets the CPU run other activities until completion).

10Hardware-software partitioning and Amdahl's Law12 min

Partitioning establishes which functions stay in software and which functions are implemented in the FPGA. A wrong choice can produce a very fast accelerator, but a slower system overall, because of the transfers and extra control.

Better suited to the CPUBetter suited to the FPGA
complex control, many branchesdata parallelism
dynamic structuresregular memory access
frequent updateshigh throughput, stable function
Computational intensity
CI = N_operations / N_bytes_transferred - a function with a high CI benefits more from acceleration; a function with a low CI may be limited by external memory, no matter how many arithmetic units the accelerator implements.
Worked exercise - the efficiency threshold of an accelerator

An operation runs on the CPU in 2 ms. An FPGA accelerator computes the result (the compute core itself) in only 0.2 ms - ten times faster - but the data transfers take 1.5 ms. What is the real, end-to-end speedup?

See the solution

T_FPGA,total = T_setup + T_compute = 0.2 + 1.5 = 1.7 ms

S = T_CPU / T_FPGA,total = 2 / 1.7 ≈ 1.18

Although the hardware core is ten times faster than the CPU, the real speedup is only ~18%, because the transfers dominate the total time. Reducing this cost requires processing more data per call, keeping the data in the FPGA between successive operations, or using DMA with double buffering.

Amdahl's Law for partial acceleration
If the proportion p of the application is accelerated by factor S_A: S_total = 1 / ((1-p) + p/S_A)
Worked exercise

60% of an application's time can be accelerated ten times over (p=0.6, S_A=10). What is the overall speedup of the complete application?

See the solution

S_total = 1 / (0.4 + 0.6/10) = 1 / (0.4 + 0.06) = 1/0.46 ≈ 2.17

Although the accelerated function is ten times faster, the complete application is only accelerated by about 2.17 times - exactly Amdahl's Law studied for parallelism in Lecture 05, but applied here to hardware acceleration. Even with an infinitely fast accelerator (S_A → ∞), the limit is S_max = 1/(1-p) = 1/0.4 = 2.5. Optimization must target the entire application flow, not just the isolated function.

Is hardware acceleration worth it? Amdahl plus the cost of transfers

11Performance, latency and resource utilization8 min

Evaluating an FPGA accelerator must distinguish between latency, throughput, initiation interval, frequency, resource utilization and bandwidth - a single "speedup" number does not fully describe the system's behavior.

Latency and throughput in a pipeline
For N_P stages at frequency f: L_pipeline = N_P/f. If the initiation interval is II: Θ_pipeline = f/II.
Worked exercise

A ten-stage pipeline runs at 200 MHz, with II=1. What is the latency of the first result, and how often does a new result appear?

See the solution

L_pipeline = 10 / (200×10⁶) = 50 ns - the first result appears after 50 ns.

With II=1, a new result can be accepted every cycle, i.e. every 1/(200×10⁶) = 5 ns - so once the pipeline is filled, even though each individual element "takes" 50 ns (latency), the system produces a new result every 5 ns (throughput). The two values describe different things and must not be confused.

Throughput is limited by the available bandwidth B_needed = Θ·D, where D is the number of bytes per result. If an accelerator produces 100 million results per second, each 8 bytes: B_needed = 100×10⁶ × 8 = 800 MB/s. A memory that effectively offers only 500 MB/s cannot feed the accelerator at the designed rate, no matter how many parallel units the circuit contains - solutions include local data reuse, Block RAM, burst transfers and reduced numeric precision.

A fair CPU vs. FPGA comparison must use the same input data, the same numeric precision, all transfers (not just the compute core) and the same definition of the measured time: S = T_CPU / T_FPGA,total. The energy for the complete task, not just the instantaneous power, is the relevant metric: E ≈ P_average · T - an FPGA may have an instantaneous power comparable to a small processor, but it can finish the operation much faster, resulting in lower total energy.

12Configuration and security of reconfigurable systems8 min

The ability to reconfigure is an important advantage, but it introduces specific risks: an attacker who can replace or modify the bitstream can change the system's very hardware structure, not just the behavior of the software running on it.

Authenticity through a digital signature
σ = Sign(K_private, H(B)), where B is the bitstream. The device checks: Verify(K_public, H(B), σ) = valid before activating the logic.
Authenticity and confidentiality are different goals The signature proves that a configuration comes from an authorized source and has not been modified. Encryption (B_encrypted = Enc(K_device, B)) protects intellectual property, preventing direct interpretation of the configuration - but encryption alone does not prove authenticity. A secure system uses both mechanisms.

Secure Boot establishes a chain of trust: boot ROM → bootloader → bitstream → software, each stage verifying the next one. A secure update goes through downloading the package, verifying the signature and version, temporary storage, controlled activation and confirming operation - and the version number must satisfy V_new > V_minimum_accepted, a direct protection against rollback attacks (reinstalling an old, correctly signed version that is known to be vulnerable).

Automatic recovery from a failed boot

A robust system keeps a recovery image protected separately from the main image. If the new configuration does not finish booting within a set time, the system automatically falls back to the recovery image, avoiding a "brick" (a completely non-functional device) following a failed update.

Configuration security does not eliminate physical attacks: an adversary with access to the device can analyze power consumption, electromagnetic emissions or operation timing (side-channel attacks), or attempt fault injection - producing controlled errors through voltage, clock or temperature variations, aiming for an error produced exactly during verification to cause an invalid configuration to be accepted. Debug interfaces (such as JTAG), useful during development, must be disabled, authenticated or restricted in the final product - protecting the bitstream alone does not help if the bootloader or the debug interface remain open.

13Frequent mistakes5 min

  • "Accelerating the slowest function tenfold means an application ten times faster." False - Amdahl's Law shows that the overall speedup is limited by the proportion of the application that was actually accelerated, not just by the speedup factor of that part. Calculate S_total = 1/((1-p) + p/S_A) before promising an overall speedup.
  • "A hardware core ten times faster guarantees an accelerator ten times faster, end-to-end." No - the CPU↔FPGA transfer time can completely dominate the actual compute time, especially for short operations or small data volumes. Always include T_setup, T_input and T_output in the real speedup calculation.
  • "A positive slack in the synthesis report guarantees correct operation on hardware." Only if all the relevant paths were correctly analyzed - incomplete or incorrect timing constraints can hide a real path that does not meet the frequency. Verify that all critical paths are covered by the timing constraints, not just that the report shows a positive slack.

14Summary and glossary5 min

The FPGA occupies an intermediate position between the flexibility of a general-purpose processor and the efficiency of a dedicated ASIC, offering spatial parallelism and pipelining configurable after manufacturing. LUT blocks implement arbitrary logic through small memories, and specialized resources (Block RAM, DSP, PLLs) avoid excessive LUT consumption for frequent operations. The hardware design flow - simulation, synthesis, place/route, timing analysis - differs fundamentally from software compilation, and a negative slack signals an error that functional simulation does not detect. Correctly partitioning an application between the CPU and the FPGA must start from profiling, include the full cost of the transfers, and account for Amdahl's Law: accelerating a function, however spectacular, is limited by the proportion that function represents within the complete application. Securing a reconfigurable configuration requires both authenticity (signature) and confidentiality (encryption), plus protection against physical attacks.

FPGA
Field-Programmable Gate Array - a circuit reconfigurable after manufacturing.
LUT
Look-Up Table - a small memory that implements an arbitrary logic function.
Bitstream
the configuration file that establishes the function and routing of an FPGA.
HLS
High-Level Synthesis - automatic generation of hardware from a high-level description (C/C++).
Slack
the timing margin of a synchronous path; negative = a violation of the required frequency.
Initiation interval (II)
the number of cycles between two successive inputs accepted by a pipeline.
Computational intensity
the number of operations performed per byte transferred.

15Self-check questions6 min

  1. What fundamental difference exists between programming a processor and configuring an FPGA?
  2. Why must profiling precede the choice of a function for hardware acceleration?
  3. How is the volume threshold N* calculated, at which the ASIC becomes more advantageous than the FPGA?
  4. Explain the principle of a LUT using the XOR function as an example.
  5. What does timing slack represent, and why does a negative slack not appear in functional simulation?
  6. Why doesn't accelerating a compute core tenfold automatically produce a tenfold overall speedup?
  7. What is the difference between the authenticity and the confidentiality of a bitstream?

16Where to go next2 min

The next lecture changes direction again: energy harvesting - how an embedded system can extract its own energy from the environment (solar, vibration, thermal, RF), and what it means to design a system that can never assume a constant energy source.

The concepts from this lecture - spatial parallelism, pipelining, hardware/software partitioning - show up, applied in practice, in any project that combines a processor with programmable logic; platforms such as Xilinx Zynq or Intel/Altera Cyclone SoC integrate exactly the CPU+FPGA architecture discussed here on a single chip.