So far, everything discussed ran on a processor executing instructions one at a time. This lecture opens a different direction: what happens when we move part of the computation directly into hardware, using an FPGA. We start from the motivation for this decision (profiling, not guessing), go through the internal architecture of an FPGA - LUTs, DSP blocks, configuration memory - through the hardware design flow, up to the central practical question: how is an application split between the processor and reconfigurable logic, and why accelerating a single function rarely accelerates the application by the same amount.
1The subject and structure of the lecture6 min
A fast processor inevitably runs into a limit: raising the clock frequency raises the power
consumed, and lowering the voltage is limited by the correct operation of the transistors. The
dynamic power of a digital circuit can be approximated by
P_dynamic = αC_LV²f, where α is the activity factor, C_L the switched capacitance, V
the voltage, and f the frequency - doubling the frequency roughly doubles the power consumed. For
this reason, modern architectures no longer chase only "faster", but "compute structures
specialized for the task at hand".
- Dynamic power depends on V², which is why DVFS reduces energy aggressively
- Hardware-accelerating an operation directly reduces the time spent at high frequency/voltage for that operation
Learning outcomes
- Explain where the FPGA sits between the general-purpose processor and a dedicated ASIC circuit
- Identify a candidate function for acceleration based on profiling, not intuition
- Describe the role of LUTs, DSP blocks and configuration memory in an FPGA
- Follow the FPGA design flow and interpret a negative timing slack
- Calculate the efficiency threshold of an accelerator, including the transfer cost
- Apply Amdahl's Law to the partial acceleration of an application
2Reconfigurable computing and spatial parallelism8 min
Reconfigurable computing occupies an intermediate position between software execution on a general-purpose processor and implementing a function in a dedicated circuit: the hardware functionality can be changed after the device is manufactured, by loading a new configuration. The main platform for this type of computing is the FPGA (Field-Programmable Gate Array) - programmable logic blocks, registers, memories and interconnects, configured to form a circuit specific to the application.
| Platform | Flexibility | Hardware efficiency |
|---|---|---|
| General-purpose processor | very high (you change the program) | limited by the general architecture |
| FPGA | high (you change the bitstream) | high, custom parallelism |
| ASIC | fixed (set at manufacturing) | very high |
y =
a·b + c·d can use, simultaneously, two parallel multipliers followed by an adder, instead
of two successive multiplications on the same arithmetic unit.The central advantage is spatial parallelism: for N independent channels, each with
processing time T_C, a sequential implementation has T_sequential ≈ N·T_C, while N
simultaneous hardware units can reach T_parallel ≈ T_C, plus the latencies of
synchronization and collecting the results.
Eight channels, each needing 10 µs: sequentially, 8 × 10 µs = 80 µs. With eight
parallel units, the time stays close to 10 µs - but the reduction is not automatic:
it requires replicating resources, enough memory ports, adequate bandwidth and the absence of data
dependencies.
A second, equally important form of parallelism is the pipeline: the algorithm is split into stages, each processing a different element simultaneously. Once the pipeline is filled, a result can be produced every initiation interval (II - the number of cycles between two successive inputs), even though the total latency remains several cycles - the latency/throughput distinction is revisited in detail later in this lecture.
3The motivation for hardware acceleration: profiling, not guessing8 min
Hardware acceleration should not be applied to the whole application. Almost always, only a small part of the code consumes a significant proportion of the total time - these portions are called critical regions, hotspots or dominant functions. The correct process starts with profiling the application, not with picking an accelerator in advance.
pᵢ = Tᵢ / T_totalAn application needs 100 ms; image filtering consumes 70 ms. What proportion does the filtering represent, and is it worth accelerating?
See the solution
p_filter = 70/100 = 0.7
70% of the total time is an excellent candidate - accelerating the filter can significantly improve the application. By contrast, accelerating a function that consumes only 2% of the time will have a small effect, no matter how efficient the accelerator built for it is.
A function is a good candidate for hardware if it contributes heavily to the total time, runs
frequently, operates on regular data, can be parallelized, and the data transfer does not cancel
the gain. This last point is easy to underestimate: the actual time of an accelerator is
T_accelerated = T_send + T_setup + T_execution + T_receive - for very short
operations, the cost of CPU-FPGA communication can exceed the time saved by hardware execution, a
topic we revisit in detail in the partitioning section.
T_application = T_control + T_compute + T_memory + T_I/O. Accelerating the
T_compute component can have a small effect if the application is dominated by memory or
input-output transfers - the entire data path must be analyzed, not just the mathematical
function in isolation.
4Processor, FPGA and ASIC: comparison and cost threshold10 min
Three main categories of compute platforms: the general-purpose processor, the FPGA and the ASIC (Application-Specific Integrated Circuit). They are not mutually exclusive - a product can use a processor, an FPGA and several integrated ASIC blocks at the same time.
| Characteristic | Processor | FPGA | ASIC |
|---|---|---|---|
| Flexibility | very high | high | very low |
| Custom parallelism | limited by the architecture | high | very high |
| Energy efficiency | moderate | high for suitable tasks | very high |
| Initial cost | low | low or moderate | very high |
| Unit cost (high volume) | depends on the platform | relatively high | low |
| Update | through software | through software and bitstream | generally impossible |
The ASIC offers maximum performance and minimum consumption for the exact function it was designed for, but at a very high non-recurring cost (NRE) and with no possibility of modification after manufacturing. The FPGA uses extra programmable resources (LUTs, routing, configuration memory) to stay flexible, which generally leads to lower area, speed and power efficiency than an equivalent ASIC.
C_total = C_NRE + N·C_unit. The volume threshold at which ASIC
and FPGA reach the same total cost: N* = (C_NRE,ASIC - C_NRE,FPGA) / (C_FPGA - C_ASIC)prototype on FPGA → validation → ASIC at high volume - but the migration is not
automatic, the hardware description, memories and DSP blocks specific to the FPGA must be adapted
to ASIC technology.5The internal architecture of an FPGA10 min
A modern FPGA is an array of configurable hardware resources, connected through a programmable network. The configuration establishes both the functions implemented in the logic blocks and the signal paths between them. The main resource categories: configurable logic blocks (CLB), registers, programmable interconnects, input-output blocks, internal memories (Block RAM), specialized arithmetic blocks (DSP) and clock-distribution resources.
The configurable logic block is the basic unit - it typically includes one or more LUTs, flip-flop registers, multiplexers, fast carry-chain logic for efficient adders, and connections to the routing network.
T_path = T_logic + T_routing, and the maximum frequency is limited by the slowest
synchronous path: f_max ≤ 1/(T_clk→q + T_logic + T_routing + T_setup). In many
implementations, the routing delay is comparable to, or even larger than, the logic delay - two
functionally equivalent descriptions can lead to different maximum frequencies after place and
route, a result often surprising to someone coming from the software world, where "equivalent"
code is usually just as fast.Configuration memory and the bitstream
The FPGA's configuration is described by a binary file called the bitstream, which establishes the contents of the LUTs, the state of the routing switches, the I/O block configuration and the DSP/memory resource modes. In SRAM-based devices, the configuration is volatile - the bitstream must be reloaded on every power-up, from external memory or through a processor.
T_config ≈ N_bitstream / R_configA 32 Mbit bitstream is loaded at a rate of 100 Mbit/s. How long does the configuration take?
See the solution
T_config ≈ 32/100 = 0.32 s
In practice, initialization, verification and internal-resource startup times are added - the real time is longer than this simplified calculation. Some FPGAs support partial reconfiguration: a region of the device can be modified without fully stopping the logic in the other regions, useful for dynamically swapping an accelerator or updating a function without stopping the whole system - at the cost of greater design and verification complexity.
6LUTs and sequential logic8 min
The fundamental logic unit of an FPGA is the LUT (Look-Up Table) - conceptually, a small memory in which the inputs form the address, and the output is the value stored at that address.
F(x₃,x₂,x₁,x₀) = LUT[x₃x₂x₁x₀] - any boolean function with at
most k inputs can be implemented directly in a LUT with k inputs.The truth table of F = A ⊕ B: (0,0)→0, (0,1)→1, (1,0)→1, (1,1)→0. These four
values, stored at addresses 00, 01, 10, 11 of a 2-input LUT, implement the function completely -
with no "real" logic gate at all, just a memory read.
If a function has more inputs than the LUT size, the synthesis tool decomposes it into several connected LUTs - the decomposition can increase the number of logic levels, the delay and the consumption, which is why the number of inputs of a LUT influences both the flexibility and the efficiency of the architecture.
Combinational vs. sequential logic
A combinational description produces an output that depends only on the current inputs:
y(t) = F(x₀(t), x₁(t), ...). A sequential description includes the previous state:
q[n+1] = F(q[n], x[n]) - registers attached to LUTs enable the implementation of
state machines, counters and pipeline registers.
-- synchronous register with reset and enable, in VHDL
process(clk)
begin
if rising_edge(clk) then
if reset = '1' then
q <= (others => '0');
elsif enable = '1' then
q <= d;
end if;
end if;
end process;
T_comb = T₁ + T₂ + T₃. With registers between stages, the clock period is limited by
the slowest individual stage: T_clk ≥ max(T₁,T₂,T₃) + T_registers - the latency
expressed in cycles increases, but the throughput can increase significantly, exactly the
trade-off studied in detail below, in the performance section.7Specialized hardware resources: BRAM, DSP, clocks8 min
Implementing every function purely through LUTs and registers would consume excessive resources. Modern FPGAs include dedicated blocks, more efficient for frequently encountered operations.
Block RAM and distributed memory
Distributed memory uses LUTs as storage elements - suitable for small tables and shift registers. Block RAM (BRAM) is a dedicated physical memory, configurable as single-port or dual-port, with variable depth and word width.
M = N_words · W_word. For 1024 words of 16 bits:
M = 1024 × 16 = 16 384 bits.A dual-port BRAM allows two accesses in the same cycle - but not an unlimited number of simultaneous users; more parallel reads require replicating the memory, splitting it into banks, or arbitrating access.
DSP blocks - Multiply-Accumulate
P = A·B + C (Multiply-Accumulate) - appears in FIR filters,
transforms, convolutions and neural networks.A FIR filter with N coefficients: y[n] = Σ h[k]x[n-k], k=0..N-1. A fully parallel
implementation can use N multipliers simultaneously; a serialized implementation reuses a smaller
number of DSP blocks, but needs more cycles - a trade-off expressed conceptually as
resources · time ≈ amount of computation. Numeric precision directly influences
resource usage: an 8-bit multiplier consumes far fewer resources than a 32-bit one.
Clocks and crossing clock domains
Clock-management blocks (PLL/MMCM) can multiply the frequency, divide it, shift the phase, or generate several independent clock domains. Signals transferred between different clock domains need Clock Domain Crossing (CDC) mechanisms - a two-register synchronizer for a single control bit, an asynchronous FIFO, or Gray code for data buses.
Without a CDC mechanism, a bus connected directly between two independent clock domains can produce metastability and incoherent data - a risk easy to overlook for someone used only to software programming, where this concept has no direct equivalent.
8The FPGA design flow8 min
Developing an FPGA application differs fundamentally from software development: compiling a program produces instructions for an existing processor, while implementing a hardware description produces the configuration of a new digital circuit.
- Architecture description - in VHDL/Verilog (RTL level) or through HLS (High-Level Synthesis, starting from a C/C++ function).
- Functional simulation - the module is stimulated through a testbench, and the outputs
are compared with the expected results:
y_HDL(x) = y_reference(x), for a set of scenarios that must include boundary values, reset and overflow situations, not just the nominal case. - Logic synthesis - the description is transformed into a network of LUTs, registers, DSPs and connections.
- Placement and routing - every logic element gets a physical position, and routing selects the paths between resources.
- Timing analysis - checks whether the circuit meets the desired frequency.
- Bitstream generation and hardware validation.
T_arrival = T_clk→q + T_logic + T_routing;
T_required = T_clk - T_setup - T_uncertainty;
Slack = T_required - T_arrival. The requirement is met if Slack ≥ 0.A negative slack indicates a timing violation: the circuit may produce correct results in functional simulation (which ignores physical delays), but may fail on hardware at the requested frequency. Fixing it involves adding extra pipeline registers, reducing the logic levels between registers, or lowering the frequency - but incorrect timing constraints are dangerous in the opposite direction too: if a real path is not analyzed at all, the tool may incorrectly report that the implementation meets the requirements.
The utilization of a resource is U_R = N_R,used / N_R,available - a utilization
close to 100% can lead to routing difficulties and a lower maximum frequency; simply fitting within
the device's capacity is not always enough for an efficient implementation.
9Heterogeneous CPU-FPGA systems8 min
A heterogeneous architecture combines a general-purpose processor with reconfigurable logic: the processor runs the control software, and the FPGA implements accelerators and interfaces with strict timing requirements. Communication is split into a control path (configuring the accelerator, writing parameters, reading status) and a data path (the large volumes being processed).
// conceptual control of a memory-mapped accelerator
accelerator->source_address = input_buffer_physical_address;
accelerator->destination_address = output_buffer_physical_address;
accelerator->data_length = number_of_elements;
accelerator->control = start_command;
while ((accelerator->status & completed_flag) == 0U) {
wait_for_interrupt_or_poll();
}
DMA - avoiding element-by-element copying
Transferring every element through CPU instructions can consume more time than the actual processing. For large volumes, DMA (Direct Memory Access) is used: the processor configures the transfer, but does not copy every word.
T_transfer = T_init + D/B_effective, where B_effective is smaller
than the theoretical bandwidth, due to arbitration and memory latency.T_iteration ≈ max(T_transfer, T_accelerator), instead of the sum of the two - a gain
that matters especially when the transfer and the computation have comparable durations.Shared memory and cache coherence
If the processor uses a cache, a modified value may not be written immediately to the memory
visible to the accelerator. Before a CPU→FPGA transfer, a flush(cache) may be needed;
after the accelerator writes the result, an invalidate(cache) may be needed before
the CPU reads it. Synchronization can use polling (simple, but consumes processor time) or
interrupts (lets the CPU run other activities until completion).
10Hardware-software partitioning and Amdahl's Law12 min
Partitioning establishes which functions stay in software and which functions are implemented in the FPGA. A wrong choice can produce a very fast accelerator, but a slower system overall, because of the transfers and extra control.
| Better suited to the CPU | Better suited to the FPGA |
|---|---|
| complex control, many branches | data parallelism |
| dynamic structures | regular memory access |
| frequent updates | high throughput, stable function |
CI = N_operations / N_bytes_transferred - a function with a high
CI benefits more from acceleration; a function with a low CI may be limited by external memory,
no matter how many arithmetic units the accelerator implements.An operation runs on the CPU in 2 ms. An FPGA accelerator computes the result (the compute core itself) in only 0.2 ms - ten times faster - but the data transfers take 1.5 ms. What is the real, end-to-end speedup?
See the solution
T_FPGA,total = T_setup + T_compute = 0.2 + 1.5 = 1.7 ms
S = T_CPU / T_FPGA,total = 2 / 1.7 ≈ 1.18
Although the hardware core is ten times faster than the CPU, the real speedup is only ~18%, because the transfers dominate the total time. Reducing this cost requires processing more data per call, keeping the data in the FPGA between successive operations, or using DMA with double buffering.
S_total = 1 / ((1-p) + p/S_A)60% of an application's time can be accelerated ten times over (p=0.6, S_A=10). What is the overall speedup of the complete application?
See the solution
S_total = 1 / (0.4 + 0.6/10) = 1 / (0.4 + 0.06) = 1/0.46 ≈ 2.17
Although the accelerated function is ten times faster, the complete application is only
accelerated by about 2.17 times - exactly Amdahl's Law studied for parallelism in Lecture 05, but
applied here to hardware acceleration. Even with an infinitely fast accelerator
(S_A → ∞), the limit is S_max = 1/(1-p) = 1/0.4 = 2.5. Optimization must
target the entire application flow, not just the isolated function.
11Performance, latency and resource utilization8 min
Evaluating an FPGA accelerator must distinguish between latency, throughput, initiation interval, frequency, resource utilization and bandwidth - a single "speedup" number does not fully describe the system's behavior.
L_pipeline = N_P/f. If the
initiation interval is II: Θ_pipeline = f/II.A ten-stage pipeline runs at 200 MHz, with II=1. What is the latency of the first result, and how often does a new result appear?
See the solution
L_pipeline = 10 / (200×10⁶) = 50 ns - the first result appears after 50 ns.
With II=1, a new result can be accepted every cycle, i.e. every 1/(200×10⁶) = 5 ns
- so once the pipeline is filled, even though each individual element "takes" 50 ns (latency), the
system produces a new result every 5 ns (throughput). The two values describe different things
and must not be confused.
B_needed = Θ·D, where D is the number of bytes per result. If an accelerator produces
100 million results per second, each 8 bytes: B_needed = 100×10⁶ × 8 = 800 MB/s. A
memory that effectively offers only 500 MB/s cannot feed the accelerator at the designed
rate, no matter how many parallel units the circuit contains - solutions include local data reuse,
Block RAM, burst transfers and reduced numeric precision.A fair CPU vs. FPGA comparison must use the same input data, the same numeric precision, all
transfers (not just the compute core) and the same definition of the measured time:
S = T_CPU / T_FPGA,total. The energy for the complete task, not just the instantaneous
power, is the relevant metric: E ≈ P_average · T - an FPGA may have an instantaneous
power comparable to a small processor, but it can finish the operation much faster, resulting in
lower total energy.
12Configuration and security of reconfigurable systems8 min
The ability to reconfigure is an important advantage, but it introduces specific risks: an attacker who can replace or modify the bitstream can change the system's very hardware structure, not just the behavior of the software running on it.
σ = Sign(K_private, H(B)), where B is the bitstream. The device
checks: Verify(K_public, H(B), σ) = valid before activating the logic.B_encrypted = Enc(K_device, B)) protects intellectual property,
preventing direct interpretation of the configuration - but encryption alone does not prove
authenticity. A secure system uses both mechanisms.Secure Boot establishes a chain of trust: boot ROM → bootloader → bitstream →
software, each stage verifying the next one. A secure update goes through downloading the
package, verifying the signature and version, temporary storage, controlled activation and
confirming operation - and the version number must satisfy
V_new > V_minimum_accepted, a direct protection against rollback attacks
(reinstalling an old, correctly signed version that is known to be vulnerable).
A robust system keeps a recovery image protected separately from the main image. If the new configuration does not finish booting within a set time, the system automatically falls back to the recovery image, avoiding a "brick" (a completely non-functional device) following a failed update.
Configuration security does not eliminate physical attacks: an adversary with access to the device can analyze power consumption, electromagnetic emissions or operation timing (side-channel attacks), or attempt fault injection - producing controlled errors through voltage, clock or temperature variations, aiming for an error produced exactly during verification to cause an invalid configuration to be accepted. Debug interfaces (such as JTAG), useful during development, must be disabled, authenticated or restricted in the final product - protecting the bitstream alone does not help if the bootloader or the debug interface remain open.
13Frequent mistakes5 min
- "Accelerating the slowest function tenfold means an application ten times faster." False - Amdahl's Law shows that the overall speedup is limited by the proportion of the application that was actually accelerated, not just by the speedup factor of that part. Calculate S_total = 1/((1-p) + p/S_A) before promising an overall speedup.
- "A hardware core ten times faster guarantees an accelerator ten times faster, end-to-end." No - the CPU↔FPGA transfer time can completely dominate the actual compute time, especially for short operations or small data volumes. Always include T_setup, T_input and T_output in the real speedup calculation.
- "A positive slack in the synthesis report guarantees correct operation on hardware." Only if all the relevant paths were correctly analyzed - incomplete or incorrect timing constraints can hide a real path that does not meet the frequency. Verify that all critical paths are covered by the timing constraints, not just that the report shows a positive slack.
14Summary and glossary5 min
The FPGA occupies an intermediate position between the flexibility of a general-purpose processor and the efficiency of a dedicated ASIC, offering spatial parallelism and pipelining configurable after manufacturing. LUT blocks implement arbitrary logic through small memories, and specialized resources (Block RAM, DSP, PLLs) avoid excessive LUT consumption for frequent operations. The hardware design flow - simulation, synthesis, place/route, timing analysis - differs fundamentally from software compilation, and a negative slack signals an error that functional simulation does not detect. Correctly partitioning an application between the CPU and the FPGA must start from profiling, include the full cost of the transfers, and account for Amdahl's Law: accelerating a function, however spectacular, is limited by the proportion that function represents within the complete application. Securing a reconfigurable configuration requires both authenticity (signature) and confidentiality (encryption), plus protection against physical attacks.
15Self-check questions6 min
- What fundamental difference exists between programming a processor and configuring an FPGA?
- Why must profiling precede the choice of a function for hardware acceleration?
- How is the volume threshold N* calculated, at which the ASIC becomes more advantageous than the FPGA?
- Explain the principle of a LUT using the XOR function as an example.
- What does timing slack represent, and why does a negative slack not appear in functional simulation?
- Why doesn't accelerating a compute core tenfold automatically produce a tenfold overall speedup?
- What is the difference between the authenticity and the confidentiality of a bitstream?
16Where to go next2 min
The next lecture changes direction again: energy harvesting - how an embedded system can extract its own energy from the environment (solar, vibration, thermal, RF), and what it means to design a system that can never assume a constant energy source.
The concepts from this lecture - spatial parallelism, pipelining, hardware/software partitioning - show up, applied in practice, in any project that combines a processor with programmable logic; platforms such as Xilinx Zynq or Intel/Altera Cyclone SoC integrate exactly the CPU+FPGA architecture discussed here on a single chip.