Previous lectures covered, in turn, the components of a microcontroller: ports, interrupts, timers, ADC, UART, SPI, I2C. An engineer who masters all of them does not, by that alone, know how to build a product. Between "I lit an LED with PWM" and "I delivered a system that has run unattended for two years in a greenhouse" lies a distance that is not closed with more knowledge about registers, but with method - the subject of this lecture, followed through a single example, from requirement to test plan.
1Scope and structure of the lecture6 min
Microcontroller projects almost always fail for reasons of method, not programming - rarely because someone misunderstood a prescaler, almost always because nobody wrote down clearly what the system had to do, because the split into modules happened on the fly, or because testing was left until the end. We will follow the entire path, from a vague requirement to a documented system, through a single reference example: an automatic watering system for a greenhouse, revisited identically in every section.
- SPI and I2C solve communication with nearby peripherals - the open question remains how to organize, as a whole, a program that uses all of them coherently (Lecture 09)
- Explain the three main causes for which microcontroller projects fail
- Turn a vague requirement into a verifiable requirement, with a numeric acceptance criterion
- Partition a system into subsystems with interfaces fixed before implementation
- Describe the logic of an application as a finite state machine, starting from a complete transition table
- Distinguish the testing levels (unit, module, system, boundary, robustness) and write a test plan linked to the specification
- Apply the three principles of systematic debugging: minimal reproducible case, measure instead of guessing, separate hardware from software
2Why a project fails: the three causes11 min
The first cause: an unclear requirement. The customer says "I want the system to react quickly" or "I want it to be safe". Nobody disagrees with such statements, because nobody knows what they mean. At acceptance, the customer notices the system reacts in 300 ms and claims it is not "fast", the designer answers that nothing else was asked for, and the discussion cannot end - there is no objective criterion either side can point to. A requirement that cannot be checked by a measurement is not a requirement, it is a wish.
The second cause: poor delimitation of subsystems. When code grows without a structure decided in advance, functions start communicating through global variables, each modifying the other's state. Such a program usually works, up until the first change - that is when it turns out that changing a threshold forces changes in the relay routine, which breaks the display, and fixing the display reintroduces the original bug. The system becomes impossible to maintain not because it is complex, but because it has no internal boundaries.
The third cause, the most costly one: testing postponed to the end. It is tempting to write the whole program and only then try it, because testing along the way seems like a waste of time. In reality, when you integrate ten untested modules and the result does not work, you have no information about where the bug is - any of the ten, or any interaction between them, could be guilty. If each module was verified separately and the system breaks when the tenth one is added, the search space has shrunk to a single module.
The two stages and the cycle between them
Development unfolds in two large stages: design, which begins with the description of the system and ends with the first code draft, and testing, which begins with simulation and ends with documentation. Showing the steps as a list is misleading, because it suggests a one-way flow - in reality there is a feedback loop: if, during testing, the system does not behave according to the specification, it is not only the code that gets corrected, but design is resumed at the level where the mistake occurred, sometimes even the specification itself. A healthy project goes through this cycle several times, with decreasing amplitude each time.
Worth remembering: the design stage produces documents, not code. Block diagrams, requirement tables, transition tables and the pin map outlive the project and determine its long-term value - a working system with no document behind it is a system whose lifetime equals the memory of its author.
3The reference example and the specification13 min
The whole chapter follows a single example: an automatic watering system for a small greenhouse. It measures soil moisture and, below a threshold, opens a solenoid valve through a relay for a limited interval; it has a manual-watering button, a stop button, two status LEDs, and it sends an event log over the serial interface. The first discussion with the customer always produces a second document, just as important as the description: the list of questions - what happens if the sensor fails and permanently indicates dry soil? If someone presses manual watering and stop at the same time? If power drops during watering? An experienced designer is recognized by the quality of this list, because every question asked now is a bug that will not be discovered later in the field.
What makes a requirement verifiable
A requirement is verifiable if, with the system in front of them and a measuring instrument, two independent people are bound to reach the same conclusion about whether it is met.
Verifiable: "the system must command the relay within at most 50 ms of the moment the measured quantity exceeds the configured threshold".
The second fixes a start moment, an end moment, a limit value, and, implicitly, a verification method.
A well-written requirement has four elements: a subject, an observable action, a triggering condition, and a numeric value with a unit and a tolerance. Words that should give pause when they appear in a specification: "fast", "stable", "safe", "friendly", "approximately" - each hides a negotiation that never took place.
Functional requirements describe what the system does; they are relatively easy to state, because they follow directly from the customer's description. Non-functional requirements describe under what conditions it must do those things, and are the source of most late surprises: average consumption decides whether the system can be powered from a solar panel; the temperature range decides whether the oscillator can be the internal one or must be a crystal; unit cost decides between 8 and 32 KB of program memory. None of these requirements appear in the customer's description, but all of them change the project if discovered late.
| Code | Requirement | Verification |
|---|---|---|
| CF1 | Measures moisture once every 1±0.1 s | timing with an oscilloscope |
| CF2 | Below threshold, opens the valve within ≤200 ms | injecting a voltage below threshold |
| CF5 | On stop, the valve closes within ≤50 ms | oscilloscope on both signals |
| CN1 | Average consumption <20 mA | ammeter, 10 minutes |
| CN4 | Thresholds changeable without reprogramming | EEPROM check |
Requirement CF5 (closing the valve within ≤50 ms of pressing the button) must be verified for an architecture where buttons are polled every 2 ms, a press is confirmed after 20 ms of a stable level, and the relay has an actuation time of 10 ms according to the datasheet. Is it achievable?
See the solution
The budget is built by adding, in the worst case, every delay along the cause-effect chain. The press can happen right after a reading: one polling interval is lost, 2 ms. Debounce confirmation: 20 ms. Decision and writing to the relay port: under 1 ms. Relay actuation: 10 ms. Total: 2+20+1+10 = 33 ms, with a margin of 17 ms against the 50 ms limit - the requirement is achievable. If, however, the debounce confirmation had been chosen as 50 ms, as often happens with no justification, the requirement would have been violated by design, and the bug would only have appeared at acceptance measurements, not here, on paper.
4System partitioning and choosing the microcontroller13 min
Partitioning means dividing the system into subsystems, each with a single responsibility, and establishing how they communicate. The fundamental rule: a subsystem's interface is fixed before the subsystem is implemented. The reason is practical - as long as the interface is not fixed, no other subsystem can be written, because nobody knows what it will receive and what it will return. From the moment the interface is written, even just as function prototypes with comments, every module can be developed in parallel.
uint16_t readSensor(void); is not a complete interface. It must specify: what
units, what range of values, with what error, what happens on invalid values, how long the
operation takes, and whether it can be called from an interrupt routine.
The second criterion of good partitioning: each subsystem has a single reason to change. If replacing the moisture sensor with another one forces a change in the code that lights the LEDs, the partitioning is wrong. The watering system is thus split into: moisture acquisition, operator interface (buttons + LEDs), valve control, time base, serial logging, and, in the middle, the application itself, which makes decisions without touching any register.
Choosing the microcontroller
The choice is not made by habit or by whatever part is in the drawer, but as a consequence of the specification and the partitioning: once it is known which signals go in and out, the pin count is determined; once the operating periods are known, the needed peripherals are determined. Only then does the catalog get opened.
| Criterion | Application requirement | Consequence |
|---|---|---|
| Pins | 1 analog, 3 digital inputs, 3 digital outputs, serial | a 28-pin package is enough |
| Peripherals | ADC, timer, USART | no SPI/I2C needed |
| Non-volatile memory | thresholds changeable in the field (CN4) | an EEPROM is mandatory |
| Consumption | under 20 mA average (CN1) | low-power modes |
| Temperature | −5...+55°C (CN2) | industrial variant or checking oscillator drift |
An ATmega328P satisfies every line of the table with margin to spare. Important: the conclusion followed from the table, it did not precede it - if the application had required eight analog channels and consumption under 100 µA, the same method would have led to a different chip. The practical rule: keep a margin of at least 20% on pins, memory and execution time - a project that uses everything it has is a project with no future, for which the first additional requirement forces a change of microcontroller and a board redesign.
5Program architecture and the finite state machine13 min
With the subsystems defined, next comes describing what each one does inside, through activity and state diagrams (the UML language). The quality criterion for an activity diagram is simple and strict: the code must be writable directly from it, with no further design decisions.
Polling loop or event-driven architecture
In the polling loop, the program endlessly walks through the same sequence: read every
input, compute, update every output - simple and predictable, but the reaction time is, in the worst
case, equal to the duration of the entire loop. In the event-driven architecture, peripherals
signal the moment of an event through interrupts, and the main loop consumes the raised flags - fast
reaction, but new problems appear: sharing variables with the ISR, volatile, critical
sections. In practice, almost every real application is a hybrid: interrupts capture fast events,
while the main loop, paced by a timer tick, makes the decisions.
The finite state machine
The natural way to organize the logic of an embedded application. The idea: the system is, at any moment, in one of a small number of well-defined states, and behavior depends not only on inputs but also on the current state - the same button press means one thing when the system is watering and another when it is waiting.
| Current state | Event | Action | Next state |
|---|---|---|---|
| WAITING | 1 s tick | starts the ADC conversion | MEASURING |
| MEASURING | average < threshold | opens the valve, starts the timer | WATERING |
| MEASURING | invalid value | lights the fault LED | FAULT |
| WATERING | maximum duration reached | closes the valve, starts the pause timer | PAUSE |
| PAUSE | manual watering button | ignores it, logs the refusal | PAUSE |
| FAULT | reset from the operator | clears the fault indication | WAITING |
The table must be complete - for every state, every possible event, including those that produce no transition; otherwise undefined corners of behavior remain, exactly where bugs appear. Translating it into code is mechanical:
typedef enum { ST_WAITING=0, ST_MEASURING, ST_WATERING, ST_PAUSE, ST_FAULT } state_t;
static state_t currentState = ST_WAITING;
void runStateMachine(event_t ev) {
switch (currentState) {
case ST_WAITING:
if (ev == EV_TICK_1S) { adcStartConversion(); currentState = ST_MEASURING; }
else if (ev == EV_MANUAL_BUTTON) { valveOpen(); currentState = ST_WATERING; }
break;
case ST_MEASURING:
if (ev == EV_INVALID_VALUE) { currentState = ST_FAULT; }
else if (ev == EV_BELOW_THRESHOLD) { valveOpen(); currentState = ST_WATERING; }
else if (ev == EV_ABOVE_THRESHOLD) { currentState = ST_WAITING; }
break;
/* ... the other states, one case per state ... */
}
}
If the state is changed from three different places in the program, the advantage of the state machine disappears completely - it becomes, once again, a maze of nested conditions, just with a nicer name.
6The main activity: initialization, loop, buttons11 min
The activity diagram of the main program describes what happens from power-up to shutdown. It invariably begins with initialization: variables, ports with their directions, the timer that provides the time base, the ADC, the serial interface, then globally enabling interrupts. The order is not indifferent.
If the direction is configured before the desired state is written to the data register, in the interval between the two instructions the output picks up the residual value in the register, and the valve can open for a few microseconds at every system power-up. The general rule: any output that commands something dangerous is brought to the safe state before being declared an output.
volatile uint8_t tickFlag = 0; /* raised by the ISR, read by the main loop */
ISR(TIMER1_COMPA_vect) { tickFlag = 1; } /* very short ISR: only marks time */
int main(void) {
valveClose(); /* safe state BEFORE configuring the output */
portsInit();
timerInit(); /* interrupt every 10 ms */
adcInit(); serialInit(9600);
parametersLoadFromEeprom();
sei();
for (;;) {
if (tickFlag) {
cli(); tickFlag = 0; sei(); /* protected clear: the ISR can write at any time */
event_t ev = collectEvent();
runStateMachine(ev);
timerTick();
logSendOneCharacter();
}
}
}
Note that the loop contains no active waiting and no long operation - anything that takes a long time, such as sending the log, is either split into pieces or left to interrupts.
Reading buttons by transition, not by level
A button is not read as a level, but as a transition: the system must react to the moment of the press, not to the fact that it is being held down. The standard method: keep the previously read value of the port and compare it with the current one; if they differ, a change has occurred, and the changed bits show which button was actuated. The same method solves the case of several buttons pressed at once, because the value of the whole port is interpreted as a code, and disallowed codes can be handled explicitly - if the user presses stop and manual watering at the same time, the program can explicitly decide that stop takes priority, a decision that must appear in the specification, not left to the whim of instruction order.
7Writing and early testing of the code11 min
With the diagrams finished, writing the code becomes largely transcription. The essential rule of this stage: the whole program is not written at once. The code is developed in layers, from the bottom up - first the actual clock frequency of the microcontroller is verified (toggling a pin in a loop, measured with the oscilloscope), then an output with an LED, then an input with a button, then the ADC with a potentiometer, then the timer, and only at the end are they brought together.
Why testing does not happen directly on the real installation
Two reasons. First: a software bug can destroy expensive hardware - commanding both arms of an H-bridge at once short-circuits the supply, a valve left open floods the greenhouse. Second, less obvious: the real installation cannot be brought on demand into the situations that matter - you cannot make the soil dry in five seconds to check the threshold, you cannot repeat a power failure a hundred times.
The solution is a cheap simulation module, which imitates the electrical interface of the installation: digital inputs from buttons, analog inputs from potentiometers, outputs on LEDs - for the watering system, a potentiometer in place of the sensor, three buttons, three LEDs. The cost is a few dollars; the gain is the ability to produce, at any time, any combination of inputs, including ones that are physically impossible, which are exactly the most interesting ones for testing.
Software simulators are invaluable for checking logic and counting cycles, but contact bounce, analog-line noise, the voltage dip when the relay engages, and coupling between the power cable and the signal cable do not show up in a simulator - and are, in practice, the source of most failures.
Code that is easy to test
It is possible to write code that lends itself to testing and code that does not - the difference lies in separating the decision from peripheral access. A function that reads the ADC itself and writes to the relay port itself cannot be verified except with hardware. The same logic, written as a pure function that receives the measured value and returns the decision, can be compiled and run on a computer, with thousands of combinations, in a few seconds.
/* hysteresis: width of the dead zone, in the same units as moisture and threshold */
uint8_t decideWatering(uint16_t moisture, uint16_t threshold,
uint16_t hysteresis, uint8_t wateringNow) {
if (wateringNow) return (moisture < threshold + hysteresis) ? 1 : 0;
return (moisture < threshold) ? 1 : 0;
}
8The test plan and its levels11 min
Once the trial module is ready, the test plan is written - not a list of things tried "whenever it comes to mind", but a document derived from the specification, in which every requirement has at least one associated test and every test has a numeric acceptance criterion. The link is made through requirement codes, so that at the end it can be shown that no requirement was left unverified.
The testing levels
Unit tests apply to purely logical functions, which touch no peripherals: converting the ADC to a percentage, averaging, the watering decision. They are compiled for the host computer and run automatically, so they can be repeated on every code change at zero effort. A good unit test includes the boundary values of the domain.
assert(decideWatering(300, 400, 50, 0) == 1); /* below threshold, closed: turns on */
assert(decideWatering(500, 400, 50, 0) == 0); /* above threshold: stays off */
assert(decideWatering(400, 400, 50, 0) == 0); /* exactly at threshold: strict comparison */
assert(decideWatering(430, 400, 50, 1) == 1); /* in the hysteresis zone: stays open */
Module tests verify an entire subsystem, with simulated hardware - a known voltage is applied and the displayed value is compared with the one measured on a voltmeter, checking not only the logic but also the peripheral configuration (reference, prescaling, alignment). System tests verify the whole through realistic usage scenarios, timed and compared against the limits in the specification. Boundary tests exercise extreme values - a 16-bit millisecond counter wraps around after about 65 s, and a carelessly written comparison gives a wrong result right there; the test must exist in the plan, otherwise the bug reaches the customer. Robustness tests check unwanted situations: sensor disconnection, unstable power, reset during watering, simultaneous button presses. The criterion is not that the system does something useful, but that it does not do something dangerous.
| Test | Requirement | How it is performed | Acceptance criterion |
|---|---|---|---|
| T1 | CF1 | toggle a pin on every measurement | period 0.9-1.1 s, 100 cycles |
| T6 | CF6 | disconnect the sensor, then short it to ground | enters FAULT, valve closed |
| T9 | CN3 | power interruption during watering | valve closed on return |
| T11 | - | pressing all buttons at once, ×50 | defined state, no lockup |
9Systematic debugging9 min
However well organized a project is, bugs will exist, and the way they are hunted makes the difference between an hour and a week. Debugging is not an activity of inspiration, but a procedure with three principles.
First: reduce to a minimal reproducible case. A bug that appears "sometimes, after about half an hour" cannot be studied. The shortest sequence of actions that triggers it every time must be found, then every part of the system not necessary for reproduction is removed, one by one - disconnect the display: does it still happen? Turn off the serial log: does it still happen? Most of the time the bug becomes clear on its own during this reduction, because the part removed at which the bug disappears points directly to the cause.
Second: measure, do not guess. Most plausible hypotheses are false, and checking them often costs less than arguing about them. If you suspect an interrupt is not firing, toggle a pin inside the ISR and look at it with the oscilloscope - in three minutes you have a certain answer.
#define TEST_PIN_HIGH (PORTD |= (1<
Third: isolate whether it is hardware or software as early as possible, because the answer completely changes the direction of the search. An LED that does not light can mean a wrongly written bit or a missing resistor; measuring the voltage on the pin immediately separates the two cases - if the pin has the correct level and the LED does not light, the problem is in the circuit; if the pin stays at zero, the problem is in the code or in the port direction configuration.
You changed the timing, you fixed nothing. Such timing-sensitive bugs almost always come from a
variable shared with an interrupt routine without volatile, or from an unprotected
access to a multi-byte variable. The correct fix protects the access, it does not keep the delay
that "makes it work".
10Documentation6 min
A large part of the documentation already exists if the previous stages were followed: the system description, the specification, the block diagrams, the transition tables, the test plan and the trial results. Final documentation means gathering them into a coherent whole.
| Pin | Direction | Signal | Notes |
|---|---|---|---|
| PC0 (ADC0) | Input | Moisture sensor | 0...5V, RC filter 10kΩ/100nF |
| PD5 | Output | Valve relay control | active high, transistor + diode |
| PD7 | Output | Test pin | debugging only, not connected in production |
| PD0, PD1 | Serial | Log RXD, TXD | 9600 bd, 8N1 |
Code documentation: every function accompanied by its purpose, its parameters with units and ranges, its return value and error conditions; comments that repeat the instruction have no value, the ones that explain why a solution was chosen are the ones that save time two years later. Finally, the decision log: a short list of the important choices and their reasons - why the hysteresis is 5% and not 2%, why the pause is 30 minutes. Every decision seems obvious on the day it is made and becomes incomprehensible a year later; without a log, the first thing whoever takes over the project does is "correct" a correct decision, and the bug reappears exactly in the situation that had originally required it.
11Common mistakes4 min
- I wrote the whole program in one piece, I will test it at the end. When integration fails, the search space for the bug includes every module and every interaction between them - debugging becomes practically unbounded. Every peripheral is verified in isolation, with a minimal program, before being integrated with the rest.
- The function that decides watering reads the ADC itself and commands the relay itself, to keep the number of functions down. Mixing the decision with peripheral access makes the function untestable without real hardware. Separate the pure logic (receives values, returns a decision) from peripheral access - the pure logic is tested on a computer, in seconds.
- I added a
delay()and the intermittent bug disappeared - it is fixed. False - usually you only changed the timing of a race on a variable shared with an interrupt; the bug reappears with a different combination of speeds. Look for the unprotected access (a missingvolatileor a missing critical section) and protect it explicitly, instead of keeping the delay that "makes it work".
12Summary and glossary5 min
A microcontroller project does not fail from a lack of knowledge about registers, but from a lack of method: unverifiable requirements, subsystems with no fixed interfaces, testing postponed to the end. The antidote is an orderly path, with a feedback loop between testing and design: the system description becomes a specification through verifiable requirements (subject, action, condition, numeric value); partitioning fixes interfaces before implementation and makes testability an explicit criterion; choosing the microcontroller follows from the requirements table, with a margin of at least 20%; application logic is organized as a finite state machine, with a complete transition table, which makes unwanted combinations impossible by construction; the code is written in layers, tested on a cheap simulation module, with pure logic separated from peripheral access; the test plan links every requirement to a test with a numeric criterion, across five levels (unit, module, system, boundary, robustness); debugging follows three principles - minimal reproducible case, measure instead of guessing, separate hardware from software; and the final documentation, including the decision log, makes the project transferable to anyone else.
13Review questions7 min
- Rewrite the requirement "the system must respond quickly to a button press" so that it becomes verifiable, also specifying the measurement method.
- Why must a subsystem's interface be fixed before it is implemented? What information must it contain, beyond the function prototype?
- A colleague proposes that the watering decision function read the ADC itself and command the relay directly. What is lost from the testing point of view?
- Compare the polling loop with the event-driven architecture from the point of view of maximum reaction time.
- Explain why the absence of a transition from the table, not a condition in the code, is the mechanism by which the state machine prevents repeated watering.
- A 16-bit millisecond counter measures a 30-minute pause. Why is the implementation wrong, and after how long does the bug appear?
14Further directions2 min
The last lecture of the discipline puts this exact method into practice, on a complete and more complex project: a robot that follows a line and avoids obstacles, from the specification to the final architecture of the program.
The integrated application of acquisition, decision and control, with no delay(), is
practiced in Laboratory 07.