LECTURE 10

Designing and Testing Microcontroller Applications

Duration: 122 min of instruction Level: undergraduate, year II - recommended after Lecture 09 Discipline: Microcontrollers and Microprocessors PDF: download the material RO versiunea română

Previous lectures covered, in turn, the components of a microcontroller: ports, interrupts, timers, ADC, UART, SPI, I2C. An engineer who masters all of them does not, by that alone, know how to build a product. Between "I lit an LED with PWM" and "I delivered a system that has run unattended for two years in a greenhouse" lies a distance that is not closed with more knowledge about registers, but with method - the subject of this lecture, followed through a single example, from requirement to test plan.

1Scope and structure of the lecture6 min

Microcontroller projects almost always fail for reasons of method, not programming - rarely because someone misunderstood a prescaler, almost always because nobody wrote down clearly what the system had to do, because the split into modules happened on the fly, or because testing was left until the end. We will follow the entire path, from a vague requirement to a documented system, through a single reference example: an automatic watering system for a greenhouse, revisited identically in every section.

Recap from Lecture 09
  • SPI and I2C solve communication with nearby peripherals - the open question remains how to organize, as a whole, a program that uses all of them coherently (Lecture 09)
  • Explain the three main causes for which microcontroller projects fail
  • Turn a vague requirement into a verifiable requirement, with a numeric acceptance criterion
  • Partition a system into subsystems with interfaces fixed before implementation
  • Describe the logic of an application as a finite state machine, starting from a complete transition table
  • Distinguish the testing levels (unit, module, system, boundary, robustness) and write a test plan linked to the specification
  • Apply the three principles of systematic debugging: minimal reproducible case, measure instead of guessing, separate hardware from software

2Why a project fails: the three causes11 min

The first cause: an unclear requirement. The customer says "I want the system to react quickly" or "I want it to be safe". Nobody disagrees with such statements, because nobody knows what they mean. At acceptance, the customer notices the system reacts in 300 ms and claims it is not "fast", the designer answers that nothing else was asked for, and the discussion cannot end - there is no objective criterion either side can point to. A requirement that cannot be checked by a measurement is not a requirement, it is a wish.

The second cause: poor delimitation of subsystems. When code grows without a structure decided in advance, functions start communicating through global variables, each modifying the other's state. Such a program usually works, up until the first change - that is when it turns out that changing a threshold forces changes in the relay routine, which breaks the display, and fixing the display reintroduces the original bug. The system becomes impossible to maintain not because it is complex, but because it has no internal boundaries.

The third cause, the most costly one: testing postponed to the end. It is tempting to write the whole program and only then try it, because testing along the way seems like a waste of time. In reality, when you integrate ten untested modules and the result does not work, you have no information about where the bug is - any of the ten, or any interaction between them, could be guilty. If each module was verified separately and the system breaks when the tenth one is added, the search space has shrunk to a single module.

The cost of a bug rises sharply with the moment it is discovered A requirement written wrong and corrected on paper costs one hour. The same requirement discovered after it has been implemented costs a rewrite of the code. Discovered after the wiring has been designed, it costs a new board. Discovered after a hundred boards have been installed in the field, it costs a trip to a hundred locations. Nothing a designer does has a better return than the time spent clarifying requirements before starting.
How much it costs to fix a bug, depending on when it is discovered

The two stages and the cycle between them

Development unfolds in two large stages: design, which begins with the description of the system and ends with the first code draft, and testing, which begins with simulation and ends with documentation. Showing the steps as a list is misleading, because it suggests a one-way flow - in reality there is a feedback loop: if, during testing, the system does not behave according to the specification, it is not only the code that gets corrected, but design is resumed at the level where the mistake occurred, sometimes even the specification itself. A healthy project goes through this cycle several times, with decreasing amplitude each time.

The design method: put the stages in the order they are followed

Worth remembering: the design stage produces documents, not code. Block diagrams, requirement tables, transition tables and the pin map outlive the project and determine its long-term value - a working system with no document behind it is a system whose lifetime equals the memory of its author.

3The reference example and the specification13 min

The whole chapter follows a single example: an automatic watering system for a small greenhouse. It measures soil moisture and, below a threshold, opens a solenoid valve through a relay for a limited interval; it has a manual-watering button, a stop button, two status LEDs, and it sends an event log over the serial interface. The first discussion with the customer always produces a second document, just as important as the description: the list of questions - what happens if the sensor fails and permanently indicates dry soil? If someone presses manual watering and stop at the same time? If power drops during watering? An experienced designer is recognized by the quality of this list, because every question asked now is a bug that will not be discovered later in the field.

The time budget: does the response requirement fit the cause-effect chain?

What makes a requirement verifiable

A requirement is verifiable if, with the system in front of them and a measuring instrument, two independent people are bound to reach the same conclusion about whether it is met.

Vague requirement vs. verifiable requirement
Vague: "the system must respond quickly when the threshold is exceeded".
Verifiable: "the system must command the relay within at most 50 ms of the moment the measured quantity exceeds the configured threshold".
The second fixes a start moment, an end moment, a limit value, and, implicitly, a verification method.

A well-written requirement has four elements: a subject, an observable action, a triggering condition, and a numeric value with a unit and a tolerance. Words that should give pause when they appear in a specification: "fast", "stable", "safe", "friendly", "approximately" - each hides a negotiation that never took place.

Functional requirements describe what the system does; they are relatively easy to state, because they follow directly from the customer's description. Non-functional requirements describe under what conditions it must do those things, and are the source of most late surprises: average consumption decides whether the system can be powered from a solar panel; the temperature range decides whether the oscillator can be the internal one or must be a crystal; unit cost decides between 8 and 32 KB of program memory. None of these requirements appear in the customer's description, but all of them change the project if discovered late.

CodeRequirementVerification
CF1Measures moisture once every 1±0.1 stiming with an oscilloscope
CF2Below threshold, opens the valve within ≤200 msinjecting a voltage below threshold
CF5On stop, the valve closes within ≤50 msoscilloscope on both signals
CN1Average consumption <20 mAammeter, 10 minutes
CN4Thresholds changeable without reprogrammingEEPROM check
Worked exercise - the time budget for CF5

Requirement CF5 (closing the valve within ≤50 ms of pressing the button) must be verified for an architecture where buttons are polled every 2 ms, a press is confirmed after 20 ms of a stable level, and the relay has an actuation time of 10 ms according to the datasheet. Is it achievable?

See the solution

The budget is built by adding, in the worst case, every delay along the cause-effect chain. The press can happen right after a reading: one polling interval is lost, 2 ms. Debounce confirmation: 20 ms. Decision and writing to the relay port: under 1 ms. Relay actuation: 10 ms. Total: 2+20+1+10 = 33 ms, with a margin of 17 ms against the 50 ms limit - the requirement is achievable. If, however, the debounce confirmation had been chosen as 50 ms, as often happens with no justification, the requirement would have been violated by design, and the bug would only have appeared at acceptance measurements, not here, on paper.

4System partitioning and choosing the microcontroller13 min

Partitioning means dividing the system into subsystems, each with a single responsibility, and establishing how they communicate. The fundamental rule: a subsystem's interface is fixed before the subsystem is implemented. The reason is practical - as long as the interface is not fixed, no other subsystem can be written, because nobody knows what it will receive and what it will return. From the moment the interface is written, even just as function prototypes with comments, every module can be developed in parallel.

A complete interface says more than a function signature

uint16_t readSensor(void); is not a complete interface. It must specify: what units, what range of values, with what error, what happens on invalid values, how long the operation takes, and whether it can be called from an interrupt routine.

The second criterion of good partitioning: each subsystem has a single reason to change. If replacing the moisture sensor with another one forces a change in the code that lights the LEDs, the partitioning is wrong. The watering system is thus split into: moisture acquisition, operator interface (buttons + LEDs), valve control, time base, serial logging, and, in the middle, the application itself, which makes decisions without touching any register.

Testability is the criterion of partitioning, not a consequence of it A well-delimited functional block is recognized by the fact that it can be tested alone. If, to verify the watering decision function, the real sensor, the 12 V supply and the solenoid valve all have to be connected, then the decision function is not a separate block - it has become mixed up with peripheral access.

Choosing the microcontroller

The choice is not made by habit or by whatever part is in the drawer, but as a consequence of the specification and the partitioning: once it is known which signals go in and out, the pin count is determined; once the operating periods are known, the needed peripherals are determined. Only then does the catalog get opened.

CriterionApplication requirementConsequence
Pins1 analog, 3 digital inputs, 3 digital outputs, seriala 28-pin package is enough
PeripheralsADC, timer, USARTno SPI/I2C needed
Non-volatile memorythresholds changeable in the field (CN4)an EEPROM is mandatory
Consumptionunder 20 mA average (CN1)low-power modes
Temperature−5...+55°C (CN2)industrial variant or checking oscillator drift

An ATmega328P satisfies every line of the table with margin to spare. Important: the conclusion followed from the table, it did not precede it - if the application had required eight analog channels and consumption under 100 µA, the same method would have led to a different chip. The practical rule: keep a margin of at least 20% on pins, memory and execution time - a project that uses everything it has is a project with no future, for which the first additional requirement forces a change of microcontroller and a board redesign.

5Program architecture and the finite state machine13 min

With the subsystems defined, next comes describing what each one does inside, through activity and state diagrams (the UML language). The quality criterion for an activity diagram is simple and strict: the code must be writable directly from it, with no further design decisions.

Polling loop or event-driven architecture

In the polling loop, the program endlessly walks through the same sequence: read every input, compute, update every output - simple and predictable, but the reaction time is, in the worst case, equal to the duration of the entire loop. In the event-driven architecture, peripherals signal the moment of an event through interrupts, and the main loop consumes the raised flags - fast reaction, but new problems appear: sharing variables with the ISR, volatile, critical sections. In practice, almost every real application is a hybrid: interrupts capture fast events, while the main loop, paced by a timer tick, makes the decisions.

The finite state machine

The natural way to organize the logic of an embedded application. The idea: the system is, at any moment, in one of a small number of well-defined states, and behavior depends not only on inputs but also on the current state - the same button press means one thing when the system is watering and another when it is waiting.

The advantage: unwanted combinations become impossible by construction If the transition from PAUSE to WATERING does not exist in the table, then it does not exist in the code either - the system cannot water twice in a row, no matter what the user presses. Safety requirements thus translate into the absence of certain transitions, a form much easier to verify than a complicated condition.
Current stateEventActionNext state
WAITING1 s tickstarts the ADC conversionMEASURING
MEASURINGaverage < thresholdopens the valve, starts the timerWATERING
MEASURINGinvalid valuelights the fault LEDFAULT
WATERINGmaximum duration reachedcloses the valve, starts the pause timerPAUSE
PAUSEmanual watering buttonignores it, logs the refusalPAUSE
FAULTreset from the operatorclears the fault indicationWAITING

The table must be complete - for every state, every possible event, including those that produce no transition; otherwise undefined corners of behavior remain, exactly where bugs appear. Translating it into code is mechanical:

typedef enum { ST_WAITING=0, ST_MEASURING, ST_WATERING, ST_PAUSE, ST_FAULT } state_t;
static state_t currentState = ST_WAITING;

void runStateMachine(event_t ev) {
    switch (currentState) {
    case ST_WAITING:
        if (ev == EV_TICK_1S) { adcStartConversion(); currentState = ST_MEASURING; }
        else if (ev == EV_MANUAL_BUTTON) { valveOpen(); currentState = ST_WATERING; }
        break;
    case ST_MEASURING:
        if (ev == EV_INVALID_VALUE) { currentState = ST_FAULT; }
        else if (ev == EV_BELOW_THRESHOLD) { valveOpen(); currentState = ST_WATERING; }
        else if (ev == EV_ABOVE_THRESHOLD) { currentState = ST_WAITING; }
        break;
    /* ... the other states, one case per state ... */
    }
}
The golden rule: a single place changes the state variable

If the state is changed from three different places in the program, the advantage of the state machine disappears completely - it becomes, once again, a maze of nested conditions, just with a nicer name.

6The main activity: initialization, loop, buttons11 min

The activity diagram of the main program describes what happens from power-up to shutdown. It invariably begins with initialization: variables, ports with their directions, the timer that provides the time base, the ADC, the serial interface, then globally enabling interrupts. The order is not indifferent.

The relay pin must be brought to the inactive state before being configured as an output

If the direction is configured before the desired state is written to the data register, in the interval between the two instructions the output picks up the residual value in the register, and the valve can open for a few microseconds at every system power-up. The general rule: any output that commands something dangerous is brought to the safe state before being declared an output.

volatile uint8_t tickFlag = 0;    /* raised by the ISR, read by the main loop */
ISR(TIMER1_COMPA_vect) { tickFlag = 1; }   /* very short ISR: only marks time */

int main(void) {
    valveClose();               /* safe state BEFORE configuring the output */
    portsInit();
    timerInit();                 /* interrupt every 10 ms */
    adcInit(); serialInit(9600);
    parametersLoadFromEeprom();
    sei();
    for (;;) {
        if (tickFlag) {
            cli(); tickFlag = 0; sei();   /* protected clear: the ISR can write at any time */
            event_t ev = collectEvent();
            runStateMachine(ev);
            timerTick();
            logSendOneCharacter();
        }
    }
}

Note that the loop contains no active waiting and no long operation - anything that takes a long time, such as sending the log, is either split into pieces or left to interrupts.

Reading buttons by transition, not by level

A button is not read as a level, but as a transition: the system must react to the moment of the press, not to the fact that it is being held down. The standard method: keep the previously read value of the port and compare it with the current one; if they differ, a change has occurred, and the changed bits show which button was actuated. The same method solves the case of several buttons pressed at once, because the value of the whole port is interpreted as a code, and disallowed codes can be handled explicitly - if the user presses stop and manual watering at the same time, the program can explicitly decide that stop takes priority, a decision that must appear in the specification, not left to the whim of instruction order.

7Writing and early testing of the code11 min

With the diagrams finished, writing the code becomes largely transcription. The essential rule of this stage: the whole program is not written at once. The code is developed in layers, from the bottom up - first the actual clock frequency of the microcontroller is verified (toggling a pin in a loop, measured with the oscilloscope), then an output with an LED, then an input with a button, then the ADC with a potentiometer, then the timer, and only at the end are they brought together.

Why testing does not happen directly on the real installation

Two reasons. First: a software bug can destroy expensive hardware - commanding both arms of an H-bridge at once short-circuits the supply, a valve left open floods the greenhouse. Second, less obvious: the real installation cannot be brought on demand into the situations that matter - you cannot make the soil dry in five seconds to check the threshold, you cannot repeat a power failure a hundred times.

The solution is a cheap simulation module, which imitates the electrical interface of the installation: digital inputs from buttons, analog inputs from potentiometers, outputs on LEDs - for the watering system, a potentiometer in place of the sensor, three buttons, three LEDs. The cost is a few dollars; the gain is the ability to produce, at any time, any combination of inputs, including ones that are physically impossible, which are exactly the most interesting ones for testing.

A software simulator models the microcontroller, not the world around it

Software simulators are invaluable for checking logic and counting cycles, but contact bounce, analog-line noise, the voltage dip when the relay engages, and coupling between the power cable and the signal cable do not show up in a simulator - and are, in practice, the source of most failures.

Code that is easy to test

It is possible to write code that lends itself to testing and code that does not - the difference lies in separating the decision from peripheral access. A function that reads the ADC itself and writes to the relay port itself cannot be verified except with hardware. The same logic, written as a pure function that receives the measured value and returns the decision, can be compiled and run on a computer, with thousands of combinations, in a few seconds.

/* hysteresis: width of the dead zone, in the same units as moisture and threshold */
uint8_t decideWatering(uint16_t moisture, uint16_t threshold,
                        uint16_t hysteresis, uint8_t wateringNow) {
    if (wateringNow) return (moisture < threshold + hysteresis) ? 1 : 0;
    return (moisture < threshold) ? 1 : 0;
}
Why hysteresis is needed Without two different thresholds for turning on and off, a value oscillating due to noise around the threshold would produce repeated relay openings and closings, wearing out the contacts quickly. The dead zone introduced means that, once watering has started, it only turns off after moisture has risen visibly above the threshold - the same problem and the same solution appear in any thermostat.

8The test plan and its levels11 min

Once the trial module is ready, the test plan is written - not a list of things tried "whenever it comes to mind", but a document derived from the specification, in which every requirement has at least one associated test and every test has a numeric acceptance criterion. The link is made through requirement codes, so that at the end it can be shown that no requirement was left unverified.

The testing levels

Unit tests apply to purely logical functions, which touch no peripherals: converting the ADC to a percentage, averaging, the watering decision. They are compiled for the host computer and run automatically, so they can be repeated on every code change at zero effort. A good unit test includes the boundary values of the domain.

assert(decideWatering(300, 400, 50, 0) == 1);  /* below threshold, closed: turns on */
assert(decideWatering(500, 400, 50, 0) == 0);  /* above threshold: stays off */
assert(decideWatering(400, 400, 50, 0) == 0);  /* exactly at threshold: strict comparison */
assert(decideWatering(430, 400, 50, 1) == 1);  /* in the hysteresis zone: stays open */

Module tests verify an entire subsystem, with simulated hardware - a known voltage is applied and the displayed value is compared with the one measured on a voltmeter, checking not only the logic but also the peripheral configuration (reference, prescaling, alignment). System tests verify the whole through realistic usage scenarios, timed and compared against the limits in the specification. Boundary tests exercise extreme values - a 16-bit millisecond counter wraps around after about 65 s, and a carelessly written comparison gives a wrong result right there; the test must exist in the plan, otherwise the bug reaches the customer. Robustness tests check unwanted situations: sensor disconnection, unstable power, reset during watering, simultaneous button presses. The criterion is not that the system does something useful, but that it does not do something dangerous.

Match each testing level to what it verifies
TestRequirementHow it is performedAcceptance criterion
T1CF1toggle a pin on every measurementperiod 0.9-1.1 s, 100 cycles
T6CF6disconnect the sensor, then short it to groundenters FAULT, valve closed
T9CN3power interruption during wateringvalve closed on return
T11-pressing all buttons at once, ×50defined state, no lockup
Tests with no associated requirement are a useful signal T11 and T12 have no requirement - they come from atypical testing, and if they uncover a bug, it means a requirement was missing from the specification. The fix is made not only in the code, but also in the specification, completed with the newly discovered requirement; then the test plan is run again, in full, because a fix is exactly the moment when other bugs get introduced.

9Systematic debugging9 min

However well organized a project is, bugs will exist, and the way they are hunted makes the difference between an hour and a week. Debugging is not an activity of inspiration, but a procedure with three principles.

First: reduce to a minimal reproducible case. A bug that appears "sometimes, after about half an hour" cannot be studied. The shortest sequence of actions that triggers it every time must be found, then every part of the system not necessary for reproduction is removed, one by one - disconnect the display: does it still happen? Turn off the serial log: does it still happen? Most of the time the bug becomes clear on its own during this reduction, because the part removed at which the bug disappears points directly to the cause.

Second: measure, do not guess. Most plausible hypotheses are false, and checking them often costs less than arguing about them. If you suspect an interrupt is not firing, toggle a pin inside the ISR and look at it with the oscilloscope - in three minutes you have a certain answer.

#define TEST_PIN_HIGH (PORTD |= (1<
"Debugging with a wire" does not stop the program Toggling a pin costs two instructions and disturbs practically nothing - unlike a breakpoint debugger, which, in a system commanding a motor or a valve, completely changes the physical behavior (sometimes dangerously) and makes exactly the timing-related bugs disappear, because time itself was altered by the measuring tool.

Third: isolate whether it is hardware or software as early as possible, because the answer completely changes the direction of the search. An LED that does not light can mean a wrongly written bit or a missing resistor; measuring the voltage on the pin immediately separates the two cases - if the pin has the correct level and the LED does not light, the problem is in the circuit; if the pin stays at zero, the problem is in the code or in the port direction configuration.

A bug that disappears when you add a delay or a print statement has not been fixed

You changed the timing, you fixed nothing. Such timing-sensitive bugs almost always come from a variable shared with an interrupt routine without volatile, or from an unprotected access to a multi-byte variable. The correct fix protects the access, it does not keep the delay that "makes it work".

10Documentation6 min

A large part of the documentation already exists if the previous stages were followed: the system description, the specification, the block diagrams, the transition tables, the test plan and the trial results. Final documentation means gathering them into a coherent whole.

The practical criterion of sufficient documentation Not the page count, but the question of whether a colleague who did not take part in the project can, using only the documents, repair a faulty unit, compile the program, load it, and check that it works. The documentation must contain: the schematic with component values; the pin map (what is connected to each pin, its direction, its active level); the communication protocol; the configuration values; the programming procedure.
PinDirectionSignalNotes
PC0 (ADC0)InputMoisture sensor0...5V, RC filter 10kΩ/100nF
PD5OutputValve relay controlactive high, transistor + diode
PD7OutputTest pindebugging only, not connected in production
PD0, PD1SerialLog RXD, TXD9600 bd, 8N1

Code documentation: every function accompanied by its purpose, its parameters with units and ranges, its return value and error conditions; comments that repeat the instruction have no value, the ones that explain why a solution was chosen are the ones that save time two years later. Finally, the decision log: a short list of the important choices and their reasons - why the hysteresis is 5% and not 2%, why the pause is 30 minutes. Every decision seems obvious on the day it is made and becomes incomprehensible a year later; without a log, the first thing whoever takes over the project does is "correct" a correct decision, and the bug reappears exactly in the situation that had originally required it.

11Common mistakes4 min

  • I wrote the whole program in one piece, I will test it at the end. When integration fails, the search space for the bug includes every module and every interaction between them - debugging becomes practically unbounded. Every peripheral is verified in isolation, with a minimal program, before being integrated with the rest.
  • The function that decides watering reads the ADC itself and commands the relay itself, to keep the number of functions down. Mixing the decision with peripheral access makes the function untestable without real hardware. Separate the pure logic (receives values, returns a decision) from peripheral access - the pure logic is tested on a computer, in seconds.
  • I added a delay() and the intermittent bug disappeared - it is fixed. False - usually you only changed the timing of a race on a variable shared with an interrupt; the bug reappears with a different combination of speeds. Look for the unprotected access (a missing volatile or a missing critical section) and protect it explicitly, instead of keeping the delay that "makes it work".

12Summary and glossary5 min

A microcontroller project does not fail from a lack of knowledge about registers, but from a lack of method: unverifiable requirements, subsystems with no fixed interfaces, testing postponed to the end. The antidote is an orderly path, with a feedback loop between testing and design: the system description becomes a specification through verifiable requirements (subject, action, condition, numeric value); partitioning fixes interfaces before implementation and makes testability an explicit criterion; choosing the microcontroller follows from the requirements table, with a margin of at least 20%; application logic is organized as a finite state machine, with a complete transition table, which makes unwanted combinations impossible by construction; the code is written in layers, tested on a cheap simulation module, with pure logic separated from peripheral access; the test plan links every requirement to a test with a numeric criterion, across five levels (unit, module, system, boundary, robustness); debugging follows three principles - minimal reproducible case, measure instead of guessing, separate hardware from software; and the final documentation, including the decision log, makes the project transferable to anyone else.

Verifiable requirement
a statement with subject, action, condition and numeric value, testable through a measurement.
Partitioning
dividing the system into subsystems with a single responsibility and interfaces fixed in advance.
Finite state machine
a model where behavior depends on the current state and the event, described by a transition table.
Hysteresis
a dead zone between the on and off thresholds, which prevents repeated switching caused by noise.
Unit test
a test of a purely logical function, with no peripheral access, run on the host computer.
Debugging with a wire
the technique of toggling a test pin to measure durations with an oscilloscope, without stopping the program.

13Review questions7 min

  1. Rewrite the requirement "the system must respond quickly to a button press" so that it becomes verifiable, also specifying the measurement method.
  2. Why must a subsystem's interface be fixed before it is implemented? What information must it contain, beyond the function prototype?
  3. A colleague proposes that the watering decision function read the ADC itself and command the relay directly. What is lost from the testing point of view?
  4. Compare the polling loop with the event-driven architecture from the point of view of maximum reaction time.
  5. Explain why the absence of a transition from the table, not a condition in the code, is the mechanism by which the state machine prevents repeated watering.
  6. A 16-bit millisecond counter measures a 30-minute pause. Why is the implementation wrong, and after how long does the bug appear?

14Further directions2 min

The last lecture of the discipline puts this exact method into practice, on a complete and more complex project: a robot that follows a line and avoids obstacles, from the specification to the final architecture of the program.

The integrated application of acquisition, decision and control, with no delay(), is practiced in Laboratory 07.