This is the first lecture and it starts from nothing: it does not assume you know what a bit is, what a protocol is, or why a network cable has eight wires in it. By the end you will understand what physically happens between the moment you press Enter and the moment the message arrives at the other end - and you will have the vocabulary in which every discussion for the rest of the semester is conducted.
1Scope and structure of the course7 min
You have two computers in a room and you want one of them to send the other a file. It sounds simple. It is not. For it to work, somebody had to answer, one at a time, some fifteen questions:
- What kind of wire connects them, and how long may it be?
- What does the digit 1 look like as an electrical voltage? And the digit 0?
- How does the receiver know where one bit ends and the next begins?
- How does it know the message has started? How does it know it has finished?
- If both of them speak at the same time, what happens?
- If a bit is corrupted along the way, how do we notice?
- If there are ten computers rather than two, how do you say who the message is for?
- What if the other computer is in a different building? In a different country?
Every one of these questions has an answer, and the answers are called protocols. This lecture answers the first five - the physical ones. The rest of the semester climbs, question by question, to the last.
A computer network is the interconnection of several computing systems so that they can exchange information. The definition is short, but it hides an important distinction.
Inside a computer, the components communicate over copper traces on the motherboard, a few centimetres long, governed by the chipset. Nobody interferes with them, nobody else is listening in, and the signal always arrives intact.
Between computers we have none of those guarantees: the distance may be metres or thousands of kilometres, the medium may be noisy, and the far end may belong to somebody else. Everything that follows in this course exists in order to make up for those three missing guarantees.
Learning outcomes
- Explain what a bit and a byte are, and why 100 Mbps does not mean 100 MB per second
- Classify a network by its reach: LAN, MAN, WAN
- State, for any piece of network equipment, which layer it works at and what decision it makes
- Explain why communication is organised in layers, and move between the OSI and TCP/IP models
- Read and draw the waveform for NRZ-L, NRZ-I, Manchester, MLT-3 and PAM-5
- Calculate the bit rate from the symbol rate and the modulation scheme
2The minimum vocabulary: bit, byte, signal11 min
Networks cannot be discussed without four words. We take them one at a time, assuming nothing.
A bit is the smallest possible unit of information: the answer to a question with two possible outcomes. Yes or no. On or off. 1 or 0. There is nothing smaller, because a question with only one possible answer conveys no information at all.
A single bit is nearly useless. The power comes from combination: with 2 bits you can express 4 different things, with 3 bits 8, with 8 bits 256. The rule is that n bits express 2n values - and practically all the arithmetic in this course follows from it.
A byte is a group of 8 bits. It can represent 28 = 256 different values,
that is, the numbers from 0 to 255. This is not an arbitrary choice: 255 is exactly the largest
number that fits in one byte of an IP address, and that is why the address
192.168.300.1 does not exist.
The same byte can be written in three ways, all meaning exactly the same thing:
- decimal - the way we count:
141 - binary - the way the equipment "thinks":
10001101 - hexadecimal - a convenient shorthand for binary:
0x8D
Binary ↔ decimal conversion is not magic and is not learned by heart. Each position in a byte has a weight - from right to left: 1, 2, 4, 8, 16, 32, 64, 128. The decimal value is simply the sum of the weights of the positions holding a 1. That is all.
Play with it until the conversion becomes reflex. You will need it in every lecture from here on, and in subnetting (lecture 5) it will be the only tool that matters. Press "new exercise" and try to guess the binary representation before it is revealed.
Base 16 uses the digits 0–9 plus the letters A–F for the values 10–15. The reason it exists is purely practical: exactly 4 bits fit into one hexadecimal digit, so a byte is always written with two hex digits, and the conversion is done in groups of four, with no calculation at all.
1000 1101 → 8 and D → 0x8D. This is why MAC
and IPv6 addresses are written in hexadecimal: they are far shorter than in binary and far easier
to convert than decimal.
A bit is an idea. For it to travel from one computer to another, the idea has to be turned into something physical: a voltage on a wire, a pulse of light in a fibre, a radio wave in the air. That "something physical" is called a signal.
The whole of the rest of this lecture deals with a single question: how do we turn a string of bits into a signal that the other end can read correctly - and back again.
The units trap
This is where the most frequent misunderstanding between a network administrator and an unhappy user arises. Network speeds are measured in bits, file sizes in bytes. Between them there is a factor of 8 - plus a second, subtler difference.
| Notation | Read as | Value | Where it appears |
|---|---|---|---|
b | bit | 1 bit | transmission speeds |
B | byte | 8 bits | file sizes |
kbps, Mbps | kilo/megabits per second | 103, 106 bits/s | bandwidth |
KB, MB | kilo/megabytes | 210, 220 bytes | storage capacity |
You have a 100 Mbps connection. At best, how long does it take to download a 500 MB file?
See the solution
500 MB = 500 × 220 bytes = 524,288,000 bytes = 4,194,304,000 bits.
100 Mbps = 108 bits per second.
4,194,304,000 / 100,000,000 ≈ 42 seconds - and that only in the ideal case, with no protocol headers, no retransmissions and nobody else on the link. In practice, allow another 15–20 %.
Notice: a user who expected "500 MB over 100 Mbps means 5 seconds" was wrong by a factor of eight - exactly the factor between a bit and a byte.
3How big is a network: LAN, MAN, WAN8 min
The oldest classification of networks is by the distance between nodes. This is not merely a convention of vocabulary: each class has developed protocols of its own, suited to its order of magnitude, because the problems really are different at 10 metres and at 10,000 kilometres.
| Distance between nodes | What it covers | Network type | Dominant protocols |
|---|---|---|---|
| 1 mm – 1 cm | on silicon, between processors | micro-networks | internal buses |
| ~1 m | personal space | PAN | Bluetooth, Zigbee |
| 10 m | a room | LAN Local Area Network | Ethernet (IEEE 802.3), Wi-Fi (IEEE 802.11) |
| 100 m | a building | ||
| 1 km | a campus | ||
| 10 km | a city | MAN | rarely seen today as a distinct entity |
| 100 – 1000 km | a country, a continent | WAN | MPLS, PPP, Frame Relay, ATM |
| ~10,000 km | the planet | Internet | the TCP/IP stack |
The course is called Local Area Networks, so the emphasis falls on the middle rows - but you will notice that they cannot be understood without what happens at the boundary.
This frontier organises the whole semester: lectures 2–4 deal with what happens inside it, lectures 5–10 with what happens at it and beyond it, and lectures 8, 11 and 12 with those who try to cross it without permission.
4The devices and the layer they work at13 min
In a shop, network devices appear to differ in port count and price. In reality the essential difference is another: how deeply each one looks into what passes through it.
- A repeater does not look at all. It receives a weakened signal, amplifies it and sends it onward. It does not know and does not care what the signal carries.
- A switch looks at the layer 2 address, the one written into the network card, and decides which port to send the frame out of.
- A router looks at the layer 3 address, the IP address, and decides which network to send the packet towards.
- A firewall looks deeper still - as far as the content itself - and decides whether it is allowed through.
The deeper it looks, the better the decisions it makes - and the more time it spends doing so. This is the trade-off that recurs throughout the rest of the course.
| Device | OSI layer | Decision based on | Effect on domains |
|---|---|---|---|
| Repeater, hub | 1 - physical | nothing; regenerates the signal | extends both collision and broadcast |
| Switch (bridge) | 2 - data link | destination MAC address | separates collision, extends broadcast |
| Router | 3 - network | destination IP address | separates both collision and broadcast |
| Firewall / IPS | 3–7 | headers and content | filters rather than merely forwards |
Do not be alarmed if the last two columns mean nothing to you yet - they are explained at length in lecture 3. For now, two sentences:
A collision domain is the group of devices that can interfere with one another if they speak at the same time. A broadcast domain is the group that hears a general announcement made by any one of them. Both are better the smaller they are.
The network interface
The term interface denotes a point of communication with a network: a computer's network card, a switch port, a router port. A computer with a single card has a single interface; one with two cards has two. A router has several, and that is precisely why it can join different networks - this is, in practice, its entire reason for existing.
In operating systems, the interface is also a software abstraction. On Linux, Ethernet cards appear
as eth0, eth1, and a virtual interface called the loopback
(lo, address 127.0.0.1) lets a computer address its own protocol stack as
though it were out on the network. It is the first thing you test when "nothing works": if even the
loopback does not answer, the problem is not in the network, it is in the computer.
5The protocol: why we must speak the same language7 min
A network protocol typically settles four things:
- Syntax - what the message looks like: which fields it has, in what order, how many bits each.
- Semantics - what each field means and what must be done with it.
- Timing - who speaks when, how long they wait, what they do if no reply comes.
- Error handling - what happens when something goes wrong.
You will meet these four components in every protocol studied this semester. When a new protocol strikes you as complicated, ask yourself which of the four is eluding you - it is usually the third.
6Why we communicate in layers13 min
Suppose we were to write an e-mail program ourselves, with nothing existing beforehand. We would have to deal with the voltage on the wire, with retransmitting lost packets, with their ordering, with the message format and with much else - all in the same body of code. It would be impossible to maintain. Worse, it would have to be rewritten from scratch every time a new kind of cable appeared.
The solution, more than forty years old, is layering: we divide the problem into layers, and each layer offers a service to the layer above and uses the services of the layer below, without knowing how they are implemented.
- A layer can be replaced without affecting the rest: moving from copper to fibre changes nothing for the application
- Problems become localised: "it does not work" becomes "it does not work at layer 3"
- Different manufacturers can implement different layers
- It can be learned one layer at a time, which is exactly what this course does
- Every layer adds a header - so there is traffic that is not useful information
- Information useful to one layer is not accessible to another; NAT, in lecture 9, is considered a violation for precisely this reason
- Passing through the layers costs processing time
The OSI model: seven layers
Click each layer on the left. Do not try to memorise the table - come back to it whenever you need it. What should stay with you is the idea: each layer has its own kind of address and its own unit of data.
The TCP/IP stack: four layers, but real ones
The OSI model is a reference model: we use it as a shared vocabulary. The TCP/IP stack is what actually runs on the Internet. They do not contradict each other - TCP/IP merges the layers that OSI separates, because in practice nobody implements session and presentation separately.
| OSI | TCP/IP | What it solves | Examples |
|---|---|---|---|
| 7. Application 6. Presentation 5. Session | Application | services for the user, data representation, control of the dialogue | HTTP, DNS, SSH, SMTP |
| 4. Transport | Transport | flow control, transmission reliability, multiplexing over ports | TCP, UDP |
| 3. Network | Internet | global addressing and choosing the path to the destination | IPv4, IPv6, ICMP |
| 2. Data link 1. Physical | Network access | access to the shared medium, framing, binary transmission | Ethernet, 802.11, PPP |
Encapsulation: how a message descends the stack
This is the most important figure in the whole lecture. Follow it step by step.
When troubleshooting a network, the first useful question is always: at which layer does it break? The rest of the semester gives you the tools with which to answer it.
7The physical layer and its four tasks9 min
The physical layer has a mission that can be stated in one sentence: to turn a bit into a signal and back again. In practice that mission breaks down into four distinct tasks.
- Turning a bit into a signal. Choosing the voltage levels, the wavelength or the frequency that represent the logical values.
- Bit synchronisation. The receiver must know where one bit ends and the next begins. It sounds like a detail; it is the reason half the encodings we are about to see exist at all.
- Speed control. Negotiating a transmission rate that both ends can sustain - the reason a gigabit card and a 100 Mbps switch nevertheless come to an understanding.
- Multiplexing. Placing several logical streams on the same physical medium. A subject treated at length in the next lecture.
Imagine I dictate a series of digits to you, but with no pause between them and without telling you how long one lasts. If I dictate ten "zero"s in a row, in an even voice, how many did you write down? You have no way of knowing - perhaps there were eight, perhaps twelve.
This is exactly the receiver's problem: if the voltage stays constant for a long time, it has no way to count the bits. The solution is either a separate clock (expensive: one more wire) or an encoding that guarantees frequent changes in the signal, from which the clock can be recovered. Almost all modern encodings choose the second.
Four combinations of data and signal
The words "digital" and "analogue" apply separately to the data and to the signal. Four situations result - all of them with real uses.
| The data is | The signal is | The operation is called | Real-life example |
|---|---|---|---|
| digital | digital | line coding | Ethernet over twisted pair |
| digital | analogue | modulation | Wi-Fi, DSL, a modem on a telephone line |
| analogue | digital | sampling and quantisation | digitised voice, VoIP |
| analogue | analogue | analogue modulation | AM and FM radio, classic telephony |
8Line coding: what a bit looks like on the wire13 min
A line code is the rule by which a string of bits becomes a waveform. It is the most concrete part of this course: here you can literally see what information looks like on the cable.
A good line code solves three problems at once:
- Distinguishability. The receiver must be able to tell a 1 from a 0 unambiguously, even with a little noise on the wire.
- Synchronisation. The receiver must be able to recover its clock from the received signal itself.
- No DC component. The average of the signal must be zero, because the isolation transformers in network cards do not pass a DC component. A signal that stays "high" for a long time does not get through them.
Line codes are classified by the number of voltage levels they use: unipolar (one level plus the absence of a signal), polar (two levels, positive and negative) and bipolar or multilevel (three or more).
What to observe, one at a time
Do not move on without actually performing the experiments below. They take five minutes and they replace an hour of reading.
| Encoding | The rule | Experiment to try | What you will observe |
|---|---|---|---|
| NRZ-L Non-Return to Zero Level | the level is the bit value | type 00000000 | a perfectly flat line - the receiver has nothing left to count |
| NRZ-I NRZ Inverted | 1 = change the level, 0 = keep it | type 11111111, then 00000000 | ones produce a transition at every bit, zeros still none at all |
| Manchester | a mandatory transition at the middle of every bit | any string | the signal changes twice as often as the bit rate |
| Differential Manchester | the mid-bit transition stays; the information is in the presence of a transition at the start | compare it with plain Manchester | it survives swapped wires |
| MLT-3 | three levels cycled through, advancing only on a 1 | type 11111111 | the cycle 0 → +1 → 0 → −1; the frequency drops fourfold |
| PAM-5 | two bits per symbol, on amplitude levels | any even-length string | half the number of transitions for the same bits |
With NRZ, an alternating string 101010… at 100 Mbps produces a level change at every
bit, that is, a fundamental frequency of 50 MHz. Cat 5 cable is certified up to 100 MHz, but at 50 MHz
it already radiates enough to disturb the neighbouring pairs and the equipment around it.
MLT-3 advances through its four-state cycle only on the 1 bits, so it needs four bits to return to its starting point. The fundamental frequency drops to 25 MHz. This is, in essence, the idea that allowed 100 Mbps to travel over the cheap telephone cable already run through buildings - and everything else followed from there.
Block coding: 4B/5B
The problem of long runs of zeros can be solved elegantly not at the signal level but before it: we rewrite the data so that it no longer contains problematic sequences.
The 4B/5B code replaces each group of 4 data bits with a group of 5 bits chosen from a specially constructed table: no code word has more than three consecutive zeros. The receiver performs the reverse translation and recovers the original data.
| data (4b) | code (5b) | data (4b) | code (5b) | control symbol | code (5b) |
|---|---|---|---|---|---|
| 0000 | 11110 | 1000 | 10010 | Q - quiet line | 00000 |
| 0001 | 01001 | 1001 | 10011 | I - idle | 11111 |
| 0010 | 10100 | 1010 | 10110 | J - start of stream | 11000 |
| 0011 | 10101 | 1011 | 10111 | K - start of stream | 10001 |
| 0100 | 01010 | 1100 | 11010 | T - end of stream | 01101 |
| 0101 | 01011 | 1101 | 11011 | R - reset | 00111 |
| 0110 | 01110 | 1110 | 11100 | S - set | 11001 |
| 0111 | 01111 | 1111 | 11101 | H - error | 00100 |
But of the 32 possible 5-bit words, only 16 are used for data. The other 16 are not wasted - they become control symbols that mark the start of a frame, its end, or the idle state of the line. In other words, the redundancy bought for synchronisation solves the problem of delimiting messages free of charge.
How they combine in real technologies
No standard uses a single technique. They are chained together, each solving a different problem:
| Standard | Medium | Coding chain | Why this way |
|---|---|---|---|
| 100BASE-TX | copper, Cat 5 | data → 4B/5B → MLT-3 | 4B/5B provides synchronisation, MLT-3 brings the frequency down |
| 100BASE-FX | multimode fibre | data → 4B/5B → NRZ-I | on fibre the frequency is not a problem, so NRZ-I suffices |
| 1000BASE-SX / LX | fibre | data → 8B/10B → NRZ | 8B/10B balances the DC component as well |
| 1000BASE-T | copper, Cat 5e | data → PAM-5, on 4 pairs at once | 250 Mbps on each pair, in both directions |
9Modulation: digital data over an analogue signal13 min
If the medium does not tolerate abrupt voltage transitions - a telephone line, the air, a radio channel - the bits cannot be transmitted directly. They have to be carried by modifying a continuous signal called the carrier.
Imagine a steady whistle, at the same pitch and the same loudness. In itself it conveys nothing - but if you make it louder and softer according to a convention agreed in advance, you have conveyed information on top of the whistle. The whistle is the carrier; modifying it is modulation.
A sine wave has exactly three parameters we can change - and consequently there are exactly three fundamental modulations.
| Parameter changed | Name | What it looks like | Where it appears |
|---|---|---|---|
| amplitude (the height of the wave) | ASK - Amplitude Shift Keying | 1 = strong wave, 0 = weak or absent wave | optical fibre, RFID |
| frequency (how often it oscillates) | FSK - Frequency Shift Keying | 1 and 0 are two distinct frequencies | old modems, classic Bluetooth |
| phase (the offset of the wave) | PSK - Phase Shift Keying | 1 and 0 differ by their position in the cycle | Wi-Fi, satellite, DVB |
Bit rate is not symbol rate
The symbol rate (baud) is the number of state changes of the signal in one second.
Always baud ≤ bps, and the ratio between them is exactly the number of bits encoded in one symbol.
This distinction is the key to the whole evolution of communications. A channel has a physical limit on its symbol rate, imposed by its bandwidth - you cannot exceed it however much you pay. The only way to transmit more bits per second is to encode more bits into each symbol.
This is where the combined schemes come in. QAM (Quadrature Amplitude Modulation) varies amplitude and phase at the same time, obtaining several distinct points in the plane - each point being one symbol.
That is why, as you move away from the Wi-Fi router, the speed drops: the equipment steps down automatically to a modulation with fewer points, which is more robust. Nothing has broken - a trade between throughput and reliability is being negotiated in real time.
Consider a line with a capacity of 2400 baud. How many data bits can be sent per second if QAM-16 is used for modulation?
See the solution
QAM-16 has 16 constellation points. The number of bits encoded in one symbol is log216 = 4.
4 bits/symbol × 2400 symbols/s = 9600 bps
The same line, with QAM-64, would carry 6 × 2400 = 14,400 bps - with no physical change at all, merely by changing the modulation scheme. Check both in the calculator above.
A Wi-Fi channel transmits 1000 symbols per microsecond using QAM-256, on a single spatial stream. What gross bit rate results? And if four MIMO streams are used at once?
See the hint
log2256 = 8 bits/symbol. 1000 symbols/µs = 109 symbols/s. That gives 8 Gbps gross per stream; with four streams, 32 Gbps.
The useful throughput is considerably lower - error-correcting codes, headers and guard intervals consume a significant share. This is the observation that returns with every technology in this course: the figure on the box is not the figure you see.
10Common mistakes4 min
- "100 Mbps means 100 MB per second" A factor of 8 between bit and byte, plus the difference between powers of 10 and powers of 2. A 100 Mbps link transfers, at best, about 11.9 MB per second. Always check whether the letter is lower case (bit) or upper case (byte).
- "A switch is a better hub" They are not in the same family. A hub blindly repeats the signal on every port; a switch reads the address and sends the frame only where it belongs. The difference is one of layer, not of quality. Remember the layer: hub = 1, switch = 2, router = 3.
- "OSI and TCP/IP are two competing things" They are not. OSI is a reference model, used as vocabulary; TCP/IP is what actually runs. They overlap, they do not contradict each other. When somebody says "layer 3", they are speaking OSI, even if the network runs TCP/IP.
- "Modulation and line coding are the same thing" Line coding produces a digital signal (voltage steps). Modulation modifies a continuous analogue signal. Ask yourself what kind of signal comes out: steps → coding; a wave → modulation.
- "The denser the modulation, the better" Only if the signal is clean enough. QAM-1024 on a noisy link produces more errors and more retransmissions than QPSK - and therefore, in the end, a lower throughput. Real throughput depends on the signal-to-noise ratio, not only on the scheme chosen.
11Summary and glossary5 min
This lecture has built the vocabulary and descended as far as the wire. Three ideas endure:
- Communication is organised in layers, each with its own address and its own unit of data. When something does not work, the first question is which layer.
- The physical layer turns bits into a signal, and its hard problem is not telling a 1 from a 0 but synchronisation.
- Throughput is obtained either by transmitting more symbols per second - which is physically limited - or by putting more bits into each symbol. The second route produced the whole evolution of the last thirty years.
12Self-check questions8 min
13Further reading and bibliography2 min
The next lecture stays at the physical layer but changes the question: not what the signal looks like, but through what we send it. Copper, fibre or air - what each can carry, how far, what degrades along the way, and how we put several conversations on a single wire.
In laboratory 1 you build your first topology in Packet Tracer and follow a real packet from one end to the other, with all the headers you have seen here.
- IEEE 802.3 - the Ethernet standard, the chapters on the physical layer (PCS, PMA)
- Andrew S. Tanenbaum, Computer Networks, chapter 2 - the physical layer
- James Kurose, Keith Ross, Computer Networking: A Top-Down Approach, chapter 1
- RFC 1122 - requirements of the TCP/IP stack for hosts
- For those who want the mathematics: Nyquist's theorem and Shannon's formula for the capacity of a noisy channel