LECTURE 01

Analogue and Digital Transmission

Duration: 113 min of teaching Level: bachelor, year III - no prior knowledge assumed Course: Local Area Networks Related lab: Laboratory 01 PDF: download the notes RO versiunea română

This is the first lecture and it starts from nothing: it does not assume you know what a bit is, what a protocol is, or why a network cable has eight wires in it. By the end you will understand what physically happens between the moment you press Enter and the moment the message arrives at the other end - and you will have the vocabulary in which every discussion for the rest of the semester is conducted.

1Scope and structure of the course7 min

You have two computers in a room and you want one of them to send the other a file. It sounds simple. It is not. For it to work, somebody had to answer, one at a time, some fifteen questions:

  • What kind of wire connects them, and how long may it be?
  • What does the digit 1 look like as an electrical voltage? And the digit 0?
  • How does the receiver know where one bit ends and the next begins?
  • How does it know the message has started? How does it know it has finished?
  • If both of them speak at the same time, what happens?
  • If a bit is corrupted along the way, how do we notice?
  • If there are ten computers rather than two, how do you say who the message is for?
  • What if the other computer is in a different building? In a different country?

Every one of these questions has an answer, and the answers are called protocols. This lecture answers the first five - the physical ones. The rest of the semester climbs, question by question, to the last.

Preliminary notions

A computer network is the interconnection of several computing systems so that they can exchange information. The definition is short, but it hides an important distinction.

Inside a computer, the components communicate over copper traces on the motherboard, a few centimetres long, governed by the chipset. Nobody interferes with them, nobody else is listening in, and the signal always arrives intact.

Between computers we have none of those guarantees: the distance may be metres or thousands of kilometres, the medium may be noisy, and the far end may belong to somebody else. Everything that follows in this course exists in order to make up for those three missing guarantees.

Learning outcomes

  • Explain what a bit and a byte are, and why 100 Mbps does not mean 100 MB per second
  • Classify a network by its reach: LAN, MAN, WAN
  • State, for any piece of network equipment, which layer it works at and what decision it makes
  • Explain why communication is organised in layers, and move between the OSI and TCP/IP models
  • Read and draw the waveform for NRZ-L, NRZ-I, Manchester, MLT-3 and PAM-5
  • Calculate the bit rate from the symbol rate and the modulation scheme

2The minimum vocabulary: bit, byte, signal11 min

Networks cannot be discussed without four words. We take them one at a time, assuming nothing.

The bit

A bit is the smallest possible unit of information: the answer to a question with two possible outcomes. Yes or no. On or off. 1 or 0. There is nothing smaller, because a question with only one possible answer conveys no information at all.

A single bit is nearly useless. The power comes from combination: with 2 bits you can express 4 different things, with 3 bits 8, with 8 bits 256. The rule is that n bits express 2n values - and practically all the arithmetic in this course follows from it.

The byte and numbers

A byte is a group of 8 bits. It can represent 28 = 256 different values, that is, the numbers from 0 to 255. This is not an arbitrary choice: 255 is exactly the largest number that fits in one byte of an IP address, and that is why the address 192.168.300.1 does not exist.

The same byte can be written in three ways, all meaning exactly the same thing:

  • decimal - the way we count: 141
  • binary - the way the equipment "thinks": 10001101
  • hexadecimal - a convenient shorthand for binary: 0x8D

Binary ↔ decimal conversion is not magic and is not learned by heart. Each position in a byte has a weight - from right to left: 1, 2, 4, 8, 16, 32, 64, 128. The decimal value is simply the sum of the weights of the positions holding a 1. That is all.

Number base converter and trainer

Play with it until the conversion becomes reflex. You will need it in every lecture from here on, and in subnetting (lecture 5) it will be the only tool that matters. Press "new exercise" and try to guess the binary representation before it is revealed.

Hexadecimal, in three sentences

Base 16 uses the digits 0–9 plus the letters A–F for the values 10–15. The reason it exists is purely practical: exactly 4 bits fit into one hexadecimal digit, so a byte is always written with two hex digits, and the conversion is done in groups of four, with no calculation at all.

1000 1101 → 8 and D → 0x8D. This is why MAC and IPv6 addresses are written in hexadecimal: they are far shorter than in binary and far easier to convert than decimal.

The signal

A bit is an idea. For it to travel from one computer to another, the idea has to be turned into something physical: a voltage on a wire, a pulse of light in a fibre, a radio wave in the air. That "something physical" is called a signal.

The whole of the rest of this lecture deals with a single question: how do we turn a string of bits into a signal that the other end can read correctly - and back again.

The units trap

This is where the most frequent misunderstanding between a network administrator and an unhappy user arises. Network speeds are measured in bits, file sizes in bytes. Between them there is a factor of 8 - plus a second, subtler difference.

NotationRead asValueWhere it appears
bbit1 bittransmission speeds
Bbyte8 bitsfile sizes
kbps, Mbpskilo/megabits per second103, 106 bits/sbandwidth
KB, MBkilo/megabytes210, 220 bytesstorage capacity
Worked example

You have a 100 Mbps connection. At best, how long does it take to download a 500 MB file?

See the solution

500 MB = 500 × 220 bytes = 524,288,000 bytes = 4,194,304,000 bits.

100 Mbps = 108 bits per second.

4,194,304,000 / 100,000,000 ≈ 42 seconds - and that only in the ideal case, with no protocol headers, no retransmissions and nobody else on the link. In practice, allow another 15–20 %.

Notice: a user who expected "500 MB over 100 Mbps means 5 seconds" was wrong by a factor of eight - exactly the factor between a bit and a byte.

3How big is a network: LAN, MAN, WAN8 min

The oldest classification of networks is by the distance between nodes. This is not merely a convention of vocabulary: each class has developed protocols of its own, suited to its order of magnitude, because the problems really are different at 10 metres and at 10,000 kilometres.

Distance between nodesWhat it coversNetwork typeDominant protocols
1 mm – 1 cmon silicon, between processorsmicro-networksinternal buses
~1 mpersonal spacePANBluetooth, Zigbee
10 ma roomLAN
Local Area Network
Ethernet (IEEE 802.3),
Wi-Fi (IEEE 802.11)
100 ma building
1 kma campus
10 kma cityMANrarely seen today as a distinct entity
100 – 1000 kma country, a continentWANMPLS, PPP, Frame Relay, ATM
~10,000 kmthe planetInternetthe TCP/IP stack

The course is called Local Area Networks, so the emphasis falls on the middle rows - but you will notice that they cannot be understood without what happens at the boundary.

Where a LAN ends The boundary between the local network and the rest of the world is the router, also called the gateway - a word that means exactly what it says. Everything on the inner side of the router belongs to the local network; everything that passes beyond it enters another network.

This frontier organises the whole semester: lectures 2–4 deal with what happens inside it, lectures 5–10 with what happens at it and beyond it, and lectures 8, 11 and 12 with those who try to cross it without permission.

4The devices and the layer they work at13 min

In a shop, network devices appear to differ in port count and price. In reality the essential difference is another: how deeply each one looks into what passes through it.

  1. A repeater does not look at all. It receives a weakened signal, amplifies it and sends it onward. It does not know and does not care what the signal carries.
  2. A switch looks at the layer 2 address, the one written into the network card, and decides which port to send the frame out of.
  3. A router looks at the layer 3 address, the IP address, and decides which network to send the packet towards.
  4. A firewall looks deeper still - as far as the content itself - and decides whether it is allowed through.

The deeper it looks, the better the decisions it makes - and the more time it spends doing so. This is the trade-off that recurs throughout the rest of the course.

Match each device with its role
PC A PC B PC C SWITCH layer 2 MAC addresses ROUTER layer 3 IP addresses another network Internet / WAN a single broadcast domain the local network boundary
Fig. 1 - The switch extends the local network; the router ends it. Every host to the left of the router hears every other one without an intermediary; to speak to anything on the right, it must ask the router.
DeviceOSI layerDecision based onEffect on domains
Repeater, hub1 - physicalnothing; regenerates the signalextends both collision and broadcast
Switch (bridge)2 - data linkdestination MAC addressseparates collision, extends broadcast
Router3 - networkdestination IP addressseparates both collision and broadcast
Firewall / IPS3–7headers and contentfilters rather than merely forwards
What "domain" means

Do not be alarmed if the last two columns mean nothing to you yet - they are explained at length in lecture 3. For now, two sentences:

A collision domain is the group of devices that can interfere with one another if they speak at the same time. A broadcast domain is the group that hears a general announcement made by any one of them. Both are better the smaller they are.

Exercise: how many domains do you see in this topology?

The network interface

The term interface denotes a point of communication with a network: a computer's network card, a switch port, a router port. A computer with a single card has a single interface; one with two cards has two. A router has several, and that is precisely why it can join different networks - this is, in practice, its entire reason for existing.

In operating systems, the interface is also a software abstraction. On Linux, Ethernet cards appear as eth0, eth1, and a virtual interface called the loopback (lo, address 127.0.0.1) lets a computer address its own protocol stack as though it were out on the network. It is the first thing you test when "nothing works": if even the loopback does not answer, the problem is not in the network, it is in the computer.

5The protocol: why we must speak the same language7 min

Protocol
A set of rules governing the way two entities exchange information. Both ends of a transmission must use the same protocol; otherwise the signal arrives but the message does not.
Analogy Two people meet. One says "早上好". The other hears every sound perfectly - the signal has arrived intact - but understands nothing, because they do not share the same convention about what the sounds mean. That is the difference between hearing and understanding, and in networking it is the difference between the physical layer and everything above it.

A network protocol typically settles four things:

  1. Syntax - what the message looks like: which fields it has, in what order, how many bits each.
  2. Semantics - what each field means and what must be done with it.
  3. Timing - who speaks when, how long they wait, what they do if no reply comes.
  4. Error handling - what happens when something goes wrong.

You will meet these four components in every protocol studied this semester. When a new protocol strikes you as complicated, ask yourself which of the four is eluding you - it is usually the third.

6Why we communicate in layers13 min

Suppose we were to write an e-mail program ourselves, with nothing existing beforehand. We would have to deal with the voltage on the wire, with retransmitting lost packets, with their ordering, with the message format and with much else - all in the same body of code. It would be impossible to maintain. Worse, it would have to be rewritten from scratch every time a new kind of cable appeared.

The solution, more than forty years old, is layering: we divide the problem into layers, and each layer offers a service to the layer above and uses the services of the layer below, without knowing how they are implemented.

Analogy You send a parcel. You wrap it and write the address - that is your layer. You take it to the post office, which puts it in a sack with other parcels for the same city - another layer. The sack goes into a lorry - another layer. The driver does not know and does not care what is in the parcel; you do not know and do not care which motorway the lorry takes. Each layer has its own "wrapping" and its own address, and at the destination everything is unwrapped in reverse order.
What we gain
  • A layer can be replaced without affecting the rest: moving from copper to fibre changes nothing for the application
  • Problems become localised: "it does not work" becomes "it does not work at layer 3"
  • Different manufacturers can implement different layers
  • It can be learned one layer at a time, which is exactly what this course does
What we lose
  • Every layer adds a header - so there is traffic that is not useful information
  • Information useful to one layer is not accessible to another; NAT, in lecture 9, is considered a violation for precisely this reason
  • Passing through the layers costs processing time

The OSI model: seven layers

Click each layer on the left. Do not try to memorise the table - come back to it whenever you need it. What should stay with you is the idea: each layer has its own kind of address and its own unit of data.

Explore the OSI stack, layer by layer

The TCP/IP stack: four layers, but real ones

The OSI model is a reference model: we use it as a shared vocabulary. The TCP/IP stack is what actually runs on the Internet. They do not contradict each other - TCP/IP merges the layers that OSI separates, because in practice nobody implements session and presentation separately.

OSITCP/IPWhat it solvesExamples
7. Application
6. Presentation
5. Session
Applicationservices for the user, data representation, control of the dialogueHTTP, DNS, SSH, SMTP
4. TransportTransportflow control, transmission reliability, multiplexing over portsTCP, UDP
3. NetworkInternetglobal addressing and choosing the path to the destinationIPv4, IPv6, ICMP
2. Data link
1. Physical
Network accessaccess to the shared medium, framing, binary transmissionEthernet, 802.11, PPP

Encapsulation: how a message descends the stack

This is the most important figure in the whole lecture. Follow it step by step.

Encapsulation, step by step
Worth remembering, even if you forget the rest Layer 2 works with MAC addresses and frames. Layer 3 with IP addresses and packets. Layer 4 with ports and segments.

When troubleshooting a network, the first useful question is always: at which layer does it break? The rest of the semester gives you the tools with which to answer it.

7The physical layer and its four tasks9 min

The physical layer has a mission that can be stated in one sentence: to turn a bit into a signal and back again. In practice that mission breaks down into four distinct tasks.

  1. Turning a bit into a signal. Choosing the voltage levels, the wavelength or the frequency that represent the logical values.
  2. Bit synchronisation. The receiver must know where one bit ends and the next begins. It sounds like a detail; it is the reason half the encodings we are about to see exist at all.
  3. Speed control. Negotiating a transmission rate that both ends can sustain - the reason a gigabit card and a 100 Mbps switch nevertheless come to an understanding.
  4. Multiplexing. Placing several logical streams on the same physical medium. A subject treated at length in the next lecture.
Why synchronisation is a real problem

Imagine I dictate a series of digits to you, but with no pause between them and without telling you how long one lasts. If I dictate ten "zero"s in a row, in an even voice, how many did you write down? You have no way of knowing - perhaps there were eight, perhaps twelve.

This is exactly the receiver's problem: if the voltage stays constant for a long time, it has no way to count the bits. The solution is either a separate clock (expensive: one more wire) or an encoding that guarantees frequent changes in the signal, from which the clock can be recovered. Almost all modern encodings choose the second.

Four combinations of data and signal

The words "digital" and "analogue" apply separately to the data and to the signal. Four situations result - all of them with real uses.

The data isThe signal isThe operation is calledReal-life example
digitaldigitalline codingEthernet over twisted pair
digitalanaloguemodulationWi-Fi, DSL, a modem on a telephone line
analoguedigitalsampling and quantisationdigitised voice, VoIP
analogueanalogueanalogue modulationAM and FM radio, classic telephony
Where the word "modem" comes from A modem is a MOdulator/DEModulator: it turns the computer's digital data into a signal suited to the medium and performs the inverse operation on reception. The word has outlived the technology that produced it - the box from your Internet provider is still called a modem, although inside it bears no resemblance at all to its ancestor on the telephone line.

8Line coding: what a bit looks like on the wire13 min

A line code is the rule by which a string of bits becomes a waveform. It is the most concrete part of this course: here you can literally see what information looks like on the cable.

A good line code solves three problems at once:

  1. Distinguishability. The receiver must be able to tell a 1 from a 0 unambiguously, even with a little noise on the wire.
  2. Synchronisation. The receiver must be able to recover its clock from the received signal itself.
  3. No DC component. The average of the signal must be zero, because the isolation transformers in network cards do not pass a DC component. A signal that stays "high" for a long time does not get through them.

Line codes are classified by the number of voltage levels they use: unipolar (one level plus the absence of a signal), polar (two levels, positive and negative) and bipolar or multilevel (three or more).

Waveform generator - try each encoding

What to observe, one at a time

Do not move on without actually performing the experiments below. They take five minutes and they replace an hour of reading.

EncodingThe ruleExperiment to tryWhat you will observe
NRZ-L
Non-Return to Zero Level
the level is the bit valuetype 00000000a perfectly flat line - the receiver has nothing left to count
NRZ-I
NRZ Inverted
1 = change the level, 0 = keep ittype 11111111, then 00000000ones produce a transition at every bit, zeros still none at all
Manchestera mandatory transition at the middle of every bitany stringthe signal changes twice as often as the bit rate
Differential Manchesterthe mid-bit transition stays; the information is in the presence of a transition at the startcompare it with plain Manchesterit survives swapped wires
MLT-3three levels cycled through, advancing only on a 1type 11111111the cycle 0 → +1 → 0 → −1; the frequency drops fourfold
PAM-5two bits per symbol, on amplitude levelsany even-length stringhalf the number of transitions for the same bits
Why MLT-3 matters - a calculation that explains an industry

With NRZ, an alternating string 101010… at 100 Mbps produces a level change at every bit, that is, a fundamental frequency of 50 MHz. Cat 5 cable is certified up to 100 MHz, but at 50 MHz it already radiates enough to disturb the neighbouring pairs and the equipment around it.

MLT-3 advances through its four-state cycle only on the 1 bits, so it needs four bits to return to its starting point. The fundamental frequency drops to 25 MHz. This is, in essence, the idea that allowed 100 Mbps to travel over the cheap telephone cable already run through buildings - and everything else followed from there.

Block coding: 4B/5B

The problem of long runs of zeros can be solved elegantly not at the signal level but before it: we rewrite the data so that it no longer contains problematic sequences.

The 4B/5B code replaces each group of 4 data bits with a group of 5 bits chosen from a specially constructed table: no code word has more than three consecutive zeros. The receiver performs the reverse translation and recovers the original data.

data (4b)code (5b)data (4b)code (5b)control symbolcode (5b)
000011110100010010Q - quiet line00000
000101001100110011I - idle11111
001010100101010110J - start of stream11000
001110101101110111K - start of stream10001
010001010110011010T - end of stream01101
010101011110111011R - reset00111
011001110111011100S - set11001
011101111111111101H - error00100
The cost of redundancy, and its upside 4B/5B introduces 25 % overhead: to carry 100 Mbps of useful data, the line must run at 125 Mbaud. It looks wasteful.

But of the 32 possible 5-bit words, only 16 are used for data. The other 16 are not wasted - they become control symbols that mark the start of a frame, its end, or the idle state of the line. In other words, the redundancy bought for synchronisation solves the problem of delimiting messages free of charge.

How they combine in real technologies

No standard uses a single technique. They are chained together, each solving a different problem:

StandardMediumCoding chainWhy this way
100BASE-TXcopper, Cat 5data → 4B/5B → MLT-34B/5B provides synchronisation, MLT-3 brings the frequency down
100BASE-FXmultimode fibredata → 4B/5B → NRZ-Ion fibre the frequency is not a problem, so NRZ-I suffices
1000BASE-SX / LXfibredata → 8B/10B → NRZ8B/10B balances the DC component as well
1000BASE-Tcopper, Cat 5edata → PAM-5, on 4 pairs at once250 Mbps on each pair, in both directions

9Modulation: digital data over an analogue signal13 min

If the medium does not tolerate abrupt voltage transitions - a telephone line, the air, a radio channel - the bits cannot be transmitted directly. They have to be carried by modifying a continuous signal called the carrier.

What a carrier is

Imagine a steady whistle, at the same pitch and the same loudness. In itself it conveys nothing - but if you make it louder and softer according to a convention agreed in advance, you have conveyed information on top of the whistle. The whistle is the carrier; modifying it is modulation.

A sine wave has exactly three parameters we can change - and consequently there are exactly three fundamental modulations.

Parameter changedNameWhat it looks likeWhere it appears
amplitude (the height of the wave)ASK - Amplitude Shift Keying1 = strong wave, 0 = weak or absent waveoptical fibre, RFID
frequency (how often it oscillates)FSK - Frequency Shift Keying1 and 0 are two distinct frequenciesold modems, classic Bluetooth
phase (the offset of the wave)PSK - Phase Shift Keying1 and 0 differ by their position in the cycleWi-Fi, satellite, DVB

Bit rate is not symbol rate

Bit rate and baud rate
The bit rate (bps) is the number of bits transmitted in one second.
The symbol rate (baud) is the number of state changes of the signal in one second.
Always baud ≤ bps, and the ratio between them is exactly the number of bits encoded in one symbol.

This distinction is the key to the whole evolution of communications. A channel has a physical limit on its symbol rate, imposed by its bandwidth - you cannot exceed it however much you pay. The only way to transmit more bits per second is to encode more bits into each symbol.

This is where the combined schemes come in. QAM (Quadrature Amplitude Modulation) varies amplitude and phase at the same time, obtaining several distinct points in the plane - each point being one symbol.

QPSK2 bits / symbol QAM-164 bits / symbol QAM-646 bits / symbol
Fig. 2 - Constellation diagrams. Each point is a distinct combination of amplitude and phase, and therefore one symbol. The more numerous and the closer together the points are, the more bits we transmit per symbol - and the less noise it takes for the receiver to confuse two neighbouring points.
The trade-off that recurs everywhere More constellation points = more bits per symbol = higher throughput. But the points get closer together, so a cleaner signal is needed to tell them apart.

That is why, as you move away from the Wi-Fi router, the speed drops: the equipment steps down automatically to a modulation with fewer points, which is more robust. Nothing has broken - a trade between throughput and reliability is being negotiated in real time.
Calculator: from symbol rate to bit rate
Worked example (an examination classic)

Consider a line with a capacity of 2400 baud. How many data bits can be sent per second if QAM-16 is used for modulation?

See the solution

QAM-16 has 16 constellation points. The number of bits encoded in one symbol is log216 = 4.

4 bits/symbol × 2400 symbols/s = 9600 bps

The same line, with QAM-64, would carry 6 × 2400 = 14,400 bps - with no physical change at all, merely by changing the modulation scheme. Check both in the calculator above.

Exercise for you

A Wi-Fi channel transmits 1000 symbols per microsecond using QAM-256, on a single spatial stream. What gross bit rate results? And if four MIMO streams are used at once?

See the hint

log2256 = 8 bits/symbol. 1000 symbols/µs = 109 symbols/s. That gives 8 Gbps gross per stream; with four streams, 32 Gbps.

The useful throughput is considerably lower - error-correcting codes, headers and guard intervals consume a significant share. This is the observation that returns with every technology in this course: the figure on the box is not the figure you see.

10Common mistakes4 min

  • "100 Mbps means 100 MB per second" A factor of 8 between bit and byte, plus the difference between powers of 10 and powers of 2. A 100 Mbps link transfers, at best, about 11.9 MB per second. Always check whether the letter is lower case (bit) or upper case (byte).
  • "A switch is a better hub" They are not in the same family. A hub blindly repeats the signal on every port; a switch reads the address and sends the frame only where it belongs. The difference is one of layer, not of quality. Remember the layer: hub = 1, switch = 2, router = 3.
  • "OSI and TCP/IP are two competing things" They are not. OSI is a reference model, used as vocabulary; TCP/IP is what actually runs. They overlap, they do not contradict each other. When somebody says "layer 3", they are speaking OSI, even if the network runs TCP/IP.
  • "Modulation and line coding are the same thing" Line coding produces a digital signal (voltage steps). Modulation modifies a continuous analogue signal. Ask yourself what kind of signal comes out: steps → coding; a wave → modulation.
  • "The denser the modulation, the better" Only if the signal is clean enough. QAM-1024 on a noisy link produces more errors and more retransmissions than QPSK - and therefore, in the end, a lower throughput. Real throughput depends on the signal-to-noise ratio, not only on the scheme chosen.

11Summary and glossary5 min

This lecture has built the vocabulary and descended as far as the wire. Three ideas endure:

  1. Communication is organised in layers, each with its own address and its own unit of data. When something does not work, the first question is which layer.
  2. The physical layer turns bits into a signal, and its hard problem is not telling a 1 from a 0 but synchronisation.
  3. Throughput is obtained either by transmitting more symbols per second - which is physically limited - or by putting more bits into each symbol. The second route produced the whole evolution of the last thirty years.
bitthe smallest unit of information: 0 or 1
bytea group of 8 bits; values from 0 to 255
protocola shared set of rules without which the message cannot be understood
LAN / MAN / WANnetworks classified by geographical reach
gatewaythe router that marks the boundary of the local network
interfacea point of connection to the network: card, port, subinterface
encapsulationthe addition, at each layer, of that layer's own header
PDUthe data unit of a layer: frame, packet, segment
line codingthe rule by which bits become voltage steps
modulationcarrying data by modifying a carrier wave
baudsymbols per second; not the same thing as bps
constellationthe map of a modulation's symbols, in amplitude and phase
Put the steps of encapsulation in order

12Self-check questions8 min

13Further reading and bibliography2 min

The next lecture stays at the physical layer but changes the question: not what the signal looks like, but through what we send it. Copper, fibre or air - what each can carry, how far, what degrades along the way, and how we put several conversations on a single wire.

In laboratory 1 you build your first topology in Packet Tracer and follow a real packet from one end to the other, with all the headers you have seen here.

  • IEEE 802.3 - the Ethernet standard, the chapters on the physical layer (PCS, PMA)
  • Andrew S. Tanenbaum, Computer Networks, chapter 2 - the physical layer
  • James Kurose, Keith Ross, Computer Networking: A Top-Down Approach, chapter 1
  • RFC 1122 - requirements of the TCP/IP stack for hosts
  • For those who want the mathematics: Nyquist's theorem and Shannon's formula for the capacity of a noisy channel