Part 1 covered how a name becomes an address and how a TCP connection gets established. Part 2 covered how that connection becomes private and verified through TLS, and how HTTP turns it into a meaningful request and response, including methods, status codes, and statelessness. This article assumes all of that and picks up with a question those first two parts never had to deal with: what happens when a single page needs dozens or hundreds of requests, not just one?

The Problem HTTP/1.1 Has in Practice

HTTP/1.1 was designed around a simple model: one request, one response, per connection, one at a time (unless the connection is reused for the next request after the first completes). Real pages don't work that way. A single page load might need the HTML document, several stylesheets, scripts, fonts, and images, each as a separate request.

Two limitations made this genuinely painful:

  • Head-of-line blocking at the HTTP layer: on a single connection, HTTP/1.1 requests are processed in order. If one response is slow, everything queued behind it on that connection waits, even if the server could otherwise have answered the later ones instantly.
  • A small limit on parallel connections per host: browsers cap how many simultaneous connections they'll open to a single domain, both to avoid overwhelming servers and because opening more connections has real overhead (each one needs its own TCP handshake and, once TLS is involved, its own TLS handshake too).

Engineers didn't wait for a protocol fix; they worked around it, and the workarounds are good evidence the pain was real rather than theoretical. Domain sharding spread assets across several subdomains specifically to get more parallel connections past the per-host limit. Spriting combined many small images into a single larger image to cut down the number of separate requests. Concatenation merged many small CSS or JavaScript files into one larger file for the same reason. All three of these are workarounds for a transport limitation, not features anyone actually wanted; they add build complexity to dodge a protocol constraint.

HTTP/2: Fixing the Application-Layer Bottleneck

HTTP/2 addressed the HTTP-layer bottleneck directly, largely by changing how requests and responses are framed, without changing HTTP's core semantics (the same methods, the same status codes, the same general request/response model from Part 2 still apply).

  • Binary framing: instead of a plain-text request line and headers, HTTP/2 breaks messages into binary frames. This makes the protocol more efficient to parse and, more importantly, lets frames from different requests be interleaved on the same connection.
  • Multiplexed streams: a single TCP connection can carry many concurrent HTTP/2 streams, each representing an independent request/response pair. Because frames are interleaved rather than sent one full response at a time, a slow response no longer blocks other responses on the same connection the way it did in HTTP/1.1. This is the direct fix for the HTTP-layer head-of-line blocking problem described above.
  • Header compression (HPACK): HTTP headers tend to repeat heavily across requests to the same host (cookies, user-agent strings, accept headers). HPACK maintains a compression context between requests on a connection so repeated header values don't need to be resent in full every time.
  • Server push: HTTP/2 also introduced the ability for a server to proactively send resources it expects the client will need, without waiting for a separate request. In practice, this didn't pan out the way it was originally pitched: it's genuinely difficult for a server to guess correctly what a client already has cached, and pushing something the client didn't need wastes bandwidth. Most major implementations have deprecated or removed server push support, and it's worth knowing about mainly to understand why it isn't something you should reach for today.
sequenceDiagram
    participant Client
    participant Server

    Note over Client,Server: Single TCP connection, single TLS session
    Client->>Server: Stream 1: GET /style.css
    Client->>Server: Stream 3: GET /app.js
    Client->>Server: Stream 5: GET /logo.png
    Server-->>Client: Stream 3 data (interleaved)
    Server-->>Client: Stream 1 data (interleaved)
    Server-->>Client: Stream 5 data (interleaved)

The Ceiling HTTP/2 Still Hits

Multiplexing at the HTTP layer solves one problem, but it doesn't remove a limitation underneath it. HTTP/2's multiple streams still ride on a single TCP connection, and TCP guarantees ordered delivery of the entire byte stream on that connection, as covered in Part 1.

That ordering guarantee is exactly where the trouble comes back in a different form. If a single packet belonging to any one HTTP/2 stream is lost, TCP won't hand any data on that connection to the application layer until the missing packet is retransmitted and arrives, even data belonging to completely unrelated streams that arrived intact. This is TCP-layer head-of-line blocking, and it's a distinct problem from the HTTP-layer one HTTP/2 fixed. HTTP/2 solved the application's ordering problem; it never touched TCP's ordering guarantee, because it's built directly on top of TCP.

This distinction matters because it explains something that otherwise looks contradictory: HTTP/2 measurably improved performance on reliable connections, but on lossy networks (weak wifi, congested mobile connections), the benefit shrinks or disappears, because a single lost packet can stall an entire multiplexed connection regardless of how well the streams above it are organized.

HTTP/3 and QUIC: Changing the Transport Itself

HTTP/3 takes a more direct approach to the problem: instead of trying to work around TCP's guarantees, it drops TCP entirely and runs on QUIC, a transport protocol built on top of UDP.

Recall from Part 1 that UDP offers none of TCP's guarantees on its own: no ordering, no retransmission, no congestion control. QUIC reimplements those guarantees itself, but does so per stream rather than for the entire connection as one ordered byte stream. That's the key structural difference: if a packet belonging to one QUIC stream is lost, only that stream stalls while its data is retransmitted; other streams on the same connection continue delivering data to the application without waiting. This directly eliminates the TCP-layer head-of-line blocking problem described above, because there's no longer a single, connection-wide ordering guarantee for QUIC to enforce.

QUIC has a couple of other properties worth knowing, both direct consequences of building a new transport rather than reusing TCP:

  • Connection migration: a TCP connection is identified by a fixed tuple of source/destination IP and port. Change networks (say, wifi to cellular) and that tuple changes, which breaks the TCP connection and forces a full reconnect, including a new TLS handshake. QUIC identifies a connection by a connection ID that's independent of the underlying network path, so a client can switch networks and keep the same QUIC connection alive without renegotiating from scratch.
  • Encryption built into the transport: TLS, as covered in Part 2, is layered on top of TCP as a separate step after the transport connection is established. QUIC integrates the equivalent of a TLS handshake directly into its own connection setup, which is why establishing a secure QUIC connection typically takes fewer round trips than establishing TCP and then TLS separately on top of it.
sequenceDiagram
    participant Client
    participant Server

    Note over Client,Server: Single QUIC connection over UDP
    Client->>Server: Stream A data
    Client->>Server: Stream B data
    Note over Server: Packet for Stream A lost
    Server-->>Client: Stream B data continues delivering
    Client->>Server: Retransmit request for Stream A
    Server-->>Client: Stream A data (recovered)

What This Actually Means in Practice

None of this changes what a request or response means; the methods, status codes, and statelessness from Part 2 are identical over HTTP/2 and HTTP/3. What changes is how efficiently the underlying connection carries many concurrent requests, and how gracefully it handles loss and network changes.

That has a very specific practical implication: the difference between HTTP/1.1, HTTP/2, and HTTP/3 is often invisible on a fast, stable, wired connection carrying a handful of requests, and increasingly significant as conditions get worse: many concurrent assets, higher packet loss, or a client switching networks mid-session. If you're debugging why a page feels slow specifically on mobile networks but fine on a office wifi, the transport layer covered in this article, not application code, is one of the first places worth looking.

Tying the Series Together

Across this series, a single page load has gone through: a name resolved to an address through a chain of DNS servers (Part 1), a TCP connection established through a three-way handshake with explicit guarantees about ordering and reliability (Part 1), a TLS handshake layered on top to make that connection private and verify identity (Part 2), an HTTP request and response exchanged with well-defined semantics on top of that secure connection (Part 2), and finally, the transport carrying that HTTP traffic evolving from a single ordered stream per connection, to multiplexed streams still bound by that same ordering guarantee, to independently ordered streams over a transport built specifically to avoid it (this part).

None of these layers are optional, and none of them are going away soon. They're also, deliberately, not tied to any particular product, browser, or vendor. They're the same handshake, the same guarantees, and the same tradeoffs whether you look at them today or in ten years. That durability is exactly what makes them worth understanding properly instead of treating them as trivia.