mobile wallpaper 1mobile wallpaper 2mobile wallpaper 3mobile wallpaper 4
3664 字
10 分钟
HTTP/1.1:持久连接
2024-04-26

在上一篇实验中,体验了 HTTP/1.0 带来的改进:请求头与响应头、状态码、多种请求方法。这些特性让 HTTP 变得可扩展、可程序化处理。然而,HTTP/1.0 在实际使用中暴露了一个严重的性能瓶颈:每个请求都需要重新建立 TCP 连接。

想象这样一个场景:你访问一个包含 20 张图片的网页。在 HTTP/1.0 中,每张图片都需要:

  1. 建立 TCP 连接(三次握手)
  2. 发送 HTTP 请求
  3. 接收 HTTP 响应
  4. 关闭 TCP 连接(四次挥手)

这意味着 21 次 TCP 连接建立和关闭!每次 TCP 握手都需要约 1-2 个 RTT(往返时间),如果客户端与服务器之间的延迟是 100ms,仅握手就消耗了 2-4 秒。这在现代网页动辄加载数十甚至上百个资源的场景下,简直是灾难。

1997 年,HTTP/1.1 作为 RFC 2068 正式发布(后由 RFC 2616 和 RFC 7230 系列更新),针对 HTTP/1.0 的性能问题做了全面优化:

持久连接(Persistent Connection):默认保持 TCP 连接打开,一个连接可以发送多个请求和响应。这被称为「keep-alive」,消除了重复建立连接的开销。

分块传输编码(Chunked Transfer Encoding):允许服务器在不知道内容总长度时就开始传输,边生成边发送,降低了首字节延迟。

请求管道化(Pipelining):允许客户端在收到前一个响应之前就发送下一个请求,进一步减少延迟(虽然实际部署中存在兼容性问题)。

Host 头要求:必须包含 Host 请求头,这为虚拟主机(一个 IP 托管多个域名)奠定了基础。

缓存增强:引入 ETagIf-None-MatchIf-Modified-Since 等条件请求头,让缓存更加智能高效。

内容协商:客户端可以通过 Accept-LanguageAccept-Encoding 等头告诉服务器自己的偏好,服务器返回最适合的内容。

一、新增特性详解#

1.1 持久连接(Keep-Alive)#

HTTP/1.0 默认在每次响应后关闭连接。虽然 HTTP/1.0 支持非标准的 Connection: keep-alive 头,但这是可选扩展,兼容性参差不齐。

HTTP/1.1 反转了默认行为:连接默认保持打开。只有当客户端或服务器显式发送 Connection: close 时,连接才会在响应后关闭。

# HTTP/1.1 默认行为(连接保持)
GET /index.html HTTP/1.1\r\n
Host: localhost:8080\r\n
\r\n
# 显式关闭连接
GET /index.html HTTP/1.1\r\n
Host: localhost:8080\r\n
Connection: close\r\n
\r\n

持久连接减少了 TCP 握手开销(一次握手,多次请求),减轻了慢启动的影响(连接复用可以充分利用已打开的拥塞窗口),同时降低了服务器负载(减少连接建立/关闭的系统调用)。

1.2 分块传输编码#

HTTP/1.0 要求服务器在发送响应前知道 Content-Length。对于动态生成的内容(如数据库查询结果、实时日志流),服务器必须先缓冲所有数据才能计算长度,这增加了延迟。

HTTP/1.1 引入了 Transfer-Encoding: chunked,允许服务器将响应分成多个块发送,每个块前标注该块的长度(十六进制):

HTTP/1.1 200 OK\r\n
Transfer-Encoding: chunked\r\n
\r\n
7\r\n
Mozilla\r\n
9\r\n
Developer\r\n
7\r\n
Network\r\n
0\r\n
\r\n

上面的响应包含三个数据块:「Mozilla」(7 字节)、「Developer」(9 字节)、「Network」(7 字节),最后以 0\r\n\r\n 表示结束。客户端收到后自动拼接成完整内容。

1.3 请求管道化#

HTTP/1.0 中,如果客户端要发送多个请求,必须等待前一个响应返回后才能发送下一个。这是典型的串行模式:

sequenceDiagram participant C as 客户端 participant S as 服务器 C->>S: 请求1 S->>C: 响应1 C->>S: 请求2 S->>C: 响应2 C->>S: 请求3 S->>C: 响应3

HTTP/1.1 允许管道化:客户端可以连续发送多个请求,服务器按顺序返回响应:

sequenceDiagram participant C as 客户端 participant S as 服务器 C->>S: 请求1 C->>S: 请求2 C->>S: 请求3 S->>C: 响应1 S->>C: 响应2 S->>C: 响应3

理论上管道化能大幅减少延迟,但实际部署中几乎没有浏览器默认开启。原因有三个:中间代理不理解管道化请求,可能乱序转发或直接丢弃;队头阻塞让管道化在丢包时比不用更糟,因为所有后续响应都被阻塞;HTTP/1.1 要求响应必须按请求顺序返回,如果第一个请求的响应很大或很慢,管道化反而增加了等待时间。Chrome 和 Firefox 从未默认启用管道化,这个特性到 HTTP/2 的多路复用方案才真正实现。

1.4 Host 头要求#

HTTP/1.0 中 Host 头是可选的。这导致一个 IP 地址只能托管一个网站:服务器无法区分请求发往哪个域名。

HTTP/1.1 要求所有请求必须包含 Host 头:

GET /index.html HTTP/1.1\r\n
Host: www.example.com\r\n
\r\n

这让虚拟主机成为可能:一个服务器 IP 可以同时托管 blog.example.comapi.example.comwww.example.com 等多个站点,服务器根据 Host 头路由到不同的应用。

1.5 缓存增强#

HTTP/1.1 引入了强大的缓存控制机制:

ETag(实体标签):服务器为资源生成唯一标识符,客户端下次请求时带上 If-None-Match: <etag>,如果资源未变化,服务器返回 304 Not Modified(无响应体),节省带宽。

If-Modified-Since:客户端告诉服务器「我上次获取这个资源的时间是 X」,如果资源自那之后没修改,服务器返回 304。

Cache-Control:更精细的缓存控制,如 max-age=3600(缓存 1 小时)、no-cache(每次使用前验证)、private(仅供单用户缓存)等。

# 首次请求
GET /resource HTTP/1.1
Host: example.com
# 首次响应
HTTP/1.1 200 OK
ETag: "abc123"
Last-Modified: Wed, 20 Mar 2026 10:00:00 GMT
Cache-Control: max-age=3600
Content-Length: 1024
[资源内容]
# 后续请求(使用 ETag 验证)
GET /resource HTTP/1.1
Host: example.com
If-None-Match: "abc123"
# 资源未变化时的响应
HTTP/1.1 304 Not Modified
ETag: "abc123"
Cache-Control: max-age=3600
[无响应体]

1.6 内容协商#

HTTP/1.1 允许客户端表达内容偏好:

  • Accept:可接受的 MIME 类型
  • Accept-Language:偏好的自然语言
  • Accept-Encoding:支持的压缩算法
  • Accept-Charset:可接受的字符集
GET /index HTTP/1.1
Host: example.com
Accept: text/html,application/xhtml+xml
Accept-Language: zh-CN,zh;q=0.9,en;q=0.8
Accept-Encoding: gzip, deflate, br

服务器根据这些头返回最合适的内容。q 值表示优先级(0-1,默认 1)。

1.7 其他重要改进#

新增方法:PUT(更新资源)、DELETE(删除资源)、OPTIONS(查询支持的方法)、TRACE(诊断)、CONNECT(代理隧道)。这些方法让 HTTP 从”文档获取协议”变成了”应用层协议”,RESTful API 的设计就建立在这些方法之上。

100 Continue:当客户端要上传一个大文件(比如 2GB 的视频)时,可以先发一个只带 Expect: 100-continue 的请求头,服务器检查完条件后回 100 Continue,客户端再开始传 body。如果服务器磁盘满了或者不允许上传,客户端不用白传 2GB 才发现。不过很多服务器实现不规范,直接忽略 100 Continue,导致客户端超时后才自己发送 body。

206 Partial Content:视频拖进度条时,客户端发 Range: bytes=5000000-10000000,服务器只传那一段。没有 206 的话,服务器必须从头开始传,前面 44 分钟的数据全浪费了。这也是 CDN 视频切片、多线程下载工具(如 aria2)的基础。

409 Conflict 和 410 Gone:409 表示”资源存在,但当前状态不允许操作”,比如试图删除一个正在被其他用户编辑的文档。410 表示”资源曾经存在,但已经被永久删除”,比如一个已经下架的商品页面。区分它们是有意义的:410 可以被缓存(资源永久不在了),409 不能(冲突可能随时解除)。

二、实验一:实现支持 Keep-Alive 的 HTTP/1.1 服务器#

现在来实现一个支持 HTTP/1.1 核心特性的服务器,重点展示持久连接和分块传输。

先看三个关键代码片段,理解核心机制后再看完整实现。

片段一:持久连接循环

# 持久连接的核心:循环处理同一连接上的多个请求
while True:
request_count += 1
# 读取请求数据...
# 检查 Connection: close 或 HTTP/1.0 → 决定是否关闭
should_close = connection == 'close' or version == 'HTTP/1.0'
# 处理请求...
if should_close:
return
continue # 继续等待下一个请求

与 HTTP/1.0 服务器不同,这里用 while True 循环持续监听同一连接上的多个请求。通过 conn.settimeout() 设置超时,空闲连接在超时后自动关闭。

片段二:分块响应构建

def build_chunked_response(status_code, status_text, headers, chunks):
# 每个数据块前标注十六进制长度
for chunk in chunks:
chunk_data = chunk if isinstance(chunk, bytes) else chunk.encode('utf-8')
result += f"{len(chunk_data):X}\r\n".encode('iso-8859-1')
result += chunk_data
result += b"\r\n"
result += b"0\r\n\r\n" # 终止块

每个数据块前标注十六进制长度,最后以 0\r\n\r\n 结束。这让服务器在不知道总长度时就能开始传输。

片段三:ETag 条件请求

# 检查 If-None-Match(ETag 验证)
if_none_match = headers.get('if-none-match')
if if_none_match and if_none_match == etag:
response = build_response(304, "Not Modified", response_headers, b'')
# 无响应体,节省带宽

ETag 让客户端在下次请求时带上 If-None-Match,服务器检查资源是否变化。未变化则返回 304(无响应体),节省带宽。

完整服务器实现(http11_server.py)
#!/usr/bin/env python3
# http11_server.py -- HTTP/1.1 server with keep-alive and chunked encoding
import socket
import threading
import os
import time
import hashlib
from datetime import datetime, timezone
HOST = '0.0.0.0'
PORT = 8080
WWW = 'www'
def parse_request(data):
"""解析 HTTP/1.1 请求,返回 (method, path, version, headers, body)"""
try:
if b'\r\n\r\n' in data:
header_part, body = data.split(b'\r\n\r\n', 1)
else:
header_part = data
body = b''
lines = header_part.decode('iso-8859-1').split('\r\n')
if not lines:
return None, None, None, {}, b''
# 解析请求行
request_line = lines[0]
parts = request_line.split()
if len(parts) < 2:
return None, None, None, {}, b''
method = parts[0].upper()
path = parts[1]
version = parts[2] if len(parts) > 2 else 'HTTP/1.0'
# 解析请求头
headers = {}
for line in lines[1:]:
if ': ' in line:
key, value = line.split(': ', 1)
headers[key.lower()] = value
return method, path, version, headers, body
except Exception as e:
print(f"Error parsing request: {e}")
return None, None, None, {}, b''
def build_response(status_code, status_text, headers, body, http_version='HTTP/1.1'):
"""构建 HTTP 响应"""
response = f"{http_version} {status_code} {status_text}\r\n"
for key, value in headers.items():
response += f"{key}: {value}\r\n"
response += "\r\n"
return response.encode('iso-8859-1') + body
def build_chunked_response(status_code, status_text, headers, chunks):
"""构建分块传输响应"""
response = f"HTTP/1.1 {status_code} {status_text}\r\n"
headers['Transfer-Encoding'] = 'chunked'
for key, value in headers.items():
response += f"{key}: {value}\r\n"
response += "\r\n"
result = response.encode('iso-8859-1')
for chunk in chunks:
chunk_data = chunk if isinstance(chunk, bytes) else chunk.encode('utf-8')
result += f"{len(chunk_data):X}\r\n".encode('iso-8859-1')
result += chunk_data
result += b"\r\n"
result += b"0\r\n\r\n"
return result
def get_mime_type(path):
"""根据文件扩展名返回 MIME 类型"""
ext = os.path.splitext(path)[1].lower()
mime_types = {
'.html': 'text/html; charset=utf-8',
'.htm': 'text/html; charset=utf-8',
'.css': 'text/css; charset=utf-8',
'.js': 'application/javascript; charset=utf-8',
'.json': 'application/json; charset=utf-8',
'.png': 'image/png',
'.jpg': 'image/jpeg',
'.jpeg': 'image/jpeg',
'.gif': 'image/gif',
'.ico': 'image/x-icon',
'.txt': 'text/plain; charset=utf-8',
}
return mime_types.get(ext, 'application/octet-stream')
def compute_etag(content):
"""计算 ETag(基于内容哈希)"""
return f'"{hashlib.md5(content).hexdigest()[:16]}"'
def format_http_date(dt):
"""格式化为 HTTP 日期格式"""
return dt.strftime('%a, %d %b %Y %H:%M:%S GMT')
def handle_conn(conn, addr):
"""处理连接(支持持久连接)"""
request_count = 0
connection_timeout = 30 # 持久连接超时时间
try:
conn.settimeout(connection_timeout)
while True:
request_count += 1
print(f"[{addr}] Waiting for request #{request_count}...")
# 读取请求数据
data = b''
while b'\r\n\r\n' not in data:
try:
chunk = conn.recv(4096)
if not chunk:
print(f"[{addr}] Client closed connection")
return
data += chunk
except socket.timeout:
print(f"[{addr}] Connection timeout after {connection_timeout}s")
return
method, path, version, headers, body = parse_request(data)
if method is None:
response = build_response(400, "Bad Request",
{"Content-Type": "text/html", "Connection": "close"},
b"<html><body><h1>400 Bad Request</h1></body></html>")
conn.sendall(response)
return
# 检查是否需要关闭连接
connection = headers.get('connection', '').lower()
should_close = connection == 'close' or version == 'HTTP/1.0'
print(f"[{addr}] {method} {path} {version} (keep-alive: {not should_close})")
# 处理 Host 头(HTTP/1.1 必需)
if version == 'HTTP/1.1' and 'host' not in headers:
response = build_response(400, "Bad Request",
{"Content-Type": "text/html", "Connection": "close"},
b"<html><body><h1>400 Bad Request</h1><p>Host header required.</p></body></html>")
conn.sendall(response)
return
response_headers = {
"Server": "HTTP11-Demo/1.0",
"Date": format_http_date(datetime.now(timezone.utc)),
}
if should_close:
response_headers["Connection"] = "close"
else:
response_headers["Connection"] = "keep-alive"
response_headers["Keep-Alive"] = f"timeout={connection_timeout}"
# 处理路径
if path == '/':
path = '/index.html'
# 演示分块传输的端点
if path == '/chunked':
chunks = [
"<html><head><title>Chunked Demo</title></head><body>\n",
f"<h1>Chunked Transfer Encoding Demo</h1>\n",
f"<p>Generated at: {datetime.now().isoformat()}</p>\n",
"<ul>\n"
]
for i in range(1, 6):
chunks.append(f"<li>Item {i}</li>\n")
time.sleep(0.1) # 模拟生成延迟
chunks.append("</ul>\n</body></html>")
response = build_chunked_response(200, "OK",
{"Content-Type": "text/html; charset=utf-8"},
chunks)
conn.sendall(response)
if should_close:
return
continue
# 演示流式日志端点
if path == '/stream':
chunks = []
for i in range(5):
chunks.append(f"[{datetime.now().isoformat()}] Log entry {i+1}\n")
time.sleep(0.2)
response = build_chunked_response(200, "OK",
{"Content-Type": "text/plain; charset=utf-8"},
chunks)
conn.sendall(response)
if should_close:
return
continue
# 安全处理:防止路径遍历攻击
safe_path = os.path.normpath(path).lstrip(os.sep)
full_path = os.path.join(WWW, safe_path)
# 处理 HEAD 请求
if method == 'HEAD':
if os.path.isfile(full_path):
with open(full_path, 'rb') as f:
content = f.read()
mime = get_mime_type(full_path)
etag = compute_etag(content)
mtime = os.path.getmtime(full_path)
last_modified = format_http_date(datetime.fromtimestamp(mtime, timezone.utc))
response_headers["Content-Type"] = mime
response_headers["Content-Length"] = str(len(content))
response_headers["ETag"] = etag
response_headers["Last-Modified"] = last_modified
response_headers["Cache-Control"] = "max-age=3600"
response = build_response(200, "OK", response_headers, b'')
else:
body = b'<html><body><h1>404 Not Found</h1></body></html>'
response_headers["Content-Type"] = "text/html"
response_headers["Content-Length"] = str(len(body))
response = build_response(404, "Not Found", response_headers, b'')
conn.sendall(response)
if should_close:
return
continue
# 处理 GET 请求(支持条件请求)
if method == 'GET':
if os.path.isfile(full_path):
with open(full_path, 'rb') as f:
content = f.read()
mime = get_mime_type(full_path)
etag = compute_etag(content)
mtime = os.path.getmtime(full_path)
last_modified = format_http_date(datetime.fromtimestamp(mtime, timezone.utc))
# 检查 If-None-Match(ETag 验证)
if_none_match = headers.get('if-none-match')
if if_none_match and if_none_match == etag:
response_headers["ETag"] = etag
response_headers["Cache-Control"] = "max-age=3600"
response = build_response(304, "Not Modified", response_headers, b'')
conn.sendall(response)
print(f"[{addr}] 304 Not Modified (ETag match)")
if should_close:
return
continue
# 检查 If-Modified-Since
if_modified_since = headers.get('if-modified-since')
if if_modified_since:
try:
client_time = datetime.strptime(if_modified_since, '%a, %d %b %Y %H:%M:%S GMT')
file_time = datetime.fromtimestamp(mtime, timezone.utc)
if file_time <= client_time.replace(tzinfo=timezone.utc):
response_headers["ETag"] = etag
response_headers["Last-Modified"] = last_modified
response_headers["Cache-Control"] = "max-age=3600"
response = build_response(304, "Not Modified", response_headers, b'')
conn.sendall(response)
print(f"[{addr}] 304 Not Modified (Last-Modified)")
if should_close:
return
continue
except:
pass
response_headers["Content-Type"] = mime
response_headers["Content-Length"] = str(len(content))
response_headers["ETag"] = etag
response_headers["Last-Modified"] = last_modified
response_headers["Cache-Control"] = "max-age=3600"
response = build_response(200, "OK", response_headers, content)
conn.sendall(response)
else:
body = b'<html><body><h1>404 Not Found</h1><p>The requested resource was not found.</p></body></html>'
response_headers["Content-Type"] = "text/html"
response_headers["Content-Length"] = str(len(body))
response = build_response(404, "Not Found", response_headers, body)
conn.sendall(response)
if should_close:
return
continue
# 处理 OPTIONS 请求
if method == 'OPTIONS':
response_headers["Allow"] = "GET, HEAD, POST, OPTIONS"
response_headers["Access-Control-Allow-Origin"] = "*"
response_headers["Access-Control-Allow-Methods"] = "GET, HEAD, POST, OPTIONS"
response = build_response(200, "OK", response_headers, b'')
conn.sendall(response)
if should_close:
return
continue
# 不支持的方法
body = b'<html><body><h1>501 Not Implemented</h1></body></html>'
response_headers["Content-Type"] = "text/html"
response_headers["Content-Length"] = str(len(body))
response = build_response(501, "Not Implemented", response_headers, body)
conn.sendall(response)
if should_close:
return
except Exception as e:
print(f"Error handling connection: {e}")
import traceback
traceback.print_exc()
finally:
conn.close()
print(f"[{addr}] Connection closed (handled {request_count} requests)")
def main():
os.makedirs(WWW, exist_ok=True)
# 创建测试文件
idx = os.path.join(WWW, 'index.html')
if not os.path.exists(idx):
with open(idx, 'w', encoding='utf-8') as f:
f.write('''<!DOCTYPE html>
<html>
<head><title>HTTP/1.1 Demo</title></head>
<body>
<h1>HTTP/1.1 Demo Server</h1>
<p>This server demonstrates HTTP/1.1 features:</p>
<ul>
<li><a href="/page1.html">Page 1</a> - Normal page</li>
<li><a href="/chunked">Chunked Demo</a> - Chunked transfer encoding</li>
<li><a href="/stream">Stream Demo</a> - Server-sent log stream</li>
<li><a href="/data.json">JSON Data</a> - JSON with caching headers</li>
</ul>
<h2>Test Keep-Alive</h2>
<p>Use nc or curl with --http1.1 to test persistent connections.</p>
<p>Run: <code>curl -v --http1.1 http://localhost:8080/</code></p>
</body>
</html>''')
page1 = os.path.join(WWW, 'page1.html')
if not os.path.exists(page1):
with open(page1, 'w', encoding='utf-8') as f:
f.write('''<!DOCTYPE html>
<html>
<head><title>Page 1</title></head>
<body>
<h1>Page 1</h1>
<p><a href="/">Back to index</a></p>
</body>
</html>''')
json_file = os.path.join(WWW, 'data.json')
if not os.path.exists(json_file):
with open(json_file, 'w', encoding='utf-8') as f:
f.write('{"name": "HTTP/1.1 Demo", "version": "1.1", "features": ["keep-alive", "chunked", "etag", "cache-control"]}')
s = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
s.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
s.bind((HOST, PORT))
s.listen(5)
print(f"HTTP/1.1 server listening on {HOST}:{PORT}")
print(f"Serving files from: {os.path.abspath(WWW)}")
print(f"Features: keep-alive, chunked encoding, ETag, conditional requests")
try:
while True:
conn, addr = s.accept()
threading.Thread(target=handle_conn, args=(conn, addr), daemon=True).start()
except KeyboardInterrupt:
print("\nShutting down.")
finally:
s.close()
if __name__ == '__main__':
main()

三、实验二:用 nc 观察持久连接和分块传输#

启动服务器后,用 nc 观察 HTTP/1.1 的特性。

实验 2.1:持久连接:同一连接发送多个请求

# 使用 nc 发送多个请求(在同一连接中)
printf 'GET /index.html HTTP/1.1\r\nHost: localhost:8080\r\n\r\nGET /data.json HTTP/1.1\r\nHost: localhost:8080\r\n\r\n' | nc localhost 8080

你会看到两个响应依次返回,服务器日志也显示这是同一个连接处理了两个请求。注意 Connection: keep-alive 响应头。

实验 2.2:显式关闭连接

printf 'GET /index.html HTTP/1.1\r\nHost: localhost:8080\r\nConnection: close\r\n\r\n' | nc localhost 8080

这次响应头会包含 Connection: close,服务器处理完这个请求后会关闭连接。

实验 2.3:观察分块传输

printf 'GET /chunked HTTP/1.1\r\nHost: localhost:8080\r\n\r\n' | nc localhost 8080

你会看到类似这样的输出:

Terminal window
HTTP/1.1 200 OK
Server: HTTP11-Demo/1.0
Date: Thu, 20 Mar 2026 10:00:00 GMT
Connection: keep-alive
Keep-Alive: timeout=30
Content-Type: text/html; charset=utf-8
Transfer-Encoding: chunked
35
<html><head><title>Chunked Demo</title></head><body>
28
<h1>Chunked Transfer Encoding Demo</h1>
30
<p>Generated at: 2026-03-20T10:00:00.123456</p>
5
<ul>
10
<li>Item 1</li>
10
<li>Item 2</li>
10
<li>Item 3</li>
10
<li>Item 4</li>
10
<li>Item 5</li>
14
</ul>
</body></html>
0

注意每个数据块前的十六进制数字(如 362b)表示该块的字节数,最后 0 表示传输结束。

实验 2.4:流式数据演示

printf 'GET /stream HTTP/1.1\r\nHost: localhost:8080\r\n\r\n' | nc localhost 8080

你会看到日志条目逐条到达,展示了分块传输对实时数据的支持。

实验 2.5:ETag 缓存验证

请求资源获取 ETag:

curl -v --http1.1 http://localhost:8080/data.json 2>&1 | grep -E '(ETag|<)'

假设获取到的 ETag 是 "abc123...",然后用这个 ETag 发送条件请求:

curl -v --http1.1 -H 'If-None-Match: "你获取的ETag值"' http://localhost:8080/data.json

如果资源未变化,你会看到 304 Not Modified 响应,没有响应体,节省了带宽。

实验 2.6:Host 头缺失测试

# 发送没有 Host 头的 HTTP/1.1 请求
printf 'GET /index.html HTTP/1.1\r\n\r\n' | nc localhost 8080

服务器会返回 400 Bad Request,因为 HTTP/1.1 要求必须有 Host 头。

实验 2.7:用 curl 验证持久连接

# curl 默认使用 HTTP/1.1 并保持持久连接
curl -v http://localhost:8080/index.html http://localhost:8080/data.json
# 观察 "Re-using existing connection" 消息
curl -v http://localhost:8080/index.html http://localhost:8080/page1.html 2>&1 | grep -i connection

四、实验三:性能对比:Keep-Alive 的实际影响#

量化持久连接带来的性能提升。

创建一个测试脚本 benchmark.sh

#!/bin/bash
# benchmark.sh - 测试持久连接 vs 非持久连接
echo "=== HTTP/1.0 风格(每次新建连接)==="
time (
for i in {1..10}; do
printf 'GET /index.html HTTP/1.0\r\n\r\n' | nc localhost 8080 > /dev/null
done
)
echo ""
echo "=== HTTP/1.1 持久连接(复用连接)==="
time (
# 使用 curl 的持久连接特性
curl -s --http1.1 http://localhost:8080/index.html \
--next --http1.1 http://localhost:8080/page1.html \
--next --http1.1 http://localhost:8080/data.json \
--next --http1.1 http://localhost:8080/index.html \
--next --http1.1 http://localhost:8080/page1.html \
--next --http1.1 http://localhost:8080/data.json \
--next --http1.1 http://localhost:8080/index.html \
--next --http1.1 http://localhost:8080/page1.html \
--next --http1.1 http://localhost:8080/data.json \
--next --http1.1 http://localhost:8080/index.html > /dev/null
)

运行测试:

chmod +x benchmark.sh
./benchmark.sh

你会观察到持久连接版本明显更快,尤其是在网络延迟较高的环境中。每次 TCP 连接建立都需要三次握手,而持久连接只需一次握手即可复用。

五、HTTP/1.1 的局限性与 HTTP/2 的动机#

虽然 HTTP/1.1 相比 1.0 有巨大改进,但它仍有不足:

队头阻塞(Head-of-Line Blocking):HTTP/1.1 要求响应必须按请求顺序返回。即使服务器已经准备好后续响应,也必须等前面的响应发送完毕。管道化理论上有帮助,但由于兼容性问题,实际很少启用。

头部冗余:每次请求都要携带完整的头部信息,对于大量小请求(如 API 调用),头部开销占比很高。

优先级缺失:无法告诉服务器「这个资源更重要,请优先传输」。

这些问题催生了 HTTP/2:二进制分帧、多路复用、头部压缩、服务器推送等特性,彻底解决了 HTTP/1.1 的队头阻塞问题。

HTTP/1.1 解决了连接复用,但没有解决请求复用。Keep-Alive 让一个 TCP 连接可以发多个请求,但请求仍然是串行的,响应必须按顺序返回。这个”连接复用但请求串行”的矛盾,就是 HTTP/2 多路复用的出发点。

六、HTTP/1.1 特性速查表#

6.1 影响评估矩阵#

特性解决了什么没有它会怎样常见误区
Keep-AliveTCP 连接重复建立的开销每请求一次握手,延迟累加以为”复用连接=并行请求”,其实还是串行
Pipeline请求无需等待响应即可发送请求仍然串行,只是减少了 RTT几乎被浏览器禁用,因为队头阻塞
Chunked Transfer服务器不知道 Content-Length 时也能传输必须等全部生成才能开始传和”压缩”是两回事,chunked 是分块不是压缩
Host 头同一 IP 部署多个网站每个网站需要独立 IP这是虚拟主机的基础,没有它就没有共享主机
ETag高效的缓存验证只能用时间戳验证,精度低ETag 不等于文件哈希,可以是任何唯一标识
Cache-Control精细的缓存控制只有 Expires,无法表达”验证后可用”等语义no-cache 不是”不缓存”,是”使用前必须验证”
详细参考数据

6.2 持久连接相关头字段#

头字段方向用途示例
Connection: keep-alive请求/响应请求/声明保持连接Connection: keep-alive
Connection: close请求/响应请求/声明关闭连接Connection: close
Keep-Alive响应配置连接超时和最大请求数Keep-Alive: timeout=30, max=100

6.3 分块传输编码格式#

sequenceDiagram participant S as 服务器 participant C as 客户端 Note over S,C: 响应头 S->>C: HTTP/1.1 200 OK\r\n S->>C: Transfer-Encoding: chunked\r\n S->>C: Content-Type: text/plain\r\n S->>C: \r\n Note over S,C: 数据块(每块前标注十六进制长度) S->>C: 7\r\nMozilla\r\n(7 字节) S->>C: 9\r\nDeveloper\r\n(9 字节) Note over S,C: 终止块 S->>C: 0\r\n\r\n(传输结束)

6.4 缓存控制指令速查#

Cache-Control 指令用途示例
max-age=<seconds>缓存有效期(秒)Cache-Control: max-age=3600
no-cache使用前必须验证Cache-Control: no-cache
no-store禁止缓存Cache-Control: no-store
public可被任何缓存存储Cache-Control: public
private仅终端用户缓存Cache-Control: private
must-revalidate过期后必须验证Cache-Control: must-revalidate
no-transform禁止转换(如压缩)Cache-Control: no-transform

6.5 内容协商头字段#

请求头用途示例
Accept可接受的 MIME 类型Accept: text/html, application/json
Accept-Language偏好的自然语言Accept-Language: zh-CN, zh;q=0.9, en;q=0.8
Accept-Encoding支持的压缩算法Accept-Encoding: gzip, deflate, br
Accept-Charset可接受的字符集Accept-Charset: utf-8
Note

q 值表示优先级(0-1),默认为 1。例如 en;q=0.8 表示英文优先级 0.8。

6.6 HTTP/1.1 新增状态码#

状态码含义典型场景
100Continue客户端可继续发送请求体
206Partial Content范围请求成功
409Conflict资源状态冲突
410Gone资源已永久删除
413Payload Too Large请求体超过服务器限制
414URI Too LongURL 过长
415Unsupported Media Type不支持的 Content-Type
417Expectation FailedExpect 头无法满足

:::

6.7 HTTP/1.1 vs HTTP/2 关键差异预览#

特性HTTP/1.1HTTP/2
传输格式文本协议二进制分帧
多路复用队头阻塞独立流,无队头阻塞
头部压缩每次完整传输HPACK 压缩
服务器推送不支持Server Push
优先级不支持流优先级
连接复用Keep-Alive(串行)多路复用(并行)

这些差异正是下一篇 HTTP/2 的核心内容。


参考资料#

  • RFC 2068 - HTTP/1.1 协议规范(初版)
  • RFC 2616 - HTTP/1.1 协议规范
  • RFC 7230 - HTTP/1.1 消息语法与路由

支持与分享

如果这篇文章对你有帮助,欢迎支持作者或分享给更多人

HTTP/1.1:持久连接
https://blog.souloss.cn/posts/web/http/http-1-1/
作者
Souloss
发布于
2024-04-26
许可协议
CC BY-NC-SA 4.0

部分信息可能已经过时