答案是:通过对比 $request_time 与 $upstream_response_time 等字段可精准定位性能瓶颈所在环节。$request_time 表示 nginx 处理请求总耗时(秒,毫秒级),含接收请求、处理、发送响应全过程;$upstream_response_time 表示 nginx 与后端交互耗时,两者差值大说明瓶颈在 nginx 本体,接近则问题在后端服务。

直接看 Nginx 日志里的 $request_time 和 $upstream_response_time,就能定位慢在哪一环——是 Nginx 自身处理慢,还是后端服务拖了后腿。
从日志中提取关键时间字段
Nginx 默认不记录详细耗时,需先在 http 或 server 块中定义带时间变量的 log_format:
-
$request_time:整个请求从接收首字节到发送完响应的总时间(单位:秒,精度毫秒),反映客户端感知的 RT -
$upstream_response_time:Nginx 向后端发起请求,到收到完整响应所花的时间(可能含多个 upstream 节点,用逗号分隔) -
$upstream_connect_time:与后端建立 TCP 连接的耗时,高值说明网络或后端连接池不足 -
$upstream_header_time:从发完请求到收到后端响应头的时间,可判断后端业务逻辑是否卡顿
示例配置:
$request_time $upstream_response_time $upstream_connect_time $upstream_header_time';
access_log /var/log/nginx/access.log timing;
识别三类典型延迟模式
分析日志时,重点关注这三种分布特征:
-
request_time 高,但 upstream_response_time 很低:问题在 Nginx 本体。常见于大文件传输未启用
sendfile on、SSL 握手开销大、或日志写入阻塞(建议异步写日志或关闭 access_log) -
upstream_response_time 明显高于 request_time:实际不可能,说明日志格式有误或变量被覆盖;正常应满足
upstream_response_time ≤ request_time。若出现反常,检查是否漏配proxy_buffering off导致响应流式阻塞 -
upstream_connect_time 突增或频繁超 100ms:后端连接池耗尽、DNS 解析慢、或未启用 keepalive 连接复用。此时应配置
keepalive 32和keepalive_requests 1000到 upstream 块中
针对性调优动作
根据日志分析结果,快速落地几项关键调整:
- 若平均
request_time> 500ms 且波动大:启用open_file_cache缓存静态文件句柄,减少系统调用 - 若大量请求
upstream_response_time在 800–1200ms 区间聚集:说明后端存在固定瓶颈(如数据库慢查询),Nginx 层可加proxy_read_timeout 10s防止长连接占用,并配合健康检查自动剔除异常节点 - 若
upstream_connect_time中位数 > 50ms:在 upstream 块中添加keepalive 64,并在 location 中配置proxy_http_version 1.1和proxy_set_header Connection '',确保复用连接 - 若小文件(request_time 普遍偏高:开启
tcp_nopush on和tcp_nodelay on平衡包合并与延迟
辅助验证与持续监控
单靠日志抽样不够,要结合实时指标交叉验证:
- 开启
ngx_http_stub_status_module,访问/nginx_status查看活跃连接与请求速率,确认是否因连接堆积导致排队延迟 - 用
awk '{sum+=$4} END {print "avg:", sum/NR}' /var/log/nginx/access.log快速算出平均request_time($4 对应日志中$request_time字段位置) - 部署 Prometheus + Grafana,通过
nginx-vts-exporter或nginx-lua-prometheus抓取分位数指标(p95/p99 request_time),比平均值更能暴露长尾问题











