诊断nginx慢请求需拆解耗时来源:$request_time(总耗时)、$upstream_response_time(后端响应)、$upstream_connect_time(建连耗时)等字段缺一不可;通过自定义日志+awk分析,结合三类归因(后端、网络、nginx/客户端)定位瓶颈,并用p95/p99替代均值评估长尾延迟。

诊断 Nginx 中响应时间相关的慢请求,关键不是只看总耗时大不大,而是拆开看时间花在哪——是后端处理慢、网络卡顿、还是 Nginx 自身或客户端拖了后腿。
确保日志能反映真实耗时
默认日志格式(如 combined)不记录任何时间字段,无法分析慢请求。必须启用含时间变量的自定义格式:
- $request_time:从收到第一个字节到发完最后一个字节的总时间(含 Nginx 处理、网络、后端等待)
- $upstream_response_time:Nginx 连上后端后,到收到第一个字节的时间(即纯后端响应耗时)
- 强烈建议补充:$upstream_connect_time(建连耗时)、$upstream_addr(具体后端地址),便于归因
示例配置:
access_log /var/log/nginx/slow.log slowlog;
用命令快速筛出异常模式
不用导入系统,几条 awk 就能定位问题方向:
- 查所有超 1 秒的请求:awk '$NF > 1' /var/log/nginx/slow.log(假设 $request_time 是最后一列)
- 看哪个后端最常超时:awk '$8 > 1 {print $9}' /var/log/nginx/slow.log | sort | uniq -c | sort -nr($8 是 upstream_response_time,$9 是 upstream_addr)
- 对比总时间和后端时间:若 $request_time 比 $upstream_response_time 高出明显(如 >300ms),瓶颈可能在 Nginx 侧(如 SSL 握手、gzip 压缩、大文件传输)或客户端弱网
按三类典型原因归因验证
光有数字不够,要结合字段组合判断责任边界:
- 后端服务慢:$upstream_response_time 高,且与 $request_time 接近 → 查对应服务日志、DB 慢查询、CPU/内存使用率
- 网络或连接问题:$upstream_connect_time 显著偏高(如 >500ms),但 $upstream_response_time 正常 → 检查 DNS 解析、后端连接池是否耗尽、网络丢包
- Nginx 或客户端侧慢:$request_time 高而 $upstream_response_time 很低(如 0.002s),但总耗时 2s+ → 考虑 client_max_body_size 限制上传、proxy_buffering off 导致流式阻塞、或客户端网络极差
避免平均值误导,关注长尾延迟
平均响应时间容易掩盖少数慢请求。更有效的方式是看分位值:
- 用 Prometheus + nginx-vts-exporter 收集带标签指标,例如:nginx_upstream_response_time_seconds{upstream="svc-order", status="200"}
- 计算 P95/P99 延迟,比均值更能反映真实用户体验
- 同一 upstream 下,对比不同 path 的 P99(如 /order/create vs /order/list),可判断是接口级问题而非服务整体退化











