要实现状态码与路径维度的秒级时序监控,需启用nginx-vts-module的location_zone统计、配置nginx-vts-exporter以1s间隔抓取/json接口、prometheus设scrape_interval:1s并重写code/location/host标签、grafana用rate()和histogram_quantile()做秒级聚合分析。

要利用 Prometheus 的 nginx-vts-exporter 实现状态码与路径维度的秒级时序监控,核心在于正确配置 Nginx 的 nginx-vts-module、暴露细粒度指标、合理配置 exporter 抓取频率,并在 Prometheus 中通过标签组合做多维聚合分析。
确保 Nginx 启用 vts 模块并开启详细统计
nginx-vts-module 需编译进 Nginx 或以动态模块加载,且必须启用 status 指令中的 upstream、server_zone 和 location_zone —— 尤其是 location_zone,它按 location 块(即路径维度)统计请求,是路径分组的关键。
- 在 Nginx 配置中为每个需监控的路径定义独立
location块,并启用vhost_traffic_status_zone;例如:location /api/users {<br> vhost_traffic_status_zone users_api;<br> proxy_pass http://backend;<br>} - 全局启用
vhost_traffic_status,并设置vhost_traffic_status_display_format html(调试用)或json(生产推荐); - 确认
nginx -V 2>&1 | grep -o with-http-vhost-traffic-status-module输出存在,表示模块已加载。
部署 nginx-vts-exporter 并对齐采集粒度
nginx-vts-exporter 默认每 5 秒拉取一次 Nginx 的 /status/format/json 接口,但秒级监控要求更短周期。需在启动参数中显式指定 -nginx.scrape-uri 和 -web.listen-address,并通过 -nginx.scrape-timeout 和 -web.telemetry-path 优化稳定性。
- 使用
--nginx.scrape-interval=1s(部分版本支持,如 v0.10+;若不支持,改用 systemd timer 或 cron 每秒调用 curl + textfile collector 间接实现); - 确保 exporter 能访问 Nginx 的 status 接口(如
http://127.0.0.1:8080/status/format/json),该接口需开放给 exporter 所在主机; - 验证 exporter 是否输出含
nginx_vts_server_request_seconds_total及带code、host、server、location标签的指标 —— 路径维度依赖location标签,状态码来自code标签。
Prometheus 配置抓取与标签重写
Prometheus 抓取间隔必须 ≤ 1s 才能支撑秒级观测,同时需通过 metric_relabel_configs 提炼关键维度,避免高基数问题。
- 在
scrape_config中设scrape_interval: 1s,并加scrape_timeout: 500ms防超时干扰; - 用
metric_relabel_configs保留必要标签:code(状态码)、location(路径)、host(虚拟主机),丢弃无意义标签如server(若未按 server 分区); - 若
location值含动态参数(如/user/123),建议在 Nginx 中用map指令归一化为模板(如/user/{id}),再由 exporter 暴露,否则会导致标签爆炸。
Grafana 查询与秒级聚合示例
在 Prometheus 中,秒级数据需用 rate() 计算单位时间请求数,结合 by (code, location) 实现多维下钻。
- 查询某路径每秒 5xx 错误率:
rate(nginx_vts_server_requests_total{code=~"5..", location="/api/orders"}[1s]) - 对比多个路径的响应延迟 P95:
histogram_quantile(0.95, sum(rate(nginx_vts_server_request_seconds_bucket{location=~"/api/.*"}[1s])) by (le, location)) - 设置告警规则时,避免直接对原始计数器告警,应基于
rate(...[1s])或increase(...[1s]),例如:1 秒内 500 错误突增 > 10 次即触发。
不复杂但容易忽略的是 Nginx location zone 的定义粒度和 exporter 的 scrape interval 对齐 —— 这两者决定你能否真正拿到“路径 + 状态码”在秒级上的真实分布。











