nginx可通过拦截错误、统一响应、状态暴露三步实现后端故障统一通告:启用proxy_intercept_errors接管502/503/504并返回定制页面或json;通过/upstream_check_module或lua暴露健康状态;结合日志与脚本/prometheus实现告警闭环。

Nginx 本身不主动广播故障,但可以通过 拦截错误 + 统一响应 + 状态暴露 三步,实现对后端集群故障的“统一通告”——即用户看到一致提示、运维能快速定位、系统可联动告警。
启用 proxy_intercept_errors 实现前端兜底
这是统一通告最直接的一环:让所有上游故障(502/503/504)不再透传原始错误,而是由 Nginx 主动接管并返回标准化页面或 JSON。
- 在代理 location 中开启:
proxy_intercept_errors on; error_page 502 503 504 =200 /_down.html;
-
/_down.html需通过独立 location 提供,例如:location = /_down.html { root /usr/share/nginx/html; internal; # 仅限内部跳转,禁止外部直连 } - 若需 API 场景返回 JSON,可改用:
error_page 502 503 504 =200 /_down.json; location = /_down.json { add_header Content-Type "application/json; charset=utf-8"; return 200 '{"status":"unavailable","message":"服务暂时不可用,请稍后再试"}'; }
暴露集群实时健康状态供外部感知
统一通告不只是给用户看,更要让监控系统、值班人员、自动化脚本能“一眼看清谁挂了”。
-
若使用
nginx_upstream_check_module(推荐),添加状态页:location /status { check_status; allow 127.0.0.1; allow 10.0.0.0/8; deny all; }访问
/status可得 HTML 或 JSON 格式结果,含每个节点的up/down、失败次数、最后检查时间等。 -
若用 OpenResty,可用 Lua 动态聚合多 upstream 状态,输出带时间戳的摘要:
content_by_lua_block { local status = { version = "1.0", updated = os.time() } -- 这里读取 shared_dict 或调用 healthcheck 模块 ngx.say(cjson.encode(status)) }
关联故障触发与日志告警闭环
真正的“通告”要延伸到人和系统。当某节点因 max_fails 被标记为 down,应立刻留下可追踪痕迹。
-
在 error_log 中记录明确信号:
error_log /var/log/nginx/health.log warn; log_format health '$time_iso8601 $upstream_addr $upstream_status $upstream_response_time'; access_log /var/log/nginx/health.log health;
-
用轻量脚本每分钟扫描:
# 检查最近 60 秒内是否有节点被连续踢出 awk -v t=$(date -d '1 minute ago' +%s) \ '$1 > t && /502|503|504/ && /192\.168\.1\.[0-9]+:8080/ {c++} END{exit c<p>返回非 0 即触发告警(如发钉钉、写 Prometheus Pushgateway)。</p> 更可靠的做法是配合
nginx-plus的/status接口,用 Prometheus 抓取nginxplus_upstream_server_state{state="unavailable"},设置count by (upstream, server) > 0告警。
不复杂但容易忽略











