nginx 本身不支持直接通过 error_log 记录后端节点健康状态变更,但可利用健康检查模块自动输出的“is up/down”日志行,配合 notice 级别配置、独立日志文件及外部工具解析,实现可观测监控。

Nginx 本身不支持直接通过 error_log 记录后端节点健康状态变更(比如“node 10.0.1.10:8080 is down”),但你可以利用 error_log 中由健康检查模块自动输出的状态提示行,配合自定义日志格式、过滤与解析,实现对节点上下线事件的可观测监控。关键不是重写 error_log 格式,而是让 error_log 承载有意义的健康事件,再通过结构化方式提取。
配置 error_log 输出带时间戳和节点标识的健康变更日志
Nginx 开源版默认 error_log 不含上游节点信息,但 nginx-upstream-check-module(或 OpenResty 的 lua-resty-upstream-healthcheck)会在节点状态切换时,自动向 error_log 写入类似这样的行:
2026/07/21 11:45:32 [notice] 12345#12345: *12345 upstream check: node 10.0.1.10:8080 is down 2026/07/21 11:46:01 [notice] 12345#12345: *12346 upstream check: node 10.0.1.10:8080 is up
要让这类日志可读、可采集、可告警,需三步:
- 确保
error_log级别设为notice或更高(warn/error会漏掉is up/down提示) - 统一使用带毫秒精度的时间戳和明确标签,便于后续解析
- 将该日志单独定向,避免被普通错误淹没
示例配置:
# 在 http 块顶部统一设置
error_log /var/log/nginx/health_events.log notice;
# 同时确保 upstream 块启用 check 并禁用被动机制
upstream app_backend {
server 10.0.1.10:8080 max_fails=0 fail_timeout=0;
server 10.0.1.11:8080 max_fails=0 fail_timeout=0;
check interval=3 rise=2 fall=3 timeout=1 type=http;
check_http_send "GET /health HTTP/1.1\r\nHost: api.example.com\r\nConnection: close\r\n\r\n";
check_http_expect_alive http_2xx;
}
✅ 注意:
max_fails=0 fail_timeout=0是必须的——否则被动失败计数会干扰主动检查结果;default_down=true可选,用于防止启动时误判。
提取并结构化健康事件日志
error_log 是文本流,不是结构化日志。要监控,得靠外部工具做轻量解析。常用做法:
- 使用
grep+awk抽取关键字段(IP、port、up/down、时间) - 用 Filebeat 或 Fluent Bit 过滤匹配
is up\|is down行,打标后发往 ELK 或 Loki - 编写简单 Python 脚本轮询日志文件,检测状态翻转后触发钉钉/企业微信推送
例如,一条可解析的典型日志行:
2026/07/21 11:45:32 [notice] 12345#12345: *12345 upstream check: node 10.0.1.10:8080 is down
可用正则提取:
- 时间:
^(\d{4}/\d{2}/\d{2} \d{2}:\d{2}:\d{2}) - 节点地址:
node ([\d.]+:\d+) - 状态:
(is up|is down)
补充:用 log_by_lua_block 增强错误上下文(OpenResty 场景)
如果你用的是 OpenResty,可在 log_by_lua_block 中捕获 $upstream_addr 和 $upstream_status,当响应异常时,主动写入带节点标识的定制 error 日志:
log_by_lua_block {
if tonumber(ngx.var.upstream_status) >= 500 then
ngx.log(ngx.NOTICE, "backend_fail: ",
"addr=", ngx.var.upstream_addr,
" status=", ngx.var.upstream_status,
" req_time=", ngx.var.request_time)
end
}
这样就把请求级失败与具体节点绑定,比单纯依赖 check 模块更细粒度。
不复杂但容易忽略











