最轻量可靠的nginx错误告警方案是用stdbuf+while read实时捕获日志,结合指纹去重与时间窗口限频,通过curl调用企微机器人api发送结构化消息,并由systemd托管实现自愈运行。

直接用 tail -f 配合 grep 捕获错误日志,再调用企业微信机器人 API 发送告警,是最轻量、最可靠的方式。关键不在“守护进程”本身,而在于日志捕获不丢行、告警不重复、脚本不退出、异常能自愈。
实时捕获错误日志(避免漏行和阻塞)
别用 tail -f | grep 简单管道——它可能因缓冲或子进程退出导致漏报。推荐用 stdbuf 控制缓冲,并用 while read 逐行处理:
-
stdbuf -oL -eL tail -n0 -f /var/log/nginx/error.log:启用行缓冲,从末尾开始实时追加 - 配合
while IFS= read -r line; do ... done安全读取每一行(支持空格、特殊字符) - 用
[[ "$line" =~ "error"|"crit"|"alert"|"emerg" ]]匹配关键等级,比单纯 grep 更可控
去重与限频(防止刷屏告警)
同一类错误高频出现时,连续发 10 条企微消息毫无意义。需做简单指纹+时间窗口控制:
- 提取错误核心片段,如
$(echo "$line" | sed -E 's/^[^]]*\] (.*)$/\1/' | cut -d',' -f1 | md5sum | cut -c1-8) - 用临时文件记录最近 5 分钟内已发过的指纹(
/tmp/nginx_alert_fingerprints),每次发送前查重 - 超过阈值(如 3 次/10 分钟)才触发,或降级为“过去 10 分钟共出现 X 次”汇总通知
调用企微机器人稳定发消息
企业微信机器人 Webhook 地址带 token,发 JSON 即可。注意三点:
- 用
curl -s -X POST -H 'Content-Type: application/json' --data-binary @- $WEBHOOK_URL发送,避免 shell 字符转义问题 - 消息体包含
msgtype: "text"和text.content,内容里带上时间、服务器 hostname、错误摘要,方便快速定位 - 加
|| echo "企微发送失败: $?" >> /var/log/nginx/alert.log记录失败,便于排查网络或 token 过期问题
作为后台服务长期运行(不依赖 nohup/screen)
用 systemd 管理最稳妥,避免挂掉后无人知晓:
- 写一个
/etc/systemd/system/nginx-error-alert.service,Type=simple,Restart=always -
ExecStart=/usr/local/bin/nginx-alert.sh,确保脚本有执行权限且路径绝对 systemctl daemon-reload && systemctl enable nginx-error-alert && systemctl start nginx-error-alert- 用
journalctl -u nginx-error-alert -f查日志,比自己写 logrotate 更省心
不复杂但容易忽略细节。核心就是:稳读日志、聪明去重、可靠发信、交给 systemd 看着它跑。











