直接用prometheus.newgaugevec会panic,因未注册到defaultregisterer或自定义registry未传给handler;必须显式调用mustregister,多实例需隔离registry,重复注册同名指标亦panic。

为什么直接用 prometheus.NewGaugeVec 会 panic?
因为没注册到默认的 prometheus.DefaultRegisterer,或者注册时用了非默认 registry 却没传给 http.Handler。常见错误是创建了指标但忘了 prometheus.MustRegister(),或者在多实例场景下误用全局 registry 导致冲突。
- 必须显式注册:创建完
prometheus.NewGaugeVec后,立刻调用prometheus.MustRegister(vec) - 若自定义 registry(比如测试或隔离场景),要确保
promhttp.HandlerFor(registry, promhttp.HandlerOpts{})用的是同一个 registry - 避免重复注册同名指标——
MustRegister遇到重名直接 panic,可用registry.MustRegister()替代并捕获 error
如何让 Go 服务暴露 /metrics 端点且不阻塞主逻辑?
别把 metrics handler 塞进主 HTTP mux 的根路径,也别用 http.ListenAndServe 单独起一个 goroutine 而不加超时控制。最稳妥做法是复用主服务的 http.ServeMux,但监听独立端口(如 :9091)更安全,避免业务路由干扰。
- 推荐开独立监听:启动一个
http.Server专门跑promhttp.Handler(),监听:9091 - 加 graceful shutdown:用
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)控制关闭等待时间 - 不要用
log.Fatal(http.ListenAndServe(...))—— 它会 kill 整个进程,应单独 handle error 并返回
prometheus.Collector 接口实现时最容易漏掉什么?
很多人只实现了 Describe 和 Collect,但忘了 Describe 必须返回 *所有可能被 Collect 发出的指标描述*,且两次调用返回的 chan *Desc 必须一致。否则 Prometheus 抓取时会报 metric family not consistent 错误。
-
Describe中的Desc必须和Collect中ch 的字段严格匹配(name、help、constLabels、variableLabels) - 如果指标带 label,
Describe里要用prometheus.NewDesc(..., nil, []string{"env", "region"}, nil)显式声明 label 名 - 不要在
Collect里动态构造新Desc—— 所有 desc 应在Describe中一次性确定
告警规则写在 alert.rules.yml 但 Prometheus 不加载?
不是文件路径写错就是 prometheus.yml 里没配 rule_files,或者 YAML 格式有隐藏空格/缩进错误。Prometheus 启动时不会报错,只会静默跳过无效 rule 文件。
- 确认
prometheus.yml中rule_files是绝对路径,或相对于配置文件所在目录的相对路径 - 用
promtool check rules alert.rules.yml验证语法,它比 Prometheus 日志更早暴露问题 - 告警规则里的
expr如果引用自定义指标,确保该指标名与 Go 代码中NewGaugeVec的name参数完全一致(包括下划线/大小写)
实际部署时,prometheus.Collector 的 Collect 方法执行耗时超过 10s 就会导致 scrape timeout;而告警触发后,alertmanager 的 route 配置若没设 group_by,相同告警可能被拆成几十条发出去——这些细节不在文档首页,但线上一出问题就卡在这里。
golang免费学习笔记(深入):立即使用
在学习笔记中,你将探索golang的核心概念和高级技巧!











