/metrics里看不到go_goroutines等默认指标,是因为promhttp.handler()虽提供采集逻辑,但需显式调用prometheus.mustregister(prometheus.newgocollector())和prometheus.mustregister(prometheus.newprocesscollector(...))才能注册运行时与进程指标;若用自定义注册器,还须用promhttp.handlerfor()绑定。

为什么/metrics里看不到go_goroutines或process_cpu_seconds_total
默认情况下,promhttp.Handler() 确实会暴露 Go 运行时和进程指标,但前提是:你得让它们被注册进去。官方 client 库不会自动注册——它只提供收集器(prometheus.NewGoCollector()、prometheus.NewProcessCollector()),不调 MustRegister 就等于没“上线”。
常见错误是只写了 http.Handle("/metrics", promhttp.Handler()),却漏了这两行:
prometheus.MustRegister(prometheus.NewGoCollector())prometheus.MustRegister(prometheus.NewProcessCollector(prometheus.ProcessCollectorOpts{}))
如果你用了自定义注册器(比如 reg := prometheus.NewRegistry()),那必须用 promhttp.HandlerFor(reg, ...),且把 collector 显式注册到 reg 上,否则 /metrics 返回空。
用 go-commons 一行启动系统指标采集
不想手动注册一堆 collector?go-commons 就是为此设计的。它把 go_goroutines、process_resident_memory_bytes、runtime_gc_pause_ns_sum 等常用指标打包进 metrics.StartDefaultMetricsServer(),开箱即用。
关键点:
- 端口独立:它默认监听
:9090,和业务服务端口(如:8080)分离,避免互相干扰 - 不侵入业务路由:不用在你的
http.ServeMux或 Gin/Echo 路由里挂/metrics,它自己起一个新 server - 自动刷新:内存、CPU、GC 等指标每 5 秒采一次,无需自己写 ticker 或 goroutine
示例代码里这行就足够:metrics.StartDefaultMetricsServer(":9090") —— 启动后直接访问 http://localhost:9090/metrics 就能看到完整系统指标。
自定义指标和系统指标混用时的注册冲突
最容易踩的坑:在多个包里 import 同一个 metrics 初始化文件,导致重复调用 prometheus.MustRegister(),panic 报错 duplicate metrics collector registration。
原因在于:Go 的 init() 函数按包加载顺序执行,如果 A 包和 B 包都 import 了 C 包(C 包里有 init(){ MustRegister(...) }),C 就会被初始化两次。
解决办法只有两个:
- 把所有
MustRegister集中在main包的init()或main()开头,其他包只定义指标变量 - 改用
promauto.NewCounter()系列(来自github.com/prometheus/client_golang/prometheus/promauto),它会在首次调用时自动注册,且线程安全,天然规避重复注册
注意:promauto 默认用的是全局注册器,若你已用自定义注册器,得传入:promauto.With(reg).NewCounter(...)。
直方图 buckets 设置不合理导致延迟监控失真
比如你用 prometheus.NewHistogramVec 记录 HTTP 延迟,但 Buckets 写成 []float64{0.1, 0.2, 0.5},而实际 P99 是 1.2 秒——那所有 >0.5s 的请求都会被塞进最后一个 bucket,根本看不出分布细节。
正确做法:
- 先用
expvar或日志抽样观察真实延迟分布,再定 bucket 边界 - 优先用
prometheus.ExponentialBuckets(0.001, 2, 12)这类函数生成合理序列(覆盖 1ms 到 ~2s) - 别为了“看着整齐”硬凑整数,比如
[0.01, 0.1, 1, 10]在微服务场景下跨度太大,中间断层严重
另外,Histogram 和 Summary 不要混用:前者适合 Prometheus pull 场景(服务端聚合),后者适合客户端分位数计算(如 gRPC 的 stats handler),选错会导致查询结果偏差明显。
golang免费学习笔记(深入):立即使用
在学习笔记中,你将探索golang的核心概念和高级技巧!











