pprof是定位go微服务性能问题的必备工具,需正确挂载路由、强制gc采样、使用debug=2获取全量goroutine,并用-base对比净增长才能精准识别内存泄漏和goroutine泄漏。

直接上 pprof,别猜。Go 微服务跑得慢,90% 的问题靠肉眼或日志根本定位不到,必须用真实采样数据说话。
pprof 接口返回 404 怎么办
常见错误现象:import _ "net/http/pprof" 加了,http.ListenAndServe(":6060", nil) 也启了,但访问 /debug/pprof/ 仍是 404。
根本原因:Gin 默认接管全部路由,nil mux 注册的 /debug/pprof/* 前缀被框架拦截丢弃。
- Gin 用户必须显式挂载:
r.GET("/debug/pprof/*pprof", gin.WrapH(http.DefaultServeMux)) - Echo 用户用:
e.Any("/debug/pprof/*", echo.WrapHandler(http.DefaultServeMux)) - 生产环境更稳妥的做法:改用
runtime/pprof手动写文件,比如收到SIGUSR1时调用pprof.WriteHeapProfile(f),避免暴露 HTTP 端点
heap profile 看不出内存泄漏?
常见错误现象:采了 /debug/pprof/heap,发现 json.Marshal 占比高,优化后 RSS 还是持续上涨。
关键遗漏:默认 heap profile 只反映 inuse_space(当前存活对象),而泄漏往往藏在 allocs(累计分配量)或两次快照的净增长里。
- 必须强制 GC 后再采样:
curl -s "http://localhost:6060/debug/pprof/heap?gc=1" > heap1.pb.gz - 压测一段时间后再采一次:
curl -s "http://localhost:6060/debug/pprof/heap?gc=1" > heap2.pb.gz - 对比净增长:
go tool pprof -base heap1.pb.gz heap2.pb.gz,重点关注 delta 值大的函数
goroutine top 显示数量很少,但 RSS 持续飙升
常见错误现象:ps aux 看 RSS 一路涨,go tool pprof http://.../goroutine 的 top 却只列出几十个 goroutine。
真相:/debug/pprof/goroutine 默认只返回状态为 running 或 runnable 的 goroutine;大量泄漏的 goroutine 实际卡在 chan receive、select 或 semacquire 上,被过滤掉了。
- 要查全量栈,加参数:
curl "http://localhost:6060/debug/pprof/goroutine?debug=2" - 配合
go tool pprof的top -cum查调用链,重点看是否大量 goroutine 堆在某个 channel recv 或 lock 上 - 典型诱因:
select缺少 default 分支、未关闭的 channel、忘记close()的 worker pool、或context.WithTimeout超时后未清理资源
真正难的是把采样时机、参数组合和结果解读串起来——比如 ?gc=1 和 ?debug=2 不是可选项,是必填项;而 go tool pprof -base 对比的不是“有没有泄漏”,而是“哪段代码在持续净增对象”。漏掉任一环,profile 就只是好看的火焰图。











