exporter内核态卡死会导致prometheus抓取超时,targets显示context deadline exceeded或up但无新数据;需通过curl测试/metrics响应、strace跟踪系统调用、ps查看d状态线程及wchan定位阻塞点,并临时禁用filesystem收集器或过滤挂载点缓解。

Exporter在采集底层系统指标时卡死在内核态,会导致Prometheus抓取请求长时间无响应,最终超时失败,Targets页面显示 context deadline exceeded 或持续处于 UP但无新数据 状态。这类故障隐蔽性强——进程仍在运行、端口可连、HTTP服务看似正常,但实际指标生成已停滞。排查需绕过应用层日志,深入到系统调用与内核行为层面。
确认Exporter是否真正在“假活”状态
不能只看进程存在或端口监听,要验证其是否能实际生成指标:
- 直接访问
http://IP:9182/metrics(Windows Exporter)或http://IP:9100/metrics(Node Exporter),观察响应时间与内容完整性;若页面加载超过5秒、返回截断内容、或反复卡在某一行(如卡在node_filesystem_...指标处),即为典型内核卡死迹象 - 执行
curl -v http://IP:9100/metrics 2>&1 | head -n 20,检查是否卡在* Connected to...后无后续,说明请求已进入内核但未返回 - 对比
ps aux | grep exporter中的 CPU% 和 RSS:若 CPU% 接近 0 但 RSS 持续缓慢上涨,常表明线程阻塞在不可中断睡眠(D状态)
检查是否存在 stuck mount 或 slow I/O 导致 filesystem collector 卡住
这是最常见诱因,尤其在挂载 NFS、CIFS、加密卷或异常设备后:
- 运行
findmnt或cat /proc/mounts,识别非常规挂载类型(如nfs4、cifs、fuse.sshfs、zfs)及非标准路径(如/mnt/backup、/data/archive) - 手动触发 statfs 测试:
strace -e trace=statfs,openat,readlink -f timeout 10 ./node_exporter --collector.filesystem.ignored-mount-points="^$" --no-collector.time,观察是否卡在某个statfs("/path", ...)调用上 - 检查对应挂载点内核状态:
cat /proc/self/mountstats | grep -A 10 "device_name",关注age:是否异常高,或出现inodes: stale类提示
定位卡在内核态的具体系统调用
需要借助 strace 和 ps 结合判断线程真实状态:
- 查出 exporter 主进程 PID:
pgrep -f 'node_exporter\|windows_exporter' - 查看线程状态:
ps -T -o pid,tid,comm,state,wchan:20 -p $PID,重点关注状态为D(uninterruptible sleep)的线程,及其wchan(等待的内核函数,如nfs_wait_bit_killable、__fuse_read、sleep_on_page) - 对可疑线程做实时跟踪:
strace -p $TID -e trace=%all -s 128 -yy 2>&1 | grep -E "(statfs|openat|readlink|getdents|ioctl)",确认是否循环阻塞在同一路径或设备上
临时规避与长期修复建议
内核卡死无法从用户态强制唤醒,只能隔离或绕过问题源:
- 立即缓解:通过启动参数跳过问题挂载点,例如
--collector.filesystem.ignored-mount-points "^/mnt/broken|^/backup"(Node Exporter)或--collectors.enabled="...,-filesystem"临时禁用整个模块 - Windows 场景下重点检查 WMI 查询:若
windows_exporter卡住,启用详细日志--log.level=debug,观察是否卡在Win32_PerfFormattedData_PerfOS_System等 WMI 类查询;可改用--collector.logical_disk.ignored-mount-points或关闭logical_disk收集器 - 长期方案:避免在生产环境挂载不稳定的远程文件系统;对关键挂载点设置
soft,nointr,timeo=10,retrans=3(NFS)等容错参数;定期巡检/proc/mounts与dmesg | tail -20中的 I/O 错误










