关键在于爬虫主动暴露指标:从首次请求即埋点,用prometheus_client按status分success/http_error/parse_error三类统计,长时进程用start_http_server,短时任务须走pushgateway,并在prometheus配置中启用honor_labels防止标签覆盖。

直接用 Prometheus 监控爬虫成功率,关键不是“搭个服务”,而是让爬虫进程主动暴露指标,并确保这些指标能真实反映业务逻辑里的“成功”定义——比如 HTTP 状态码 200 且解析出有效数据,而不是单纯发出去就算成功。
如何在爬虫代码里暴露 prometheus_client 指标
别等整个爬虫写完再加监控。从第一个请求开始就该埋点。用官方 prometheus_client 库最稳妥,它支持多进程(需配合 MultiProcessCollector)和异步(AsyncGauge 等)。
常见错误是只统计“请求次数”,却没区分失败类型。建议至少拆成三类:
-
scrape_total{status="success"}:页面下载 + 解析都 OK -
scrape_total{status="http_error"}:如 404、503 -
scrape_total{status="parse_error"}:HTTP 200 但 XPath/JSONPath 找不到关键字段
示例片段:
from prometheus_client import Counter, start_http_server
import time
<p>scrape_counter = Counter('scrape_total', 'Total scrape attempts', ['status'])</p><p>def fetch_and_parse(url):
try:
resp = requests.get(url, timeout=10)
if resp.status_code == 200:
data = parse_content(resp.text) # 你自己的解析逻辑
if data:
scrape_counter.labels(status='success').inc()
return data
else:
scrape_counter.labels(status='parse_error').inc()
else:
scrape_counter.labels(status='http_error').inc()
except requests.RequestException:
scrape_counter.labels(status='http_error').inc()</p>
为什么不能只依赖 metrics_path 和静态配置?
很多人把爬虫当普通 HTTP 服务配进 Prometheus,用 static_configs + metrics_path: /metrics,结果发现指标永远是空的——因为 prometheus_client 默认不自动启动 HTTP server,也不会自动注册所有指标。
必须显式调用 start_http_server(8000),且这个端口要和 Prometheus 的 scrape_config 对齐。更关键的是:如果爬虫是短生命周期任务(比如每小时跑一次的脚本),不能靠 HTTP 暴露,得改用 PUSHGATEWAY。
- 长时运行爬虫(如常驻协程)→ 用
start_http_server - 短时任务(如 Cron 启动的单次抓取)→ 必须走
push_to_gateway - 多进程爬虫(如用
multiprocessing)→ 必须设环境变量PROMETHEUS_MULTIPROC_DIR并用MultiProcessCollector
prometheus.yml 中怎么写爬虫 job 才不会漏指标?
别直接抄官网 demo 里的 static_configs。爬虫服务 IP 经常变(尤其是 Docker/K8s 环境),推荐用服务发现。如果实在只能静态配,务必加 honor_labels: true,否则你代码里打的 status 标签会被覆盖。
一个安全的最小配置:
scrape_configs: - job_name: 'crawler' honor_labels: true static_configs: - targets: ['localhost:8000'] # 对应 start_http_server 的地址 metrics_path: '/metrics'
如果你用了 PUSHGATEWAY,job 配置就完全不一样:
- job_name: 'pushgateway' static_configs: - targets: ['pushgateway:9091'] metrics_path: '/metrics'
这时所有指标实际来自 pushgateway 的内部存储,不是你的爬虫进程直连。
成功率计算容易被忽略的陷阱
Prometheus 里算成功率不能只写 rate(scrape_total{status="success"}[1h]) / rate(scrape_total[1h])。问题在于:rate() 基于样本变化率,而短时任务 push 的指标可能被 pushgateway 清理(默认 2 小时 TTL),或因重试导致重复计数。
更稳的做法是用 increase() + 显式时间窗口对齐:
(increase(scrape_total{status="success"}[2h]) + 1e-10)
/
(increase(scrape_total[2h]) + 1e-10)
加 1e-10 是防除零;窗口长度(如 [2h])必须大于最大可能的 push 间隔,否则 increase 会丢数据。另外,如果爬虫有重试机制,记得在指标 label 里带上 retry_count,否则“第一次失败、第三次成功”会被记成三次独立事件。
Python免费学习笔记(深入):立即使用
在学习笔记中,你将探索 Python 的核心概念和高级技巧!











