requests同步爬虫慢因串行阻塞i/o,cpu空等;aiohttp需复用clientsession、await调用、控制连接数与超时,避免ssl验证开销和异常中断。

为什么 requests 同步爬虫慢得明显
同步请求本质是串行:发一个、等响应、再发下一个。哪怕网络延迟平均只有 200ms,100 个链接就要卡 20 秒以上,中间 CPU 其实空着——它在干等 TCP 握手、SSL 解密、远端服务器处理这些 I/O 操作。
requests 本身不支持异步,硬套 threading 或 multiprocessing 又容易触发 GIL 争抢或连接数爆炸,反而更卡。
常见错误现象:
- 用
time.sleep()模拟等待后发现总耗时线性增长 - 开了 10 个线程,但
requests.get()还是排队执行(没真正并发) -
ConnectionResetError或TooManyRedirects突然变多(服务端反爬或连接池打爆)
aiohttp + asyncio 怎么写才不翻车
核心就两点:所有 HTTP 调用必须用 aiohttp.ClientSession,所有等待必须用 await,且整个流程得跑在 event loop 里。
使用场景:
- 批量抓取静态页面、API 接口、JSON 数据
- 不适合需要执行复杂 JS 渲染的页面(得换 Playwright 或 Pyppeteer)
实操建议:
-
ClientSession必须复用,不能每个请求都新建——否则 DNS 查询和 TCP 连接开销全白费 - 设置
connector=TCPConnector(limit=100)控制并发连接数,避免被目标封 IP 或触发服务端限流 - 加
timeout=ClientTimeout(total=10),防止某个请求卡死拖垮整批任务
简短示例:
import asyncio import aiohttp <p>async def fetch(session, url): async with session.get(url) as resp: return await resp.text()</p><p>async def main(): connector = aiohttp.TCPConnector(limit=50) timeout = aiohttp.ClientTimeout(total=10) async with aiohttp.ClientSession(connector=connector, timeout=timeout) as session: tasks = [fetch(session, url) for url in urls] results = await asyncio.gather(*tasks) return results</p><h1>别忘了这句才能跑起来</h1><p>asyncio.run(main())</p><div class="aritcle_card flexRow artxards"> <div class="artcardd flexRow"> <a class="aritcle_card_img" rel="nofollow" href="/xiazai/skill7351" title="testing-python"><img src="https://img.php.cn/upload/skill/000/000/081/179143938488980.jpg" alt="testing-python" onerror="this.onerror='';this.src='/static/lhimages/moren/morentu.png'" ></a> <div class="aritcle_card_info flexColumn"> <a rel="nofollow" href="/xiazai/skill7351" title="testing-python" class="overflowclass">testing-python</a> <p class="overflowclass">使用pytest编写和评估有效的Python测试。适用于编写测试、审查测试代码、调试测试失败或提高测试覆盖率。</p> </div> <a rel="nofollow" href="/xiazai/skill7351" title="testing-python" class="aritcle_card_btn flexRow flexcenter"><b></b><span>下载</span> </a> </div> </div>
requests 和 aiohttp 的参数差异容易漏看
很多习惯写 requests.get(url, headers={}, params={}) 的人,直接套用到 aiohttp 会出错。
关键区别:
-
params写法一样,但headers必须是dict,不能是CaseInsensitiveDict(requests.utils.default_headers()返回的那种) -
data和json不能同时传;想发 JSON,直接用session.post(url, json={...}),它会自动设Content-Type: application/json并序列化 -
verify=False在aiohttp里对应connector=TCPConnector(ssl=False),不是传给get()的参数
性能影响:
- 忘关 SSL 验证(开发调试时常见)会让每次连接多花 50–100ms
-
json=...比手动data=json.dumps(...)+headers更快,因为内置了编码复用逻辑
异步爬虫卡住不动?先查这几个点
最常被忽略的是 event loop 状态和异常吞没。
容易踩的坑:
- 在 Jupyter 或某些 IDE 的交互式环境里直接跑
asyncio.run(),可能报RuntimeError: asyncio.run() cannot be called from a running event loop -
asyncio.gather(*tasks)中某个请求抛异常,默认会中断整个批次;加return_exceptions=True才能继续跑完其他请求 - 目标网站返回 4xx/5xx 时,
aiohttp不会自动 raise 异常,得自己检查resp.status,否则看似“成功”实则拿到空页或错误 HTML
兼容性注意:
- Python asyncio.run(),得手动
loop = asyncio.get_event_loop()+loop.run_until_complete() -
aiohttp3.9+ 默认禁用 HTTP/1.0,如果目标服务只支持老协议,得显式加connector=TCPConnector(force_close=True)
事情说清了就结束
Python免费学习笔记(深入):立即使用
在学习笔记中,你将探索 Python 的核心概念和高级技巧!










