asyncio程序不能直接用cprofile.run(),因其不感知协程调度,仅记录线程栈帧;应改用cprofile.profile()在async main函数体内启用/禁用,专注测量await间cpu耗时,并用pstats过滤干扰项。

asyncio 程序不能直接用 cProfile.run()
cProfile.run() 会把整个 async def 函数当同步代码执行,结果里只看到 run_until_complete 或 run_forever 占满 100% 时间,真实协程耗时完全不可见。根本原因是 cProfile 基于 CPython 的 C API 钩子,它不感知事件循环调度和协程挂起/恢复,只记录线程栈帧——而 asyncio 大量时间花在 select、epoll 等系统调用上,这些不会被 cProfile 捕获为 Python 函数调用。
实操建议:
- 改用
asyncio.run()启动主协程,并在入口处手动启动cProfile.Profile() - 必须在事件循环启动前启用 profiler,在所有协程结束后显式
disable()和dump_stats() - 避免在
async with或await中间穿插enable()/disable(),否则 profile 数据会断裂
用 profile.Profile() 包裹 async main 函数体
关键不是“测整个脚本”,而是“测协程内部的 Python 执行耗时”。把 profile.Profile() 的生命周期严格限定在协程函数体内部,才能捕获 await 之间的真实 CPU 消耗(比如 JSON 解析、正则匹配、数据转换等)。
示例结构:
import asyncio
import cProfile
<p>async def main():
pr = cProfile.Profile()
pr.enable() # ⚠️ 必须在 await 前启用
await do_work() # 这里才是你要测的部分
pr.disable() # ⚠️ 必须在 await 后、协程返回前禁用
pr.dump_stats("main.prof")</p><p>if <strong>name</strong> == "<strong>main</strong>":
asyncio.run(main())</p>
注意:do_work() 内部的 await asyncio.sleep(1) 不会出现在 profile 结果中——这正是对的,它不占 CPU;但 json.loads() 或 pd.DataFrame() 构造就会清晰暴露。
stats 文件需用 pstats 加载并过滤协程调用
生成的 .prof 文件默认包含大量 asyncio 底层函数(如 _run_once、_step),干扰判断。直接 pstats.Stats("main.prof").sort_stats("cumtime").print_stats(20) 会看到一堆无意义的事件循环调度开销。
实操建议:
- 用
filter排除asyncio/和selector_events.py相关路径:stats.filter_stats("your_module_name") - 优先看
tottime(函数自身耗时),而非cumtime(含子调用),因为协程跳转会让cumtime失真 - 若发现
__await__或send()耗时高,说明有协程对象反复创建或未复用,不是 CPU 瓶颈,而是对象分配问题
真正瓶颈常在 awaitable 类型混用上
profile 显示某个 await 行总耗时很长,但该行调用的函数本身 tottime 很低?大概率是 await 的对象类型不对。比如用 await asyncio.to_thread(os.path.getsize, path) 测文件大小,比直接 await aiofiles.os.stat(path) 慢 5–10 倍——前者要跨线程切换+序列化参数,后者走纯异步 syscall。
常见陷阱:
- 误把
requests.get()包进to_thread当“异步方案”,其实应换aiohttp.ClientSession - 用
await asyncio.sleep(0)做 yield,但 profile 里看不到它——它不占 CPU,但可能暴露调度延迟问题,需结合asyncio.current_task().get_coro()手动打点 -
concurrent.futures.ThreadPoolExecutor提交任务后 await Future,profile 会显示wait()耗时,但真实瓶颈在 worker 线程里,此时得去查线程级 profile
asyncio 的性能评估,本质是分清“CPU-bound”和“IO-bound + 调度开销”。profile 只管前者;后者得靠 asyncio.create_task() 数量、sys.getsizeof() 协程对象、以及 loop.slow_callback_duration 配置来交叉验证。
Python免费学习笔记(深入):立即使用
在学习笔记中,你将探索 Python 的核心概念和高级技巧!











