torch.compile有时更慢是因为输入shape不稳导致频繁重编译,需固定batch_size和max_length、启用fullgraph=true、分离编译纯张量计算路径,并避免python标量操作和动态控制流。

torch.compile 为什么有时加了反而更慢?
不是所有模型加 torch.compile 都能提速,关键看输入是否稳定。如果每个 batch 的 shape 都不同(比如 NLP 中变长序列 padding 不一致),Dynamo 会为每个新 shape 重新编译,开销远超收益。你看到的“第一次慢、后面快”,只在 shape 固定时成立。
实操建议:
- 训练前用典型 shape(如
(8, 512))预热一次torch.compile模型 - 固定
batch_size和max_length,避免动态 padding 破坏 shape 稳定性 - 启用
torch._dynamo.config.verbose = True,观察日志里是否频繁出现"compiling new graph" - A100 上 Inductor 生成的 Triton kernel 比 V100 快 2–3×,硬件差异直接影响收益
fullgraph=True 是必须显式写的,不是默认值
torch.compile(model) 默认 fullgraph=False,意味着遇到不可追踪分支就 fallback 到 eager 模式,编译形同虚设。真正触发 Inductor 图融合,必须写成 torch.compile(model, fullgraph=True)。
常见报错和修复:
-
torch._dynamo.exc.Unsupported: call_function aten._local_scalar_dense→ 检查是否用了loss.item()、print(loss)或.cpu().numpy();这些 Python scalar 操作必须移出编译范围 -
if x.shape[0] > 1:这类运行时 shape 判断 → 提前在外层判断,或改用torch.where替代 - 调试时的
print、list.append→ 全部删掉,或移到compiled_model调用之外
梯度累积和 compile 要分开编译,不能整个 train_step 一起编译
把 optimizer.step()、loss.backward()、print、指标统计等逻辑塞进 torch.compile,大概率失败。Inductor 只适合纯张量计算路径。
正确做法是分层编译:
compiled_model = torch.compile(model, fullgraph=True, dynamic=False)compiled_loss_fn = torch.compile(loss_fn, fullgraph=True)- 保持
train_step函数为 eager 模式:只调用编译后的 model 和 loss_fn,其余逻辑(zero_grad、step、log)不编译 - 注意
loss.backward()前后不能有 .item() 或任何 CPU 转换操作
mode='max-autotune' 不等于“一定更快”
mode='max-autotune' 会让 Inductor 尝试更多 kernel 组合,编译时间可能长达 2–3 分钟,且对小模型或简单网络往往得不偿失。它适合 ResNet-50、ViT、LLM 这类计算密集型模型。
选 mode 的实际策略:
- 开发调试阶段用
mode='default'或干脆不设(默认值),快速验证逻辑 - 正式训练前跑一轮
mode='max-autotune',但只做一次,结果会 cache - 如果显存紧张,加
dynamic=False强制静态 shape,避免 runtime recompile - 首次运行务必留出 30–120 秒编译时间,别误判为卡死











