
本文介绍使用 guzzlehttp 实现 php 异步并发 http 请求的最佳实践,帮助你将原本需 19 小时的 20 万次 tfl api 调用压缩至分钟级,显著提升地理距离批量计算效率。
本文介绍使用 guzzlehttp 实现 php 异步并发 http 请求的最佳实践,帮助你将原本需 19 小时的 20 万次 tfl api 调用压缩至分钟级,显著提升地理距离批量计算效率。
在处理大规模地理服务调用(如 TfL 公共交通路线规划)时,串行请求(file_get_contents 或单个 cURL)会因网络延迟叠加导致性能崩溃——正如你当前遇到的近 19 小时耗时问题。根本症结在于 I/O 阻塞:每次请求必须等待响应返回后才发起下一次。解决方案是并发非阻塞请求,即同时发起多个 HTTP 请求并统一等待全部完成,充分利用网络带宽与 API 服务端并发能力。
推荐采用 GuzzleHttp + Promises 方案,它基于 cURL multi 接口封装,提供简洁、健壮、可调试的异步编程模型,远优于手动管理 curl_multi_* 函数。
✅ 正确实现并发请求的核心步骤
-
安装 Guzzle(确保 PHP ≥ 7.2,启用 curl 扩展):
composer require guzzlehttp/guzzle
-
构建批量 Promise 队列:
避免一次性创建 20 万 Promise(内存溢出风险),应分批处理(例如每批 50–100 个请求):use GuzzleHttp\Promise; use GuzzleHttp\Client; $client = new Client([ 'base_uri' => 'https://api.tfl.gov.uk/', 'timeout' => 10.0, 'headers' => ['User-Agent' => 'PHP-GeoBatch/1.0'] ]); $allResults = []; $batchSize = 80; // 每批并发请求数,根据服务器资源和 TfL 限流策略调整 // 假设 $universities 和 $hosts 已从数据库加载为二维数组 foreach (array_chunk($universities, ceil(count($universities) / 10)) as $uniChunk) { // 外层按大学分组降压 foreach (array_chunk($hosts, $batchSize) as $hostBatch) { $promises = []; foreach ($uniChunk as $uni) { foreach ($hostBatch as $host) { $url = "journey/journeyresults/{$uni['Postcode']}/to/{$host['Postcode']}"; $promises[] = $client->getAsync($url, [ 'query' => ['app_key' => 'a59c7dbb0d51419d8d3f9dfbf09bd5cc'], 'http_errors' => false // 防止 4xx/5xx 抛异常中断整个批次 ])->then( function ($response) use ($uni, $host) { if ($response->getStatusCode() === 200) { $data = json_decode($response->getBody(), true); $durations = array_column($data['journeys'] ?? [], 'duration'); $minDuration = !empty($durations) ? min($durations) : null; return [ 'uni_postcode' => $uni['Postcode'], 'host_postcode' => $host['Postcode'], 'duration_min' => $minDuration, 'status' => 'success' ]; } return [ 'uni_postcode' => $uni['Postcode'], 'host_postcode' => $host['Postcode'], 'duration_min' => null, 'status' => 'api_error', 'code' => $response->getStatusCode() ]; }, function ($reason) use ($uni, $host) { return [ 'uni_postcode' => $uni['Postcode'], 'host_postcode' => $host['Postcode'], 'duration_min' => null, 'status' => 'network_error', 'error' => (string)$reason ]; } ); } } // 并发执行本批次所有请求 $results = Promise\settle($promises)->wait(); // 收集结果并写入数据库(建议使用批量 INSERT) $batchInserts = []; foreach ($results as $result) { if ($result['state'] === 'fulfilled') { $data = $result['value']; if ($data['duration_min'] !== null) { $batchInserts[] = [ $data['uni_postcode'], $data['host_postcode'], $data['duration_min'] ]; } } } if (!empty($batchInserts)) { // 使用 PDO 批量插入(示例伪代码) $stmt = $pdo->prepare("INSERT INTO travel_times (uni_postcode, host_postcode, duration_min) VALUES (?, ?, ?)"); foreach ($batchInserts as $row) { $stmt->execute($row); } } echo "✅ Batch completed: " . count($batchInserts) . " valid records saved.\n"; usleep(100000); // 可选:轻度限流,避免触发 TfL 频率限制 } }
⚠️ 关键注意事项
- TfL API 限制:TfL 免费 Key 有 严格调用配额(通常 1000 次/小时)。务必在生产环境添加 Retry-Middleware 和 RateLimiting 逻辑,或申请商业 Key。
- 错误容错:使用 Promise\settle() 而非 Promise\unwrap(),确保单个失败请求不影响整批执行;对 429 Too Many Requests、503 Service Unavailable 等需指数退避重试。
- 内存优化:20 万次请求对象不可全量加载进内存。始终分批(array_chunk)+ 流式处理 + 及时 unset() 中间变量。
- 数据库写入:避免逐条 INSERT。改用事务包裹的批量插入(INSERT ... VALUES (...), (...), (...)),性能可提升百倍。
- 本地缓存:对重复 uni→host 组合(如多用户查询同一学校)引入 Redis 缓存,TTL 设为 24h,大幅降低实际调用量。
✅ 最终收益
经合理配置(80 并发 + 分批 + 错误重试 + 批量入库),原 19 小时任务可压缩至 15–45 分钟内完成,且后续“查某大学最近 10 个住宿点”只需一条 SQL 查询(ORDER BY duration_min LIMIT 10),真正实现毫秒级响应。
异步不是银弹,但它是大规模地理数据服务化的必经之路——关键在于平衡并发数、错误恢复与外部 API 约束。从今天开始,用 Guzzle Promise 重构你的循环吧。
php免费学习视频:立即使用
踏上前端学习之旅,开启通往精通之路!从前端基础到项目实战,循序渐进,一步一个脚印,迈向巅峰!











