kratos微服务中需设http读写超时、grpc客户端deadline透传及服务端响应截断。http服务端配置readtimeout=3s、writetimeout=5s;grpc客户端用context.withdeadline或timeout.client中间件;服务端校验deadline并流式检查ctx.done()。

在Kratos微服务中正确设置请求超时与Deadline,是防止级联空耗、避免GPU算力焚化炉式浪费的关键操作。网关3秒超时后若下游未及时感知,大模型仍会完成12秒无效计算,显存和带宽被彻底锁死。
HTTP服务端全局超时控制
在internal/server/http.go中配置http.Server的ReadTimeout和WriteTimeout字段:
将srv := http.NewServer()替换为显式构造http.Server实例:
srv := &http.Server{ Addr: ":8000", Handler: router, ReadTimeout: 3 * time.Second, WriteTimeout: 5 * time.Second }
这一步必须手动写死超时值,Kratos默认不设限;不设ReadTimeout会导致慢连接长期占用线程,不设WriteTimeout则响应卡住时无法释放goroutine。
gRPC客户端Deadline透传
所有gRPC调用必须基于context.WithDeadline或context.WithTimeout生成子Context,且不能复用上游原始Context——因为上游Deadline已扣除链路前序耗时。
方法一:手动计算剩余预算
在编排服务中接收请求后,立即读取req.Context().Deadline(),减去当前已耗时(如向量库+重排耗时),再用context.WithDeadline生成新Context传给大模型客户端。
方法二:使用Kratos内置中间件自动截断
在internal/client/grpc.go中注册timeout.Client中间件:
conn, err := grpc.Dial( "discovery:///llm.service", grpc.WithMiddleware( timeout.Client( timeout.WithTimeout(2 * time.Second), // 此处填预估最大允许耗时 ), ), )
【注意】此中间件仅对Unary RPC生效,Streaming接口必须手动在Handler内调用ctx, cancel := context.WithTimeout(req.Context(), 2*time.Second)并确保defer cancel()
gRPC服务端强制响应截断
第一步:在internal/server/grpc.go的ServerOption中启用grpc.KeepaliveParams
grpc.KeepaliveParams(keepalive.ServerParameters{ MaxConnectionAge: 30 * time.Minute, MaxConnectionAgeGrace: 5 * time.Minute })
第二步:为每个RPC方法添加context.Deadline()校验
打开internal/service/llm_service.go,在Generate方法开头插入:
if d, ok := ctx.Deadline(); ok && time.Until(d)
第三步:在流式响应循环中每产出一个Token就检查ctx.Done()
若select { case 触发,立即终止<code>Send()并返回,否则SSE连接断开后模型仍在后台生成。











