AWS Machine Learning Blog 中文 关注 使用 Amazon SageMaker Inference 的基于前缀的路由来降低 LLM 延迟 Amazon SageMaker Inference 现已提供前缀感知路由(prefix-aware routing),这是一种将具有相同提示前缀的请求路由到同一实例的策略,以保持 KV 缓存处于热状态。在 Llama 3.1 70B 上的基准测试显示,该策略将 P50 首 token 延迟降低了高达 77%,并将 KV 缓存命中率从约 25% 提升至 80% 以上。 Reduce LLM latency with prefix-aware routing on Amazon SageMaker Inference aws.amazon.com AWS Machine Learning Blog 中文 RSS thenote.app