AWS Machine Learning Blog 中文 关注 使用模型缓存减少 Amazon SageMaker HyperPod 的推理冷启动 Amazon SageMaker HyperPod 现已支持推理模型缓存,该功能可将模型权重和容器镜像预加载至集群节点,使 Pod 从本地 NVMe 存储读取数据,而非通过网络下载。了解模型缓存如何将冷启动时间从数十分钟缩短至数秒,其工作原理以及启用方法。 Reduce inference cold starts on Amazon SageMaker HyperPod with model caching aws.amazon.com AWS Machine Learning Blog 中文 RSS thenote.app