Prompt Caching Doesn't Save Mo... Note
DZone.com

Prompt Caching Doesn't Save Money on Turn One

I went looking for a clean way to show what prompt caching actually saves an agent, and the first thing I found was a fact that's easy to miss if you only read the "up to 90% savings" headline. Caching a fresh conversation's first turn costs more than not caching it. There's no cache to read from yet, so you pay the input price on the content plus a 25% premium to write to the cache, and get nothing back. The savings arrive starting on turn two, once there's something to read. Where This Lives in deepagentsDeep Agents ships AnthropicPromptCachingMiddleware from langchain-anthropic in its default middleware stack. It doesn't decide what to cache by guessing. It tags exactly two things on every model call: