DEV Community
Follow
Your agent loads every tool schema by default. Decide which ones it should see.
Before modifying agent configurations, it's crucial to understand prompt token usage, especially the allocation for tool schemas. The author highlights that most teams cannot answer these basic questions due to a lack of ownership. Pi 1.0 introduced deferred tool loading and Codemode, features that shift tool visibility to a per-tool setting, a significant change. Previously, Pi resisted MCP, but version 1.0 added native support due to the need for metadata regarding tool exposure. This metadata dictates whether a tool is directly visible to the model, loaded on demand, or callable only from Codemode.Tool schemas incur a cost, consuming tokens in the system prompt or tool block on every request, regardless of relevance. This cost manifests as financial expense, model attention drain, and reduced reproducibility. A vendor example shows a request's prompt tokens dropping by approximately 40% with these changes. Pi's new metadata allows tools to be directly exposed, deferred, or Codemode-only. Direct exposure is for frequently used tools, deferred for rarely used ones, and Codemode only when output needs filtering or for tool combinations.The Codemode approach, where the model writes code to call tools within a sandbox, is particularly interesting as it alters what enters the context window. However, the author cautions that Codemode doesn't fix server-side issues where servers might inefficiently return text blobs instead of structured data. Before tuning, an audit is recommended, measuring cold start prompt tokens with and without tools loaded. Sorting tools into the three exposure categories clarifies usage patterns and identifies unnecessary schema bloat.The decision of whether a tool needs to be chosen by name determines its placement; if not, it belongs in Codemode or deferred. Frequency of use dictates between direct (frequent) and deferred (rare). While Codemode can reduce token counts, server-side efficiency remains a concern. The author emphasizes measuring token counts with the provider's tokenizer and performing a before-and-after comparison for fixed tasks. Visibility is a safety measure as tools a model cannot see cannot be mistakenly chosen. Finally, the audit helps identify underutilized connectors that inflate prompt size.