Google Cloud Blog
Follow
Claude at scale on Google Cloud: Frontier AI, built for enterprise production
Running frontier AI like Claude in production is complex, requiring careful management of accelerators and consistent latency across global regions. Claude on Google Cloud offers a solution by integrating Anthropic's advanced AI models with Google's robust infrastructure. This partnership leverages Claude's reasoning capabilities with Google Cloud's managed services, worldwide reach, and compliance features. The integration makes using Claude operationally identical to other Google Cloud services, simplifying development and reducing infrastructure management overhead.The platform provides managed infrastructure, freeing engineers to focus on building features rather than managing clusters for compute provisioning, auto-scaling, and load balancing. Claude is available as a Model-as-a-Service through Agent Platform's Model Garden, accessible via standard REST APIs. This approach ensures consistent operational experience with existing Google Cloud security and observability tools. Global endpoints are offered to maintain low latency and high uptime for a worldwide user base. Regional endpoints ensure data residency and low latency for specific geographic areas, while multi-region endpoints offer data sovereignty without single-region dependency.Claude on Google Cloud inherits Google Cloud's comprehensive security posture, including FedRAMP High and HIPAA compliance, enabling deployment in regulated industries. VPC Service Controls create secure perimeters, and IAM-native access control manages endpoint access, eliminating the need for separate API keys. Visibility into usage, errors, and latency is provided through Cloud Logging and Cloud Monitoring. The service is optimized for both cost and performance at scale, incorporating Claude's native features like prompt caching and streaming responses.Google Cloud's serving infrastructure complements Claude's capabilities with features such as batch prediction for large-scale offline workloads and provisioned throughput for guaranteed performance. The platform also handles memory management and scheduling for long-context requests. This unified approach allows teams to optimize at both the model and infrastructure levels. Furthermore, the same infrastructure powers the agent layer of Agent Platform, enabling Claude to orchestrate multi-step tasks and delegate to other agents through the Agent2Agent protocol. This allows for the creation of powerful, auditable, and unified AI agents within the Google Cloud ecosystem.