Google Cloud Blog
Follow
Scaling agentic AI: How UiPath built its high-performance GPU platform on AI Hypercomputer
UiPath is transitioning to agentic AI, deploying autonomous agents that reason and make decisions to orchestrate complex business processes. This shift requires significant computational power and reliable infrastructure, especially for large enterprises. Orchestrating hundreds of GPUs efficiently is crucial for global AI platforms, balancing training and inference without increasing costs or latency. UiPath re-architected its infrastructure on Google Cloud, moving to a shared GPU fleet for intelligent document processing. They now use A3 VM instances for training and G4 VM instances for inference, addressing spiky workloads and ensuring predictable costs. This partnership with Google Cloud provides the necessary scale, flexibility, and specialized AI capabilities for UiPath's ambitious agentic AI initiatives. The company's previous approach of provisioning GPUs on demand became unmanageable due to spiky workloads, supply bottlenecks for high-end chips, and operational overhead. By treating GPUs as a shared strategic resource, UiPath's ML Services platform prioritizes work and balances demand across workflows, optimizing utilization. Google Cloud's AI Hypercomputer architecture, integrating hardware, software, and flexible consumption models, further supports UiPath's growing scale. UiPath leveraged Google Kubernetes Engine and Google Cloud's Dynamic Workload Scheduler (DWS) to secure GPU capacity in advance, enabling proactive planning. They now use A3 VM instances for training and cost-effective G4 VM instances for inference, optimizing performance and price. This new infrastructure allows UiPath to bring advanced models into production, such as those used by Omega Healthcare and Thermo Fisher Scientific, achieving remarkable accuracy and processing efficiencies. Key lessons learned include decoupling capacity, scheduling compute in advance, and right-sizing silicon for different workloads.