To do so, UiPath re-architected its infrastructure to support high-scale intelligent document processing (IDP) using UiPath IXP and moved from isolated clusters to a shared Google Cloud GPU fleet, balancing A3 VM instances (NVIDIA H100 GPUs) for training with G4 VM instances (NVIDIA RTX Pro 6000) for inference. This architecture lets UiPath solve its “spiky workload” problem and count on predictable costs and open-source patterns that the company’s engineering teams can use to replicate this architecture themselves.
“Realizing the full potential of enterprise agentic AI requires an infrastructure that matches our ambition. Google Cloud provides the scale and flexibility we need to train specialized models and deploy them globally. This partnership allows us to deliver high-precision intelligent document processing and autonomous agents that don’t just chat, but actively drive business outcomes for our customers.” – Raghu Malpani, Chief Technology Officer, UiPath
The context: heavy-duty math
UiPath has run its full-stack automation platform on Google Cloud for years, but as its agentic AI initiatives expanded, it faced a series of new infrastructure challenges.
Core capabilities like IDP, computer vision, and LLM-powered reasoning require heavy-duty math, so UiPath’s engineering team utilizes LLAMA model grounding that allows its robots to “see” interfaces with human-like clarity. And with specialized document models built on the Qwen architecture, the team can extract valuable data from messy, real-world paperwork.
These models live on the UiPath cloud infrastructure, where cutting every possible millisecond of latency is essential. Moving from a “cool demo” to a reliable production tool without exploding costs meant the team had to rethink its underlying silicon.
The challenge: more demand than supply
In the past, when a team at UiPath needed to train a new model or run inference, it provisioned GPU nodes on demand and scaled up or down depending on whether the workloads were spiking or slowing.
This was a functional strategy when cloud capacity was cheap, abundant, and perfectly elastic. But as its AI ambitions grew, UiPath found this approach could no longer keep up with its operational complexity. It now faced three new challenges:
-
Spiky workloads: To ensure it had sufficient power for peak demand, UiPath often had to buy extra capacity that sat idle during quieter periods, wasting expensive headroom. The company needed intelligent, on-demand scaling that didn’t require paying for silicon that wasn’t crunching numbers.
-
Supply bottlenecks: For large-scale fine-tuning, the price-to-performance ratio on gold standard high-end A3 VM instances with 8-cluster H100s is unbeatable. But global demand for those chips has outstripped supply, making it nearly impossible to scale training efforts at the speed UiPath desired just by adding nodes.
-
Operational overhead: UiPath was also struggling with geographical inefficiency because stable inference demand still meant maintaining dedicated clusters in multiple regions to ensure low latency for international customers. Further, managing GPU infrastructure for both training and inference added inefficient layers of operational overhead.
The solution: a shared GPU fleet
With all of that in mind, UiPath decided to treat its GPUs as a shared strategic resource instead of a product-centric elastic infrastructure.
As a result, its engineering team designed a platform-level shared GPU fleet managed by its machine learning services (MLS) platform, which prioritizes work across teams and time windows while balancing demand across workflows. During the day, the fleet serves real-time inference and latency-sensitive workloads, and at night or during off-peak hours, it automatically switches to batch training and long-running jobs.
By coordinating workloads at the fleet level, MLS lets UiPath maximize utilization while reducing contention, all without relying on per-instance elasticity. It also enables the company to schedule capacity in advance, which improves predictability for both research and production use cases.
Why Google Cloud: AI Hypercomputer architecture
To support its growing scale, UiPath leveraged Google Cloud AI Hypercomputer, which offers a system-level approach integrating performance-optimized hardware, open software, and flexible consumption models into a unified environment. AI Hypercomputer also minimizes the friction between hardware and software layers, which allows engineering teams to focus on model performance rather than infrastructure management.
Source Credit: https://cloud.google.com/blog/topics/customers/how-uipath-built-its-high-performance-gpu-platform/
