Oct 06, 2026
/
By Bruno S.
/
Cloud GPU pricing ranges from under $0.50 per hour for an entry-level card to over $12 for the most powerful ones. For the same GPU, the price can still vary by as much as sevenfold depending on which provider you rent from. That gap comes from how providers package the hardware, not from the chip itself.
What you pay depends heavily on how that access is structured: some providers rent you just the GPU; others bundle it inside a full server with dozens of processors and terabytes of memory, charge for everything, and require you to rent several GPUs at once.
The pricing model adds a second layer: on-demand, interruptible, or reserved access can halve or double the rate for the same card.
The advertised rate also doesn’t capture the full cost: storage fees, data transfer charges, and idle GPUs running between jobs can push a real monthly bill well above what the hourly rate suggests.
How much do cloud GPUs cost per hour?
Cloud GPUs cost between $0.14 and $12.29 per GPU-hour (the cost of renting one GPU for one hour) at on-demand (pay-as-you-go) list prices, as of September 2026.
The comparison table covers seven GPU models from seven providers across US regions. The VRAM column shows each GPU’s on-chip memory, which determines how large a model the GPU can run.
|
GPU |
VRAM |
Hostinger |
RunPod |
Vast.ai |
Lambda |
AWS |
GCP |
Azure |
|
RTX 4090 |
24 GB |
$0.38 |
$0.74 |
from $0.14 |
– |
– |
– |
– |
|
RTX A6000 |
48 GB |
– |
$0.53 |
from $0.28 |
$1.09 |
– |
– |
– |
|
RTX PRO 6000 |
96 GB |
$0.60 |
$2.09 |
from $0.89 |
– |
– |
$4.50† |
– |
|
L40S |
48 GB |
$0.92* |
$0.99 |
from $0.47 |
– |
– |
– |
– |
|
A100 80GB |
80 GB |
$1.43 |
$1.39 |
from $0.44 |
from $1.99§ |
$3.43† |
$3.67§ |
$3.67† |
|
H100 |
80 GB |
– |
$2.89 |
from $1.33 |
$3.29 |
$6.88 |
$11.06† |
$12.29† |
|
B200 |
192 GB |
$4.50 |
$6.79 |
from $4.38 |
from $6.69 |
– |
$16.11‡ |
– |
On-demand rates, normalized per GPU-hour, US regions, as of September 2026. RunPod: Secure Cloud tier. Vast.ai: marketplace range; prices change daily. (–) = not offered.
† Whole-VM or minimum 8-GPU node; per-GPU rate derived from instance total.
‡ Reservation required; no standard on-demand.
* Availability varies.
§ Lambda $1.99 and GCP $3.67 reflect A100 40GB variants (cheapest option). Lambda 80GB SXM is approximately $2.79; GCP a2-ultragpu 80GB is approximately $5.03.
Beyond the per-GPU rate, the tier you pick – shared community hardware or dedicated data-center capacity –, regional availability, and uptime commitments all affect the real cost of a workload. On those dimensions, comparing GPU cloud providers often changes the final choice more than a small rate difference would.
Important
GPU cloud pricing updates monthly and sometimes weekly; marketplace platforms like Vast.ai can change hourly. The rates above were verified in September 2026 – re-check primary pricing pages before you deploy anything.
How does cloud GPU pricing work?
Cloud GPU pricing runs on three models: on-demand pay-as-you-go, interruptible spot capacity at a significant discount, and reserved terms for predictable long-running workloads.
On-demand
On-demand gives you immediate access at a fixed rate with no commitment. You pay for the whole time the instance (your running virtual server) is running, whether the GPU is processing a job or sitting idle. It’s the default model for variable or short-term workloads and the basis for every rate in the comparison table.
Spot and interruptible
Spot pricing cuts costs from 50 to 91 percent depending on the provider, but the instance can be reclaimed with little notice when the provider needs the capacity back. Google Cloud Platform (GCP) gives 30 seconds of warning; Amazon Web Services (AWS) gives two minutes.
Spot only makes sense with proper checkpointing: saving your job’s progress at intervals so it can resume from the last save point after an interruption, rather than from scratch. A job with no checkpoints that gets evicted mid-run can end up costing more than an on-demand run would have.
Spot prices are market rates, not guaranteed discounts. On bid-based platforms, a supply crunch can push the spot rate close to or above on-demand. Always check the live spot rate before assuming it’s cheaper than on-demand.
|
Provider |
Spot/interruptible model |
Typical discount |
Notes |
|
AWS |
Spot Instances |
~57% |
~$2.96/GPU-hr for H100 (p5.48xlarge); fluctuates |
|
GCP |
Spot VMs |
60–91% |
30-second eviction notice |
|
Azure |
Spot VMs |
~80% |
30-second eviction notice |
|
RunPod |
Community Cloud |
~50% |
No published spot rate card |
|
Vast.ai |
Interruptible |
50%+ |
Host-set; varies daily |
Reserved and committed
Reserved pricing locks in a GPU for a fixed term in exchange for a lower effective hourly rate. Discounts range from 25 to 65 percent, depending on the provider and commitment length.
|
Provider |
Discount |
Term options |
|
AWS |
25–45% |
Capacity Blocks (variable), Savings Plans |
|
GCP |
up to 65% |
1 or 3 years (Committed Use Discounts) |
|
Azure |
up to 63% |
1 or 3 years; 5 years on ND H100 v5 (Reserved Instances) |
|
Vast.ai |
Up to 50% |
1, 3, or 6 months |
|
Lambda |
1-Click Clusters from $5.54/GPU/hr |
2 weeks – 1 year; 16+ GPUs minimum; 1yr+ contact sales |
|
RunPod |
Contact sales |
Enterprise agreements |
Billing granularity
Providers charge in different units, and the difference matters most for short or interrupted jobs. RunPod and Vast.ai bill per second; Lambda bills per-minute; AWS, GCP, and Azure bill per second with a one-minute minimum. A 16-minute job on RunPod costs roughly 27 percent of the hourly rate; on Lambda it costs the full hour.
Hostinger bills hourly using a prepaid credit system (1 credit = $0.01), with the first full hour deducted at the moment of deployment. Destroying an instance mid-hour does not return unused time.
Cloud GPU pricing by model
Cloud GPU models fall into several pricing tiers, from under $0.14/hr for the RTX 4090 to $6.79/hr for the B200 on standard on-demand – and higher still at hyperscalers, where a reserved B200 runs $16.11/hr. Each tier offers different VRAM capacity and suits different jobs.
RTX 4090 pricing per hour
The RTX 4090 costs from $0.14 per hour on Vast.ai and $0.38 per hour on Hostinger, as of September 2026. RunPod lists it at $0.34 on Community Cloud and $0.74 on Secure Cloud.
Community Cloud uses hardware contributed by independent hosts that RunPod has screened; Secure Cloud uses data-center-grade infrastructure with higher reliability guarantees.
No hyperscaler carries the RTX 4090, since GPUs originally built for gaming and workstation use aren’t available on AWS, GCP, or Azure.
With 24 GB of VRAM, the 4090 handles most 7-billion-parameter models at full 16-bit precision and up to roughly 30 billion parameters at 4-bit quantization, a compression method that reduces memory usage at a small cost to accuracy.
Ollama, a free tool for running open-source AI models on cloud or local hardware, is a natural fit for this card. At the 24 GB tier, choosing a GPU for Ollama means most popular open-weight models run without additional configuration.
The 4090 is also a popular choice for Stable Diffusion (an open-source image generation model) and ComfyUI (a visual workflow tool for running it), where generating images quickly matters more than raw training speed.
RTX A6000 and RTX PRO 6000 pricing
The RTX A6000 starts at $0.28 per hour on Vast.ai, $0.53 on RunPod, and $1.09 on Lambda, as of September 2026. Built on NVIDIA’s older Ampere architecture with 48 GB of VRAM, it keeps NVLink, which lets multiple GPUs share memory directly – a feature the newer L40S drops.
The RTX PRO 6000 (96 GB, built on NVIDIA’s newer Blackwell architecture) is its successor. Hostinger lists it at $0.60 per hour, while RunPod charges $2.09. GCP’s G4 series bundles the same card with 48 vCPUs (virtual CPU cores) and 180 GB of RAM into a single virtual machine from approximately $4.50 per hour.
The memory difference is the decision point. At 48 GB, the A6000 can hold a model of roughly 20 billion parameters at full FP16 precision (16-bit floating point, the standard format for full-accuracy AI model weights) once you leave room for the context window; larger models require quantization to fit.
With 96 GB, the PRO 6000 can handle most 70-billion-parameter models on a single card at 4-bit quantization, without needing to link multiple GPUs.
L40S pricing per hour
The L40S costs $0.92 per hour on Hostinger and $0.99 on RunPod Secure Cloud, as of September 2026. It’s built on NVIDIA’s Ada Lovelace architecture with 48 GB of VRAM and native FP8 support – FP8 is an 8-bit number format that reduces memory usage and can speed up AI inference workloads.
That FP8 support makes it the better option for inference (running a trained AI model to generate outputs) in the 48 GB tier compared to the older A6000.
GCP’s G2 series offers the L4 (24 GB) from $0.71 per hour. The L4 and L40S share the same Ada Lovelace architecture, but the L4 is a low-power inference card with half the VRAM and a fraction of the throughput – it draws 72 W against the L40S’s 350 W. If both appear in GCP’s pricing for an inference workload, the L4 is not a cheaper L40S: it is a much smaller card.
Important
L40S availability on Hostinger fluctuates. If the card doesn’t appear in the deploy picker, it’s simply out of stock, not unavailable in your region. Hostinger shows live inventory before you commit to a deploy, so what you see is what’s actually available.
A100 80GB pricing per hour
The A100 80GB starts at $1.39 per hour on RunPod and $1.43 per hour on Hostinger, rising to $3.67–$5.03 per hour at hyperscalers.
RunPod’s $1.39 rate is the PCIe variant; the SXM variant (faster for multi-GPU training) is $1.59. At hyperscalers, the gap widens sharply: Azure’s NC24ads A100 v4 is $3.67 per hour for a single-GPU VM; the per-GPU rate on AWS’s p4de (8× A100 80GB) works out to roughly $3.43, and GCP’s a2-highgpu-1g (A100 40GB) runs $3.67.
The A100’s 80 GB of high-bandwidth memory (HBM2e) handles fine-tuning (adapting an existing model for a specific task) and inference for most models up to 70 billion parameters in quantized form. At the specialist-cloud rate of $1.39–1.43 per hour, it’s the default step up from the 48 GB tier for training runs that need more memory.
H100 pricing per hour
An H100 costs $2.89 per hour on RunPod and up to $12.29 on Azure, as of September 2026. That is a more than 4x spread between the cheapest specialist cloud and the most expensive hyperscaler for the same chip. Hostinger doesn’t offer an H100; the B200 covers its top tier.
- RunPod: $2.89 (PCIe); $3.29 (SXM); $3.19 (NVL).
- Lambda: $3.29 (PCIe); from $3.99 (SXM; cheapest per-GPU at the 8-GPU config, single GPU available at $4.29).
- Nebius: $3.85.
- DigitalOcean: $4.41.
- AWS p5.4xlarge: $6.88 (single H100 SXM5; no minimum node).
- Azure NC H100 v5: $6.98 (NVL form factor, single GPU).
- GCP A3 (a3-highgpu-8g): from $11.06 per GPU (8-GPU minimum on-demand).
- Azure ND H100 v5: $12.29 per GPU ($98.32 for the 8-GPU node).
PCIe, SXM, and NVL are the form factors the H100 ships in. PCIe fits a standard server slot; SXM uses NVIDIA’s proprietary high-power socket with NVLink interconnects that let multiple GPUs share memory directly, which matters for distributed training (splitting a job across several GPUs in parallel). NVL pairs two cards and carries more memory per GPU.
The SXM premium at specialist clouds is roughly 3–10%. For single-card inference, PCIe is the cheaper option with no performance difference at that scale.
GCP’s A3 on-demand requires a minimum of 8 GPUs per node – if you need two H100s, you still pay for eight. Lambda sells H100 SXM in 1×, 2×, 4×, and 8× configurations.
B200 and H200 pricing
The B200 costs $4.50 per hour on Hostinger, $6.79 on RunPod, from $4.38 on Vast.ai, and from $6.69 on Lambda. The H200 starts at $3.87 on Vast.ai and $4.59 on RunPod. GCP’s A4 series starts around $16.11 per GPU-hour but requires a reservation. There’s no standard on-demand option.
The H200 (141 GB HBM3e, a high-bandwidth memory type) carries roughly 75% more VRAM than the H100; RunPod lists it at $4.59. GCP’s A3 Ultra and Azure’s ND H200 v5 also require reservations and aren’t available on standard on-demand.
Both cards sit above the H100, but in different ways: the B200 is a new architecture with higher throughput, while the H200 keeps the H100’s compute and adds memory. As these cards become more widely available, H100 on-demand rates are likely to keep falling. A100 pricing has already dropped below $1 per hour on some marketplace providers as the market shifts toward newer GPU generations.
Why do hyperscalers charge more for the same GPU?
Hyperscalers (AWS, Google Cloud and Azure) charge significantly more than specialist GPU clouds for on-demand access because they rent you a whole virtual machine, not just the GPU. That package bundles compute, memory, and storage, often with a minimum number of GPUs per rental, plus the compliance certifications and support tooling most AI workloads don’t need.
Forced bundling. AWS’s p5.4xlarge pairs one H100 with 16 vCPUs (virtual CPU cores) and 256 GB of system RAM. Azure’s ND H100 v5 bundles eight GPUs with 96 vCPUs and 1.9 TB of RAM. You can’t rent the GPU on its own – the processors and memory come attached in fixed ratios, whether you need them or not.
On Hostinger, the same B200 GPU is available in tiers from $4.50 to $6.00/hr, a 33% spread from smallest to largest.
Quota lead time. AWS and GCP both cap how many high-end GPU instances a new account can launch – these limits are called quotas, and they default to zero for GPU instances.
Getting AWS P5 access requires a written justification and typically takes 3–7 business days for approval; GCP’s A3 H100 quota can take longer. That wait is a real cost: teams blocked on quota often end up renting more capacity than they need on whatever they can access immediately..
What hyperscalers give you in return. They run high-speed networking between GPU nodes, up to 3.2 Tbps (terabits per second) on AWS P5 instances, which matters for distributed training across dozens of GPUs. Enterprise compliance certifications are built in.
Regional coverage spans every major geography. For regulated industries or workloads that need 64+ GPU clusters, the hyperscaler premium buys infrastructure you can’t assemble elsewhere.
For developers running inference, fine-tuning, or image generation, most of that bundle is unnecessary. Specialist GPU clouds strip it out, and alternatives to RunPod for GPU workloads in the same tier typically come without quota approval or minimum node sizes.
The price gap is most visible when you only need one GPU, where specialist clouds consistently undercut hyperscaler rates.
What you actually pay for cloud GPUs
The hourly GPU rate is only part of the cost: storage, data transfer fees, and idle time can push the real bill up by 20–40%.
Data transfer costs
Moving data out of a provider’s network is free on some platforms and costs up to $0.12 per GB on others, which adds up fast on large workloads. Providers call this egress rate.
|
Provider |
Egress rate |
Notes |
|
AWS |
$0.09/GB |
First 100 GB free per month |
|
GCP |
$0.08–0.12/GB |
Varies by destination and volume |
|
Azure |
~$0.087/GB |
First 5 TB from East US |
|
Vast.ai |
Host-set |
Charged per byte, both directions |
|
Hostinger |
$0 |
No egress fees |
|
RunPod (Pods) |
$0 |
Zero egress on Pod workloads |
|
Lambda |
$0 |
No egress fees |
The difference scales with volume: moving a 140 GB Llama-70B model checkpoint out of AWS costs around $12.60, while moving a 10 TB training dataset costs roughly $900. On Hostinger, RunPod Pods, or Lambda, the same transfers cost nothing.
Storage costs
Cloud GPU storage typically runs $0.04–0.20/GB/month, and on most platforms it keeps billing even after you stop your instance.
- RunPod. $0.10/GB/month while running, $0.20/GB/month when stopped. A 200 GB volume left idle for a month costs $40.
- Vast.ai. Rates are host-set and vary by listing. Storage bills continuously even when stopped, typically at a higher rate than when running. Deleting the instance is the only way to stop the charge.
- Lambda. $0.20/GB/month for persistent file storage, which bills even with no GPU instance attached.
- AWS. $0.08/GB/month (EBS), billed separately on top of compute.
- GCP. $0.04–0.17/GB/month (standard to SSD persistent disk), billed separately on top of compute.
- Azure. $0.17/GB/month (Premium SSD), billed separately on top of compute.
- Hostinger. Included with each instance tier. No separate storage charge.
Idle GPU costs
Idle GPUs bill at the full hourly rate, the same as when they are running a job. A GPU left running after a job finishes costs exactly as much as one that was computing the whole time.
This is true of every provider: a session forgotten overnight costs a full night’s compute everywhere. Billing granularity only changes the rounding at the edges. Hostinger deducts the full first hour at deploy, so a 10-minute session costs the same as a 60-minute one. On RunPod or Vast.ai, that same 10-minute session costs exactly 10 minutes.
How much does Hostinger GPU hosting cost?
Hostinger GPU hosting starts at $0.38 per hour for an RTX 4090 and reaches $7.08 per hour for a dedicated B200, billed hourly in credits with no egress fees, as of September 2026.
|
GPU |
VRAM |
From |
|
RTX 4090 |
24 GB |
$0.38/hr |
|
RTX PRO 6000 (Server) |
96 GB |
$0.60/hr |
|
L40S |
48 GB |
$0.92/hr |
|
A100 80GB PCIe |
80 GB |
$1.43/hr |
|
B200 |
192 GB |
$4.50/hr |
|
B200 (Dedicated) |
192 GB |
$7.08/hr |
Every instance gets a dedicated NVIDIA GPU – the card isn’t shared between customers – with full root and SSH access. The separate B200 (Dedicated) tier goes further, giving you the whole underlying server rather than a share of it. One-click apps for Ollama, ComfyUI, and Jupyter are available at deploy time alongside a base Ubuntu 24.04 option. Deployment takes a few minutes once you select a GPU and instance size.
Each GPU comes in multiple size tiers. The B200’s base tier (Nano) gives you 1 core, 1 GB RAM, and 10 GB storage at 450 credits per hour ($4.50). The largest shared tier (32 cores, 64 GB RAM, 500 GB storage) costs 600 credits per hour ($6.00). The spread across the shared tiers is 33 percent.
Hostinger GPU hosting uses credits purchased in advance. Three pack sizes are available:
- 500 credits. $5.00 ($0.0100/credit, base rate).
- 2,200 credits. $20.00 ($0.0091/credit).
- 6,000 credits. $50.00 ($0.0083/credit).
The full first hour is deducted at deploy. Destroying an instance mid-hour does not return the unused portion.
Hostinger GPU hosting is US-only (Los Angeles and Dallas, region depending on the GPU) and on-demand only. No spot or reserved tiers exist. There is no GPU switching after deploy; changing to a different card requires a new instance and a data migration.
How to estimate your monthly cloud GPU bill
To estimate a monthly cloud GPU bill, multiply your planned hours by the hourly rate, then add storage costs and egress fees. On hyperscalers both are billed separately and grow with the amount of data you move; on Hostinger they add nothing, while RunPod and Lambda charge for storage but not egress.
This example uses a single A100 80GB for LLM (Large Language Model) inference using a model such as Llama or Mistral, 8 hours per day over 30 days (240 GPU-hours), with 10 GB of model weights in storage and 200 GB of data transferred out per month.
|
Cost component |
Hostinger |
RunPod (Secure) |
Azure (NC24ads) |
|
Compute (240 hr) |
240 × $1.43 = $343.20 |
240 × $1.39 = $333.60 |
240 × $3.67 = $880.80 |
|
Storage (10 GB/mo) |
Included |
10 GB × $0.07 = $0.70 |
10 GB × $0.17 = $1.70 |
|
Egress (200 GB/mo) |
$0 |
$0 |
200 GB × $0.087 = $17.40 |
|
Monthly total |
$343.20 |
$334.30 |
$899.90 |
At 8 hours per day, Hostinger and RunPod land within $9 of each other. Azure costs 2.7 times more, with the compute rate driving almost all the difference.
Applying Hostinger’s $50 credit pack (effective $0.0083/credit) reduces the Hostinger total to approximately $285. At that rate, the gap over RunPod widens in Hostinger’s favor for this workload.
For production LLM inference, how you size the GPU to the model matters as much as which provider you pick.
Deploying LLMs on GPU infrastructure with vLLM requires matching GPU memory to model size, which determines whether an A100 or a smaller card is the right choice for a given workload.
Set a shutdown reminder or a budget alert for any long-running instance. On hourly billing, a two-day accidental run costs the same as two days of actual work.
How to choose the right cloud GPU pricing model for your workload
The right cloud GPU pricing model depends on three factors: how long your jobs run, whether they can be interrupted, and how often you need the GPU. For most workloads, that means on-demand access at a specialist cloud.
|
Workload |
Recommended model |
Typical cost |
|
Experiments and development |
On-demand |
From $0.38/hr (RTX 4090) |
|
Multi-hour batch jobs |
Spot / interruptible |
From $0.34/hr (RTX 4090) |
|
Always-on production service |
Reserved |
From $1.36/hr (A100, Azure 3-year) |
If your jobs take several hours and you don’t need to watch them run, spot pricing cuts costs by 50% or more. RunPod’s Spot Pods and Vast.ai’s interruptible tier offer the steepest discounts, though the provider can reclaim the GPU mid-job, so save your progress regularly.
If you need a GPU running continuously for a production AI service with consistent user traffic, on-demand pricing stops making sense. Azure’s 3-year A100 reservation drops from $3.67 to $1.36 per hour, a 63% reduction for a workload that was going to run anyway. Most teams reach this point when they’re serving a live product.
Before picking a provider, check the data transfer rate: moving 10 TB out of GCP or AWS can add $800–1,000 per month. The same transfer is free on Hostinger, RunPod Pods, and Lambda. On data-heavy workloads, the transfer fee can outweigh any savings from a cheaper hourly rate.
One approach worth trying: keep your application and data on a VPS and spin up a GPU only when you need to run something. You pay for the GPU only while it runs.
All of the tutorial content on this website is subject to
Hostinger’s rigorous editorial standards and values.
Apply for Premium Hosting
Source Credit: https://www.hostinger.com/in/tutorials/cloud-gpu-pricing/
