
NVIDIA / Blackwell
RTX 5090
| Provider / plan | Price / GPU-hour | Rental type | Terms & configuration | Last checked | Provider pricing link |
|---|---|---|---|---|---|
| SaladCloudCommunity / batch | $0.25Lowest flexible rateGPU only · CPU/RAM extra | ServerlessActive GPU time | 1 GPU minimumConfirm with providerView detailsPublic USD per GPU-hour; billed per second while running. High priority avoids priority preemption but nodes may disconnect.
|
2026-10-03Auto-collected | View pricing ↗ |
| SaladCloudCommunity / low | $0.3330GPU only · CPU/RAM extra | ServerlessActive GPU time | 1 GPU minimumConfirm with providerView detailsPublic USD per GPU-hour; billed per second while running. High priority avoids priority preemption but nodes may disconnect.
|
2026-10-03Auto-collected | View pricing ↗ |
| SaladCloudCommunity / medium | $0.4170GPU only · CPU/RAM extra | ServerlessActive GPU time | 1 GPU minimumConfirm with providerView detailsPublic USD per GPU-hour; billed per second while running. High priority avoids priority preemption but nodes may disconnect.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 44173830 | $0.4704 | On-demandNo reservation | 1 GPU minimumSouth Korea, KRView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 45669124 | $0.4704 | On-demandNo reservation | 1 GPU minimumSouth Korea, KRView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 49024423 | $0.4704 | On-demandNo reservation | 1 GPU minimumSouth Korea, KRView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 53294509 | $0.4704 | On-demandNo reservation | 1 GPU minimumSouth Korea, KRView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 46753302 | $0.4763 | On-demandNo reservation | 1 GPU minimumUnited States, USView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| SaladCloudCommunity / high | $0.50GPU only · CPU/RAM extra | ServerlessActive GPU time | 1 GPU minimumConfirm with providerView detailsPublic USD per GPU-hour; billed per second while running. High priority avoids priority preemption but nodes may disconnect.
|
2026-10-03Auto-collected | View pricing ↗ |
| CloudRift3-month reservation | $0.51 | ReservedCommitment required | 1 GPU minimumSee provider3 months, upfrontView detailsPer-GPU published list rate; capacity and exact server configuration must be checked with the provider.
|
2026-10-03Auto-collected | View pricing ↗ |
| CloudRift1-month reservation | $0.54 | ReservedCommitment required | 1 GPU minimumSee provider1 month, upfrontView detailsPer-GPU published list rate; capacity and exact server configuration must be checked with the provider.
|
2026-10-03Auto-collected | View pricing ↗ |
| CloudRiftOn-demand | $0.60 | On-demandNo reservation | 1 GPU minimumSee providernoneView detailsPer-GPU published list rate; capacity and exact server configuration must be checked with the provider.
|
2026-10-03Auto-collected | View pricing ↗ |
| RunpodSecure Cloud Pods | $0.99 | On-demandNo reservation | 1 GPU minimumSee providerView detailsSecure Cloud Pods published GPU-hour rate. CPU/RAM allocation varies; storage/networking extra. Not live capacity.
|
2026-10-03Auto-collected | View pricing ↗ |
Prices are dated listings, not live availability. Compare minimum GPU counts and commitments. Storage, networking and taxes may add charges; serverless rates cover GPU time only.
RTX 5090 is a Blackwell GeForce GPU with 32 GB of GDDR7. It suits single-GPU development, image generation and models that benefit from more memory than a 24 GB card. Choose a rental image with explicit support for this GPU generation.
CHOOSING THIS GPU
Is it right for your workload?
32 GB GeForce RTX 5090 reference configuration. Board-vendor cooling, clock and power settings may differ; this is not an RTX PRO model.
Best suited to
- Single-GPU image generation
- Quantized language-model inference
- Rendering and video creation
Strengths
- 32 GB GDDR7
- Blackwell Tensor and RT hardware
- AV1 encode and decode support
Things to consider
- No NVLink
- 575 W reference graphics power
- 32 GB remains a limit for larger models and high concurrency
HARDWARE DETAILS
Inside the GPU
- Manufacturer
- NVIDIA
- Architecture
- Blackwell
- VRAM (GB)
- 32
- Memory type
- GDDR7
- Form factor
- PCIe
- Memory bandwidth (GB/s)
- Not verified for this edition
- Maximum board power (W)
- 575
- CUDA cores
- 21760
- Tensor cores
- 5th generation
- Ray tracing cores
- 4th generation
- CUDA capability
- 12.0
- Memory ECC
- Not verified for this edition
- Host interface
- PCIe Gen5
- GPU interconnect
- Not supported
- Hardware partitioning (MIG)
- Not verified for this edition
- Cooling
- Active; board design varies
Hardware capabilities are not a guarantee of access in a cloud instance. Check the exact edition, GPU allocation and server topology in the rental offer.
MODEL MEMORY
Plan your model size
Compare 8B/14B FP16 and larger quantized models using the weight-only estimates below. A 70B 4-bit model has about 35 GB of raw weights before quantization metadata and runtime memory; nominal capacity alone does not establish a fit.
| Example model class | Weight format | Raw weights ≈ | Minimum GPUs for weights only |
|---|---|---|---|
| Llama 3.1 8B ↗ | 4-bit | 4 GB | 1Runtime needs more memory |
| Llama 3.1 8B ↗ | FP16 | 16 GB | 1Runtime needs more memory |
| Qwen2.5 14B ↗ | 4-bit | 7 GB | 1Runtime needs more memory |
| Qwen2.5 32B ↗ | 4-bit | 16 GB | 1Runtime needs more memory |
| Llama 3.1 70B ↗ | 4-bit | 35 GB | 2Runtime needs more memory |
| Llama 3.1 70B ↗ | 8-bit | 70 GB | 3Runtime needs more memory |
| Llama 3.1 70B ↗ | FP16 | 140 GB | 5Runtime needs more memory |
Arithmetic lower bound: rounded parameter count × bits per weight ÷ 8, in decimal GB. Excludes quantization metadata, KV cache, activations, CUDA workspaces and training states. Actual model sizes differ from their rounded names. No context length, batch size or concurrency is guaranteed. GPU counts assume supported model sharding and can be higher in practice. Quantized checkpoints and compatible kernels are required for 4-bit/8-bit execution.
SOFTWARE & PERFORMANCE
Before you deploy
Software compatibility
Use an NVIDIA CUDA-enabled framework/container compatible with this GPU and the host driver. Check PyTorch build and inference-engine kernel requirements before deployment. vLLM documents NVIDIA compute capability 7.5+ as a baseline; support for each model and quantization method still needs checking. Select builds that explicitly support Blackwell; a container that runs on Hopper may not include the required kernels.
Measured benchmarks
A reproducible, comparable application benchmark has not yet been recorded for this edition. Peak TFLOPS above are not measured tokens per second or image-generation speed.
Compare tests using the same model, precision, GPU count, engine version, input/output length and batch or concurrency. A provider’s server configuration can change the result.
COMPARE YOUR OPTIONS
Alternatives to consider
BUDGET YOUR RUN
GPU cost estimator
GPU compute only: hourly rate × billable hours × GPU count. CPU, RAM, storage, egress, taxes, billing increments and minimum terms may add charges. Serverless hours mean active GPU task time; spot capacity can be interrupted. Reserved commitments and quote-only offers are excluded from this simple hourly estimator.