
NVIDIA / Ampere
RTX 3090
| Provider / plan | Price / GPU-hour | Rental type | Terms & configuration | Last checked | Provider pricing link |
|---|---|---|---|---|---|
| SaladCloudCommunity / batch | $0.09Lowest flexible rateGPU only · CPU/RAM extra | ServerlessActive GPU time | 1 GPU minimumConfirm with providerView detailsPublic USD per GPU-hour; billed per second while running. High priority avoids priority preemption but nodes may disconnect.
|
2026-10-03Auto-collected | View pricing ↗ |
| SaladCloudCommunity / low | $0.1170GPU only · CPU/RAM extra | ServerlessActive GPU time | 1 GPU minimumConfirm with providerView detailsPublic USD per GPU-hour; billed per second while running. High priority avoids priority preemption but nodes may disconnect.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 50484110 | $0.1385 | On-demandNo reservation | 1 GPU minimumCalifornia, USView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 53871134 | $0.1430 | On-demandNo reservation | 1 GPU minimumJiangsu, CNView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| SaladCloudCommunity / medium | $0.1430GPU only · CPU/RAM extra | ServerlessActive GPU time | 1 GPU minimumConfirm with providerView detailsPublic USD per GPU-hour; billed per second while running. High priority avoids priority preemption but nodes may disconnect.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 35599501 | $0.1516 | On-demandNo reservation | 1 GPU minimumQuebec, CAView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 25325627 | $0.1563 | On-demandNo reservation | 1 GPU minimumCzechia, CZView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 41028909 | $0.1696 | On-demandNo reservation | 1 GPU minimumUnited States, USView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| SaladCloudCommunity / high | $0.17GPU only · CPU/RAM extra | ServerlessActive GPU time | 1 GPU minimumConfirm with providerView detailsPublic USD per GPU-hour; billed per second while running. High priority avoids priority preemption but nodes may disconnect.
|
2026-10-03Auto-collected | View pricing ↗ |
| RunpodSecure Cloud Pods | $0.50 | On-demandNo reservation | 1 GPU minimumSee providerView detailsSecure Cloud Pods published GPU-hour rate. CPU/RAM allocation varies; storage/networking extra. Not live capacity.
|
2026-10-03Auto-collected | View pricing ↗ |
Prices are dated listings, not live availability. Compare minimum GPU counts and commitments. Storage, networking and taxes may add charges; serverless rates cover GPU time only.
RTX 3090 is an Ampere GeForce card with 24 GB of GDDR6X. It remains a candidate for development and quantized model experiments when the rental price suits the workload. An NVLink connector is available, but a rental must include the bridge and compatible topology.
CHOOSING THIS GPU
Is it right for your workload?
RTX 3090, not RTX 3090 Ti. Power is 350 W for the reference RTX 3090; add-in-card specifications can vary.
Best suited to
- Development and model prototyping
- Quantized inference within 24 GB
- Rendering and supported two-GPU experiments
Strengths
- 24 GB memory
- Optional two-GPU NVLink bridge
- Ampere CUDA ecosystem
Things to consider
- No native FP8 Tensor acceleration
- 24 GB limits larger models
- Check bridge installation and provider power limits
HARDWARE DETAILS
Inside the GPU
- Manufacturer
- NVIDIA
- Architecture
- Ampere
- VRAM (GB)
- 24
- Memory type
- GDDR6X
- Form factor
- PCIe
- Memory bandwidth (GB/s)
- Not verified for this edition
- Maximum board power (W)
- 350
- CUDA cores
- 10496
- Tensor cores
- 3rd generation
- Ray tracing cores
- 2nd generation
- CUDA capability
- 8.6
- Memory ECC
- Not verified for this edition
- Host interface
- PCIe Gen4
- GPU interconnect
- Two-card bridge supported; verify compute peer access
- Hardware partitioning (MIG)
- Not verified for this edition
- Cooling
- Active; board design varies
Hardware capabilities are not a guarantee of access in a cloud instance. Check the exact edition, GPU allocation and server topology in the rental offer.
MODEL MEMORY
Plan your model size
Start with an 8B-class model in a supported 4-bit format. FP16 8B weights alone are about 16 GB before runtime memory; reserve room for KV cache and activations. Adapter fine-tuning needs a separate training memory estimate.
| Example model class | Weight format | Raw weights ≈ | Minimum GPUs for weights only |
|---|---|---|---|
| Llama 3.1 8B ↗ | 4-bit | 4 GB | 1Runtime needs more memory |
| Llama 3.1 8B ↗ | FP16 | 16 GB | 1Runtime needs more memory |
| Qwen2.5 14B ↗ | 4-bit | 7 GB | 1Runtime needs more memory |
| Qwen2.5 32B ↗ | 4-bit | 16 GB | 1Runtime needs more memory |
| Llama 3.1 70B ↗ | 4-bit | 35 GB | 2Runtime needs more memory |
| Llama 3.1 70B ↗ | 8-bit | 70 GB | 3Runtime needs more memory |
| Llama 3.1 70B ↗ | FP16 | 140 GB | 6Runtime needs more memory |
Arithmetic lower bound: rounded parameter count × bits per weight ÷ 8, in decimal GB. Excludes quantization metadata, KV cache, activations, CUDA workspaces and training states. Actual model sizes differ from their rounded names. No context length, batch size or concurrency is guaranteed. GPU counts assume supported model sharding and can be higher in practice. Quantized checkpoints and compatible kernels are required for 4-bit/8-bit execution.
SOFTWARE & PERFORMANCE
Before you deploy
Software compatibility
Use an NVIDIA CUDA-enabled framework/container compatible with this GPU and the host driver. Check PyTorch build and inference-engine kernel requirements before deployment. vLLM documents NVIDIA compute capability 7.5+ as a baseline; support for each model and quantization method still needs checking.
Measured benchmarks
A reproducible, comparable application benchmark has not yet been recorded for this edition. Peak TFLOPS above are not measured tokens per second or image-generation speed.
Compare tests using the same model, precision, GPU count, engine version, input/output length and batch or concurrency. A provider’s server configuration can change the result.
COMPARE YOUR OPTIONS
Alternatives to consider
BUDGET YOUR RUN
GPU cost estimator
GPU compute only: hourly rate × billable hours × GPU count. CPU, RAM, storage, egress, taxes, billing increments and minimum terms may add charges. Serverless hours mean active GPU task time; spot capacity can be interrupted. Reserved commitments and quote-only offers are excluded from this simple hourly estimator.