
NVIDIA / Hopper
H100 NVL
| Provider / plan | Price / GPU-hour | Rental type | Terms & configuration | Last checked | Provider pricing link |
|---|---|---|---|---|---|
| Vast.aiOn-demand offer 13527066 | $2.5476Lowest flexible rate | On-demandNo reservation | 1 GPU minimumJapan, JPView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 35408170 | $2.5876 | On-demandNo reservation | 1 GPU minimumTexas, USView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 24548924 | $2.6015 | On-demandNo reservation | 1 GPU minimumBulgaria, BGView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 54024228 | $2.9222 | On-demandNo reservation | 1 GPU minimumIreland, IEView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 54019006 | $3.0156 | On-demandNo reservation | 1 GPU minimumIowa, USView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| RunpodSecure Cloud Pods | $3.19 | On-demandNo reservation | 1 GPU minimumSee providerView detailsSecure Cloud Pods published GPU-hour rate. CPU/RAM allocation varies; storage/networking extra. Not live capacity.
|
2026-10-03Auto-collected | View pricing ↗ |
Prices are dated listings, not live availability. Compare minimum GPU counts and commitments. Storage, networking and taxes may add charges; serverless rates cover GPU time only.
H100 NVL is a Hopper PCIe configuration with 94 GB of memory per GPU. A two-GPU NVLink-connected pair has 188 GB in total, but using that capacity requires software that distributes the model. This profile and rental price unit are per GPU.
CHOOSING THIS GPU
Is it right for your workload?
94 GB per GPU, not 188 GB per card. NVIDIA lists 350–400 W configurable power. Confirm whether the offer supplies one card or a bridged pair.
Best suited to
- Memory-intensive inference
- Models distributed across an NVLink-connected pair
- Hopper deployments requiring a PCIe form factor
Strengths
- 94 GB HBM3 per GPU
- 3.9 TB/s memory bandwidth
- NVLink bridge support
Things to consider
- A two-card memory total is not a single automatic memory pool
- Per-GPU pricing can differ from the minimum pair cost
- NVL peak rates differ from H100 SXM
HARDWARE DETAILS
Inside the GPU
- Manufacturer
- NVIDIA
- Architecture
- Hopper
- VRAM (GB)
- 94
- Memory type
- HBM3
- Form factor
- PCIe / NVL
- Memory bandwidth (GB/s)
- 3900
- Maximum board power (W)
- 400
- CUDA cores
- Not verified for this edition
- Tensor cores
- 4th generation
- Ray tracing cores
- Not verified for this edition
- CUDA capability
- 9.0
- Memory ECC
- Not verified for this edition
- Host interface
- PCIe Gen5
- GPU interconnect
- 600 GB/s aggregate bidirectional; bridge required
- Hardware partitioning (MIG)
- Up to 7 hardware instances; provider must enable it
- Cooling
- Passive; server airflow required
Hardware capabilities are not a guarantee of access in a cloud instance. Check the exact edition, GPU allocation and server topology in the rental offer.
MODEL MEMORY
Plan your model size
70B 8-bit weights alone are about 70 GB; a full runtime needs additional memory. Compare quantized single-GPU serving with supported multi-GPU execution. Full fine-tuning requires substantially more than inference weights.
| Example model class | Weight format | Raw weights ≈ | Minimum GPUs for weights only |
|---|---|---|---|
| Llama 3.1 8B ↗ | 4-bit | 4 GB | 1Runtime needs more memory |
| Llama 3.1 8B ↗ | FP16 | 16 GB | 1Runtime needs more memory |
| Qwen2.5 14B ↗ | 4-bit | 7 GB | 1Runtime needs more memory |
| Qwen2.5 32B ↗ | 4-bit | 16 GB | 1Runtime needs more memory |
| Llama 3.1 70B ↗ | 4-bit | 35 GB | 1Runtime needs more memory |
| Llama 3.1 70B ↗ | 8-bit | 70 GB | 1Runtime needs more memory |
| Llama 3.1 70B ↗ | FP16 | 140 GB | 2Runtime needs more memory |
Arithmetic lower bound: rounded parameter count × bits per weight ÷ 8, in decimal GB. Excludes quantization metadata, KV cache, activations, CUDA workspaces and training states. Actual model sizes differ from their rounded names. No context length, batch size or concurrency is guaranteed. GPU counts assume supported model sharding and can be higher in practice. Quantized checkpoints and compatible kernels are required for 4-bit/8-bit execution.
SOFTWARE & PERFORMANCE
Before you deploy
Software compatibility
Use an NVIDIA CUDA-enabled framework/container compatible with this GPU and the host driver. Check PyTorch build and inference-engine kernel requirements before deployment. vLLM documents NVIDIA compute capability 7.5+ as a baseline; support for each model and quantization method still needs checking.
Measured benchmarks
A reproducible, comparable application benchmark has not yet been recorded for this edition. Peak TFLOPS above are not measured tokens per second or image-generation speed.
Compare tests using the same model, precision, GPU count, engine version, input/output length and batch or concurrency. A provider’s server configuration can change the result.
COMPARE YOUR OPTIONS
Alternatives to consider
BUDGET YOUR RUN
GPU cost estimator
GPU compute only: hourly rate × billable hours × GPU count. CPU, RAM, storage, egress, taxes, billing increments and minimum terms may add charges. Serverless hours mean active GPU task time; spot capacity can be interrupted. Reserved commitments and quote-only offers are excluded from this simple hourly estimator.