NVIDIA / Blackwell
B200
| Provider / plan | Price / GPU-hour | Rental type | Terms & configuration | Last checked | Provider pricing link |
|---|---|---|---|---|---|
| ModalGPU Tasks | $6.2496Lowest flexible rateGPU only · CPU/RAM extra | ServerlessActive GPU time | 1 GPU minimumSee providerView detailsGPU time only; CPU, RAM, storage, region and non-preemptible surcharges are extra.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 50255933 | $6.2521 | On-demandNo reservation | 1 GPU minimumOregon, USView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| RunpodSecure Cloud Pods | $6.79 | On-demandNo reservation | 1 GPU minimumSee providerView detailsSecure Cloud Pods published GPU-hour rate. CPU/RAM allocation varies; storage/networking extra. Not live capacity.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 53976631 | $7.5104 | On-demandNo reservation | 1 GPU minimumVirginia, USView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 53384349 | $7.8146 | On-demandNo reservation | 1 GPU minimum, USView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 53943901 | $8.6222 | On-demandNo reservation | 1 GPU minimumMaryland, USView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 51061898 | $9.3854 | On-demandNo reservation | 1 GPU minimum, USView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
Prices are dated listings, not live availability. Compare minimum GPU counts and commitments. Storage, networking and taxes may add charges; serverless rates cover GPU time only.
B200 is a Blackwell data-center accelerator for large-model inference and distributed training. This profile covers the 180 GB HGX configuration. Its large HBM capacity and high memory bandwidth make it a candidate when smaller GPUs require model sharding.
CHOOSING THIS GPU
Is it right for your workload?
180 GB per GPU in the cited HGX B200 configuration; this is not the memory of an entire eight-GPU server. Power is configurable up to 1,000 W per GPU in that reference platform.
Best suited to
- Large language model serving with large weight and KV-cache footprints
- Distributed training on an NVLink-connected HGX platform
- Memory-intensive AI workloads that exceed an 80 GB accelerator
Strengths
- 180 GB HBM3e per GPU
- Up to 8 TB/s local memory bandwidth
- Fifth-generation NVLink for supported multi-GPU systems
Things to consider
- Requires a compatible data-center platform
- The provider must expose the required NVLink topology
- Use a software image that explicitly supports Blackwell
HARDWARE DETAILS
Inside the GPU
- Manufacturer
- NVIDIA
- Architecture
- Blackwell
- VRAM (GB)
- 180
- Memory type
- HBM3e
- Form factor
- SXM
- Memory bandwidth (GB/s)
- 8000
- Maximum board power (W)
- 1000
- CUDA cores
- Not verified for this edition
- Tensor cores
- 5th generation
- Ray tracing cores
- Not verified for this edition
- CUDA capability
- 10.0
- Memory ECC
- Not verified for this edition
- Host interface
- PCIe Gen5
- GPU interconnect
- 5th generation; up to 1,800 GB/s per GPU (aggregate bidirectional)
- Hardware partitioning (MIG)
- Not verified for this edition
- Cooling
- Server platform dependent
Hardware capabilities are not a guarantee of access in a cloud instance. Check the exact edition, GPU allocation and server topology in the rental offer.
MODEL MEMORY
Plan your model size
70B FP16 weights alone are about 140 GB. Allow additional room for KV cache, activations and the runtime; 141 GB nominal memory is not enough evidence of a practical single-GPU fit. Use quantization or model sharding when needed.
| Example model class | Weight format | Raw weights ≈ | Minimum GPUs for weights only |
|---|---|---|---|
| Llama 3.1 8B ↗ | 4-bit | 4 GB | 1Runtime needs more memory |
| Llama 3.1 8B ↗ | FP16 | 16 GB | 1Runtime needs more memory |
| Qwen2.5 14B ↗ | 4-bit | 7 GB | 1Runtime needs more memory |
| Qwen2.5 32B ↗ | 4-bit | 16 GB | 1Runtime needs more memory |
| Llama 3.1 70B ↗ | 4-bit | 35 GB | 1Runtime needs more memory |
| Llama 3.1 70B ↗ | 8-bit | 70 GB | 1Runtime needs more memory |
| Llama 3.1 70B ↗ | FP16 | 140 GB | 1Runtime needs more memory |
Arithmetic lower bound: rounded parameter count × bits per weight ÷ 8, in decimal GB. Excludes quantization metadata, KV cache, activations, CUDA workspaces and training states. Actual model sizes differ from their rounded names. No context length, batch size or concurrency is guaranteed. GPU counts assume supported model sharding and can be higher in practice. Quantized checkpoints and compatible kernels are required for 4-bit/8-bit execution.
SOFTWARE & PERFORMANCE
Before you deploy
Software compatibility
Use an NVIDIA CUDA-enabled framework/container compatible with this GPU and the host driver. Check PyTorch build and inference-engine kernel requirements before deployment. vLLM documents NVIDIA compute capability 7.5+ as a baseline; support for each model and quantization method still needs checking. Select builds that explicitly support Blackwell; a container that runs on Hopper may not include the required kernels.
Measured benchmarks
A reproducible, comparable application benchmark has not yet been recorded for this edition. Peak TFLOPS above are not measured tokens per second or image-generation speed.
Compare tests using the same model, precision, GPU count, engine version, input/output length and batch or concurrency. A provider’s server configuration can change the result.
COMPARE YOUR OPTIONS
Alternatives to consider
BUDGET YOUR RUN
GPU cost estimator
GPU compute only: hourly rate × billable hours × GPU count. CPU, RAM, storage, egress, taxes, billing increments and minimum terms may add charges. Serverless hours mean active GPU task time; spot capacity can be interrupted. Reserved commitments and quote-only offers are excluded from this simple hourly estimator.