
NVIDIA / Ada Lovelace
L40S
| Provider / plan | Price / GPU-hour | Rental type | Terms & configuration | Last checked | Provider pricing link |
|---|---|---|---|---|---|
| CloudRift3-month reservation | $0.54 | ReservedCommitment required | 1 GPU minimumSee provider3 months, upfrontView detailsPer-GPU published list rate; capacity and exact server configuration must be checked with the provider.
|
2026-10-03Auto-collected | View pricing ↗ |
| CloudRift1-month reservation | $0.57 | ReservedCommitment required | 1 GPU minimumSee provider1 month, upfrontView detailsPer-GPU published list rate; capacity and exact server configuration must be checked with the provider.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 52429317 | $0.6156Lowest flexible rate | On-demandNo reservation | 1 GPU minimumQuebec, CAView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 38574939 | $0.8015 | On-demandNo reservation | 1 GPU minimumTaiwan, TWView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 38575013 | $0.8015 | On-demandNo reservation | 1 GPU minimumTaiwan, TWView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| Vast.aiOn-demand offer 38575675 | $0.8015 | On-demandNo reservation | 1 GPU minimumTaiwan, TWView detailsObserved marketplace offer, not a provider-wide list rate. Total hourly quote for one GPU with the API default disk allocation; bandwidth and additional storage can cost extra. Up to 5 cheapest matching offers per catalog GPU; availability can change.
|
2026-10-03Auto-collected | View pricing ↗ |
| RunpodSecure Cloud Pods | $1.09 | On-demandNo reservation | 1 GPU minimumSee providerView detailsSecure Cloud Pods published GPU-hour rate. CPU/RAM allocation varies; storage/networking extra. Not live capacity.
|
2026-10-03Auto-collected | View pricing ↗ |
| ModalGPU Tasks | $1.9512GPU only · CPU/RAM extra | ServerlessActive GPU time | 1 GPU minimumSee providerView detailsGPU time only; CPU, RAM, storage, region and non-preemptible surcharges are extra.
|
2026-10-03Auto-collected | View pricing ↗ |
| CloudRiftOn-demand quote | Contact sales | On-demandNo reservation | 1 GPU minimumSee providerView detailsNo public on-demand price. Contact provider.
|
2026-09-28Manually checked | Request quote ↗ |
Prices are dated listings, not live availability. Compare minimum GPU counts and commitments. Storage, networking and taxes may add charges; serverless rates cover GPU time only.
L40S is a 48 GB Ada data-center GPU that combines AI compute with ray tracing and video engines. It is a candidate for inference, adapter fine-tuning and rendering on the same hardware. It does not provide NVLink or MIG.
CHOOSING THIS GPU
Is it right for your workload?
L40S is distinct from L40: identical nominal VRAM does not imply identical Tensor throughput or power limits.
Best suited to
- LLM inference and adapter fine-tuning
- Image generation and rendering
- Mixed graphics, video and AI services
Strengths
- 48 GB GDDR6 with ECC
- Dedicated RT cores and AV1 media engines
- Supports FP8 Tensor operations
Things to consider
- No NVLink for GPU-to-GPU scaling
- No MIG hardware partitioning
- 350 W passive card needs suitable server cooling
HARDWARE DETAILS
Inside the GPU
- Manufacturer
- NVIDIA
- Architecture
- Ada Lovelace
- VRAM (GB)
- 48
- Memory type
- GDDR6
- Form factor
- PCIe
- Memory bandwidth (GB/s)
- 864
- Maximum board power (W)
- 350
- CUDA cores
- 18176
- Tensor cores
- 568 / 4th generation
- Ray tracing cores
- 142 / 3rd generation
- CUDA capability
- 8.9
- Memory ECC
- Supported
- Host interface
- PCIe Gen4 x16
- GPU interconnect
- Not supported
- Hardware partitioning (MIG)
- Not supported
- Cooling
- Passive; server airflow required
Hardware capabilities are not a guarantee of access in a cloud instance. Check the exact edition, GPU allocation and server topology in the rental offer.
MODEL MEMORY
Plan your model size
Compare 8B/14B FP16 and larger quantized models using the weight-only estimates below. A 70B 4-bit model has about 35 GB of raw weights before quantization metadata and runtime memory; nominal capacity alone does not establish a fit.
| Example model class | Weight format | Raw weights ≈ | Minimum GPUs for weights only |
|---|---|---|---|
| Llama 3.1 8B ↗ | 4-bit | 4 GB | 1Runtime needs more memory |
| Llama 3.1 8B ↗ | FP16 | 16 GB | 1Runtime needs more memory |
| Qwen2.5 14B ↗ | 4-bit | 7 GB | 1Runtime needs more memory |
| Qwen2.5 32B ↗ | 4-bit | 16 GB | 1Runtime needs more memory |
| Llama 3.1 70B ↗ | 4-bit | 35 GB | 1Runtime needs more memory |
| Llama 3.1 70B ↗ | 8-bit | 70 GB | 2Runtime needs more memory |
| Llama 3.1 70B ↗ | FP16 | 140 GB | 3Runtime needs more memory |
Arithmetic lower bound: rounded parameter count × bits per weight ÷ 8, in decimal GB. Excludes quantization metadata, KV cache, activations, CUDA workspaces and training states. Actual model sizes differ from their rounded names. No context length, batch size or concurrency is guaranteed. GPU counts assume supported model sharding and can be higher in practice. Quantized checkpoints and compatible kernels are required for 4-bit/8-bit execution.
SOFTWARE & PERFORMANCE
Before you deploy
Software compatibility
Use an NVIDIA CUDA-enabled framework/container compatible with this GPU and the host driver. Check PyTorch build and inference-engine kernel requirements before deployment. vLLM documents NVIDIA compute capability 7.5+ as a baseline; support for each model and quantization method still needs checking.
Measured benchmarks
A reproducible, comparable application benchmark has not yet been recorded for this edition. Peak TFLOPS above are not measured tokens per second or image-generation speed.
Compare tests using the same model, precision, GPU count, engine version, input/output length and batch or concurrency. A provider’s server configuration can change the result.
COMPARE YOUR OPTIONS
Alternatives to consider
BUDGET YOUR RUN
GPU cost estimator
GPU compute only: hourly rate × billable hours × GPU count. CPU, RAM, storage, egress, taxes, billing increments and minimum terms may add charges. Serverless hours mean active GPU task time; spot capacity can be interrupted. Reserved commitments and quote-only offers are excluded from this simple hourly estimator.