
NVIDIA / Turing
T4
| Provider / plan | Price / GPU-hour | Rental type | Terms & configuration | Last checked | Provider pricing link |
|---|---|---|---|---|---|
| ModalGPU Tasks | $0.5904Lowest flexible rateGPU only · CPU/RAM extra | ServerlessActive GPU time | 1 GPU minimumSee providerView detailsGPU time only; CPU, RAM, storage, region and non-preemptible surcharges are extra.
|
2026-10-03Auto-collected | View pricing ↗ |
Prices are dated listings, not live availability. Compare minimum GPU counts and commitments. Storage, networking and taxes may add charges; serverless rates cover GPU time only.
T4 is a Turing inference accelerator with 16 GB of GDDR6 and a 70 W power envelope. It suits smaller models, video pipelines and development workloads that fit its memory budget. Use precision formats and inference kernels that explicitly support Turing.
CHOOSING THIS GPU
Is it right for your workload?
NVIDIA T4, 16 GB. NVIDIA lists memory bandwidth as 320+ GB/s; 320 GB/s is the reference value shown here.
Best suited to
- Small-model inference
- Video analytics and transcoding
- Quantized development workloads within 16 GB
Strengths
- 70 W power envelope
- FP16 and INT8 Tensor acceleration
- Compact inference-oriented design
Things to consider
- No native BF16 or FP8 Tensor acceleration
- 16 GB sharply limits larger models and context
- New inference kernels may impose additional architecture requirements
HARDWARE DETAILS
Inside the GPU
- Manufacturer
- NVIDIA
- Architecture
- Turing
- VRAM (GB)
- 16
- Memory type
- GDDR6
- Form factor
- PCIe
- Memory bandwidth (GB/s)
- 320
- Maximum board power (W)
- 70
- CUDA cores
- 2560
- Tensor cores
- 320 / 2nd generation
- Ray tracing cores
- 40 / 1st generation
- CUDA capability
- 7.5
- Memory ECC
- Supported
- Host interface
- PCIe Gen3 x16
- GPU interconnect
- Not verified for this edition
- Hardware partitioning (MIG)
- Not verified for this edition
- Cooling
- Passive; server airflow required
Hardware capabilities are not a guarantee of access in a cloud instance. Check the exact edition, GPU allocation and server topology in the rental offer.
MODEL MEMORY
Plan your model size
Start with an 8B-class model in a supported 4-bit format. FP16 8B weights alone are about 16 GB before runtime memory; reserve room for KV cache and activations. Adapter fine-tuning needs a separate training memory estimate.
| Example model class | Weight format | Raw weights ≈ | Minimum GPUs for weights only |
|---|---|---|---|
| Llama 3.1 8B ↗ | 4-bit | 4 GB | 1Runtime needs more memory |
| Llama 3.1 8B ↗ | FP16 | 16 GB | 1Runtime needs more memory |
| Qwen2.5 14B ↗ | 4-bit | 7 GB | 1Runtime needs more memory |
| Qwen2.5 32B ↗ | 4-bit | 16 GB | 1Runtime needs more memory |
| Llama 3.1 70B ↗ | 4-bit | 35 GB | 3Runtime needs more memory |
| Llama 3.1 70B ↗ | 8-bit | 70 GB | 5Runtime needs more memory |
| Llama 3.1 70B ↗ | FP16 | 140 GB | 9Runtime needs more memory |
Arithmetic lower bound: rounded parameter count × bits per weight ÷ 8, in decimal GB. Excludes quantization metadata, KV cache, activations, CUDA workspaces and training states. Actual model sizes differ from their rounded names. No context length, batch size or concurrency is guaranteed. GPU counts assume supported model sharding and can be higher in practice. Quantized checkpoints and compatible kernels are required for 4-bit/8-bit execution.
SOFTWARE & PERFORMANCE
Before you deploy
Software compatibility
Use an NVIDIA CUDA-enabled framework/container compatible with this GPU and the host driver. Check PyTorch build and inference-engine kernel requirements before deployment. vLLM documents NVIDIA compute capability 7.5+ as a baseline; support for each model and quantization method still needs checking. Prefer FP16-compatible or supported quantized kernels; do not select BF16-only or FP8-only execution.
Measured benchmarks
A reproducible, comparable application benchmark has not yet been recorded for this edition. Peak TFLOPS above are not measured tokens per second or image-generation speed.
Compare tests using the same model, precision, GPU count, engine version, input/output length and batch or concurrency. A provider’s server configuration can change the result.
COMPARE YOUR OPTIONS
Alternatives to consider
BUDGET YOUR RUN
GPU cost estimator
GPU compute only: hourly rate × billable hours × GPU count. CPU, RAM, storage, egress, taxes, billing increments and minimum terms may add charges. Serverless hours mean active GPU task time; spot capacity can be interrupted. Reserved commitments and quote-only offers are excluded from this simple hourly estimator.