NVIDIAVendor documented
A100 80GB SXM
Ampere · Ampere · SXM4 · 2020
The prior-generation standard. Still highly capable and widely available; lacks FP8 so newer inference kernels are less efficient.
LLM trainingLLM inferenceFine-tuningHPCVision
Precision fingerprint
6432t3216BF168i8
Memory
80 GB
HBM2e
Bandwidth
2.039 TB/s
peak
TDP
400 W
air or liquid
Max model
~30B
FP16, planning est.
Compute throughput
| FP64 | 19.5 TFLOPS |
| FP32 | 19.5 TFLOPS |
| TF32 | 156 TFLOPS |
| FP16 | 312 TFLOPS |
| BF16 | 312 TFLOPS |
| FP8 | — |
| INT8 | 624 TFLOPS |
Platform & software
InterconnectNVLink 3 — 600 GB/s
PCIePCIe 4.0 x16
Coolingair or liquid
MIGSupported
PartitioningUp to 7× MIG
VirtualizationvGPU, MIG
FrameworksCUDA, TensorRT, Triton, vLLM
AvailabilityAWS, Azure, GCP, OCI, bare-metal
Known limitations
- ·No FP8 / Transformer Engine — slower on modern LLM stacks
- ·Prior generation, but broadly available and cost-effective