nvidia-logo

Nvidia A40

Ampere-based data center GPU designed for AI inference, virtual workstations, and professional visualization workloads.

Release

2020

GPU Class

Data Center / Visualization & AI Accelerator

Architecture

Ampere

PRICE SNAPSHOT

Loading price comparison...
Loading live GPU prices...

On-premise Module

~$4k–$8k*

Turnkey System

~$20k–$80k†

Cloud Pricing
(per GPU/hr)

~$0.70–$3.00/hr‑

chip identity

Nvidia A40

On-premise Module

A40

GPU Class

Data Center / Visualization & AI Accelerator

Release

2020

Architecture

Ampere

Target Workload

  • AI inference workloads
  • Virtual desktop infrastructure (VDI)
  • 3D rendering and visualization
  • Computer vision pipelines
  • Simulation and engineering workloads

Compatible Platforms

  • NVIDIA A40 PCIe accelerator cards
  • OEM GPU servers (Dell, HPE, Lenovo, Supermicro)
  • GPU virtualization infrastructure for enterprise workloads

Interconnect: PCIe Gen4

Software Stack: CUDA, TensorRT, NVIDIA AI Enterprise, NVIDIA vGPU

Ideal Buyer Profile

Organizations that commonly deploy A40 GPUs include:

  • Enterprises running GPU-accelerated virtual desktops
  • Engineering firms performing simulation and rendering workloads
  • Companies deploying AI inference pipelines
  • Cloud providers offering GPU virtualization services

The GPU is especially attractive for buyers seeking a versatile accelerator capable of supporting multiple enterprise workloads.

Availability Notes

The A40 is widely available through OEM server vendors and enterprise infrastructure providers. It is commonly deployed in environments that require GPU virtualization or a combination of AI and graphics workloads.

Because it has been on the market for several years, the A40 is often available at competitive pricing in enterprise GPU marketplaces.

Recent Developments

  • Continued use of A40 GPUs in enterprise virtualization environments.
  • Increasing migration toward newer Ada-based GPUs such as the L40S.
  • Ongoing support through NVIDIA’s enterprise software ecosystem.

overview

The NVIDIA A40 is a data center GPU built on the Ampere architecture, designed to support a combination of AI acceleration, professional visualization, and virtual workstation workloads. It provides a flexible compute platform for enterprise environments that require GPU acceleration across multiple types of workloads.

The accelerator includes second-generation Tensor Cores, enabling acceleration of deep learning inference tasks and machine learning workflows. With 48 GB of GDDR6 memory, the A40 can support larger datasets and complex models compared to many earlier data center GPUs.

Unlike many training-focused accelerators, the A40 emphasizes versatility, making it suitable for AI inference, 3D rendering, engineering simulations, and GPU-accelerated virtualization environments.

Key specifications

A40 GPU Specifications

A40 GPU Specifications

Specification Value
Architecture NVIDIA Ampere
CUDA Cores 10,752
Tensor Cores 336 (2nd Gen)
Memory 48 GB GDDR6
Memory Bandwidth ~696 GB/s
Interconnect PCIe Gen4
Form Factor PCIe
Max TGP ~300 W
Precision Support FP64, FP32, TF32, FP16, BF16, INT8
Typical AI Compute ~150 TFLOPS (FP16 Tensor)
Process Node Samsung 8 nm
Transistor Count ~28 Billion
MIG Support Not Supported
NVLink (Peer) Supported (2-GPU bridge)

Performance Summary

  • AI Inference: The A40 provides strong inference performance using Tensor Core acceleration.
  • Visualization Workloads: The GPU is optimized for rendering and professional visualization environments.
  • Memory Capacity: With 48 GB of memory, the A40 supports larger datasets and moderately sized AI models.
  • Virtualization: Support for NVIDIA vGPU technology enables GPU sharing across multiple virtual machines.

Compared with newer GPUs such as the L40S, the A40 delivers lower AI performance but remains widely deployed due to its versatility and enterprise software ecosystem.

primary use case

  • AI inference pipelines
  • Virtual workstation deployments
  • 3D rendering and visualization
  • Computer vision workloads
  • Engineering simulations

The GPU is particularly suited for organizations requiring GPU acceleration across multiple enterprise workloads.

Alternatives & Upgrade Path

Comparable NVIDIA GPUs:

  • A10G: Ampere GPU commonly used for inference and cloud workloads.
  • L40S: Newer Ada Lovelace GPU offering higher AI performance.
  • RTX 6000 Ada: Workstation GPU with similar memory capacity.

Competitor Accelerators:

  • AMD data center GPUs
  • Intel AI accelerators

These alternatives provide different tradeoffs between AI performance, memory capacity, and enterprise features.

Related Chips & Providers

Related NVIDIA GPUs:

  • A10G
  • L40S
  • RTX 6000 Ada

Complementary Infrastructure:

  • NVIDIA vGPU virtualization platform
  • CUDA software ecosystem

SUMMARY

The NVIDIA A40 is a versatile Ampere-based data center GPU designed to support a wide range of enterprise workloads, including AI inference, professional visualization, and virtual workstation environments. With 48 GB of GDDR6 memory and Tensor Core acceleration, it provides a flexible platform for organizations deploying GPU-accelerated infrastructure.

Although newer architectures offer improved AI performance, the A40 remains widely used due to its balance of compute capability, memory capacity, and enterprise virtualization support.

The next important GPU to add is NVIDIA T4.

Why this one matters:

  • One of the most widely deployed inference GPUs ever
  • Extremely common in Google Cloud, AWS, and enterprise inference servers
  • Frequently compared with L4 and A10G
  • Huge installed base across AI inference and video processing workloads

Below is the page written in the same structure as your other GPU pages.

Newsletter

Stay Ahead in Cloud
& Data Infrastructure

Get early access to new tools, insights, and research shaping the next wave of cloud and storage innovation.