Enterprises and organizations requiring:
Typical adopters include cloud providers offering alternative AI instances, enterprises building private AI clusters on Ethernet, and teams optimizing for throughput-per-dollar at multi-node scale.
Gaudi 3 is delivered through OEM platforms and select cloud providers rather than as a turnkey Intel-branded system. It is commonly deployed in 8-accelerator OAM baseboards for dense data center configurations and as PCIe add-in cards for broader server compatibility. Availability and pricing vary by OEM integration, cooling configuration, and volume, with procurement often tied to complete platform builds rather than single-card retail channels.
Gaudi 3 is delivered through OEM platforms and select cloud providers rather than as a turnkey Intel-branded system. It is commonly deployed in 8-accelerator OAM baseboards for dense data center configurations and as PCIe add-in cards for broader server compatibility. Availability and pricing vary by OEM integration, cooling configuration, and volume, with procurement often tied to complete platform builds rather than single-card retail channels.
The Intel Gaudi 3 is a data center AI accelerator built for scalable training and inference with an architecture optimized around matrix engines, large on-package high-bandwidth memory, and integrated high-speed Ethernet networking. It targets generative AI workloads by combining strong BF16/FP8 performance with native scale-out connectivity, aiming to reduce reliance on proprietary GPU fabrics for multi-node scaling.
Gaudi 3 is delivered primarily as an OCP OAM 2.0 mezzanine module for dense 8-accelerator baseboards and as a PCIe add-in card for broader server compatibility. Its integrated 200GbE RDMA ports enable direct accelerator-to-accelerator networking, making it well suited to Ethernet-based AI clusters where cost, openness, and scale-out deployment patterns are primary considerations.
| Specification | Gaudi 3 AI Accelerator |
|---|---|
| Architecture | Intel Gaudi (5th-gen heterogeneous architecture) |
| Memory | 128Β GB HBM2e |
| Memory Bandwidth | ~3.7Β TB/s |
| Interconnect | 24xΒ 200GbE RDMA (RoCEΒ v2) integrated networking (platform-dependent usage) |
| Form Factor | OCP OAMΒ 2.0 mezzanine (HL-325L), PCIe Gen5Β x16 add-in card (HL-338) |
| Max TGP | Up to ~900Β W (OAM air-cooled); higher with liquid-cooled configurations; up to ~600Β W (PCIe) |
| Precision Support | FP32, TF32-equivalent workflows (framework-dependent), FP16, BF16, FP8, INT8, FP64 (workload-dependent support) |
| Typical AI Compute | ~1.8Β PFLOPS (FP8/BF16 peak, vendor-claimed) |
| Process Node | TSMCΒ 5Β nm |
| Transistor Count | TBD |
| MIG Support | TBD (no NVIDIA MIG equivalent; partitioning/virtualization is platform and software dependent) |
| NVLink (peer) | Not supported |
Compared to NVIDIA H100-class systems, Gaudi 3βs positioning emphasizes scale-out economics and open networking alongside competitive LLM throughput on supported frameworks, while ecosystem maturity and software portability remain key decision factors.
Comparable AI accelerators:
For organizations prioritizing open networking and scale-out economics, Gaudi 3 is typically evaluated against H100/H200-class deployments on a cost-per-throughput basis. Buyers requiring maximal software portability or tightly coupled multi-GPU scale-up may favor NVIDIA platforms, while those optimizing for Ethernet-native clusters may view Gaudi 3 as a scalable alternative.
Intel accelerators:
Competitor GPUs:
Intel Gaudi 3 is a data center AI accelerator optimized for cost-efficient LLM training and inference with an Ethernet-native scale-out model. It combines 128 GB of HBM2e, ~3.7 TB/s memory bandwidth, and integrated 24x 200GbE RDMA networking to support large distributed deployments without a proprietary GPU fabric. Gaudi 3 is typically deployed in dense 8-accelerator OEM platforms or PCIe form factors, appealing most to organizations prioritizing open networking, throughput-per-dollar, and an alternative software stack to CUDA.