인피니언 6월23일부터
Physical AI HBM Smart Factory SDV AIoT Power Semicon 특수 가스 정정·반론보도 모음 e4ds plus

AMD Expands GPU Options for Enterprise AI Inference

Google 우선 소스 기사입력2026.05.08 13:53


'MI350P PCIe' GPU Unveiled for Infrastructure Changes

As the adoption of artificial intelligence (AI) by enterprises accelerates, AMD has unveiled an enterprise GPU that leverages existing data center infrastructure. The product targets the demand for running AI inference in on-premises environments without the need for large-scale equipment replacement.

On the 7th, AMD unveiled the 'AMD Instinct MI350P PCIe' GPU card for enterprise AI. Designed as a standard PCIe-based card, the product allows for immediate use in existing air-cooled server racks and power/cooling environments.

The MI350P PCIe card features a dual-slot PCIe form factor. Its key feature is that it can be installed in existing servers to handle AI inference workloads without the need to build a separate dedicated GPU accelerator platform. AMD explained that this can serve as an alternative between the costs associated with migrating to the cloud, data privacy concerns, and the burden of large-scale on-premises upgrades.

According to AMD, the card can be utilized in air-cooled system configurations equipped with up to eight accelerators. Accordingly, it is designed to support inference tasks for various AI models ranging from small to large, as well as Search Augmented Generation (RAG) pipelines.

The MI350P PCIe is primarily targeted at enterprise environments where CPUs alone are insufficient but the adoption of large GPU platforms is burdensome. AMD stated that this product offers flexible options depending on the stage of AI adoption in enterprise environments.

Key specifications include native support for low-precision computing formats such as MXFP6 and MXFP4 to provide high throughput, and the ability to utilize sparsity acceleration even in precision formats such as INT8 and BF16. AMD presented peak performance of up to 4,600 TFLOPS based on MXFP4, approximately 144GB of HBM3E memory, and memory bandwidth of up to 4TB/s.

AMD emphasized openness not only in hardware but also in software. The MI350P PCIe natively supports major AI frameworks such as PyTorch and is designed to integrate with existing enterprise AI stacks, including Kubernetes GPU operators and AI inference microservices.

AMD plans to provide its partners with an open-source-based enterprise AI reference stack with no licensing costs. The company explained that this will enable relatively fast operation of inference workloads even in on-premises environments, while increasing code transparency and lowering operating costs.

AMD stated that the MI350P PCIe is designed to minimize power and cooling requirements in standard air-cooled data center environments by supporting various precision levels, including FP8, MXFP8, and MXFP4. This enables enterprises to migrate existing AI pipelines without code modifications and gradually increase the number of users and models.

AMD stated, “Adopting AI does not mean a complete overhaul of infrastructure,” adding that it is “a realistic option for companies looking to run AI inference in existing data center environments.”