Techday
This page was machine-translated and may differ from the original. View original

Intel Proves Inference Performance in ML Commons Results

Google 우선 소스Published2023.09.12 16:32

▲ Intel Gaudi 2 (Photo: Intel)
ML Commons Benchmark Results Confirm Overall AI Competitiveness

MLPerf, generally the most reliable metric for evaluating AI performance, enables fair and repeatable performance comparisons. With high MLPerf metrics serving as a competitive advantage for AI semiconductors, news has emerged that Intel has achieved superior results compared to competing products in the ML Commons benchmark.

On the 11th (local time), MLCommons announced the results of 'MLPerf Inference 3.1' for computer vision and natural language processing models, including GPT-J, a massive language model with 6 billion parameters.

Here, Intel submitted performance measurement results for the Havana Gaudi 2 accelerator, 4th Gen Intel Xeon Scalable processors, and Intel Xeon CPU Max Series products. Through this, Intel showcased its competitiveness in AI inference as well as its efforts to enhance AI accessibility at various scales across AI workloads, from clients and the edge to networks and the cloud.

“As demonstrated by the ML Commons results, Intel has a powerful and competitive portfolio of AI products designed to meet customer demands for high-performance, high-efficiency deep learning inference and training,” said Sandra Rivera, Senior Vice President and General Manager of Intel’s Data Center and AI Group. “Our AI product family spans the entire spectrum of AI models, ranging from the smallest to the largest, and offers excellent value for money.”

This announcement is an extension of the AI training and Hugging Face performance results from ML Commons last June, which showed that Gaudi 2 products could outperform Nvidia's H100 in the latest vision language models. It also highlights that Intel provides the only alternative to Nvidia H100 and A100 products to meet AI computing demands.

Customers have different considerations, and Intel supports the implementation of AI for any use case through products capable of handling inference and learning across AI workloads. Intel's AI products provide customers with the flexibility and choice to select the optimal AI solution based on performance, efficiency, and cost goals, while simultaneously helping them break away from closed ecosystems.

In the GPT-J inference performance results for Havana Gaudi2, the Gaudi2 inference performance for GPT-J-99 and GPT-J-99.9 was recorded at 78.58 times per second for each server query and 84.08 times per second for offline samples.

Intel stated that Gaudi 2 demonstrated a slight advantage over the NVIDIA H100, providing approximately 9% higher performance in server mode and approximately 28% higher performance in offline mode. It added that Gaudi 2 delivered 2.4 times higher performance in server mode and 2 times higher performance in offline mode compared to the NVIDIA A100, and achieved 99.9% accuracy on new data types using FP8.

Intel plans to continuously provide performance improvements and an expanded range of models in the MLPerf benchmark through Gaudi 2 software updates released every 6 to 8 weeks.


▲ 4th Gen Intel Xeon Scalable Processor (Photo: Intel)

Intel has submitted results for all seven inference benchmarks, including GPT-J, on 4th Gen Intel Xeon Scalable processors. These results demonstrate excellent performance for various general AI workloads, including vision, language processing, and speech and audio translation models, as well as for much larger models such as DLRM v2 recommendation and Chat GPT-J models. In addition, Intel is still the only company submitting open CPU results using industry-standard deep learning ecosystem software.

4th Gen Intel Xeon Scalable processors are ideal for building and deploying common AI workloads using the most popular AI frameworks and libraries. 4th Gen Intel Xeon processors summarized two paragraphs per second in offline mode and one paragraph per second in real-time server mode for the GPT-J 100-word summary task of news articles approximately 1,000 to 1,500 words long.

Intel has submitted MLPerf results for the Intel Xeon CPU Max Series for the first time, offering up to 64GB of high-bandwidth memory. For GPT-J, it is the only CPU capable of achieving 99.9% accuracy, playing a critical role in applications where the highest accuracy is a key performance requirement.

Intel collaborated with OEMs to enable them to submit results directly. Through this, Intel demonstrated the AI performance scalability and broad availability of general-purpose servers based on Intel Xeon processors that can meet customer service level agreements (SLAs).

Meanwhile, Intel announced that it plans to submit new AI training performance results in the next MLPerf benchmark.
본 기사에 대한 정정·반론·추후보도 청구는 보도 청구 안내를, 그간 게재된 보도문은 정정·반론보도 모아보기를 참고해 주세요.
명세환 기자
명세환 기자