Techday
This page was machine-translated and may differ from the original. View original

NVIDIA's deep learning computing platform's performance has improved tenfold in just six months.

Google 우선 소스Published2018.03.29 13:22
Tesla V100 32GB GPU with double the memory
NVSwitch Fabric enhances performance with a comprehensive software stack.


At GTC 2018, NVIDIA unveiled a series of performance improvements to its deep learning computing platform, announcing a 10x performance improvement over the previous generation in just six months for deep learning workloads.

Key enhancements to the NVIDIA platform include doubling the memory of its NVIDIA Tesla V100 data center GPU and NVIDIA NVSwitch, a groundbreaking GPU interconnect fabric that allows up to 16 Tesla V100 GPUs to communicate simultaneously at a record-breaking 2.4 terabytes per second. NVIDIA also announced updates and optimizations to its software stack.

With the launch of the NVIDIA DGX-2, NVIDIA has unveiled the first single server capable of delivering two petaflops of compute power to deep learning computing. The DGX-2 delivers the deep learning processing power of 300 servers occupying 15 racks in a data center, yet is 60 times smaller and 18 times more power efficient.

“Many of these improvements build on NVIDIA’s deep learning platform, which has quickly become the standard worldwide,” said Jensen Huang, founder and CEO of NVIDIA, when announcing the news at GTC 2018. “We are dramatically enhancing the performance of this platform at a rate that far exceeds Moore’s Law, creating breakthroughs that will drive transformational change in healthcare, transportation, scientific exploration, and countless other areas,” he said.



Tesla V100 with double the memory
The Tesla V100 GPU has double the memory to handle memory-intensive deep learning and high-performance computing workloads.

Data scientists can enhance the quality and quantity of deep learning model training with the Tesla V100 GPU, equipped with 32GB of memory, while also improving accuracy. Furthermore, it can improve the performance of memory-constrained HPC applications by up to 50% compared to the previous 16GB version.

The Tesla V100 32GB GPU is available immediately across the entire NVIDIA DGX system portfolio. Major computer manufacturers Cray, Hewlett Packard Enterprise, IBM, Lenovo, Supermicro, and Tyan have announced that they will begin shipping systems featuring the new Tesla V100 32GB GPU within the second quarter. Oracle Cloud Infrastructure also announced plans to offer the Tesla V100 32GB GPU in its cloud within the first half of this year.

NVSwitch: A Breakthrough Interconnect Fabric
NVSwitch delivers five times the bandwidth of the best PCIe switches, enabling developers to build systems with more GPUs hyper-connected. This will allow developers to overcome previous system limitations and run larger datasets. It also opens up the possibility of executing complex, large-scale workloads, such as parallel training modeling of neural networks.

NVSwitch extends the technological innovations achieved with NVIDIA NV Link, NVIDIA's first high-speed interconnect technology. NVSwitch enables system designers to build advanced systems that flexibly connect any topology of NV Link-based GPUs.

Advanced GPU-accelerated deep learning and HPC software stack
NVIDIA's deep learning and HPC software stack updates are available free of charge to NVIDIA's developer community. The NVIDIA developer community now boasts over 820,000 registered members, a significant increase from 480,000 a year ago.

This release includes new versions of NVIDIA CUDA, TensorRT, NCCL, and cuDNN, as well as the new Isaac Software Development Kit for Robotics. Furthermore, through close collaboration with industry-leading cloud service providers, ongoing optimizations are underway to ensure all major deep learning frameworks fully leverage the diverse benefits of NVIDIA's GPU computing platform.

NVIDIA DGX-2: The World's First 2 Petaflop System
NVIDIA's DGX-2 system delivers 2 petaflops by integrating several leading technology advancements that NVIDIA has driven at every level of the computing stack.

The DGX-2 is the first system to feature NVSwitch, which allows all 16 GPUs in the system to share a unified memory space. Developers now have access to deep learning training performance that can handle the largest datasets and most complex deep learning models.

DGX-2, featuring optimized and updated NVIDIA deep learning software, is designed for data scientists challenging the limits of deep learning research and computation.

On the DGX-2, training FAIRSeq, a state-of-the-art neural network-based machine translation model, takes less than two days. This represents a tenfold performance improvement compared to the Volta architecture-based DGX-1 introduced in September.
본 기사에 대한 정정·반론·추후보도 청구는 보도 청구 안내를, 그간 게재된 보도문은 정정·반론보도 모아보기를 참고해 주세요.
김학준 기자