This page was machine-translated and may differ from the original. View original
Achieved dramatically improved performance per watt through CNN algorithm-based cloud data center acceleration using Altera FPGAs
Hong Kong, March 3, 2015 — Altera (NASDAQ: ALTR) announced that Microsoft (NASDAQ: MSFT) has achieved dramatically improved performance per watt through data center acceleration based on convolutional neural network (CNN) algorithms by adopting Altera Arria® 10 field programmable gate arrays (FPGAs). CNN algorithms are widely used in image classification, image recognition, and natural language processing.
Microsoft researchers are conducting studies to improve cloud technology, and by utilizing the Arria 10 developer kit and Arria 10 FPGA engineering samples, they achieved performance reaching 40 GFLOPS per watt. This represents the industry's best level of data center performance. Furthermore, compared to using GPGPU, this FPGA performance offers a power-to-performance ratio that is more than three times better for a CNN platform. This performance was achieved by coding the Arria 10 FPGA and its IEEE754 hard floating-point DSP (digital signal processing) blocks using open software development languages such as OpenCL or VHDL.

“Our research team was able to achieve significant improvements in CNN performance and power efficiency by using Arria 10 engineering samples,” said Doug Burger, Director of Client and Cloud Apps at Microsoft Research. “The precision hard floating-point operations of the DSP blocks being integrated into this silicon are one of the factors that enabled the achievement of such leapfrog performance results,” he said. If you visit Microsoft’s blog (http://bit.ly/1MMMzvG), you can find an article in which Director Burger examines the infrastructure challenges facing data centers and explains how Microsoft was able to solve these challenges by replacing existing CPUs with reprogrammable FPGAs.
Michael Strickland, Director of Computation and Storage at Altera, stated, “FPGAs are advantageous for use in neural network algorithms at an architectural level because they can perform convolve and pooling very efficiently by utilizing flexible data paths. This allows many OpenCL kernels to transfer data directly to each other without needing to go to external memory. Another architectural advantage of the Arria 10 is that it supports hard floating-point for both multiplication and addition.” "By doing so, this hard floating-point allows for the utilization of more redundant logic than existing FPGA products and enables faster clock speeds," he said.
Altera previously announced that it has adopted its Stratix V FPGA to accelerate search using the innovative Catapult board for use in servers of Microsoft's first Bing data center, which is scheduled to be completed in the second half of this year.
Industry evaluation
Achieved outstanding performance and power efficiency by utilizing Altera 20nm FPGAs with integrated hard floating-point DSPs.
Many companies are achieving dramatically improved performance per watt by using Altera Arria® 10 FPGA products that integrate onboard hard floating-point DSPs. Altera is developing solutions for high-performance computing (HPC), data center acceleration, and financial systems in close collaboration with various customers and partners.
Microsoft – Doug Burger, Director of Client and Cloud Apps, Microsoft Research
“Our research team was able to see a dramatic improvement in CNN performance and power efficiency by using Arria 10 engineering samples. The precision hard floating-point capabilities of the DSP blocks integrated into this silicon were one factor in achieving these leapfrog performance results,” said Doug Burger, Director of Client and Cloud Apps at Microsoft Research. Refer to the Microsoft blog (http://bit.ly/1MMMzvG).
Bittware - Jeff Milrod, President/CEO, Bittware
Jeff Milrod, President and CEO of Bittware, stated, “Altera’s Arria 10 will be a true game changer. By providing a native floating-point engine onboard, these devices enable system designers to utilize massive floating-point resources with FPGAs with extreme ease and outstanding power efficiency. Classic signal processing applications can now interface analog signals directly to the Arria 10 and process them as floating-point numbers. For HPC and acceleration applications, there is no longer a need to port FPGA algorithms to fixed-point or implement them inefficiently through floating-point emulation. The Arria 10’s native floating-point achieves over 40 GFLOPS/W using a higher Fmax while consuming only one-third of the logic resources. This makes it easier to use, lowers power consumption, is faster, and reduces resource usage compared to other previous solutions.”
Gidel - Reuven Weintraub, Founder/CTO, Gidel
“We cannot hide our excitement over the Altera Arria 10’s unprecedented power-to-flops performance. FPGAs have long achieved outstanding power-to-performance ratios for bit, byte, and integer processing.” “The Altera Arria 10’s remarkably leap in power versus floating-point performance will enable Gidel products to be used in an even wider range of HPC and DSP applications,” he said.
Nallatech - Allan Cantle, President/Founder, Nallatech
Allan Cantle, President and Founder of Nallatech, said, “Nallatech is porting production code for several customers that require floating-point operations using Altera’s OpenCL compiler. To achieve this, we are utilizing the new Arria 10 FPGA, which provides a dedicated floating-point DSP, to reduce logic resource usage, increase clock frequencies, and further improve the performance index per watt. This makes Nallatech’s new Arria 10-based accelerators suitable for use in an even wider range of applications.”
ReFLEX CES - Yann Casteignau, Senior Engineer, ReFLEX CES
Yann Casteignau, Senior Engineer at ReFLEX CES, stated, “The FPGA boards recently released by ReFLEX CES based on the Altera Arria 10 FPGA enable many advantages thanks to the new floating-point DSP blocks implemented in this 10th-generation FPGA family. Our goal is to allow customers to significantly improve GFLOPS/W performance (expecting a 3x improvement) while simultaneously reducing the logic required for complex floating-point operations, thereby leaving as much free space as possible for implementing custom designs. Many of our customers use ReFLEX CES boards for high-performance computing, and power consumption is a critical challenge above all else.” “By utilizing the Arria 10 FPGA, it has become possible to achieve higher computing performance while reducing power consumption. The new hard-implemented DSP floating-point operations of the Arria 10 play a crucial role in enabling the ReFLEX CES board to improve performance, reduce logic requirements, and maximize GFLOPS/W performance,” he said.
본 기사에 대한 정정·반론·추후보도 청구는 보도 청구 안내를, 그간 게재된 보도문은 정정·반론보도 모아보기를 참고해 주세요.














