This page was machine-translated and may differ from the original. View original

▲CEVA launches enhanced NeuPro-M NPU IP product family (Image: CEVA)
Provides 350 TOPS/Watt, maximizing cost and power efficiency
The use of low-power, low-cost inference processing using NPUs is increasing to realize AI inference technology from the cloud to the edge.
CEVA, a custom system-on-chip (SoC) company, today announced the launch of its enhanced NeuPro-M NPU.
NeuPro-M NPU is expected to meet the requirements of next-generation generative AI with industry-leading performance and power efficiency for all AI inference workloads from the cloud to the edge.
The NeuPro-M NPU architecture and tools have been redesigned to support transformer networks and future machine learning inference models, in addition to convolutional neural networks (CNNs) and other networks.
This enables NeuPro-M NPUs to seamlessly develop and run applications that optimize the capabilities of generative and traditional AI in communications gateways, optical networks, automobiles, laptops and tablets, AR/VR headsets, smartphones, and other cloud or edge environments.
The enhanced NeuPro-M architecture integrates The Vector Processing Unit (VPU) can be used for a variety of purposes and can support network layers that will be developed in the future. It also supports all activations and data flows, and can accelerate performance by up to 4x through true sparsity of data and weights.
Consumers can also address a variety of applications and markets with a single NPU product. NeuPro-M adds new NPM12 and NPM14 NPU products, which include two and four NeuPro-M engines respectively, to increase scalability required across a variety of AI markets.
Therefore, easy migration to high-performance AI workloads has become possible, and it currently consists of a total of four NPU product groups: △NPM11 △NPM12 △NPM14 △NPM18.
NeuPro-M delivers peak performance of 350 TOPS/W at the 3nm process node and is capable of processing over 1.5 million tokens per second on transformer-based LLM inference workloads.
The architecture is based on the NeuPro-M parallel processing engine and is supported by an innovative, comprehensive development toolchain based on CEVA’s architecture-aware network AI compiler, CDNN, to maximize the performance of consumer AI applications.
The CDNN software includes a memory manager for memory bandwidth optimization and optimal load balancing algorithms, and is compatible with common open source frameworks including TVM and ONNX.
“Transformer-based networks that power generative AI will require a dramatic increase in compute and memory resources,” said Ran Snir, vice president and general manager, Vision Business Unit at CEVA. “Meeting this increased compute and memory demand will require new approaches and optimized processing architectures.”
“The performance gains enabled by this architecture open up incredible possibilities for generative AI across use cases, from cost-sensitive edge devices to high-efficiency cloud computing,” he said.
ABI Research predicts that edge AI shipments will reach 2.4 billion units by 2023It is predicted that the number of edge applications will increase to 6.5 billion by 2028, at a compound annual growth rate (CAGR) of 22.4%. Generative AI is expected to play a key role in supporting this growth, and more sophisticated and intelligent edge applications are demanding more powerful and efficient AI inference technology.
In particular, large-scale language models (LLMs) and vision and audio converters used in generative AI can revolutionize products and industries, but they also pose new challenges in terms of performance, power, cost, latency, and memory when running on edge devices.
“Today’s hardware market for generative AI is highly concentrated, driven by a small number of vendors,” said Reece Hayden, senior analyst at ABI Research. “To make AI inference technology a reality, a clear strategy is needed for low-power, low-cost inference processing in the cloud and at the edge, using smaller models and more efficient hardware.”
“CEVA’s NeuPro-M NPU IP provides a proposition for deploying generative AI on devices with minimal power consumption,” he said. “Furthermore, the scalability of NeuPro-M allows it to be applied to use cases that demand ultra-high performance in network appliances or beyond.”
Advances in inference and modeling technologies are enabling new capabilities that leverage small-scale, domain-specific LLMs, vision transformers, and other generative AI models operating at the edge device level, transforming applications in the infrastructure, industrial, mobile, consumer, automotive, PC, and mobile markets.
본 기사에 대한 정정·반론·추후보도 청구는 보도 청구 안내를, 그간 게재된 보도문은 정정·반론보도 모아보기를 참고해 주세요.















