This page was machine-translated and may differ from the original. View original
NVIDIA's Blackwell Platform Reduces AI Inference Costs by Up to 10x Per Token
Accelerating open-source cost innovation across healthcare, gaming, and customer service.
NVIDIA's full-stack strategy spanning computing, networking, and software is dramatically improving the economics of AI inference and accelerating AI adoption across industries.
NVIDIA announced on the 20th that it is accelerating the spread of AI across industries by reducing the cost per token of AI inference services by up to 10 times through its next-generation AI computing platform, 'NVIDIA Blackwell'.
Major inference service providers, including Baseten, DeepInfra, Fireworks AI, and Together AI, have adopted optimized inference stacks based on Blackwell to achieve both efficiency and scalability.
Various AI interactions, such as AI-based medical diagnoses, interactive games, and customer service agents, operate based on the same unit of intelligence called a “token.”
Tokenomics, which allows companies to afford more tokens, is emerging as a key challenge as they expand their AI services.
A recent MIT study found that the cost of cutting-edge AI inference is decreasing by up to tenfold annually, driven by improvements in infrastructure and algorithmic efficiency.
The NVIDIA Blackwell platform makes these cost savings a reality through tight co-design of hardware and software.
By combining optimization techniques such as the low-precision NVFP4 data format, TensorRT-LLM, and the Dynamo inference framework, we have enabled processing of significantly more tokens with the same infrastructure cost.
In the healthcare sector, Baseten and Sulli.ai have reduced AI inference costs by 90% by leveraging Blackwell-based open-source models.
Response times for repetitive tasks like creating medical records and writing codes were improved by 65%, giving healthcare professionals more than 30 million minutes of time back.
In the gaming space, DeepInfra and Latitude have achieved a fourfold reduction in cost per token with Blackwell-based inference, while also delivering a stable user experience even in large-scale AI-native gaming environments.
Success continued in the areas of agentic chat and customer service.
Fireworks AI and Sentient Labs achieved up to 50% cost-effectiveness improvements with the Blackwell-optimized inference stack, while Together AI and Decagon achieved sub-400 millisecond response times while reducing cost per query by 6x in voice AI customer support.
본 기사에 대한 정정·반론·추후보도 청구는 보도 청구 안내를, 그간 게재된 보도문은 정정·반론보도 모아보기를 참고해 주세요.















