This page was machine-translated and may differ from the original. View original
Gartner: “Corporate Burden Continues Even as AI Token Prices Drop”… Proliferation of Agents Changes Cost Structures
Despite falling inference unit costs, the challenge of managing total corporate AI costs persists.
In an analysis released on March 30, Gartner predicted that the cost of inference for massive language models with one trillion parameters would fall by more than 90% by 2030 compared to 2025. Tokens are the basic units used by generative AI to process sentences and data, and in this analysis, they were defined as data at the level of approximately 3.5 bytes.
Gartner cited improved semiconductor and infrastructure efficiency, changes in model design, enhanced chip utilization, the expansion of inference-specialized semiconductors, and the increased application of edge devices in certain areas as the background for this cost reduction. Accordingly, it projected that LLM in 2030 could demonstrate up to 100 times higher cost efficiency compared to an initial model of the same size in 2022.
However, Gartner noted that a drop in unit costs does not immediately lead to the "popularization of AI." In particular, since AI agents can use 5 to 30 times more tokens per task than traditional chatbots, the overall cost of inference could actually increase even if the price of individual tokens decreases. The explanation is that even if basic functions become cheaper, computing resources for processing complex reasoning remain limited.
This analysis compared costs based on a 'Frontier' scenario using state-of-the-art semiconductors and a 'Legacy Mix' scenario using a combination of various existing semiconductors. Gartner explained that the Mix scenario showed higher costs than the Frontier scenario due to relatively lower computational performance. Ultimately, this means that the type of semiconductor and infrastructure on which AI is operated has a direct impact on the cost structure.
Gartner predicted that going forward, corporate AI competitiveness will be determined not by adopting a single large-scale model, but by operational strategies for deploying and coordinating multiple models across different tasks. A more realistic alternative was suggested: assigning repetitive, high-frequency tasks to smaller or domain-specific models, while deploying costly, frontier-level models only for complex, high-value-added work. Ultimately, the ability to design which models to connect to which tasks is increasingly likely to determine corporate AI profitability, rather than the reduction in token prices itself.
본 기사에 대한 정정·반론·추후보도 청구는 보도 청구 안내를, 그간 게재된 보도문은 정정·반론보도 모아보기를 참고해 주세요.















