From Semiconductors to Tokens: The Current State of the AI Market
Jordan Nanos of semiconductor research firm SemiAnalysis analyzed the impact of semiconductor manufacturing constraints, data center expansion, and network bottlenecks on AI software architecture. Based on the company's research data, he explains the current state of the AI market from chip manufacturing to model inference, including GPU performance scalability and the economics of token processing costs (tokenomics).

From semiconductor manufacturing constraints to data center expansion and token processing for AI-generated text, various bottlenecks are now emerging throughout the AI infrastructure supply chain. Jordan Nanos of semiconductor research firm SemiAnalysis presented an analysis of these challenges and the current market situation.
With the acceleration of the AI boom, demand for AI semiconductors including GPUs is rapidly expanding. However, semiconductor manufacturing (fabs) faces high technical and physical constraints such as circuit miniaturization and securing manufacturing equipment, creating a structure where supply struggles to keep pace with surging demand. This "semiconductor constraint" can be seen as one of the fundamental factors determining the expansion speed of the entire AI infrastructure.
Data center expansion also presents complex challenges that go beyond simple server expansion. To efficiently operate large numbers of GPUs bundled together, high-speed networking technology connecting servers is essential, but Nanos points out that this network component often becomes a bottleneck. To fully leverage GPU processing power, not only the performance of individual chips but also the development of "communication infrastructure" for data exchange is required simultaneously.
Important evaluation metrics for GPU performance include benchmark scores and scaling efficiency—the efficiency when multiple GPUs are deployed at larger scales. Based on SemiAnalysis research data, Nanos analyzed the current market situation from these perspectives. Even when benchmark values are high, parallel operation at large scales can show reduced efficiency, and accurately assessing scaling characteristics directly impacts adoption decisions.
Particularly notable is the concept of "tokenomics." Tokens are the minimum units AI uses to process text, and tokenomics refers to the cost and economic viability of generating and processing those tokens. Understanding the complete cost structure—from chip manufacturing costs to data center power and cooling costs, through to the per-token price when models perform inference tasks like answering questions—is becoming an essential perspective for considering the profitability of AI services.
What this analysis reveals is the reality that AI competitive advantage is moving beyond being determined solely by model accuracy. Semiconductor procurement capability, data center network design, and token-level economic efficiency—these three elements are now intricately intertwined, determining both the practical implementation costs and competitiveness of AI services. The ability to optimize the entire infrastructure as a "system" is positioned as a source of differentiation for companies and nations.
A key point to watch going forward is how much the pace of semiconductor supply improvement and network technology evolution contribute to reducing AI service costs. If cost per token decreases, more industries and users can easily adopt AI. Conversely, if infrastructure bottlenecks persist, the gap between early adopters who secured equipment and latecomers will widen. The perspective of understanding the "behind-the-scenes" of AI—from chip manufacturing sites to the moment of token generation—will become increasingly important.
This article is an original work independently written and edited by the AI issue editorial team based on factual reporting. © AI issue. Unauthorized reproduction, redistribution, or use for AI training is prohibited.