AI Progress Shifts From Moore’s Law to 'Full Stack' Efficiency
Industry leaders are moving beyond transistor scaling to a composite of hardware, architecture, and model design to drive economic viability.
The era of relying on Moore’s Law as the sole barometer for computing progress has ended, replaced by a compounding 'full stack' of technological curves. As AI evolves toward 2026, the industry is shifting its focus from raw transistor counts to the integrated efficiency of chips, system architecture, model design, and security.
This transition is evidenced by a dramatic collapse in the cost of running AI. According to Andreessen Horowitz, AI inference costs at a constant performance level fell roughly tenfold annually between 2021 and 2024, representing a 1,000x decrease in just three years. This acceleration is not the result of a single breakthrough but the synergy of multiple layers of the stack. For example, hardware partnerships are pushing the boundaries of power efficiency; AMD and Cerebras recently announced an inference system pairing AMD's Helios platform with the Cerebras Wafer-Scale Engine, which is projected to deliver up to five times more tokens per second per watt than setups using Cerebras hardware alone.
The Architecture Shift
For decades, the industry viewed computing performance through the narrow lens of silicon scaling. However, modern AI progress is now a composite of hardware, software architecture, and system-level optimizations involving networking and memory. This shift allows performance to scale faster than transistors alone could permit, opening the door for more complex autonomous agents and embodied AI.
Model design is playing an equally critical role in these efficiency gains. Moonshot AI’s Kimi K3 model demonstrates this by activating only 16 of its 896 specialist components per token, resulting in scaling efficiency approximately 2.5 times better than its predecessor, Kimi K2. By optimizing how the model utilizes its parameters, developers are achieving higher intelligence without a linear increase in compute requirements.
Redefining Economic Value
This technical evolution is forcing a fundamental change in how the market measures value. The industry is moving away from 'cost per token' toward metrics like 'cost per qualified task' or 'intelligence per dollar.' This shift signifies that the economic viability of AI is no longer about raw compute power, but about the ability to complete useful, reliable work cheaply.
As a result, the primary bottleneck is shifting from chip availability to the ability to deploy and secure agents. This expands the critical market beyond traditional chipmakers to include the entire infrastructure stack. This includes specialized AI cloud providers like Nebius Group and security initiatives such as Nvidia's Open Secure AI Alliance, both of which are essential for the reliable deployment of AI workflows.
The Path Forward
As ETF Trends notes, while Moore’s Law taught investors to watch one curve, AI requires monitoring how chips, architecture, models, data, and security improve in tandem. The coming years will likely see further integration between specialized silicon and sparse model architectures. The industry will be watching whether these compounding gains can maintain their current velocity to make fully autonomous agents economically sustainable across the enterprise.