Chinese AI Firms Use Software Optimization to Bridge Nvidia Performance Gap
Companies are adopting heterogeneous computing to maintain scale as domestic hardware struggles to match restricted Nvidia processors.
Chinese AI companies are increasingly relying on software optimization and heterogeneous computing strategies to maintain operations as domestic hardware struggles to match Nvidia's performance. This shift comes as firms attempt to scale inference capabilities amid strict US export controls on high-end semiconductors.
To manage surging demand, firms are splitting workloads between domestic hardware, such as the Huawei Ascend 910B, and restricted Nvidia processors like the H20. This approach allows companies to route basic tasks to local chips while reserving scarce Nvidia architecture for complex, high-tier requirements. The urgency is driven by a massive spike in activity; according to the National Data Administration and the South China Morning Post, China's average daily token calls exceeded 140 trillion in March, representing a more than 1,000-fold increase since the start of 2024.
The Performance Divide
This technical pivot is a direct response to a growing performance gap in "high-quality" AI tasks. While domestic chips can handle basic inference, they currently lack the architecture necessary for complex operations, specifically coding. This has created a bipolarization of demand where high-tier tokens are in acute shortage because they cannot be efficiently processed on local hardware.
Guan Jiawei, vice-president of Approaching.AI, highlighted the commercial risk of this divide, stating, "If we rely solely on domestic chips for inference, they can only handle the low-quality tier – the tier with weak demand and weak monetisation."
A Strategy for Survival
To survive these constraints, the industry is moving toward "heterogeneous P/D disaggregation." Approaching.AI is utilizing this technique to distribute tasks across different chip architectures via software, effectively using code to bridge the gap between disparate hardware capabilities.
This strategy is critical for the commercial viability of advanced Chinese AI agents. Without the ability to execute complex tasks at scale, the industry risks a ceiling on the sophistication of its deployed models. By decoupling the processing and delivery of tokens across a mixed fleet of chips, firms can maximize the utility of their limited Nvidia supply while still integrating domestic silicon into their infrastructure.
The Path Forward
As the demand for AI inference continues to grow exponentially, the reliance on software-based workarounds remains a temporary bridge. The industry's long-term success depends on whether domestic chipmakers can close the architectural gap for high-tier tasks or if software optimization can sufficiently mask the hardware deficit. For now, the ability to orchestrate a hybrid environment of Huawei and Nvidia chips is the primary mechanism allowing Chinese AI firms to scale despite ongoing sanctions.