Chinese LLMs Top US Rivals in Call Volume on OpenRouter for 20th Straight Week(Yicai) Sept. 17 -- The combined weekly call volume of Chinese large language models exceeded that of US models on application programming interface aggregation platform OpenRouter for the 20th week in a row despite lacking in intelligent capabilities.
Global LLM call volume on OpenRouter reached 127 trillion tokens from Sept. 7 to 13, with Chinese models accounting for nearly 61.2 trillion tokens, compared with 21.8 trillion tokens for their US rivals, according to the latest data released by the platform. The figure for Chinese LLMs rose 7.9 percent week on week, while that for US models surged 32 percent.
GPT-5.6 Luna from US developer OpenAI ranked first with 18.2 trillion, followed by Hy4 preview with 16.8 trillion and GLM-5.3-Flash with 11.9 trillion by China’s Tencent Holdings and Z.AI. Other Chinese models in the top 10 included DeepSeek’s V4-Flash-0731, V4.1-Flash, and V4-Flash-0423, which ranked fourth, sixth, and seventh, respectively, as well as Xiaomi’s MiMo-V2.5 in fifth.
DeepSeek’s V4.1 Flash processed about one trillion tokens within 24 hours of its Sept. 10 release. “It’s interesting that it just took AI users hours rather than days to use DeepSeek’s new model with their workflows and applications,” a developer pointed out.
LLMs used to be upgraded yearly, then every few months, and now almost every week. Seven leading global developers upgraded their models between Sept. 1 and Sept. 11, including US’ Anthropic launching Claude Fable 5.1 and OpenAI releasing its flagship model GPT-6 Astra.
After the latest round of upgrades, the global LLM competitive landscape shows two distinct characteristics, according to OpenRouter. In terms of capabilities, Chinese developers still miss the top five, with only Z.AI and Moonshot AI making the top 10. However, seven Chinese models rank in the top 10 by call volume, up from only two a year ago.
Chinese models are not leading in technological capabilities but have become the choice of more and more developers. Amid AI’s large-scale implementation, enterprises and users are gradually shifting their focus from capabilities to higher input-output ratios, with “Token ROI” becoming a key term in discussions, Xiong Wei, an internet industry analyst at UBS Securities in China, told Yicai at a recent exchange event.
After a round of price increases and drops, DeepSeek still plays the role of a “price butcher” in the global LLM market. About 90 percent of V4.1-Flash’s call volume comes from cache reads, the cost of which is around 0.6 US cents per million tokens based on market pricing, about 80 percent lower than similar models such as GLM-5.3-Flash, OpenRouter noted.
“Chinese models are already capable of doing most tasks of the application layer,” the developers of GitHub’s open-source project Xiaobei told Yicai, adding that unless a task is complex at the underlying level, there is no need to pay tens of times the cost for top models.
Speaking with overseas investors, Xiong said he has found that some of the world’s leading information technology companies have begun to consider incorporating Chinese open-source models into their workflows. The trend significantly accelerated from May to June, with more and more companies and developers starting to use such models to build workflows or perform specific tasks, he noted.
In addition, although Chinese model developers offer higher cost-effectiveness, they are not paying to boost adoption of their models, according to a UBS survey. Their gross profit margin is mostly in the range of 20 percent to 40 percent, meaning that they improve profitability and maintain cost-effectiveness through innovative model architecture, optimized training, and inference efficiency, the findings showed.
Editor: Martin Kadiev
