For months, cheaper token prices have helped Chinese AI models build a reputation as the budget-friendly alternative to OpenAI and Anthropic. But a new study suggests enterprises chasing the lowest sticker price could end up paying more.
Research from the AI-powered market intelligence platform AlphaSense found that OpenAI's GPT-5.6 Sol and Anthropic's Opus 4.8 outperformed Chinese rivals Kimi K3 and GLM-5.2 on both cost and quality in complex financial analysis tasks.
The findings challenge the widely held assumption that lower token prices automatically translate into lower AI costs.
The Cheapest Model Isn't Always the Cheapest to Use
At first glance, Chinese models appear significantly cheaper. Moonshot charges $15 per one million output tokens for Kimi K3, compared with $25 for Anthropic's Opus 4.8 and $30 for OpenAI's GPT-5.6 Sol.
However, AlphaSense found that token pricing tells only part of the story.
After testing 246 financial analysis tasks—including analyzing earnings transcripts, SEC filings, analyst estimates and acquisition activity—the company concluded that OpenAI's GPT-5.6 Sol delivered answers with roughly 20% higher quality while costing about 13% less than Kimi K3 on a median basis. Anthropic's Opus 4.8 performed even better, generating responses that scored around 13% higher in quality at roughly half the overall cost of Kimi K3.
The reason, according to the study, is that more capable models often require fewer tokens and fewer processing steps to complete the same task, reducing the total cost despite charging higher prices per token.
"Some of the more expensive models, the ones that look more expensive based on just their price per token, actually ended up being less costly because they were more efficient in using tokens," AlphaSense CEO Jack Kokko said.
AI Buyers May Need a New Way to Measure Cost
The findings arrive as enterprises increasingly weigh whether to adopt frontier AI models from companies like OpenAI and Anthropic or to opt for lower-priced, open-weight alternatives from Chinese developers.
Rather than focusing solely on token prices, AlphaSense argues businesses should evaluate the total cost of completing a task alongside the quality of the output. In knowledge-intensive workloads such as financial research, more intelligent models may justify their premium by delivering accurate answers more efficiently.
At the same time, the report doesn't dismiss open models altogether.
Companies running AI workloads on their own infrastructure can avoid per-token charges entirely, while less demanding tasks—such as email summarization—may not require frontier models.
Instead, AlphaSense says the most cost-effective strategy may be using multiple models together. Its AI search platform employs a routing system that automatically assigns different parts of a query to different models—for example, using a more capable model to plan a response before handing execution to a smaller, cheaper model.
The broader takeaway is that as enterprises move beyond benchmark scores and token pricing, the AI industry's next pricing battle may be decided less by who charges the least—and more by who delivers the lowest cost per completed task.