# From Selling Intelligence to Selling Workflows: How AI Companies Are Redefining Their Value
In July 2026, MiniMax released M3. If you look only at its benchmark scores, you might think this is just another round in the endless model parameter race.
But look closer at what MiniMax actually presented, and you'll see something different—not a smarter model, but a new narrative about how AI companies create value.
This article unpacks why M3's launch isn't just a routine version update. It's a signal that the entire industry's value anchor is shifting.
01 The Evaluation Has Changed: From Test Scores to Getting Things Done
When MiniMax launched M3, the headline metrics weren't MMLU (a test of broad knowledge), GSM8K (math reasoning), or HumanEval (code generation). Instead, MiniMax devoted significant space to a different set of benchmarks: BrowserComp (browser operation capability), SWE-Bench (fixing real project bugs), Terminal Bench (command-line operation), and OSWorld (operating-system level tasks).
These metrics measure the same thing: **can the model complete an actual job? Fix a real bug in a real project? Navigate a web page independently? Call a development environment? Connect to an enterprise system?** [1]
36Kr's Qidianpai put it precisely: the industry used to evaluate "Intelligence" (knowledge). Today it evaluates "Task Completion." [1]
**It's like hiring. Companies used to look at your transcript. Now they watch you work. The AI industry's "exam system" is shifting from written tests to on-the-job assessments.**
MiniMax isn't alone in this shift. Around the same time, LMArena launched Agent Mode—evaluating models on real task completion rates rather than chat quality. [5] Vercel unveiled its skills ecosystem, with installable AI skill packages hitting 2.3 million installs. [7]
The trend is unmistakable: in consumer and general enterprise scenarios, nearly everyone is moving from "how smart is your model?" to "how much work can your model actually do for me?"
02 Why MiniMax—And Why It Matters
MiniMax's M3 launch simultaneously advanced four capabilities: Browser operation, Code fixing, Terminal environments, and MCP (Model Context Protocol—a standard for connecting to enterprise systems). Put together, MiniMax wasn't just showing off a model. It was showcasing a complete workflow. [1]
Why? Because in consumer and general enterprise scenarios, most models' "knowledge benchmark" scores have converged to within a few percentage points. Meanwhile, enterprise AI spending is shifting from "validation and exploration" to "production deployment." As Qidianpai notes, citing a Gartner forecast, global AI software spending is projected to reach $135 billion in 2026—with the bulk flowing to products that integrate into real workflows. [1]
One important caveat: **"model capability convergence" applies more to consumer and general enterprise scenarios. In frontier reasoning, long-context, and multimodal domains, meaningful gaps remain between top-tier models and the rest.** M3's workflow focus is partly a strategic choice driven by China's market reality—under API price wars, workflow is the key differentiation escape.
This reveals two interesting China-US differences in motivation:
**In the US** — OpenAI, Google, and Anthropic's driving logic is "general AGI as foundation → enterprise agents as extension." Per Qidianpai's citation of a 2026 Sequoia AI investment report, Agent/Workflow funding share jumped from 18% in 2024 to 42%. [1] These companies build the moat first, then extend to the workflow layer.
**In China** — Baidu's ERNIE API price has dropped over 90% in the past 18 months. ByteDance's Doubao is giving away service for free, bundling with scenarios to capture market share. [1] For MiniMax and peers, workflow strategy is a strategic escape from a commodity pricing trap.
Both roads lead to Workflow, but for completely different reasons. For US incumbents, it's a natural extension of building capabilities. For Chinese players, it's survival in a prisoner's dilemma.
03 From Token to Workflow: Restructuring the Business Model
Tokens—the basic billing unit for most AI APIs, roughly equivalent to a character or a few English characters. You use more, you pay more.
But more and more companies aren't buying AI for more "answers." They're buying it to complete more "jobs." Fix a bug. Summarize a meeting. Process a ticket. The value unit is shifting from Token to Task.
How serious is this shift? Look at Arena.
Arena (formerly Chatbot Arena)—a platform that builds no models, sells no APIs. It only evaluates. Eight months after launching commercial services, it hit $100 million in annualized revenue at a $1.7 billion valuation. Its clients include OpenAI, Google, Anthropic, and Meta—every model maker pays Arena to evaluate them. [2][3]
**Capital isn't betting on AI models themselves. It's betting on whoever defines the infrastructure for "what AI can do."**
Another supporting signal: the explosion of Vibe Coding—describing what you want in natural language and letting AI write the code. A fintech company called Slash spent roughly $80,000 in token fees on AI-powered coding to build a pixel game. The game went viral. It brought attention and new clients. That translated into $100 million in new Assets Under Management (AUM). [6]
The logic chain: **$80K to make a game → game goes viral → attracts clients → clients convert to AUM → $100M AUM growth.**
This isn't a "make games and get rich" formula. The real point: AI has lowered the barrier to shipping a complete product. Individuals can now complete a commercial task at a fraction of the traditional cost. And "completing commercial tasks" is the core driver of AI's shift from the Token era to the Workflow era.
As Qidianpai notes: Workflow moats run deeper than technology moats. Technology can be caught up—a better model launches tomorrow and your ranking drops. But workflow binds data accumulation, employee habits, and system integration. That creates stickier user relationships and more stable business models. [1]
**A company's data doesn't accumulate in chat logs. It accumulates in the workflows that happen every day.**
04 Industry-Wide Evidence of the Value Shift
Put M3 in the broader picture, and four dimensions show the same shift happening simultaneously:
**Model layer** — Per developer community reports (not officially confirmed by OpenAI), GPT-5.6 abandoned the traditional Mini/Standard/Ultra tiering. Instead, it introduced Sol (hard reasoning and code), Terra (daily production), and Luna (batch, high-frequency)—priced by capability tier, not by model size. [4] Not selling "the smartest brain" anymore, but selling precision productivity tools at different levels.
**Tool layer** — Vercel's skills ecosystem turns AI capabilities into reusable, downloadable modules. But Snyk's security audit found that over 30% of 3,984 skills had security defects, 13.4% rated severe. [7] More capable tools also mean more risk—a reality enterprises must face when deploying.
**Evaluation layer** — Arena expanded from "humans judge who chats better" to "machines automatically evaluate task completion rates." Whoever defines this new standard holds commercial leverage. [2][3]
**Enterprise layer** — A 36Kr investigation titled "The First Big Tech Layoffs by AI" revealed a sobering picture: at Ctrip, Meituan, and ByteDance, the "high-salary, high-performance, high-rank" survival shield has been shattered. AI-driven efficiency gains, stagnant legacy business growth, and the cash pressure of investing in new AI initiatives form a triple bind. A Stanford study found that since ChatGPT's launch, employment for software developers aged 22-25 has dropped nearly 20% from its late 2022 peak. [9]
**Every signal converges on the same direction: AI's unit of value is moving from Token to Task.**
05 What This Means for Knowledge Workers
At this point, you might think this is just an industry inside story. How does it affect you directly?
It changes how you choose and use AI tools.
Used to be: check the benchmark leaderboard, pick whoever's smartest. Now, the standard is different. Don't just ask how many questions it gets right. Ask whether it can help you finish today's work.
Here's a simple three-step self-check to evaluate any AI tool you're considering:
**Step one: Can it complete a specific task?**
- Don't ask "what can you do." Give it a real task: write an email, organize a page of notes, draw a flowchart, analyze a dataset.
- Evaluation: Did it deliver? To what degree?
**Step two: Can it plug into your workflow?**
- Can it access your calendar, documents, codebase, data sources?
- Does it support API integration, file imports, batch processing?
- Evaluation: Is it a standalone tool, or can it become part of your workflow loop?
**Step three: What does it leave behind after the job?**
- What did it output after completing the task? A one-time result, or a reusable capability?
- Is your data accumulated and available next time?
- Evaluation: Did it help you build your own "work capacity reservoir"?
These three questions don't require AI expertise. They only require you to understand one thing: **AI's value, ultimately, depends on how much work it helps you get done.**
When model capabilities converge, an AI company's value anchor isn't how many test questions it answers correctly. It's how many of your workflows it can enter.
That shift has just begun.
References
- **"MiniMax M3:一家AI公司,为什么开始重新定义自己的价值?"** — 奇点湃 (Qidianpai), 36Kr, 2026-07-05. https://www.36kr.com/p/3882467938040710
- **LMArena lands $1.7B valuation four months after launching its product** — Julie Bort, TechCrunch, 2026-01-06. https://techcrunch.com/2026/01/06/lmarena-lands-1-7b-valuation-four-months-after-launching-its-product/
- **Arena, the AI leaderboard everyone uses, is now a $100M business** — Marina Temkin, TechCrunch, 2026-06-29. https://techcrunch.com/2026/06/29/arena-the-ai-leaderboard-everyone-uses-is-now-a-100m-business/
- **"硅谷无间道:GPT-5.6'偷跑'截杀Claude,最快明天端出三大模型"** — DIYIXinsheng, 36Kr, 2026-07-06. https://www.36kr.com/p/3883839281099016
- **Arena launches Agent Mode** — Cited in TechCrunch Arena report (reference [3])
- **"Vibe Coding误烧55万元token,他靠嘲笑声反手赚了7个亿"** — 量子位 (Liangziwei), 36Kr, 2026-07-06. https://www.36kr.com/p/3883634355695878
- **"狂揽2.4万星标:一行命令,AI会自己找技能了"** — 新智元 (Xinzhiyuan), 36Kr, 2026-07-06. https://www.36kr.com/p/3883329457483784
- **Gartner Forecasts Worldwide AI Software Spending to Reach $135 Billion in 2026** — Gartner, January 2026. Data cited via reference [1]
- **"AI砍掉的第一批大厂人:高薪,高绩效,高P|深氪"** — 任彩茹 (Ren Cairu), 36Kr, 2026-07-06. https://www.36kr.com/p/3883456791163138
- **Canaries in the Coal Mine? The Impact of AI on the Employment of Software Developers** — Stanford University, 2025. Cited via reference [9]