September 2026
English translation of the Korean original, prepared with AI assistance. Korean original
This month, the contest in AI moved from the performance ceiling to the cost of deployment. Cost per task and Tokens Per Watt are becoming the tests of adoption. The market is now going to those who enter the enterprise stack first as a default layer, with a portfolio of models for different uses rather than a single flagship.
The three gates an AI capability must pass to reach commercial deployment are the regulatory gate (source tracing and governance), the reliability gate (domain verification) and the economic gate (benefit against operating cost). This month the first two became routine work through vertical products in law, finance and healthcare and through domain benchmarks. The third, the economic gate, became the only bottleneck, and its pass criterion was redefined from price per token to unit-economics measures: cost per task, Tokens Per Watt and latency.
- Key AI trends of the month
First, competition has shifted fully from scale to unit cost. Anthropic cut cost per task by 30% with Claude Sonnet 5.5. OpenAI lowered the price of voice AI to $0.05, and Alibaba cut the price of its voice tools by up to 95%. Second, agents have become the drivers of workflow change and a core enterprise capability, and they are now being asked to verify sources, as with Source-Aware Verification. Third, AI has moved off the screen and into the physical world. GPT-6 Astra’s tests of robot and drone control signal the start of the embodied-intelligence stage.
- The foundation model race
The top tier has narrowed to three leaders: OpenAI’s GPT-6 Astra, Anthropic’s Claude Opus 5.5 and Google’s Gemini 4 Pro. Astra raised the performance ceiling with a 100% score on a cybersecurity benchmark. Google differentiated itself on interaction, with Gemini 3.8 Live offering voice conversation in 97 languages and real-time reasoning. The notable change is the splitting of product lines. OpenAI separated performance from running cost with GPT-6 Sol and Luna and quoted 300 tokens per second for GPT-6.1 Sol. This marks a shift from a contest between single flagships to a contest between portfolios built for different uses. The start of Gemini 4 training and the teaser for Fable 5.2 show that the next round is already under way.
- Chips, infrastructure and open source
With Vera Rubin and NVL72, NVIDIA put Tokens Per Watt forward as the key metric. It has redefined the basis of infrastructure competition from compute to power efficiency. NVIDIA’s on-site AI acceleration at IFA 2026 and its real-time processing for broadcast and sports are aimed at demand outside the data centre. In the open-weight camp, lightweight MoE models such as DeepSeek-V4.1-Flash and Altar-1, the one-million-token multimodal capability of Qwen3.8-Omni-Flash and the on-device operation of MiniCPM5-2B have quickly narrowed the performance gap. However, moves by the US government to restrict Chinese models create an adoption risk that is separate from technical advantage.
- AI products and the start-up ecosystem
Monetisation has moved from hypothesis to results. Salesforce’s Claude CRM integration, Snowflake Cortex AI’s link to Astra and AWS’s autonomous task-handling feature show that models have become a default layer of the enterprise stack, not an optional feature. ChatGPT advertising revenue has grown to about $1bn a year, which is evidence that the consumer revenue model is also becoming self-sustaining. Vertical strategies have also sharpened. Specialist chatbots for finance, law and research, EHR integration and Proaction’s use in vehicle fleet monitoring suggest that combining domain knowledge, rather than general-purpose performance, now decides the contest. Perplexity’s RTX-based on-device processing is an attempt to differentiate by combining personalisation with privacy.
- Outlook for next month
Three things deserve attention. First, the floor under price competition. A 30% cut in cost per task and a 95% price cut came within a single month, so the price per token itself is likely to drop out as a differentiator, and competition is likely to shift to latency and throughput per watt. Second, whether the overlap of Gemini 4’s full rollout and the Fable 5.2 launch will shorten the flagship replacement cycle to a quarter. Third, the growing real-world use of embodied intelligence and autonomous agents will expose debates on safety and international governance as lagging behind the pace of the technology. If restrictions on Chinese models take concrete form, the regional fragmentation of the open-weight ecosystem could accelerate.
What to do
- Because model competition has been reorganised around efficiency, evaluate frontier models such as GPT-6 Astra and lightweight models such as DeepSeek V4.1 Flash and MiniCPM5-2B separately, using benchmarks. Set a routing policy by task difficulty, and track the inference cost per unit of equivalent quality every month.
- Agents have become a core capability for end-to-end work automation. Standardise the agent runtime, including tool calls, state management and failure recovery. Combine it with prompt caching, batch processing and quantised serving (such as GGUF) to cut API costs structurally.
- As the Salesforce-Claude and Snowflake Cortex-GPT-6 cases show, value is moving to the platform touchpoint. Avoid the contest over general-purpose models, and reposition your product as a vertical solution that combines domain knowledge and workflow, in areas such as finance and healthcare (EHR integration).
- On-device operation and on-site AI acceleration are rising together. Design a cloud-edge hybrid inference architecture, and compare NVIDIA acceleration solutions, TPU optimisation and Samsung-Mistral-style partnerships, so as to secure several compute procurement options for next quarter.
- Assume a more complex regulatory environment and more demanding cybersecurity benchmarks. Build an internal governance system ahead of time, covering model evaluation, red-teaming and audit logs, and state it in your product descriptions as a trust differentiator for enterprise customers.
Based on 1022 items over 29 days