Jul 19 – Jul 25, 2026
English translation of the Korean original, prepared with AI assistance. Korean original
Not investment advice — these are resources for learning and forming your own view.
The strategist's view
The money in AI now comes not from who has the smartest model but from who can run a single inference most cheaply.
Taken separately, this week's news looks like unrelated events: the rise of China's Kimi, Google's cloud profits and in-house chips, Nvidia's next-generation inference accelerator, and the White House debate over open weights. Taken together, they tell one story. Model performance is converging upward fast, so the premium from simply owning a model is thinning. Profit is flowing to the infrastructure and operating capability that runs inference cheaply and reliably. AI differs decisively from traditional software in that inference cost rises along with revenue. As scale grows, cost efficiency becomes a matter of survival.
So the question leaders should ask changes. It is not where our model ranks on benchmarks. It is how much inference costs us each time we sell, and whether we hold the levers to push that cost down faster than revenue grows. Google is building that answer by pushing upstream cost down to a utility with its own chips. Those adopting open-weight models are finding the same answer by another route, eliminating license fees entirely.
The practical conclusions are three. First, split workloads by accuracy requirement and run a dual setup of closed and open models. Second, for cost items tied to external GPUs and APIs, restore bargaining power through multiple vendors, and concentrate in-house resources only on the data and industry workflows where real differentiation remains. Third, treat models as parts you can swap at any time, and dig your moat in the distribution, trust and regulatory compliance built on top. Organizations that cannot switch models will pay the highest price.
Cases through a framework 3
Who’s Afraid of Chinese Models?
The rapid rise of open-weight models and Chinese players is redefining the core cost structure of the AI market. Unlike traditional software, AI incurs cost at the inference stage that keeps growing in proportion to revenue. This has shown that the ability to build good models alone does not protect profit.
Value migration applies when customer needs or technology change and the source of profit moves. The key fact of this article is that as open-weight models catch up on performance for free, money once earned from the model itself is draining away. Profit is starting to flow toward the operating capability and infrastructure that run inference cheaply.
This news is usually read as a performance-ranking contest in which Chinese models threaten American ones. Through the value migration lens, though, the real event is not the ranking. It is that profit is evaporating from the model layer. The accurate reading is that even closed-LLM companies with seemingly healthy revenue have already begun to hollow out.
Decision prompt — If your business relies only on closed-model license revenue, then even if sales are strong now, shift your investment priorities to serving optimization, dedicated chips and caching that lower inference cost in proportion to revenue. Redesign your pricing so that margin is recovered not from the model itself but from the workflow, data and distribution layered on top.
About the framework
Industry profit does not stay in one place. It flows to the business models that better meet what customers really want. When technology or customer needs change, places that once made money empty out and value moves to entirely different places. So a company from which value has already begun to drain hollows out slowly even while its revenue still looks healthy. The contest is decided not by where you earn now but by reading where profit is flowing.
Stratechery
Google justifies its massive AI spending with a booming cloud business
Google has shown that it is recouping its huge AI investment through record results in its cloud business. At the same time, it has begun developing a dedicated AI chip to raise Gemini's compute efficiency. It is pushing costs down across its own value chain, from models to chips to data centers.
Wardley Mapping is used to judge what to build in-house and what to hand off as a commodity by looking at the maturity of each component in the value chain. Google is moving from buying GPUs to its own chips, and pushing model serving down into a cloud utility. This is exactly that landscape judgment: push down the layers about to become commodities and capture margin on top of them.
This news is usually read as an earnings story: Google Cloud is doing well, so its AI investment is justified. Through Wardley Mapping, though, the point is not the earnings. It is a shift in the landscape. Google is trying to turn its upstream cost, dependence on Nvidia, into a utility with its own chips, and so to break the cost structure its competitors pay.
Decision prompt — Draw a map of which layers of the AI stack will become commodities (utilities whose prices fall) within the next 12 to 18 months. For upstream cost items tied to external GPUs and APIs, regain bargaining power through multiple vendors and in-house optimization. Concentrate in-house development resources only on the data and industry workflow layers where real differentiation remains.
About the framework
To build a strategy, you must first see the landscape. Wardley Mapping places customer value at the top and links the components needed to deliver it in a chain beneath. It then positions each component along a horizontal axis by its maturity (genesis, custom-built, product, utility). This makes it visible on the map what will soon become a commodity and lose value, and where to invest in-house and where to outsource. You can then act on the landscape rather than on instinct.
TechCrunch
OpenAI is scared of open-weight models. Should the US be?
Leading closed-LLM companies are worried about the surge and spread of open-weight models. There is even discussion, on national security grounds, of restricting models from particular regions. Open models may not be the top performers, but they are rising quickly into the mainstream along a different axis: free, open and self-deployable.
Disruptive innovation applies when a new technology is inferior on the established performance measures but superior on other axes such as price and accessibility, and when incumbents, optimized for high-margin customers, find it hard to defend the low end. Open-weight models are now rising from the low-cost, self-hosted niche on the axis of slightly lower performance but free and runnable on my own server. That fits the condition exactly.
This news is usually read through a political and security frame: OpenAI is trying to block competition through regulation. Through the disruptive innovation lens, though, the call for regulation is itself the classic defensive reaction of an incumbent to an inferior technology climbing up from the low end. Whether or not regulation passes, the structural upward erosion by fast-improving open models should be expected to continue.
Decision prompt — Do not make top closed-model performance your only standard. Run a dual setup in which high-volume, low-cost workloads that tolerate slightly lower accuracy go to open-weight models, lowering vendor lock-in and inference cost together. Lay down an abstraction layer in advance so that models can be swapped, in case of regulatory scenarios such as a ban on models from certain regions.
About the framework
What topples incumbents is usually not a better technology but a worse one. A cheap, simple, lower-performing product starts in a low-end niche that mainstream customers ignore, and climbs upward as its performance improves. Christensen's paradox is that incumbents do not lose through incompetence. They lose because they managed rationally, listening faithfully to their most profitable customers, so they abandoned the low end and were slow to respond.
TechCrunch
What to watch
- Measured comparisons of inference TCO and power efficiency between the AMD Helios rack system and Nvidia Vera Rubin, to see whether the GPU cost structure is actually diversifying
- Progress of the AI Kill Switch Act and the debate over restricting open-weight models, a regulatory risk that could forcibly disturb the ability to swap models
- Actual deployment of Google's in-house AI chip and other hyperscalers' responses with their own silicon, which sets the speed of the race to push inference cost down to a utility
Based on 84 items over 6 days