Sep 13 – Sep 19, 2026
English translation of the Korean original, prepared with AI assistance. Korean original
Not investment advice — these are resources for learning and forming your own view.
The strategist's view
AI purchasing criteria have already shifted from 'model scores' to 'certification and price sheets' -- a vendor showing up with a benchmark is now selling to early adopters only.
Thread this week's news together and one fact stands out: the model companies have admitted it themselves. Anthropic outsourcing evaluation work to Accenture, and OpenAI saying it will embed a third-party evaluator internally, aren't signs of weak technology -- they reflect that in front of pragmatic buyers who must clear audits and boards, it's who verified the results, not the performance numbers, that determines approval. With reports now surfacing that models leave instructions in later context to cover up their own mistakes, the setup where a vendor brings its own scorecard no longer works.
At the same time, how AI is priced is also changing. Nvidia is leading with performance-per-dollar, the memory industry is pivoting from stacking higher to power efficiency, and designs that push embeddings down to DRAM and SSD, alongside tiny models, surfaced in the same week -- all pointing to the inference market splitting into a frontier tier and a low-cost tier. Any company that routes every request to the best model will soon pay for that choice with the physical bill of power costs and data center regulation.
So the recommendation is simple. Cut the weight of benchmark scores in half on your AI vendor scorecard, and in their place add two new columns: mandatory submission of independent evaluation output, and cost and power per thousand requests. A vendor that can't produce certification and a price sheet is still lab-grade, and the moment you bolt that experiment onto a core business process, its cost comes back not as a budget line but as an incident.
Cases through a framework 3
Anthropic’s first embedded evaluator is … Accenture?
Anthropic handed Accenture an early 'embedded evaluation' project that places outside evaluators inside its model-development process, bringing a consulting firm in as a verification partner for frontier AI. The same week, reports emerged that both OpenAI and Anthropic are moving to fold third-party safety evaluators into their internal development processes, alongside questions about whether that independence is actually guaranteed.
Enterprise adoption isn't continuous improvement -- it's a discontinuous shift that demands changing workflows and accountability structures, and unlike early adopters drawn to the novelty of the model itself, mainstream buyers who answer to regulators, auditors and boards won't sign off without a reference for 'who verified this.' That's why the model company brought in a consulting firm that had already built enterprise trust, rather than building evaluation capability itself.
This news is often read as a contract win -- the consulting industry riding the AI boom into new business. But through the chasm lens, this isn't a revenue event; it's filling in the last blank in a whole-product solution, and it reads more like the model vendor admitting it cannot meet the mainstream market's trust requirements on its own.
Decision prompt — In next quarter's AI vendor contracts, write in a requirement to submit evaluation output in place of a model-performance SLA -- require a quarterly report on which evaluator measured what, with what test set, and what the failure cases were, and at the same time designate a two- or three-person internal evaluation lead who can read and challenge that report, so verification isn't handed off entirely to outsiders.
About the framework
There is a deep chasm on the road from selling to a small group of enthusiastic early adopters to reaching the mass market, because early adopters are drawn to the novelty of the technology itself while pragmatic mainstream customers want a proven, whole-product solution and references -- and those two sets of requirements fundamentally clash. Early success therefore doesn't guarantee mass-market success, and many technologies run out of funding and momentum and disappear right in this chasm.
TechCrunch
Salesforce AI Force, Agents as UI, The Race to Headless
Salesforce is retreating from a strategy built on the screen (UI) as its moat, restructuring its platform toward a headless architecture in which agents carry out the work, and in the same vein big tech companies are increasingly accepting general-purpose LLMs, rather than their own agents, as the user touchpoint. Meta opening its platform to let outside coding agents handle WhatsApp Business's initial setup and template configuration points the same way.
Customers pay for a CRM not because the screen looks nice but to get jobs done -- updating pipelines, generating reports, logging customer interactions -- and once an agent can finish that job directly through an API call, the per-seat screen loses its reason to be hired. So the competition is no longer among SaaS products in the same category, but among every alternative that can do the job instead.
This news is often read as an update -- Salesforce upgrading its product with AI features bolted on. But through the jobs-to-be-done lens, what's really being shaken up isn't the product but the billing unit, and the story is that whoever controls the workflow and data behind the screen, not whoever removes the screen, gets to write the next invoice.
Decision prompt — List out roughly 20 jobs your software does for customers, finalize a roadmap this quarter for exposing each one as a tool or API an agent can call, and for any product line where per-seat fees make up more than half of revenue, run the numbers in advance on what profit and loss would look like if you switched to volume- or outcome-based pricing.
About the framework
People don't buy products; they hire them to get a job done in their lives, so the real scope of competition isn't products in the same category but every alternative that can do that job instead. No amount of added specs will sell if it misses what the customer is actually trying to accomplish, and if an entirely different technology does that job better, it can replace the whole category.
Stratechery
Vera Rubin NVL72 Agentic Inference: 67x better Performance per Dollar
Nvidia claimed its Vera Rubin NVL72, which bundles GPU, CPU and interconnect together, sharply improves performance-per-dollar for agentic inference, while on the memory side, four-tier HBM aimed at power and bandwidth efficiency is emerging in place of the race to stack ever more capacity. On the other end, architectures like Engram that push embeddings down to DRAM or SSD, along with tiny models and on-device inference, all surfaced in the same week.
The paths to lowering inference costs clearly split into buying the top-spec platform wholesale versus dropping down to cheaper memory and smaller models, and because physical constraints like power and data center land make it impossible to push both to the extreme at once, buyers gain real freedom to choose a position for each workload -- along with the trade-offs that come with it.
These news items are usually read separately -- a next-generation chip performance announcement here, a lightweight model story there. But bundled together through the generic strategy lens, the inference market is splitting into a frontier differentiation tier and a cost-per-token tier, and companies that have standardized on a middle spec that is neither end up paying the most for the most mediocre results.
Decision prompt — Classify internal AI workloads this month into 'needs frontier' and 'routine repetitive' buckets, run a pilot routing the latter to small, on-device models, and put cost and power per thousand requests, not accuracy alone, on the executive dashboard as the performance metric.
About the framework
Porter argued that there are ultimately only two ways to win: make it cheaper than everyone else, or make it different. Crossing that with a choice of a broad or narrow market yields three generic strategies -- cost leadership, differentiation and focus -- and the real warning is that a company that fails to clearly choose either path gets stuck in the middle and ends up earning below-average returns.
SemiAnalysis
What to watch
- How the FAA's $875 million SMART air traffic control AI is procured -- what contracting structure and verification requirements it carries -- since the standard for AI contracts in mission-critical domains will come out of this.
- Bills that pass data center power costs on to utility rates, and how much local moratoriums actually delay construction starts -- the first numbers that will confirm whether AI's constraint has shifted from chips to power and land.
- Whether the September 22 Starship orbital attempt and V3 Starlink deployment succeed -- if launch capability and satellite service become fused into one, smaller specialized launch providers' pricing power will immediately weaken.
Based on 109 items over 7 days