The Grid Is the New GPU

Why energization rights, interconnection capacity and firm power are becoming strategic assets in the AI economy.

Solten & Co. Research Report
Published/updated: August 20, 2026
Research universe: AI Infrastructure / Energy / Data Centers / Capital & Deals

Research Snapshot

Research type: Flagship Research Report 

Evidence cut-off: August 20, 2026

Estimated reading time: ~30 minutes

Executive Summary

The AI infrastructure race is moving into a new phase. For the past several years, the scarcest strategic inputs were advanced accelerators, high-bandwidth memory, networking equipment and the capital required to buy them. Those constraints have not disappeared. But a more physical bottleneck is moving upstream: the ability to energize a data center at the scale and on the timetable that AI developers now require.

The change is visible in the numbers. Lawrence Berkeley National Laboratory’s June 2026 update estimates that U.S. data centers could consume 11.8% of total U.S. electricity in 2030 in its reference case, with sensitivity scenarios ranging from 9.5% to 15.3%. The same report estimates roughly 148 GW of interconnection capacity could be required for data centers by 2030 under a 50% average utilization assumption. AI servers are the dominant driver of the increase.

That load is arriving much faster than the infrastructure built to serve it. An advanced data center can be developed in roughly two to three years, while Berkeley Lab’s latest national interconnection analysis shows that the median path from a generation interconnection request to commercial operation now exceeds five years in the regions with available data. The result is a structural timing mismatch: compute capital can be committed faster than reliable power can be connected.

This changes what “capacity” means in AI infrastructure. A site with land, fiber and a data-center shell is not necessarily a usable AI asset. The economically scarce object is increasingly an energization right: a credible path to a specific quantity of power, at a specific location, by a specific date, with acceptable reliability, cost and curtailment terms. In markets where grid capacity is tight, queue position, executed utility agreements, transmission access and generation arrangements can become as strategically important as the servers themselves.

Virginia offers a revealing leading indicator. Dominion Energy says it has assigned energization dates through 2031 to 25 GW of new data-center projects for which it has current or soon-to-be-connected capacity. Another 45 GW of proposed new data-center projects do not yet have future connection dates. At the same time, the Virginia State Corporation Commission has created a separate large-load rate class that includes long-term minimum obligations, minimum monthly transmission and distribution charges, and potential collateral requirements. The power contract is beginning to resemble an infrastructure commitment rather than a simple utility bill.

Grid operators are adapting as well. PJM spent 2025–2026 developing special frameworks for large load additions, including “connect-and-manage” concepts and possible pathways for customers that bring new generation or accept curtailment. ERCOT introduced a batch process for loads of 75 MW or greater so the system can assess large projects collectively rather than one by one. In other words, data-center interconnection policy is becoming part of AI industrial policy.

The strategic response from technology companies is increasingly visible in power procurement. Microsoft’s 20-year agreement with Constellation supports the restart of the 835 MW Crane Clean Energy Center. Google’s agreement with Kairos Power creates a path to up to 500 MW of advanced nuclear capacity by 2035, with the first deployment targeted for 2030. Meta signed a 20-year agreement supporting 1,121 MW at Constellation’s Clinton plant while separately pursuing new nuclear projects. These transactions should not be read simply as sustainability initiatives. They are long-duration attempts to secure firm power and reduce future energy optionality risk.

The central Solten & Co. conclusion is that the AI infrastructure stack is being reordered. Chips remain essential, but chips depreciate quickly and can be procured from multiple vendors over time. A credible high-voltage interconnection, transmission pathway and firm-power arrangement can take many years to create and may be much harder to replicate. In constrained regions, the most durable infrastructure moat may therefore move from ownership of compute equipment toward control of time-to-power.

Key Findings

  1. Power is becoming a first-order AI input. Data centers are no longer a marginal electricity category in the United States; the 2030 reference case from Berkeley Lab implies roughly one-eighth of national electricity consumption.
  2. The critical scarcity is not electricity in the abstract but deliverable electricity at a specific place and time. Grid connection, transmission capacity, transformers, substations, firm generation and queue position determine whether nominal power supply can actually become usable AI capacity.
  3. AI development and grid development operate on different clocks. Data centers can be built faster than generation and transmission can be interconnected, creating a growing premium on pre-secured energization.
  4. Virginia shows how energization rights can acquire economic value. Dominion reports 25 GW of new data centers with assigned energization dates and another 45 GW without dates.
  5. Utility contracts are becoming strategic obligations. Virginia’s new large-load framework requires long contract periods and minimum payments designed to reduce the risk that ordinary ratepayers finance infrastructure for projects that later underutilize or abandon capacity.
  6. Grid policy is becoming part of AI competitive strategy. PJM and ERCOT are redesigning large-load processes around reliability, curtailment, generation contribution and project readiness.
  7. Firm power is creating a new class of technology-energy transactions. Nuclear restarts, life extensions, advanced nuclear development, onsite generation and storage are increasingly linked to technology-company load growth.
  8. Time-to-power may become a valuation variable. Data-center land with an executed path to hundreds of megawatts should not be valued like otherwise similar land with uncertain energization.
  9. The power bottleneck can reshape geography. AI clusters may migrate toward regions with faster interconnection, surplus generation, expandable transmission, fuel access and flexible large-load rules rather than simply toward historic cloud regions.
  10. The counter-thesis is meaningful. Better hardware efficiency, lower energy per AI task, load flexibility, grid reform and slower-than-expected project realization could reduce the severity of the bottleneck. The correct thesis is not “the U.S. will run out of electricity,” but that timely, location-specific, reliable capacity is becoming scarce enough to influence AI economics.

Table of Contents

  1. From GPU Scarcity to Power Scarcity
  2. The Demand Curve Has Crossed a System Threshold
  3. The Real Asset Is an Energization Right
  4. The Time-Scale Mismatch
  5. Virginia as a Leading Indicator
  6. Grid Rules Are Becoming AI Industrial Policy
  7. Power Contracts Are Becoming Strategic Obligations
  8. Why Firm Power Is Back
  9. Behind-the-Meter Power: Escape Valve, Not Free Lunch
  10. Flexibility Becomes a Currency
  11. The Capital Structure of AI Power
  12. Time-to-Power as a Valuation Variable
  13. Winners, Losers and Second-Order Effects
  14. Solten & Co. Thesis, Counter-Thesis and Falsification
  15. What to Watch Next

Sources & Evidence

Methodological Note

About Solten & Co.

Scope & Methodology

Research question. This report examines whether electricity infrastructure is becoming a binding strategic constraint on AI development in the United States, and what that shift means for data-center economics, hyperscaler strategy, utilities, power developers, infrastructure investors and adjacent technology markets.

Scope. The analysis focuses on U.S. data-center electricity demand, generator and large-load interconnection, selected regional examples, firm-power procurement and the financial implications of time-to-power. It does not forecast individual power prices, recommend securities or assume that every announced data-center project will be built.

Evidence hierarchy. Priority is given to Lawrence Berkeley National Laboratory, the International Energy Agency, FERC, PJM, ERCOT, state utility regulators, utilities and direct corporate announcements. Company energy agreements are treated as strategic evidence, not as proof that announced capacity will be delivered on schedule.

Critical limitation. Data-center project pipelines contain duplication, speculative requests and projects that will never reach operation. Electricity-demand forecasts also depend heavily on hardware efficiency, model architecture, utilization, deployment pace and AI adoption. The analysis therefore distinguishes announced load, interconnection capacity, contracted power and realized consumption.

What Changed

The power constraint is not new, but its strategic importance changed sharply in 2025–2026. The IEA estimates that capital expenditure by five major technology companies exceeded $400 billion in 2025 and is expected to rise another 75% in 2026. Its satellite-based tracking indicates that cutting-edge “AI factory” capacity more than tripled in roughly eighteen months. This is industrial expansion at a speed the electricity system was not designed to mirror.

At the same time, the energy intensity of individual AI tasks is falling quickly. The IEA reports at least an order-of-magnitude annual decline in energy per AI task in recent years. That would normally relieve infrastructure pressure. But the mix of work is changing toward reasoning, video and agentic tasks that can consume hundreds or thousands of times more energy than simple text queries, while usage itself is growing rapidly. Efficiency is therefore fighting a rebound effect rather than simply reducing total load.

The newest development is institutional. Large-load connection is no longer being treated as routine utility service. PJM, ERCOT, Virginia regulators and utilities are creating specialized rules around data centers and other large loads because a single project can now resemble a traditional power plant in scale while arriving on a technology-company timetable.

1. From GPU Scarcity to Power Scarcity

The first phase of the generative-AI infrastructure boom was dominated by accelerator scarcity. Access to NVIDIA GPUs became a competitive advantage, cloud capacity was rationed, and startups raised capital partly to secure compute. That framing remains useful, but it is incomplete in 2026.

A GPU cluster is economically useless without power. More importantly, the marginal difficulty of adding power is increasing as cluster size rises. Large AI campuses require substations, transformers, transmission upgrades, cooling systems, backup systems and generation resources capable of supporting loads measured in hundreds of megawatts or multiple gigawatts. The infrastructure problem therefore moves upstream from the server rack into the electricity system.

The IEA’s 2026 analysis highlights the physical intensity of this transition. Between 2020 and 2025, AI-server power density increased roughly eleven-fold, and by 2027 it is expected to rise another four-fold. An advanced rack could reach peak demand comparable to 65 U.S. households. Higher density improves the economics of scarce data-center floor space, but it concentrates power and thermal requirements into a much smaller physical footprint.

The strategic implication is straightforward: the AI industry can manufacture more computational density faster than the grid can manufacture new delivery capacity.

Exhibit 1 — U.S. Data Centers Are Moving From a Large Load to a System-Level Load

Berkeley Lab’s 2026 update moves the discussion beyond anecdotes. Its reference case reaches 11.8% of total U.S. electricity consumption by 2030, with a 9.5%–15.3% sensitivity range. The report estimates that AI servers account for 84% of projected server energy use and 55% of all projected data-center electricity use by 2030.

That does not mean AI will consume a fixed share regardless of price or infrastructure constraints. It means data centers are now large enough to influence generation planning, transmission investment, rate design and regional resource adequacy.

2. The Demand Curve Has Crossed a System Threshold

Electric systems routinely absorb new industrial loads. What makes AI different is the combination of size, concentration, speed and uncertainty.

Size matters because a single campus can require hundreds of megawatts. Concentration matters because data centers cluster around fiber, existing cloud regions, skilled labor and tax regimes. Speed matters because technology companies can finance and build a campus much faster than a transmission corridor or major generating plant. Uncertainty matters because announced projects can be delayed, resized, duplicated across utility queues or cancelled.

The IEA emphasizes this distinction: globally, data centers remain a modest share of electricity use, but they create outsized local integration challenges because demand is geographically concentrated. This is why national electricity abundance does not solve a local interconnection shortage.

For AI investors, the right question is therefore not “Does the United States have enough electricity?” It is “Can this specific site receive the required power, on the required date, under terms that remain economic if utilization is lower than planned?”

3. The Real Asset Is an Energization Right

In conventional data-center analysis, investors often focus on land, building cost, fiber connectivity, cooling and server economics. The current market adds another asset class: a credible energization pathway.

An energization right is not a formal legal category. It is an analytical way to describe the bundle of permissions, infrastructure and contractual commitments that allow a site to draw meaningful power. It can include utility service agreements, queue position, completed studies, transmission upgrades, substation capacity, generation contracts, transformer procurement and regulatory approvals.

This distinction matters because two sites with identical acreage can have radically different economic value. One may have an executable path to 500 MW in 2028. Another may have theoretical access to a large regional grid but no credible connection date before 2032. The second site is not merely “later.” It may miss an entire hardware and model cycle.

In AI infrastructure, time has unusually high option value. A two-year delay can mean several generations of accelerators, lower model costs, different cooling architectures and changed demand assumptions. That makes a secured energization date economically closer to a scarce real option than to a routine utility connection.

4. The Time-Scale Mismatch

Exhibit 2 — AI Infrastructure Is Being Built on a Faster Clock Than the Grid

The mismatch is visible in development timelines. The IEA notes that a data center can be operational in roughly two to three years. Berkeley Lab’s July 2026 interconnection update shows that the median duration from generator interconnection request to signed interconnection agreement was well above three years in 2025, while the path to commercial operation exceeded five years in regions with available data.

These are not perfectly comparable processes: one describes building a load asset, the other adding generation. But that is precisely the point. Demand can materialize faster than supply can be studied, permitted, financed, connected and commissioned.

The queue itself is enormous. Berkeley Lab counted 2,061 GW of generation and storage actively seeking U.S. interconnection at the end of 2025. More than 750 GW of requests were withdrawn during the year, and historically only 13% of capacity submitted from 2000–2020 had reached operation by the end of 2025. A large queue therefore does not equal a large pipeline of near-term power.

This is why nominal resource announcements can mislead AI infrastructure investors. A region may have hundreds of gigawatts “in queue” while still lacking enough firm, deliverable capacity for a data-center campus on the required timetable.

5. Virginia as a Leading Indicator

Northern Virginia became the world’s most important data-center market because it combined fiber, cloud-network effects, business density, land development expertise and historically reliable power access. Its current constraints therefore deserve attention as a preview of what can happen elsewhere.

Exhibit 3 — In Virginia, the Scarce Asset Is Increasingly an Energization Date

Dominion Energy reports that it has assigned energization dates through 2031 to 25 GW of new data centers where capacity is currently available or expected to be connected. For another 45 GW of proposed new data-center projects, future connection dates have not yet been offered.

The striking point is not that all 70 GW will be built. They almost certainly will not be. The point is that the request pipeline is large enough that the utility must explicitly ration certainty about when projects can connect.

Virginia regulators have also changed the commercial relationship. The State Corporation Commission established a separate GS-5 rate class for large loads. New qualifying customers contracting from 2027 are subject to a minimum 14-year service obligation. Large-load customers are generally required to pay at least 85% of the transmission and distribution costs incurred to serve them each month regardless of actual usage, and some customers may need to provide collateral covering a substantial portion of minimum charges.

This is economically important. Utilities are effectively saying: if the grid makes long-lived investments for an AI campus, the customer must assume part of the stranded-asset risk.

6. Grid Rules Are Becoming AI Industrial Policy

When connection to the grid determines where AI infrastructure can be built, interconnection rules become a competitive-policy variable.

PJM, the largest U.S. regional transmission organization, initiated an accelerated process in 2025 to address large load additions and spent 2026 developing reliability-focused solutions. One direction is a “connect-and-manage” framework in which certain large loads could connect before all necessary system upgrades are complete if they accept curtailment or otherwise operate within reliability limits. PJM has also explored distinctions between loads that bring new generation and those that do not.

ERCOT moved in a different but related direction. In June 2026, Texas regulators approved “Batch Zero,” a process that groups qualified large loads of 75 MW or greater so ERCOT can assess their combined system impact, allocate available capacity and identify transmission needs. The logic resembles modern generator-queue reform: study projects as a portfolio rather than allowing a flood of individually evaluated requests to overwhelm the system.

These changes create strategic questions for AI operators. Is a flexible connection better than waiting years for fully firm service? Is building or contracting new generation worth faster connection? How much curtailment can a training cluster tolerate? Can inference workloads be shifted geographically or temporally? The answer will vary by workload.

Grid design is therefore becoming part of systems architecture.

7. Power Contracts Are Becoming Strategic Obligations

AI companies are already accustomed to large, long-term compute commitments. The energy layer is developing similar characteristics.

A multi-year power arrangement can lock in access to scarce capacity, support financing of new generation and improve certainty for a data-center build. But it also creates risk if model economics, hardware efficiency or demand growth change faster than expected.

The Virginia GS-5 structure makes the analogy explicit. Minimum payment obligations and long contract terms transform electricity from a fully variable operating cost into something closer to a quasi-fixed strategic commitment. The customer is not legally issuing debt, but economically it may be assuming a long-duration obligation tied to infrastructure that was built for its anticipated load.

This mirrors a broader pattern already visible in AI compute contracts: the industry is trading future flexibility for present capacity.

8. Why Firm Power Is Back

Renewables remain essential to data-center power procurement, but very large AI loads place a premium on firm capacity: electricity that can be delivered reliably across hours and seasons, not merely matched annually through renewable-energy certificates.

The strategic value of firm power helps explain the new technology-company interest in nuclear energy.

Microsoft / Constellation. A 20-year power purchase agreement supports the restart of the former Three Mile Island Unit 1 as the Crane Clean Energy Center, expected to add approximately 835 MW of carbon-free generation to the grid.

Google / Kairos Power. Google and Kairos created a multi-plant development agreement for up to 500 MW of advanced nuclear generation by 2035, with the first deployment targeted for 2030.

Meta / Constellation. Meta signed a 20-year agreement supporting continued operation of the 1,121 MW Clinton Clean Energy Center beginning in 2027, while separately pursuing new nuclear capacity through a broader request-for-proposals process.

These transactions are structurally different and should not be added together as if they were comparable delivered capacity. Some preserve existing plants, some restart retired assets, and some attempt to commercialize new reactor designs. What they share is the willingness of technology companies to make long-duration commitments to power infrastructure because future firm capacity has strategic value.

The energy company is becoming part of the AI supply chain.

9. Behind-the-Meter Power: Escape Valve, Not Free Lunch

Slow grid connections are pushing some developers toward onsite or behind-the-meter generation. The attraction is obvious: if the grid cannot deliver power quickly enough, build generation closer to the load.

The IEA’s 2026 work identifies onsite natural-gas generation as an emerging U.S. data-center response and notes that a meaningful share of tracked projects has already begun land clearing or construction. But onsite generation does not eliminate infrastructure constraints; it changes them.

A reliable isolated system needs redundancy. The IEA estimates that reliable onsite gas generation for critical, variable data-center load may require 30%–70% overbuilding relative to demand. Developers then face turbine availability, gas-pipeline capacity, emissions permitting, maintenance, noise, local opposition and financing requirements.

Behind-the-meter power is therefore not “free from the grid.” It is a substitution of one infrastructure stack for another.

10. Flexibility Becomes a Currency

The easiest data-center load for a power system to serve is not necessarily the smallest. It is the load that can adapt when the system is stressed.

AI infrastructure has several potential flexibility levers: delaying non-urgent training runs, shifting training across regions, modulating batch inference, using batteries for short-duration support, operating backup or onsite generation, and designing software to move work between facilities.

The IEA estimates that 20–25 GW of battery storage could be installed in data centers globally by 2030. If appropriately controlled and compensated, some of that storage can support both the facility and the wider grid.

This creates a new trade: faster grid access in exchange for operational flexibility. The emerging PJM connect-and-manage concept illustrates the direction. A data center that can curtail during limited system events may be economically easier to connect than one demanding fully firm, inflexible capacity from day one.

For AI architects, resilience and electricity-market participation may therefore become part of workload orchestration.

11. The Capital Structure of AI Power

Power scarcity changes capital requirements across the AI value chain.

First, data-center developers may need to finance substations, transmission upgrades, generation assets and long-lead electrical equipment earlier in the project cycle.

Second, utilities require greater certainty that large-load customers will actually materialize. Minimum contracts, collateral and readiness milestones transfer more development risk back to the customer.

Third, generation developers gain a new class of anchor offtaker. A creditworthy technology company willing to sign a long-term agreement can make a nuclear restart, gas plant, renewable project, storage installation or advanced-energy demonstration financeable.

Fourth, private infrastructure capital moves closer to AI economics. Investors financing generation and grid assets increasingly need a view on model demand, data-center utilization and hyperscaler capital spending because those variables determine whether the load supporting the infrastructure remains durable.

The separation between “technology investing” and “energy infrastructure investing” is becoming less useful.

12. Time-to-Power as a Valuation Variable

The most important investment implication may be a change in how AI infrastructure assets are valued.

A traditional data-center valuation can emphasize leased megawatts, utilization, rent, replacement cost, tenant quality and cap rates. In the AI era, investors may need to add a more explicit measure: time-to-power certainty.

Consider two otherwise similar development sites:

  • Site A has land, fiber, permits and an executed utility pathway to 300 MW by 2028.
  • Site B has cheaper land and excellent fiber but no credible energization date before 2031.

The difference is not simply three years of lost rent. Site A can host hardware generations, model launches and customer demand that Site B may never capture. It can also give the tenant negotiating leverage in compute procurement because the limiting input has already been secured.

This suggests a new hierarchy of AI infrastructure assets:

  1. Powered and operating capacity.
  2. Contracted capacity with high-confidence energization dates.
  3. Advanced-stage capacity with identified upgrades and generation.
  4. Speculative powered-land claims without firm delivery dates.
  5. Land with only conceptual access to future grid capacity.

Markets that fail to distinguish these categories risk overvaluing “gigawatts” that are not economically deliverable.

13. Winners, Losers and Second-Order Effects

Potential winners

Existing firm generation. Nuclear plants, efficient gas generation, hydro and other reliable assets gain strategic value when power becomes scarce.

Utilities and transmission developers that can add capacity quickly. Regions able to plan and build credible infrastructure can attract high-value AI investment.

Electrical equipment suppliers. Transformers, switchgear, power electronics, turbines, cooling and storage become critical links in the AI supply chain.

Data-center developers with real interconnection rights. A credible power position can differentiate otherwise commoditized real estate.

Flexible AI operators. Companies able to shift workloads, use storage or tolerate curtailment can monetize flexibility through earlier or cheaper connections.

Potential losers

Speculative data-center land. Sites marketed around theoretical future power may be repriced as buyers demand harder evidence of energization.

Inflexible large loads. Customers requiring fully firm service at all times may face higher connection costs and longer delays.

Ratepayers if cost allocation is poorly designed. If utilities build expensive infrastructure for projects that later disappear, stranded costs can shift to other customers. This risk is precisely why new large-load tariffs are emerging.

AI projects with mismatched commitments. A company can over-contract both compute and power if demand or monetization fails to scale.

Second-order effects

Power constraints can alter AI geography. Regions with available generation, fuel, transmission and faster permitting may gain share from established hubs. They can also change semiconductor economics: efficiency per watt becomes more valuable when the bottleneck is electricity rather than chip supply alone.

Third-order effects reach industrial policy. Governments increasingly face trade-offs among data centers, manufacturing, household affordability, electrification and grid reliability. Decisions about who pays for new generation and transmission can influence where the next AI clusters are built.

14. Solten & Co. Thesis, Counter-Thesis and Falsification

Solten & Co. Thesis

The durable bottleneck in frontier AI infrastructure is moving upstream from accelerator procurement toward the ability to secure and energize large quantities of power on a predictable timetable. As this happens, queue position, utility contracts, transmission access, firm generation and load flexibility acquire strategic and financial value. In constrained regions, time-to-power becomes a competitive moat.

Counter-Thesis

The market may be extrapolating peak infrastructure demand too aggressively. Energy efficiency per AI task is improving extremely quickly. Specialized chips, better cooling, higher utilization, workload routing and lower-cost models can reduce power per unit of useful output. Many announced data-center projects will never be built. Grid reform can accelerate interconnection, and flexible loads can reduce the requirement for new firm capacity. If AI monetization grows more slowly than infrastructure commitments, today’s perceived power scarcity could turn into localized overcapacity.

Falsification Criteria

The thesis would weaken materially if several of the following occur:

  • Berkeley Lab’s U.S. data-center electricity-demand trajectory is repeatedly revised sharply downward because AI-server deployments or utilization disappoint.
  • Median generator and large-load interconnection timelines fall enough that power no longer constrains site delivery.
  • Major data-center markets develop large surplus generation and transmission capacity without materially higher customer costs.
  • AI efficiency gains consistently outpace growth in model usage and capability, reducing aggregate electricity demand.
  • Large-load queues experience very high cancellation rates without corresponding executed projects, revealing much of the apparent scarcity as speculative duplication.
  • Corporate firm-power agreements fail to expand beyond a small number of flagship transactions.
  • Utilities stop requiring special tariffs, minimum obligations or collateral because stranded-asset risk proves immaterial.

15. What to Watch Next

Data-center electricity share. Track the 2026 Berkeley Lab reference case against actual server shipments, utilization and regional load.

Energization queues. The ratio of projects with assigned connection dates to projects merely requesting capacity may become more informative than headline gigawatts.

Large-load tariff design. Watch minimum contract terms, collateral, cost allocation and curtailment rules across Virginia, Texas, PJM states and other fast-growing markets.

Bring-your-own-generation frameworks. A wider use of customer-supported generation would strengthen the thesis that power procurement is moving inside AI infrastructure strategy.

Nuclear execution. Restarts, uprates and advanced-reactor milestones matter more than announcement volume. Delivery schedule is the key evidence.

Onsite gas and storage. Growth would indicate that developers are willing to pay a premium to bypass grid timing constraints.

Transformer and turbine lead times. If these equipment bottlenecks remain severe, nominal generation investment may still fail to translate into rapid energization.

Geographic migration. Watch whether new AI campuses shift toward regions with faster time-to-power even when those regions are less established as cloud hubs.

Power intensity per useful AI outcome. The long-term balance depends on whether efficiency gains outrun the growth in reasoning, agentic and multimodal workloads.

Sources & Evidence

Primary / authoritative research and system sources

Lawrence Berkeley National Laboratory — United States Data Center Energy Usage Report: 2025 Update, June 2026

https://datacenters.lbl.gov/publications/united-states-data-center-energy-2025

Lawrence Berkeley National Laboratory — Queued Up / U.S. generator interconnection update, July 1, 2026

https://emp.lbl.gov/news/backlog-power-plants-seeking-transmission-grid-connection-eased-somewhat-2025-amidst

International Energy Agency — Key Questions on Energy and AI, April 16, 2026

https://www.iea.org/reports/key-questions-on-energy-and-ai

International Energy Agency — Energy and AI, April 2025

https://www.iea.org/reports/energy-and-ai

Federal Energy Regulatory Commission — Order No. 2023, Generator Interconnection Reforms

https://www.ferc.gov/explainer-interconnection-final-rule

PJM — Critical Issue Fast Path: Large Load Additions

https://www.pjm.com/committees-and-groups/cifp-lla

PJM — Connect and Manage Senior Task Force

https://www.pjm.com/committees-and-groups/task-forces/camstf

ERCOT — Large Load Integration / Batch Zero

https://www.ercot.com/services/rq/large-load-integration

ERCOT — PUCT Approves Batch Zero Process, June 18, 2026

https://www.ercot.com/news/release/06182026-puct-approves-ercots

Virginia State Corporation Commission — Data Center Initiatives / GS-5 large-load framework

https://www.scc.virginia.gov/about-the-scc/scc-facts/

Dominion Energy Virginia — Meeting the Demands for Large-Load Customers, 2026

https://sustainability.dominionenergy.com/GS-5%20Large%20Load%20Rate%20Class%20Report.pdf

Selected corporate firm-power signals

Constellation — Microsoft / Crane Clean Energy Center, September 20, 2024

https://investors.constellationenergy.com/news-releases/news-release-details/constellation-launch-crane-clean-energy-center-restoring-jobs

Google — Kairos Power advanced nuclear agreement, October 14, 2024

https://blog.google/company-news/outreach-and-initiatives/sustainability/google-kairos-power-nuclear-energy-agreement/

Kairos Power — Google partnership / 500 MW deployment pathway

https://www.kairospower.com/google

Meta — Constellation nuclear agreement, June 3, 2025

https://about.fb.com/news/2025/06/meta-constellation-partner-clean-energy-project/

Methodological Note

Electricity-demand forecasts, data-center pipelines and interconnection queues should not be treated as committed realized capacity. This report intentionally separates electricity consumption, requested interconnection capacity, announced projects, assigned energization dates, executed power contracts and operating generation. Quantities from different categories are not added together.

The term energization right is a Solten & Co. analytical concept, not a regulatory or accounting classification. It describes the economic value of having a credible, sufficiently advanced path to deliver a defined quantity of power to a site by a useful date.

About Solten & Co.

Solten & Co. is an independent research and analysis firm focused on the AI economy, with deeper research emphasis on AI infrastructure, software AI, Physical AI, robotics and autonomous systems. We study the technologies, companies, markets, transactions and capital structures shaping the next phase of AI.

Need an independent perspective?

Solten & Co. provides independent research and analytical support for investors and decision-makers evaluating companies, markets and investment opportunities across the AI economy.

Discuss a research question · Discuss an investment opportunity · Request independent analysis

soltenco.com

Etched: The $21 Billion Bet on Specialized AI Inference

What a $700 million financing round reveals about the economics of inference, the limits of GPU generality, and the emerging contest to reshape AI infrastructure

Solten & Co. Deal Analysis 
Published / updated: August 20, 2026 
Research universe: AI Infrastructure / Semiconductors / Capital & Deals

Research Passport

Company: Etched
Headquarters: San Jose, California
Founded: 2022
Founders: Gavin Uberti, Robert Wachen, Chris Zhu

Transaction: $700 million financing
Announced valuation: $21 billion
Date announced: August 18, 2026
Lead investor: Jane Street
Other disclosed participants: Kleiner Perkins, Sequoia, Andreessen Horowitz, Tiger Global, Bain Capital Ventures, Neo, Primary, Stripes, Positive Sum, and Blackstone

Company-stated total funding: $1.9 billion
Evidence cut-off: August 20, 2026
Research type: Deal Analysis
Estimated reading time: ~35 minutes

Executive Summary

 

Etched’s latest financing is not interesting primarily because a three-year-old semiconductor company raised $700 million. It is interesting because investors are assigning a $21 billion value to a company whose commercial history is still extremely short, whose first customer rack was delivered only last month, whose publicly identified customer base remains minimal, and whose most important performance claims are not yet supported by broad independent benchmark evidence.

 

At the same time, dismissing the valuation as simple AI exuberance misses what investors may actually be underwriting. Etched sits at the intersection of several powerful structural forces: inference is becoming the dominant recurring compute burden in AI; model-serving economics increasingly depend on tokens per dollar and tokens per watt; hyperscalers and model labs are actively seeking alternatives to a single-vendor GPU stack; and the AI infrastructure market is beginning to reward systems designed around specific workload economics rather than general-purpose programmability.

 

The company has also changed materially. In 2024 Etched publicly presented itself as a radical transformer-only ASIC company. The proposition was intentionally narrow: hardwire the dominant model architecture into silicon and sacrifice generality for extraordinary efficiency. By mid-2026, Etched’s public positioning had broadened into “frontier inference clusters” — co-designed chips, memory, interconnect, racks, cooling, software and manufacturing. Its current materials emphasize Low Voltage Inference and Cluster Scale Memory, and the company says its systems can run large mixture-of-experts and non-transformer designs. That evolution reduces one of the original thesis risks, but also means the company should no longer be analyzed simply as “the transformer ASIC startup.”

 

The current round contains an unusually strong strategic signal: Jane Street is simultaneously the lead investor and Etched’s first disclosed customer. Jane Street says it tested the chip, received the first rack in July and is deploying it in production workloads. This matters because Jane Street is not a passive financial sponsor. Earlier in 2026 it committed approximately $6 billion to CoreWeave for AI cloud capacity and invested $1 billion in CoreWeave equity. Its investment in Etched is therefore consistent with a broader strategy of controlling access to high-performance compute for latency- and research-intensive workloads.

 

The central valuation question is severe. A $21 billion entry valuation means that, before accounting for future dilution or preference terms, investors need an eventual company value of roughly $42 billion for 2x, $63 billion for 3x and $105 billion for 5x. If Etched experiences 20% future dilution, those thresholds rise to roughly $52.5 billion, $78.8 billion and $131.3 billion. This is not impossible in a market as large as AI infrastructure, but it requires Etched to become much more than a successful chip startup. It likely requires the company to establish a durable platform position in large-scale inference, capture significant system-level economics, and survive successive NVIDIA and hyperscaler product cycles.

 

The strongest Solten & Co. interpretation is that the Etched round is an early institutional bet on a market structure in which inference fragments away from a universal GPU architecture. The strongest counter-thesis is that NVIDIA’s software ecosystem, scale, systems integration, pace of product improvement and financing power continue to compress the available window for specialized challengers faster than Etched can convert technical advantage into a durable commercial moat.

 

Key Findings

 

  1. The $700 million round is best understood as a bet on the future structure of inference, not simply on one chip. Etched is attempting to own an integrated inference system spanning silicon, memory, interconnect, racks, cooling, software and production.

 

  1. The valuation has moved much faster than publicly demonstrated commercial maturity. Etched went from a reported $5 billion post-money valuation in December 2025 to $10.3 billion in July 2026 and $21 billion in August 2026. The latest valuation more than doubled in less than one month.

 

  1. Jane Street is the most strategically important participant in the round because it is both lead investor and first disclosed customer. That dual role provides stronger demand validation than a conventional venture syndicate, but also introduces concentration and signaling questions.

 

  1. The company’s “more than $1 billion in customer contracts” is meaningful evidence of demand, but it is not equivalent to recognized revenue, recurring revenue or even necessarily fully binding backlog. Public disclosure remains insufficient to determine contract quality, customer concentration, delivery schedules or gross-margin economics.

 

  1. Etched’s product thesis has broadened materially since 2024. The original transformer-only framing exposed the company to architecture obsolescence. The 2026 system is presented as a broader inference platform using Low Voltage Inference and Cluster Scale Memory and is said to run MoE and non-transformer designs.

 

  1. NVIDIA remains the reference competitor, but Etched’s real competitive set is wider: NVIDIA, AMD, Google TPU, AWS Trainium, Microsoft Maia, Meta MTIA, Cerebras and other specialized inference architectures. The market is evolving toward heterogeneous compute rather than a simple NVIDIA-versus-startup contest.

 

  1. The deal is strategically important for the semiconductor ecosystem because success would validate a merchant specialized-inference business model distinct from both general-purpose GPUs and vertically integrated hyperscaler ASICs.

 

  1. The most important unresolved question is no longer whether Etched can produce working silicon. It can. The question is whether it can repeatedly manufacture, deploy and support systems at scale while delivering independently verifiable cost, latency and power advantages after software, networking, utilization and customer migration costs are included.

 

Table of Contents

 

Executive Summary

Key Findings

Why This Deal Matters

Scope & Methodology

The Company: From Transformer ASIC to Frontier Inference Systems

What Etched Actually Builds

Manufacturing and Production Strategy

Commercial Evidence: $1 Billion in Contracts Is Not $1 Billion in Revenue

Jane Street: Why the Lead Investor Matters

Financing History

Current Transaction Anatomy

Valuation: What Must Be True at $21 Billion

The Investor Coalition

Competitive Landscape

Market Structure: The Inference Economy Is Becoming Its Own Industry

Industry Impact: First-, Second- and Third-Order Effects

Broader Economic Implications

Risks to Etched

Risks to Investors

Risks to the Industry

Scenario Analysis

Solten & Co. Thesis

Counter-Thesis

Falsification Criteria

Key Unknowns

What to Watch Next

Research Exhibits

Sources & Evidence

Methodological Note

About Solten & Co.

 

Scope & Methodology

 

Research question. This report asks what Etched’s August 2026 financing reveals about the company’s emerging business, the economics of specialized AI inference, investor expectations embedded in a $21 billion valuation, and the likely competitive effects on the broader AI infrastructure market.

 

Scope. The analysis covers Etched’s corporate history, product evolution, disclosed financing history, current investor coalition, customer evidence, valuation implications, competitive landscape, industry structure and potential first-, second- and third-order effects. It is not a full technical audit of Etched silicon and does not constitute an investment recommendation.

 

Evidence hierarchy. Priority is given to Etched disclosures, official investor and partner statements, SEC filings and other primary sources. Reuters, TechCrunch and other high-quality specialist reporting are used where private-company terms are not publicly disclosed. Academic accelerator research is used to test the general validity of performance comparisons.

 

Evidence cut-off. August 20, 2026.

 

Material limitations. Etched is private. Detailed financial statements, cap-table data, preferred-stock terms, production yields, customer contracts, pricing, gross margins and normalized third-party benchmark results are not publicly available. Therefore ownership, return and valuation analyses are explicitly illustrative where required.

 

Evidence classes. DISCLOSED FACT identifies primary-source facts. REPORTED TERM identifies credible but externally reported information. ESTIMATE identifies calculations based on incomplete public data. SOLTEN & CO. INTERPRETATION identifies analytical synthesis.

 

What Changed

 

Etched is no longer best understood through its original 2024 description as a transformer-only chip company. Its public 2026 strategy has moved upward in the stack toward complete frontier inference systems and toward a broader architecture story built around Low Voltage Inference and Cluster Scale Memory. The company now says its systems are running massive MoE models and non-transformer designs.

 

The commercial evidence has also changed. In June, the company had working silicon and more than $1 billion in customer contracts but no disclosed production customer. By August, Jane Street had received the first rack and was actively deploying it. This materially improves the evidence base, although it does not resolve questions around normalized performance, contract quality, revenue conversion or margins.

 

The financing context changed just as quickly. Etched’s valuation moved from $5 billion in the financing disclosed for December 2025 to $10.3 billion in July 2026 and $21 billion in August. The market is therefore not merely rewarding technical execution; it is rapidly capitalizing an expected future position in the inference value chain.

 

Why This Deal Matters

 

Etched announced on August 18 that it had raised $700 million at a $21 billion valuation in a round led by Jane Street. The company said the financing included Kleiner Perkins, Sequoia, Andreessen Horowitz, Tiger Global, Bain Capital Ventures, Neo, Primary, Stripes, Positive Sum and Blackstone. It also disclosed that Jane Street was its first customer, had received the company’s first shipped rack in July and was actively deploying the system.

 

This financing deserves attention for three reasons.

 

First, the speed of valuation expansion is extreme even by current AI standards. Etched was valued at $10.3 billion in a $300 million Series C announced July 23. Less than four weeks later, the valuation was $21 billion. That is approximately a 104% increase in headline valuation in 26 days.

 

Second, the round arrives at the moment Etched is crossing the line from technical promise to commercial execution. The company emerged from stealth in June with working A0 silicon on TSMC N4P, a team of more than 400, over $1 billion in customer contracts and a plan to ship its first racks during the summer. By August, the first disclosed rack had reached Jane Street. The investment therefore prices not merely a design concept but a very early production system.

 

Third, the deal tests a larger industry hypothesis: whether the economics of AI inference are now large enough to support highly specialized merchant hardware companies alongside GPUs and hyperscaler custom silicon.

 

The Company: From Transformer ASIC to Frontier Inference Systems

 

Etched was founded in 2022 by Gavin Uberti, Robert Wachen and Chris Zhu, three Harvard dropouts who later became Thiel Fellows. Early reporting described Uberti and Zhu as the technical founders; by 2024 Primary Venture Partners publicly described all three as co-founders.

 

Uberti’s background includes compiler work and development of a Cortex-M backend for TVM. Zhu has a mathematics and high-performance-computing background. Wachen’s role has been more commercially oriented. Etched’s current leadership team has been deliberately built around experienced semiconductor operators, including former Cypress CTO Mark Ross, former NVIDIA platform leader Brian Loiler, former NVIDIA/Auradine architect Saptadeep Pal, former Google TPU software leader David Munday and production executives with experience in high-volume consumer hardware.

 

This composition matters because semiconductor startups often fail not at architecture design but at production, packaging, supply, system validation, developer tooling and customer deployment. Etched’s hiring strategy appears designed to close precisely that execution gap.

 

The original product thesis was unusually simple and unusually risky. In 2024 the company described Sohu as an ASIC designed specifically for transformer inference. CEO Gavin Uberti openly acknowledged the binary nature of the bet: if transformers disappeared, the company’s original architecture would be impaired; if they remained dominant, specialization could create a major performance advantage.

 

By 2026, however, the public product definition had changed. Etched no longer leads with Sohu or with a “transformer-only” identity. Its current category is “frontier inference clusters.” The company says it co-designs chips, racks, software and manufacturing methods to optimize throughput, latency, cost and power efficiency across both prefill and decode workloads.

 

That shift is strategically significant. It suggests Etched is moving from an architecture-specific chip thesis to a system-level inference thesis.

 

What Etched Actually Builds

 

Etched’s current system claims two central technical differentiators.

 

Low Voltage Inference (LVI). Etched argues that conventional AI chips cannot sustain peak floating-point throughput because power draw and thermal limits force clock reductions. Its architecture is designed to run math blocks at materially lower voltage, allowing higher compute density within a fixed power envelope. The company says this enables large sparse mixture-of-experts models to operate at over 80% of peak FLOPS without thermal throttling.

 

Cluster Scale Memory (CSM). Etched argues that decode workloads are frequently constrained by memory access and inter-chip communication rather than raw FLOPS. Its system combines HBM and SRAM with a proprietary interconnect to create a lower-latency shared memory pool across a scale-up domain. The intended effect is to reduce the trade-off between high-capacity HBM systems and very fast but capacity-limited SRAM-centric designs.

 

These are company claims, not independent performance conclusions. The company has not yet published a broad third-party benchmark suite that allows investors to normalize performance across batch size, context length, model architecture, precision, concurrency, power, rack configuration and software stack.

 

That absence is currently one of the most important analytical constraints.

 

Independent academic work on AI accelerators reinforces why such normalization matters. A 2026 comparative study of NVIDIA, AMD, Cerebras, SambaNova, Gaudi and TPU systems found that the optimal hardware platform varies significantly with model size, sequence length, batch size and workload characteristics. It also found that high utilization is critical to realizing energy-efficiency benefits. A headline tokens-per-second figure therefore does not by itself establish superior production economics.

 

Manufacturing and Production Strategy

 

Etched’s first A0 silicon returned from TSMC on the N4P process in early 2026. The company says it achieved first-pass silicon success in under three years from its seed round.

 

It has also chosen a higher degree of operational integration than many fabless semiconductor startups. Etched says it has opened a Taiwan factory, built a data center, test house and NPI prototyping lab in San Jose, and opened an approximately 80,000-square-foot facility near its headquarters to accelerate production and prototyping. It describes “production is the product” as an operating principle.

 

This strategy has two interpretations.

 

The positive interpretation is that Etched understands the commercialization bottleneck. A custom AI accelerator is not useful merely because the chip works. Customers buy complete systems, qualification, networking, thermals, firmware, software and reliable supply. Vertical integration may let Etched iterate faster and reduce the hand-offs that slow traditional semiconductor development.

 

The counterpoint is capital intensity. Bringing more system design, validation and manufacturing processes inside the company increases fixed costs, working-capital needs and organizational complexity. The $700 million round therefore appears partly designed to finance an industrial scaling problem, not just semiconductor R&D.

 

Commercial Evidence: $1 Billion in Contracts Is Not $1 Billion in Revenue

 

Etched says it has secured more than $1 billion in customer contracts across public and private frontier AI companies and cloud providers. The current announcement adds one major validation point: Jane Street is now a named customer with a rack physically deployed.

 

This is meaningful progress. In June, outside reporting correctly noted that Etched had not publicly named customers, disclosed contract terms or provided independent performance benchmarks. By August, one customer had moved from anonymous demand to actual deployment.

 

But the public data still leave major diligence gaps.

 

We do not know:

 

  • how much of the $1 billion is binding purchase commitments versus reservations, milestones or conditional orders;
  • how concentrated the contracts are among a small number of customers;
  • the delivery schedule;
  • cancellation rights;
  • gross margin on initial systems;
  • recognized revenue to date;
  • how much revenue depends on customer acceptance testing;
  • whether contracts cover chips, racks, services or future generations;
  • how much capital Etched must spend before it can recognize the contracted revenue.

 

This distinction is critical because the company’s $21 billion valuation is roughly 21 times the headline $1 billion contract figure. That is not a revenue multiple. It should not be treated as one.

 

The correct interpretation is that investors are valuing expected future economics far beyond currently disclosed commercial realization.

 

Jane Street: Why the Lead Investor Matters

 

Jane Street is unusually important to the Etched story.

 

It is not simply a hedge fund or a venture investor adding an AI hardware position. Jane Street is one of the world’s most technology-intensive trading firms. Its core business rewards extremely high-performance research and computation, and it has become a major direct buyer of AI infrastructure.

 

In April 2026, Jane Street committed approximately $6 billion to CoreWeave for AI cloud capacity across multiple facilities, including NVIDIA Vera Rubin technology, and separately invested $1 billion in CoreWeave equity. Reuters described the transaction as part of Jane Street’s broader effort to scale machine learning across its global markets research.

 

That makes Jane Street’s Etched investment much more informative than a conventional VC endorsement. Jane Street has access to NVIDIA-based infrastructure through CoreWeave, has the technical capacity to evaluate alternatives, and has an economic incentive to improve inference performance for highly latency-sensitive and compute-intensive workloads.

 

According to Etched, Jane Street tested the chip before leading the round and is now running a rack in its own data center.

 

This creates a powerful proof point — but it also requires analytical caution.

 

Jane Street is simultaneously customer, investor and signal provider. Those roles can reinforce each other. A customer that owns equity may tolerate early product friction or adopt infrastructure partly because it expects strategic upside. Conversely, the fact that Jane Street already has access to large-scale NVIDIA systems makes its decision to deploy Etched more significant, not less.

 

The most useful conclusion is therefore not “Jane Street proves Etched wins.” It is that a technically sophisticated, capital-rich buyer with access to alternative infrastructure believes Etched is sufficiently credible to test in production and sufficiently valuable to lead a major financing.

 

Financing History

 

March 2023 — Seed

Etched raised approximately $5.4 million at a reported $34 million valuation. Reuters later confirmed the valuation when covering the Series A.

 

June 2024 — Series A

Etched raised $120 million, co-led by Primary Venture Partners and Positive Sum. The round included Peter Thiel and a long list of technology founders and investors. The company said the capital would support chip development and manufacturing. A company-level valuation was not publicly disclosed at the time.

 

December 2025 / disclosed June 2026 — $500 million financing

When Etched emerged from stealth, it disclosed that its latest prior financing had been $500 million at a $5 billion post-money valuation. It also said it had raised $800 million across multiple unannounced financings by that point.

 

July 23, 2026 — Series C

Etched raised $300 million at a $10.3 billion valuation. Sequoia led, with Andreessen Horowitz, Jane Street, Diffusion and SK Hynix among participants. The company said proceeds would accelerate production and customer deployments.

 

August 18, 2026 — $700 million round

Jane Street led at a $21 billion valuation, joined by Kleiner Perkins, Sequoia, Andreessen Horowitz, Tiger Global, Bain Capital Ventures, Neo, Primary, Stripes, Positive Sum and Blackstone.

 

Etched’s current announcement says the company has raised $1.9 billion in total. Publicly identifiable round amounts do not reconcile perfectly to that figure because several financings were never individually announced and historical totals were rounded. This is a useful example of why venture funding databases should not be treated as primary evidence when the company itself describes undisclosed rounds.

 

Current Transaction Anatomy

 

DISCLOSED FACT: Etched raised $700 million at a $21 billion valuation.

 

DISCLOSED FACT: Jane Street led the round.

 

DISCLOSED FACT: Kleiner Perkins, Sequoia, Andreessen Horowitz, Tiger Global, Bain Capital Ventures, Neo, Primary, Stripes, Positive Sum and Blackstone participated.

 

UNKNOWN: Whether the $21 billion figure is explicitly pre-money or post-money in the legal financing documents. Public coverage generally presents it as the round valuation without disclosing the security terms.

 

UNKNOWN: Secondary component, if any.

 

UNKNOWN: Liquidation preference, participation rights, anti-dilution protection, board rights and other preferred-stock terms.

 

ESTIMATE: If $21 billion is a post-money valuation and the entire $700 million is primary equity, the round would represent approximately 3.3% of post-money ownership before other option-pool or security effects. This is only an illustrative estimate and should not be treated as a disclosed cap-table fact.

 

The striking point is that Etched is raising large absolute amounts with relatively modest headline dilution because valuation has accelerated so quickly.

 

Valuation: What Must Be True at $21 Billion

 

A $21 billion private valuation changes the analytical standard. At this level, the central question is no longer whether Etched can become a meaningful semiconductor company. It is whether Etched can become a very large infrastructure platform.

 

Ignoring future dilution, new investors require approximately:

 

  • $42 billion eventual equity value for a 2x gross multiple;
  • $63 billion for 3x;
  • $105 billion for 5x.

 

With 20% future dilution, the required exit values increase to approximately:

 

  • $52.5 billion for 2x;
  • $78.8 billion for 3x;
  • $131.3 billion for 5x.

 

Those values are possible only if Etched reaches a scale comparable with major public semiconductor or infrastructure companies.

 

Because current recognized revenue is not publicly disclosed, conventional revenue-multiple analysis is impossible. The $1 billion-plus contract figure is insufficient because contract value, recognized revenue, gross margin and duration are unknown.

 

A more useful valuation framework is therefore milestone-based.

 

To support the current valuation, Etched likely needs to demonstrate several of the following simultaneously:

 

  1. Production-scale reliability across multiple customers.
  2. Independent evidence of superior total cost of ownership on important frontier inference workloads.
  3. A software and deployment layer that makes migration materially easier than adopting a typical custom ASIC.
  4. Multi-generation product execution, not a single architecture win.
  5. A sufficiently large customer base to prevent Jane Street or any one AI lab from dominating revenue.
  6. Gross margins consistent with attractive merchant semiconductor or integrated-system economics.
  7. Access to foundry, advanced packaging and memory supply at scale.
  8. Continued relevance as model architectures evolve.
  9. A market large enough for both hyperscaler custom silicon and merchant specialized systems.
  10. Evidence that NVIDIA cannot eliminate Etched’s cost/performance advantage through rapid platform iteration or pricing.

 

The Investor Coalition

 

The syndicate tells its own story.

 

Jane Street brings customer validation and direct experience buying large-scale compute.

 

Sequoia led the July Series C and returned immediately in the August financing. Its follow-on participation is a strong signal of continued conviction after receiving access to non-public diligence during the prior round.

 

Andreessen Horowitz is another repeat investor, reinforcing the view that Etched is being underwritten as a major AI infrastructure platform rather than a conventional semiconductor startup.

 

Kleiner Perkins is notable because its managing partner publicly framed the opportunity around inference economics — specifically tokens per dollar and tokens per watt. The firm also has exposure across AI labs and infrastructure companies, giving it a broad view of demand.

 

SK Hynix, which participated in the July round, is strategically relevant as a major memory supplier even though it was not named in the August investor list. Its presence in the cap table illustrates how Etched’s financing network intersects with the semiconductor supply chain.

 

VentureTech Alliance, a fund linked to the semiconductor ecosystem and disclosed in earlier financings, further deepens the strategic investor layer.

 

The broader coalition — Blackstone, Tiger Global, Bain Capital Ventures, trading firms and major technology angels — suggests that Etched has moved beyond early venture financing into crossover-style capital formation.

 

Competitive Landscape

 

Exhibit 4 — The Competitive Battlefield Is Not Simply Etched vs. NVIDIA

This framework separates merchant general-purpose platforms, hyperscaler-owned custom silicon and merchant specialized architectures. Etched’s strategic opening exists only if a meaningful customer segment wants hyperscaler-like specialization without owning a hyperscaler-scale internal silicon program.

 

The lazy framing is “Etched versus NVIDIA.” The real market is more complicated.

 

NVIDIA remains the dominant reference platform because it combines high-performance accelerators with CUDA, networking, systems, enterprise software, developer tooling, financing relationships and an enormous installed base. NVIDIA reported fiscal 2026 revenue of approximately $216 billion, with data-center growth driven by AI. Its gross margins remain above 70%, which demonstrates the economic value of platform control.

 

But NVIDIA is facing competition from several directions.

 

Hyperscaler custom silicon. Google TPUs, AWS Trainium, Microsoft Maia and Meta MTIA are increasingly important because the largest buyers have enough workload volume to justify architecture-specific optimization. Google and Broadcom have extended TPU collaboration through 2031, and Google is expanding external TPU distribution. AWS has made Trainium central to major OpenAI and Anthropic compute relationships.

 

Merchant accelerators. AMD remains the largest conventional GPU alternative. Cerebras uses wafer-scale processors and is explicitly targeting fast inference. Its newly announced CS-4 system illustrates how quickly the specialized-inference market is moving.

 

Specialized architectures. Groq demonstrated demand for low-latency inference and ultimately entered a major technology relationship with NVIDIA, a reminder that incumbent responses can include acquisition, licensing and ecosystem absorption rather than only price competition.

 

Other startups and new architectures. Tenstorrent, SambaNova, d-Matrix, Furiosa and others continue to pursue different combinations of programmability, memory architecture, inference efficiency and system design.

 

Etched is therefore competing on multiple dimensions:

 

  • tokens per dollar;
  • tokens per watt;
  • latency;
  • throughput;
  • programmability;
  • model compatibility;
  • memory capacity and bandwidth;
  • interconnect performance;
  • rack-level deployment complexity;
  • supply availability;
  • software maturity;
  • customer switching cost;
  • production scale.

 

The important market question is not whether one architecture wins universally. Independent accelerator research increasingly suggests that workload characteristics determine the optimal hardware. The likely market structure is heterogeneous.

 

Etched’s opportunity is to own a sufficiently valuable portion of that heterogeneity.

 

Market Structure: The Inference Economy Is Becoming Its Own Industry

 

The shift from training to inference is central to the Etched thesis.

 

Training creates frontier capability, but inference monetizes that capability. Every user query, coding-agent step, enterprise workflow and autonomous tool call creates recurring inference demand. As agentic systems perform longer reasoning chains and call other models or software repeatedly, the number of inference operations per unit of useful work can rise dramatically.

 

NVIDIA itself now describes inference as the dominant AI workload. The company’s fiscal 2026 annual materials state that inference has overtaken training as the primary workload and emphasize the changing economics of AI as models move into large-scale use.

 

This shift matters because inference rewards different optimization choices than training.

 

Training values flexibility, enormous distributed scale and rapid support for new model operations.

 

Inference can reward specialization when workloads are repeated at high volume. Once a model architecture is stable enough, purpose-built silicon can theoretically remove general-purpose overhead and optimize memory, scheduling and power around the real workload.

 

That creates a structural opening for companies like Etched.

 

But the same economics also motivate every hyperscaler to build custom silicon internally. Etched must therefore prove that a merchant specialized platform can compete not only with NVIDIA but with customers’ own chips.

 

Industry Impact: First-, Second- and Third-Order Effects

 

First-order effect: more credible competition in inference accelerators.

 

A working Etched system with a production customer increases pressure on NVIDIA and other accelerator vendors to compete on inference-specific economics rather than only headline training performance. It also gives AI labs and cloud providers another procurement option.

 

Second-order effect: greater bargaining power for large compute buyers.

 

Even if Etched never takes dominant market share, credible alternatives can influence NVIDIA pricing, supply terms and roadmap priorities. Large buyers gain leverage when they can move marginal inference workloads to specialized systems.

 

Second-order effect: more fragmentation in the software stack.

 

Every new accelerator creates integration cost. Model runtimes, compilers, kernels, observability, orchestration and deployment tooling must support heterogeneous systems. This creates opportunity for software layers that abstract hardware differences — but also raises switching costs for customers.

 

Second-order effect: more pressure on custom silicon economics.

 

If Etched can offer hyperscaler-like specialization without requiring a customer to design its own chip, it creates a new strategic option between buying NVIDIA and building an internal ASIC program.

 

Third-order effect: semiconductor value may shift from chips toward integrated inference systems.

 

Etched’s own evolution suggests that the economic unit is moving upward from the accelerator die to the rack or cluster. Memory, networking, cooling, power delivery, software and manufacturing become part of competitive differentiation.

 

Third-order effect: capital intensity spreads downstream.

 

If specialized inference becomes a major infrastructure category, large venture-funded companies may need hundreds of millions or billions of dollars to finance fabrication, inventory, systems and customer deployments before revenue scales. The distinction between venture capital and industrial project finance begins to blur.

 

Broader Economic Implications

 

Etched is one small company relative to the global semiconductor market, so macro claims should be restrained. But the financing contributes to several larger patterns.

 

AI is becoming more industrial. Capital is flowing not only into software but into fabs, packaging, memory, data centers, power, cooling, networking and purpose-built systems.

 

Inference efficiency has downstream economic significance. If specialized hardware materially lowers the cost per useful token, then applications that are currently uneconomic can become viable. That can affect pricing for AI software, enterprise automation and agentic workloads.

 

Hardware competition may also shift capital allocation. The stronger the evidence that inference supports multiple architectures, the more venture and growth capital may move toward specialized silicon, memory systems, networking and infrastructure software.

 

At the same time, the market risks overinvestment. High private valuations and massive compute commitments can create excess capacity if model monetization fails to scale at the pace assumed by infrastructure investors.

 

Risks to Etched

 

Technology risk. Performance claims may narrow under independent testing or production conditions. Different workloads may reduce the theoretical advantage of specialization.

 

Architecture risk. Although Etched’s 2026 system appears broader than the original transformer-only thesis, rapid changes in model architecture can still invalidate hardware assumptions.

 

Software risk. NVIDIA’s most durable moat is not only silicon. CUDA, libraries, tooling, developer familiarity and system integration create powerful switching costs.

 

Manufacturing risk. First-pass silicon is a major achievement but does not guarantee high-yield volume production across multiple generations.

 

Supply-chain risk. Etched depends on advanced foundry capacity, packaging, memory and other constrained semiconductor inputs.

 

Execution risk. Scaling from a few racks to gigawatt-scale deployments is an industrial challenge substantially harder than producing working A0 silicon.

 

Customer concentration risk. Only Jane Street is currently publicly identified. More than $1 billion in contracts could still be concentrated in a small number of counterparties.

 

Capital risk. Building inventory and production capacity may require further large financings before cash generation becomes self-sustaining.

 

Valuation risk. At $21 billion, execution disappointments can create severe private-market repricing even if the company remains technologically viable.

 

Incumbent-response risk. NVIDIA has the resources to respond through product acceleration, pricing, financing, software integration, partnerships or acquisition/licensing strategies.

 

Hyperscaler risk. The largest potential customers may prefer their own TPU, Trainium, Maia or MTIA systems rather than buying from a merchant startup.

 

Risks to Investors

 

The current round has unusually high expectations embedded in the entry price.

 

A successful product launch is not sufficient. Investors need Etched to sustain a major competitive position over multiple product generations.

 

Future dilution could materially raise required exit values. A capital-intensive company may need several more large rounds.

 

Private-market headline valuation does not disclose preference terms. Investors entering at different rounds may have very different downside protection.

 

The customer-investor overlap is both strength and risk. Strategic investors can accelerate adoption, but commercial relationships may make market demand appear stronger than it would under purely arm’s-length procurement.

 

Liquidity is uncertain. An eventual IPO must support a valuation far above current levels for venture-style returns, while strategic acquisition at extreme valuations becomes difficult because only a small set of buyers could finance it and antitrust considerations may constrain some obvious acquirers.

 

Risks to the Industry

 

Large financings can accelerate competitive innovation, but they can also distort market behavior.

 

Capital crowding. Well-funded hardware companies can bid aggressively for scarce semiconductor talent and supply capacity, raising costs for smaller entrants.

 

Supply concentration. More accelerator designs still depend on a small number of foundries, advanced packaging providers and memory suppliers.

 

Architecture fragmentation. Heterogeneous hardware can improve efficiency but increase software and operational complexity across the ecosystem.

 

Overcapacity. If inference demand disappoints, capital-intensive accelerator and data-center investments could create stranded or underutilized assets.

 

Synchronized assumptions. Many current AI infrastructure investments depend on the same premise: rapidly rising inference demand and sustained willingness to pay. Correlated error in that assumption would affect labs, clouds, chipmakers, power developers and infrastructure financiers simultaneously.

 

Scenario Analysis

 

Base Case

 

Etched successfully ramps production, converts a meaningful portion of its current contracts into recognized revenue and demonstrates superior inference economics on selected frontier workloads. It becomes a credible second-source or specialized accelerator supplier for several AI labs, clouds and high-performance enterprises. NVIDIA remains dominant overall, but Etched captures a defensible high-value niche and grows into a major infrastructure company.

 

Signals supporting this case: several named production customers; independent benchmark validation; repeat orders; evidence of improving gross margins; second-generation silicon delivered on schedule; expansion beyond Jane Street without sacrificing performance.

 

Upside Case

 

Inference becomes substantially larger than training in total compute spend, and workloads increasingly reward specialization. Etched’s architecture demonstrates a durable cost and power advantage while its system integration reduces migration friction. Hyperscalers use Etched as a merchant alternative to internal ASIC programs, AI labs adopt the systems at multi-gigawatt scale, and the company becomes an independent platform with economics closer to a leading accelerator vendor than a niche chip supplier.

 

In this case, a valuation above $100 billion becomes plausible.

 

Signals supporting this case: multi-gigawatt customer commitments, very strong gross margins, broad model compatibility, major cloud distribution, sustained performance-per-watt advantage over two NVIDIA generations, rapid third-generation product execution.

 

Downside Case

 

Etched’s initial systems work but deliver a narrower advantage than expected once real-world utilization, software costs and new NVIDIA platforms are considered. The $1 billion-plus contracts convert slowly, customers retain Etched as experimental capacity rather than core production infrastructure, and hyperscalers prioritize internal silicon. High production spending forces further financing, compressing returns for current investors.

 

Signals supporting this case: delayed deliveries, limited independent benchmarking, cancellations or contract slippage, continued customer anonymity, heavy pricing discounts, rising inventory, repeated financing before meaningful revenue, slower roadmap execution.

 

Solten & Co. Thesis

 

The Etched financing is evidence that the AI accelerator market is moving from a single dominant architecture toward a portfolio of workload-specific compute systems.

 

The important economic variable is not “Can Etched beat NVIDIA?” in the abstract. It is whether the cost of inference becomes large and repetitive enough that customers are willing to accept hardware and software fragmentation in exchange for materially better tokens per dollar, tokens per watt and latency on specific high-value workloads.

 

If that threshold has been crossed, Etched does not need to replace NVIDIA. It needs to become economically indispensable for a sufficiently large subset of inference.

 

The round suggests sophisticated investors believe that subset may be very large.

 

Counter-Thesis

 

The strongest counter-thesis is that the market is overestimating the value of specialized silicon while underestimating the value of general-purpose platform integration.

 

NVIDIA can amortize R&D across a vastly larger installed base, improve hardware every generation, bundle networking and systems, subsidize adoption through financing and maintain developer lock-in through CUDA. Hyperscalers can optimize internally for their own workloads without paying a merchant supplier margin.

 

Etched therefore occupies a difficult middle position: more specialized and less ecosystem-rich than NVIDIA, but less vertically integrated with end-customer demand than Google or AWS custom silicon.

 

The company must show that its system-level efficiency advantage is large enough to overcome that structural disadvantage.

 

Falsification Criteria

 

The Solten & Co. thesis would weaken materially if several of the following occur:

 

  • independent production benchmarks show only modest total-cost advantage over NVIDIA or hyperscaler alternatives;
  • customer contracts fail to convert into deployments and recognized revenue;
  • the majority of demand remains concentrated in Jane Street or one or two counterparties;
  • new model architectures materially reduce the efficiency of Etched’s hardware approach;
  • NVIDIA’s next two platform generations close most of the tokens-per-watt or latency gap;
  • Etched requires repeated large financings without corresponding commercial scale;
  • software migration becomes a persistent barrier to production adoption;
  • hyperscalers refuse to adopt merchant specialized inference hardware because internal chips are economically superior.

 

Key Unknowns

 

The public investment case remains constrained by missing information. The highest-value unanswered questions are:

 

  1. What is Etched’s recognized revenue today?
  2. How much of the $1 billion-plus contract value is legally binding and non-cancellable?
  3. What is customer concentration?
  4. What is the expected delivery schedule for those contracts?
  5. What are current and expected gross margins per rack?
  6. What is actual production yield?
  7. What is the all-in system cost per token under realistic production workloads?
  8. How does Etched compare with NVIDIA Blackwell and Vera Rubin under identical model, precision, latency and batch conditions?
  9. What percentage of workloads can run without material software modification?
  10. What foundry, HBM and packaging capacity is contractually secured?
  11. How much additional capital will be required before positive free cash flow?
  12. What are the preference and governance terms in the latest round?
  13. How much ownership did Jane Street obtain?
  14. Which frontier AI companies and cloud providers are under contract?
  15. What proportion of current contracts relate to first-generation versus future-generation systems?

 

What to Watch Next

 

The next six to eighteen months should provide far more information than the financing announcement itself.

 

The critical signals are:

 

  • additional named customers;
  • independent benchmark publication;
  • production shipment volume;
  • contract conversion to revenue;
  • second and third hardware generation milestones;
  • hyperscaler adoption;
  • broader software ecosystem availability;
  • pricing disclosure or customer total-cost evidence;
  • manufacturing scale and yield;
  • additional fundraising;
  • changes in NVIDIA inference pricing and roadmap;
  • competitive launches from Cerebras, AMD, AWS, Google and other accelerator providers;
  • any evidence that Etched is becoming a standard merchant alternative rather than a specialist system for a narrow set of customers.

 

Research Exhibits

 

Exhibit 1 — Etched Valuation Trajectory

The valuation trajectory illustrates how rapidly market expectations have changed. The Series A valuation is intentionally omitted because Etched did not publicly disclose it at the time.

 

Exhibit 2 — Known Financing Rounds

Public financing records do not fully reconcile to Etched’s stated $1.9 billion total because several rounds were unannounced and historical totals were rounded. The correct approach is to preserve that uncertainty rather than manufacture precision.

 

Exhibit 3 — What a $21 Billion Entry Valuation Implies

This is an illustrative return-hurdle model, not a valuation forecast. It demonstrates how future dilution raises the exit value required for current investors to achieve venture-style returns.

 

Sources & Evidence

 

Primary / company sources

 

Etched — Frontier Inference Clusters, June 30, 2026

https://www.etched.com/progress/frontier-inference-clusters

 

Etched — company website and leadership materials

https://www.etched.com/

 

Etched — $300M financing at $10.3B valuation, July 23, 2026

https://www.globenewswire.com/news-release/2026/07/23/3332366/0/en/Etched-raises-300M-at-a-10-3B-Valuation-to-Scale-Production-of-Frontier-Scale-Inference-Hardware.html

 

Etched — $700M financing at $21B valuation / first Jane Street delivery, August 18, 2026, company release syndicated by AIwire

https://www.hpcwire.com/aiwire/2026/08/18/etched-raises-700m-at-21b-valuation-and-completes-1st-customer-delivery-to-jane-street/

 

CoreWeave — Jane Street $6B cloud agreement and $1B equity investment, April 15, 2026

https://www.coreweave.com/news/jane-street-signs-6-billion-ai-cloud-agreement-with-coreweave

 

NVIDIA — FY2026 annual results

https://investor.nvidia.com/news/press-release-details/2026/NVIDIA-Announces-Financial-Results-for-Fourth-Quarter-and-Fiscal-2026/

 

NVIDIA — FY2026 Form 10-K

https://www.sec.gov/Archives/edgar/data/1045810/000104581026000021/nvda-20260125.htm

 

Google Cloud — Anthropic expands TPU usage, April 6, 2026

https://www.googlecloudpresscorner.com/2026-04-06-Anthropic-Expands-Use-of-Google-Cloud-and-TPUs

 

High-quality reporting / research

 

Reuters — AI chip startup Etched doubles valuation to $21 billion in under a month, August 18, 2026

https://www.reuters.com/technology/ai-chip-startup-etched-valued-21-billion-latest-funding-round-2026-08-18/

 

Reuters — Etched raises $120 million to develop specialized chip, June 25, 2024

https://www.reuters.com/technology/artificial-intelligence/ai-startup-etched-raises-120-million-develop-specialized-chip-2024-06-25/

 

Reuters — Jane Street signs $6B AI cloud deal with CoreWeave, April 15, 2026

https://www.reuters.com/legal/transactional/jane-street-signs-6-billion-ai-cloud-deal-with-coreweave-boosts-stake-2026-04-15/

 

Reuters — Cerebras launches new inference server system, August 19, 2026

https://www.reuters.com/technology/cerebras-launches-new-server-chip-system-designed-speed-ai-chatbots-2026-08-19/

 

Reuters — Broadcom signs long-term Google custom AI chip agreement, April 6, 2026

https://www.reuters.com/business/broadcom-signs-long-term-deal-develop-googles-custom-ai-chips-2026-04-06/

 

TechCrunch — Etched Series C at $10.3B valuation, July 23, 2026

https://techcrunch.com/2026/07/23/ai-chip-startup-etched-defies-skeptics-hits-10-3b-valuation-from-big-name-investors/

 

TechCrunch — Etched Series A / transformer-only strategy, June 25, 2024

https://techcrunch.com/2024/06/25/etched-is-building-an-ai-chip-that-only-runs-transformer-models/

 

Primary Venture Partners — Etched Series A thesis / company history

https://www.primary.vc/articles/etcheds-series-a-to-revolutionize-ai-hardware

 

Academic research — The xPU-athalon: Quantifying the Competition of AI Acceleration, 2026

https://arxiv.org/abs/2604.10852

 

Methodological Note

 

This report distinguishes disclosed facts, reported terms, estimates and Solten & Co. interpretation.

 

DISCLOSED FACT — directly supported by a company, regulatory or authoritative primary source.

 

REPORTED TERM — reported by a credible secondary source but not independently disclosed in full by the relevant company or counterparty.

 

ESTIMATE — an analytical calculation based on incomplete public data. Assumptions are stated where material.

 

SOLTEN & CO. INTERPRETATION — our synthesis or inference from the evidence.

 

The report does not treat contract value as revenue, valuation as enterprise value, or headline round size as sufficient evidence of ownership. Private-company financial statements, cap-table terms, contract schedules and detailed customer economics are not publicly available. Valuation return scenarios are illustrative and do not represent an investment recommendation.

 

About Solten & Co.

 

Solten & Co. is an independent research and analysis firm focused on the AI economy, with deeper research emphasis on AI infrastructure, Physical AI, robotics and autonomous systems. We study the technologies, companies, markets, transactions and capital structures shaping the next phase of AI.

 

Need an independent perspective?

 

Solten & Co. provides independent research and analytical support for investors and decision-makers evaluating companies, markets and investment opportunities across the AI economy.

 

Discuss a research question · Discuss an investment opportunity · Request independent analysis

soltenco.com

The New SaaS Math

How Generative AI Is Rewriting Software Margins, Pricing and Operating Leverage

Solten & Co. Research Report
Published / updated: August 20, 2026
Research universe: Software AI / AI Infrastructure / Economics & Benchmarks

Executive Summary

 

The most common financial shorthand about generative-AI software is that traditional SaaS enjoys 70–90% gross margins while AI-native products fall to 30–60% because every user action consumes tokens, GPUs, or model-API capacity. The direction of that claim is useful; the conclusion is too simple.

 

Generative AI does change the financial architecture of software. It converts part of software delivery from a largely fixed-cost infrastructure problem into a workload-sensitive production cost. A conventional SaaS user may log in more often without materially changing the vendor’s cost to serve. An AI user can create a very different economic outcome: a longer context window, a more capable model, an autonomous agent loop, repeated tool calls or an image/video workload can multiply compute consumption while the customer continues paying the same monthly fee.

 

That makes usage itself a financial variable. It also explains why AI pricing is rapidly moving away from the old idea of an unlimited seat. Credits, metered usage, model tiers, rate limits, and hybrid seat-plus-consumption structures are not merely product packaging. They are mechanisms for allocating inference-cost risk between vendor and customer.

 

But lower gross margins today do not prove that AI software is structurally inferior to SaaS. The cost of intelligence is falling extraordinarily quickly. Stanford’s AI Index found that the cost of querying a model with GPT-3.5-equivalent benchmark performance fell from about $20 per million tokens in November 2022 to $0.07 by October 2024 — more than a 280-fold decline. OpenAI cut the API price of GPT-5.6 Luna by 80% in July 2026 and Terra by 20%, illustrating that model price compression continues. Routing, caching, quantization, speculative decoding, smaller specialized models, and owned infrastructure can all lower unit cost further.

 

The complication is that demand is evolving just as quickly. Agentic software is not simply “chat with more tokens.” Microsoft Azure Research analyzed 13 million GitHub Copilot coding-agent sessions from June 2026 and found 761 million LLM calls and 775 million tool invocations. User archetypes differed by roughly 50× in token consumption; tool failures generated retry loops that materially amplified compute. As software moves from answering questions to performing multi-step work, a single customer task can become dozens of inference events.

 

This produces the central economic tension of AI software: cost per unit of intelligence is falling, but the number of units consumed per customer outcome may rise even faster.

 

Traditional SaaS therefore remains an important benchmark, but not a universal destination. Salesforce generated a fiscal-2026 total gross margin of approximately 77.7%; ServiceNow’s 2025 subscription gross margin was 80%; Datadog’s 2025 gross margin was 80%. Yet even cloud-native, consumption-heavy software already shows lower structural margins: Snowflake’s fiscal-2026 product gross margin was 72%, with the company explicitly citing third-party cloud costs, including AI inference, as a major cost-of-revenue driver. Software margins were never determined only by whether the product was “software”; they reflect where compute, data, support, and infrastructure costs sit in the value chain.

 

At the frontier-model layer, the contrast is sharper. Reuters reported that OpenAI’s adjusted gross margin fell to approximately 33% in 2025 from 40% in 2024 as inference costs increased fourfold. Reuters Breakingviews cited a PitchBook estimate of approximately 44% gross margin for Anthropic in 2026. Those figures are not directly comparable to public SaaS accounting and should not be generalized across all AI applications, but they demonstrate the economic burden of operating frontier models at scale.

 

The most important conclusion is therefore not “AI has lower margins than SaaS.” It is this:

 

AI changes software from a business in which usage is usually desirable into one in which usage must be economically designed.

 

The quality of an AI business increasingly depends on four variables operating together: the value created per unit of inference cost; the vendor’s ability to price that value; the variance and growth of usage; and the rate at which model and infrastructure costs decline. A 55% gross-margin AI product that replaces $1,000 of human labor with $50 of compute may be economically stronger than an 85% gross-margin SaaS tool that saves a customer $20. Conversely, an AI wrapper selling $20 subscriptions while exposing itself to $30 of unbounded monthly inference cost is not a software business with temporarily bad margins; it is a structurally mispriced compute reseller.

 

For investors, this means SaaS multiples should not be applied mechanically to AI revenue. The right question is whether the company has a credible path from raw intelligence consumption to durable gross profit. For founders, it means product design, model architecture and pricing architecture have become inseparable from financial architecture.

 

Key Findings

 

  1. Generative AI does not eliminate SaaS economics; it makes them workload-dependent. The key change is not “software becomes expensive to serve,” but that cost-to-serve varies materially with user behavior, model choice, and task complexity.

 

  1. The widely repeated 30–60% gross-margin range for AI should be treated as a category estimate, not a law. Battery Ventures’ 2025 State of AI framework estimated 0–30% gross margins for some AI application businesses and 30–60% for model inference, but actual company outcomes vary dramatically by workload, pricing, and vertical value.

 

  1. Traditional SaaS itself spans a broad margin range. Public benchmarks show Salesforce near 78% total gross margin, ServiceNow at 80% subscription gross margin, Datadog at 80%, and Snowflake at 72% product gross margin. Consumption intensity already mattered before generative AI.

 

  1. Frontier-model economics remain materially below classic SaaS benchmarks. OpenAI’s reported adjusted gross margin fell to roughly 33% in 2025; Anthropic’s 2026 margin has been estimated around 44%. These are model-provider economics, not universal AI-software margins.

 

  1. Falling model prices create a powerful margin-expansion opportunity, but cost deflation does not automatically become profit. Competition may pass the savings to customers, and agentic products can increase the amount of inference consumed per task.

 

  1. Agentic AI makes the “unlimited seat” structurally dangerous for high-variance workloads. Microsoft’s production-scale GitHub Copilot traces show highly heterogeneous usage and dozens of LLM/tool events per agent session. Heavy users can have radically different costs under the same seat price.

 

  1. The market is already responding through pricing redesign. GitHub Copilot now combines subscription tiers with AI credits; Microsoft increasingly meters agent workloads through Copilot Credits. Credits are not merely monetization tactics — they cap vendor exposure to the long tail of usage.

 

  1. The most important economic metric for AI applications is not gross margin in isolation. It is value created per inference dollar, combined with pricing power and retention. High-value labor substitution can support excellent economics even at gross margins below historic SaaS norms.

 

  1. “Would the product be useful without AI?” is a weak viability test. Many genuinely valuable AI-native products would not exist without AI. The stronger test is whether the workflow creates durable customer value after fully loaded inference, infrastructure, support and deployment costs.

 

  1. AI-native software will likely split into multiple economic classes rather than converge on one margin benchmark: low-margin intelligence resellers, high-margin workflow software, consumption platforms, outcome-priced automation and integrated model/infrastructure businesses.

 

Table of Contents

 

  1. The SaaS Benchmark — What Actually Created the 80% Gross-Margin Ideal
  2. AI Changes COGS, Not the Definition of Software
  3. The Margin Evidence: What We Know and What We Do Not
  4. Why the 30–60% “AI Margin” Rule Is Too Crude
  5. The Unlimited-Usage Trap
  6. Agentic AI Makes Usage Variance a First-Class Financial Problem
  7. Model-Cost Deflation: The Most Important Counterforce
  8. Why Cheaper Intelligence May Not Improve Margins
  9. Pricing Architecture Is Now Part of Systems Architecture
  10. Consumer Freemium: Why Conversion Alone Does Not Diagnose PMF
  11. The Wrapper Question — When Is an AI Product Just Compute Resale?
  12. Vertical AI and the Economics of Labor Substitution
  13. The New AI Software Economic Taxonomy
  14. Valuation: When Does AI Software Deserve a SaaS Multiple?
  15. Risks to Founders and Investors
  16. Industry-Level Effects
  17. Solten & Co. Framework: Value per Inference Dollar
  18. Thesis, Counter-Thesis and Falsification
  19. What to Watch Next
  20. Sources & Evidence

 

Scope & Methodology

 

Research question. This report examines how generative and agentic AI alter the gross-margin structure, pricing design, operating leverage, and valuation logic of software businesses.

 

Scope. The report compares public SaaS and cloud-software benchmarks with reported frontier-model economics, current model API pricing, production-scale agent workload data, and emerging software pricing models. It is not an accounting standard or a forecast for every AI company.

 

Evidence hierarchy. Priority is given to SEC filings, company pricing pages, company disclosures, and academic or production-scale technical research. Reuters and other high-quality financial reporting are used where private-company financial information is not publicly filed. Investor frameworks are treated as estimates and analytical lenses, not audited facts.

 

Evidence classes. DISCLOSED FACT refers to primary or authoritative evidence. REPORTED TERM refers to credible reporting not independently disclosed in full. ESTIMATE refers to calculations or third-party analytical estimates. SOLTEN & CO. INTERPRETATION refers to our synthesis.

 

Critical limitation. Gross margins across the companies cited are not perfectly comparable. Salesforce and Datadog report consolidated gross margin; ServiceNow separately reports subscription gross margin; Snowflake reports product gross margin; OpenAI and Anthropic figures are private-company estimates or reported adjusted measures. The comparison is therefore directional and structural, not an accounting league table.

 

What Changed

 

The popular 2023–2025 framing of AI economics focused on the obvious problem: every generation costs money. By 2026, the more important issue is no longer simply token cost. Three things have changed.

 

First, inference has become substantially cheaper per unit of capability. Model providers now actively segment price-performance tiers and push customers toward smaller models, caching, and batch processing. Raw model cost is becoming an optimization surface, not a fixed tax.

 

Second, software is becoming agentic. A single user instruction may trigger long contexts, repeated reasoning, tool execution, file reads, searches, retries, and subagents. The economic unit is moving from “request” toward “completed workflow.”

 

Third, pricing is adapting. GitHub’s current Copilot structure explicitly combines fixed monthly plans with AI credits whose consumption varies by model and task complexity. That is a structural break from the unlimited-seat intuition inherited from SaaS.

 

The emerging question is therefore not whether AI margins are low today. It is who can convert rapidly falling intelligence costs into durable gross profit before competition passes those savings to customers or expanding usage consumes them.

 

1. The SaaS Benchmark — What Actually Created the 80% Gross-Margin Ideal

 

The mythology of SaaS begins with a true observation: once software has been built, the marginal cost of delivering another copy is small relative to the selling price. Cloud delivery introduced hosting, support, and third-party infrastructure costs, but the economics remained attractive because the vendor could serve many customers from a common code base.

 

This is visible in current public-company filings. Salesforce reported fiscal-2026 revenue of $41.5 billion and gross profit of $32.3 billion, implying total gross margin of about 77.7%. Its subscription and support business was even stronger: $39.4 billion of revenue against $6.8 billion of related cost, or roughly 82.7% gross margin.

 

ServiceNow reported 80% subscription gross margin for full-year 2025. Datadog reported 80% consolidated gross margin. These economics create powerful operating leverage because incremental revenue can finance sales, R&D, and administration while leaving a substantial contribution pool.

 

But the important historical detail is that “software” has never guaranteed 80% margins. Snowflake’s fiscal-2026 product gross margin was 72%. The company’s cost of product revenue rose by $268 million, driven primarily by $248 million of additional third-party cloud infrastructure expense, including costs related to AI inference. Consumption-heavy software already carried a more visible infrastructure burden.

 

The real SaaS advantage is therefore not zero marginal cost. It is predictable marginal cost that is low relative to price.

 

That distinction becomes central in AI.

 

Exhibit 1 — Software Gross Margins Are Becoming More Workload-Sensitive

2. AI Changes COGS, Not the Definition of Software

 

When a SaaS customer uses a CRM dashboard ten times instead of five, the vendor rarely experiences a proportionate increase in direct cost. In AI-native software, use can be directly connected to compute.

 

The relevant cost stack may include model API charges, owned GPU depreciation, cloud inference, retrieval infrastructure, embeddings, vector databases, tool execution, external search, code sandboxes, moderation, storage, observability, and human fallback. Some of these costs are small; some can dominate.

 

More importantly, the cost is not merely per user. It depends on what the user does.

 

A short classification task sent to a small model may cost fractions of a cent. A long-context reasoning task using a frontier model can cost orders of magnitude more. Video generation, voice, and image workloads add different cost curves. Agents can call models repeatedly before returning one visible answer.

 

The result is a financial system in which engagement can improve retention while simultaneously compressing gross margin.

 

This is the opposite of the classic SaaS instinct that “more product usage is always good.” For AI, more usage is good only when the economics of that usage are controlled.

 

3. The Margin Evidence: What We Know and What We Do Not

 

The public evidence supports a meaningful margin gap between mature SaaS and frontier-model providers, but not a universal 30–60% rule for AI applications.

 

Reuters reported in February 2026 that OpenAI’s adjusted gross margin fell to 33% in 2025 from 40% in 2024 after inference costs increased fourfold. OpenAI’s planned compute spending through 2030 is extraordinarily large, illustrating that frontier-model economics remain infrastructure intensive.

 

For Anthropic, Reuters Breakingviews cited a PitchBook estimate of approximately 44% gross margin in August 2026. Anthropic’s revenue run rate had exceeded $65 billion by the end of July, but that growth continues to require massive infrastructure expansion. The margin estimate is useful but should remain explicitly labeled as a third-party estimate rather than an Anthropic disclosure.

 

Battery Ventures’ State of AI 2025 report provided a broader category framework: it estimated gross margins of roughly 0–30% for some AI application businesses and 30–60% for model inference, compared with 80%+ for traditional SaaS applications. The report also argued that margins should improve as token costs fall, routing improves, and pricing moves toward value or outcomes.

 

The important analytical correction is that these are different layers of the stack.

 

A frontier-model company owning training and inference infrastructure is not economically equivalent to an AI application buying a low-cost API. An AI coding agent with long autonomous sessions is not equivalent to a document-classification product. An image generator is not equivalent to compliance software that runs a model once per document and charges for a high-value regulated workflow.

 

There is no single “GenAI gross margin.”

 

4. Why the 30–60% “AI Margin” Rule Is Too Crude

 

Gross margin is an outcome of at least six design choices: workload intensity, model choice, price architecture, customer behavior, infrastructure ownership, and workflow value.

 

Consider two companies with identical $30 monthly subscriptions.

 

Company A performs a few lightweight classifications and spends $2 per active user on inference. Company B provides an agent that reads a repository, reasons over files, executes tools, and retries failures, generating $18 of monthly model and infrastructure cost for a heavy user. They are both “AI software.” Their economics are fundamentally different.

 

Exhibit 2 — $30 Subscription: How Usage-Driven Inference Cost Changes Gross Margin

The illustrative sensitivity model in Exhibit 2 assumes $3 per user of non-inference COGS. At $2 of inference cost, gross margin is about 83%. At $8, it falls to about 63%. At $20, gross margin is roughly 23%.

 

The difference is not category. It is unit economics.

 

This matters because founders and investors can misdiagnose a product by applying a category average. A 45% gross-margin business may be temporarily inefficient and rapidly improving. An 80% gross-margin product may achieve that number only because usage is weak. The direction of margin must be connected to usage, retention and customer value.

 

5. The Unlimited-Usage Trap

 

Fixed subscriptions were ideal for SaaS because the vendor could sell predictability without exposing itself to extreme marginal cost. AI weakens that assumption.

 

The risk is a heavy-user tail. A customer paying $20 or $30 per month can consume far more than that in frontier-model compute if the product allows long contexts, premium models or autonomous agent execution without limits.

 

This is why historical anecdotes such as the widely repeated estimate that ChatGPT once cost hundreds of thousands of dollars per day to operate are less useful than current pricing behavior. Whatever the precise 2023 number, the strategic response across the industry is visible: quotas, credits, model multipliers, rate limits and premium tiers.

 

The market is learning to price variance.

 

GitHub Copilot’s current individual plans show this clearly. The $10 Pro plan includes a defined monthly pool of AI credits; Pro+, Max and enterprise tiers provide progressively larger allowances. GitHub explicitly states that chat, agent mode, code review, cloud agents and CLI interactions consume credits, with cost depending on the model and complexity of the task. Code completion remains unlimited because its cost profile is different.

 

That is not a cosmetic billing change. It is a decomposition of the product into economically distinct workloads.

 

6. Agentic AI Makes Usage Variance a First-Class Financial Problem

 

The economics become more difficult as AI moves from assistant to agent.

 

Microsoft Azure Research’s 2026 production study of GitHub Copilot coding agents is unusually useful because it describes real workload behavior rather than synthetic benchmarks. The dataset covered roughly 13 million sessions from more than 3.2 million users, with 761 million LLM calls, 95 trillion tokens and 775 million tool invocations.

 

That is roughly 58 LLM calls and 60 tool invocations per sampled session on average. The study also found a roughly 50× token range across five user archetypes. Tool failures occurred in about 9% of turns and could drive approximately four times the compute through retry loops. Model switching severely reduced cache reuse.

 

Exhibit 3 — Agentic Software Turns One User Task Into Many Compute Events

This changes product finance in three ways.

 

First, users are not equal-cost seats. A fixed $30 agent license may contain users whose workload differs by an order of magnitude or more.

 

Second, reliability becomes a margin variable. A failed tool call is not merely a UX problem; it can trigger another model loop and another compute bill.

 

Third, systems engineering directly affects gross margin. Cache retention, model routing, context compaction, scheduler design and tool orchestration move from backend optimization to financial strategy.

 

The future CFO of an AI software company will need to understand inference architecture in a way the CFO of a classic SaaS company usually did not.

 

7. Model-Cost Deflation: The Most Important Counterforce

 

The bearish margin story becomes incomplete if it ignores the falling cost of intelligence.

 

Stanford’s 2025 AI Index estimated that the cost of querying a model with GPT-3.5-equivalent MMLU performance fell from $20 per million tokens in November 2022 to $0.07 by October 2024. Depending on workload, the report observed annual inference-price declines ranging from roughly 9× to 900×.

 

The trend has continued through product segmentation and efficiency. OpenAI’s July 30, 2026 update cut GPT-5.6 Luna’s API price by 80% and Terra’s by 20%. Current pricing spans a huge range: GPT-5.6 Luna at $0.20 per million input tokens and $1.20 output; Terra at $2 and $12; Sol at $5 and $30. Anthropic’s current pricing similarly ranges across model tiers, with Sonnet 5 offered at $2 input / $10 output per million tokens through August 2026 before moving to standard pricing.

 

Exhibit 4 — Inference Cost Deflation Is Real — But It Does Not Guarantee Margin Expansion

This creates a path to margin expansion unavailable in many physical businesses. If a workflow generates constant customer value while its inference cost falls 80%, gross profit can expand rapidly without raising price.

 

But there are two catches.

 

The first is competition. If every vendor gains access to cheaper models, customers may capture the savings through lower prices.

 

The second is induced demand. When intelligence becomes cheaper, products may use much more of it.

 

8. Why Cheaper Intelligence May Not Improve Margins

 

The natural assumption is that cheaper tokens mean higher gross margins. That is only true if usage does not expand faster than unit cost falls.

 

AI may exhibit a form of Jevons effect: efficiency lowers the cost of intelligence, which encourages products to consume more intelligence. A support assistant becomes an autonomous resolution agent. A coding autocomplete becomes a coding agent. A document summarizer becomes a full diligence workflow. A chatbot becomes a multi-agent research system.

 

The economic unit therefore changes.

 

Suppose model cost per token falls by 70%, but the product moves from one model call per task to ten calls because autonomy improves customer value. Total inference expense per completed workflow can still rise.

 

This means investors should track cost per customer outcome, not cost per token alone.

 

The relevant metric is not “How cheap is the model?” but “How much intelligence must the product buy to produce one unit of billable value?”

 

9. Pricing Architecture Is Now Part of Systems Architecture

 

Traditional SaaS pricing primarily captured willingness to pay. AI pricing must also manage cost variance.

 

Four structures are emerging.

 

Fixed subscription. The vendor accepts usage risk. This can work when tasks are low-cost, predictable or heavily cached, but it is dangerous when the usage distribution is long-tailed.

 

Subscription plus credits. The customer buys predictable access while the vendor caps extreme consumption. GitHub Copilot is a current example.

 

Pure usage-based pricing. Cost risk passes more directly to the customer. This is natural for APIs and infrastructure but can weaken budget predictability and make software feel like a utility.

 

Seat plus usage. Enterprise customers pay a platform or governance fee plus metered intelligence consumption. This may become the dominant compromise for high-value agentic software.

 

Outcome-based pricing. The vendor charges for completed work, savings or business results rather than tokens. This can create exceptional economics when customer value is large relative to compute, but it shifts execution and measurement risk to the vendor.

 

Exhibit 5 — AI Pricing Is Becoming a Risk-Allocation Mechanism

This is one reason “SaaS versus AI” is the wrong framing. AI is forcing software companies to rediscover pricing as a form of risk management.

 

10. Consumer Freemium: Why Conversion Alone Does Not Diagnose PMF

 

A common argument is that if only a small percentage of hundreds of millions of free AI users convert to paid plans, the business has weak product-market fit. That conclusion is analytically unsound without more information.

 

Freemium businesses deliberately maximize free reach. Low percentage conversion can coexist with tens of millions of paying customers, strong retention, and enormous absolute revenue. Free users may also create distribution, brand, data, or enterprise lead generation.

 

The right questions are cohort-specific: What percentage of high-intent users convert? What is paid-user retention? What is ARPU? What does a free user cost? How much free usage is subsidized by enterprise revenue or strategic distribution? Does the free tier improve acquisition efficiency?

 

For AI, one additional metric becomes essential: contribution margin by user cohort.

 

A low-engagement free user may cost almost nothing. A sophisticated unpaid user running long agent loops may be expensive. “Free MAU” is therefore not one economic category.

 

11. The Wrapper Question — When Is an AI Product Just Compute Resale?

 

“GPT wrapper” is often used as an insult rather than an analytical category. The relevant issue is not whether a company calls an external model. Many valuable software businesses depend on infrastructure they do not own.

 

The economic question is whether the product adds a defensible layer between raw model output and customer value.

 

A weak wrapper has little proprietary workflow, low switching cost, minimal data advantage and prices close to model consumption. Its gross margin is exposed both to upstream model pricing and downstream competition. If the model provider adds the feature directly, the application can disappear.

 

A strong AI application may also use third-party models, but it owns workflow integration, proprietary data context, distribution, compliance, user experience, evaluation systems, business process orchestration or outcome accountability. The model is an input, not the product.

 

The distinction is value capture.

 

If a company buys $1 of model inference and sells $1.30 of undifferentiated output, it is a low-margin reseller. If it buys $10 of inference to automate $500 of legal, accounting or engineering work, the model cost can be economically trivial even if accounting gross margin is only 60% during early scaling.

 

12. Vertical AI and the Economics of Labor Substitution

 

This is why vertical AI may produce some of the strongest businesses despite relatively heavy compute.

 

A legal drafting agent, insurance claims system, healthcare coding workflow, accounting close product or cybersecurity investigation agent can be priced against labor, delay, error rates and regulatory risk rather than against a generic software seat.

 

The economic denominator changes from “software budget” to “cost of work.”

 

A traditional SaaS product might charge $50 per user to make an employee 10% more productive. An AI agent that performs 40% of the employee’s workflow can potentially charge hundreds or thousands of dollars while consuming tens of dollars of compute.

 

This does not guarantee a good business. Vertical products may require implementation, human review, compliance, support and domain-specific data work. But the higher value pool gives the vendor room to absorb inference cost.

 

The strongest AI companies may therefore look less like software licenses and more like digitally delivered labor with software-like scalability.

 

13. The New AI Software Economic Taxonomy

 

Solten & Co. sees at least five emerging economic classes.

 

  1. Intelligence resellers. Products with limited differentiation that mainly package third-party models. Their margins are vulnerable to model commoditization and feature bundling.

 

  1. Workflow software with AI COGS. Products where AI is one variable cost inside a sticky workflow. These may converge toward traditional SaaS-like margins as models become cheaper.

 

  1. Agentic software. Products that perform multi-step work and carry high usage variance. Their economics require credits, usage controls or outcome pricing.

 

  1. AI infrastructure and platforms. Model APIs, inference clouds and data platforms. These are structurally more capital and compute intensive, with margins closer to cloud infrastructure than classic application SaaS.

 

  1. Outcome-priced digital labor. Products priced against completed work or avoided labor cost. These may have lower accounting gross margins than SaaS but much stronger absolute gross profit per customer.

 

The mistake is valuing all five as one category.

 

14. Valuation: When Does AI Software Deserve a SaaS Multiple?

 

High SaaS multiples historically reflected a specific combination: recurring revenue, high gross margin, strong retention, low capital intensity and predictable operating leverage.

 

AI software should deserve comparable valuation only when it demonstrates comparable economic durability, not because revenue is called ARR.

 

Investors should ask:

 

  • Is revenue recurring because customers are contractually committed, or because consumption happened to be high this quarter?
  • Does higher usage expand gross profit or compress it?
  • Is inference cost falling faster than price?
  • Can the company route across models without materially degrading quality?
  • Does the product have enough workflow value to pass compute cost through to customers?
  • Are gross margins expanding with scale?
  • Is customer retention strong after usage limits and pricing changes?
  • How much capital must be invested in infrastructure to support growth?
  • Is the company exposed to a model provider that can vertically integrate into the application?

 

A company with 50% gross margin, 150% net revenue retention and rapidly falling cost-to-serve may deserve more confidence than a company reporting 80% gross margin because customers barely use the product.

 

AI valuation should therefore incorporate margin trajectory and usage economics, not static margin alone.

 

15. Risks to Founders and Investors

 

Pricing risk. A fixed plan can become uneconomic as users discover higher-value or more compute-intensive workflows.

 

Provider risk. Upstream model vendors can change API prices, rate limits, model availability or product scope.

 

Commoditization risk. Falling model prices may reduce costs but also lower barriers to entry and compress application pricing.

 

Usage-growth risk. Greater engagement can raise COGS faster than revenue when pricing is not consumption-aware.

 

Quality-cost risk. Customers may prefer premium models, forcing the vendor to choose between margin and perceived product quality.

 

Agent-loop risk. Tool failures, retries and longer reasoning can create invisible compute leakage.

 

Infrastructure lock-in. Moving between models or clouds can destroy caching, require re-evaluation and create technical switching costs.

 

Capital risk. Companies that own model or inference infrastructure can require extraordinary financing before reaching positive cash flow.

 

Measurement risk. AI companies can present ARR or bookings without disclosing the cost required to produce that revenue. Revenue growth alone can hide deteriorating contribution economics.

 

16. Industry-Level Effects

 

The margin problem will reshape the software stack beyond individual companies.

 

First, model routing becomes financially valuable. Software that automatically chooses the cheapest model capable of completing a task can capture gross-margin expansion without visible product change.

 

Second, caching and context management become economic moats. The ability to reuse context or avoid repeated inference can create meaningful cost advantage.

 

Third, pricing complexity will rise. Buyers will face more credits, consumption units, outcome fees and blended contracts, making procurement and cost forecasting harder.

 

Fourth, AI will blur software and services. Outcome-priced agents may compete directly with labor budgets, business-process outsourcing and professional services rather than only software vendors.

 

Fifth, infrastructure ownership will become a strategic choice. At sufficient scale, companies may move from public APIs to reserved capacity, custom models or owned inference in order to improve margins.

 

Sixth, the gross-margin premium historically attached to software may fragment. Public markets may eventually distinguish high-margin workflow AI from infrastructure-heavy intelligence businesses in the same way they distinguish application SaaS from cloud data platforms today.

 

17. Solten & Co. Framework: Value per Inference Dollar

 

We propose a simple primary lens for AI application economics:

 

Value per Inference Dollar = Customer Economic Value Created / Fully Loaded AI Delivery Cost

 

The numerator should include measurable value: labor displaced, time saved, revenue generated, errors avoided, risk reduced or cycle time compressed.

 

The denominator should include more than API tokens: model calls, retrieval, embeddings, tool execution, hosting, human review, retries, support and any infrastructure directly attributable to the AI workflow.

 

This ratio should then be considered alongside three additional dimensions.

 

Pricing capture. What percentage of created value can the vendor monetize?

 

Usage variance. How widely does cost differ across customers performing ostensibly the same product action?

 

Cost trajectory. Is the fully loaded cost per successful workflow declining over time?

 

An attractive AI business has high customer value, meaningful pricing capture, controllable usage variance and a declining cost curve.

 

This framework is more useful than asking whether a product could exist without AI.

 

18. Thesis, Counter-Thesis and Falsification

 

Solten & Co. Thesis

 

Generative AI will not permanently destroy software gross margins. It will split software into new economic categories and force vendors to price intelligence explicitly. The strongest AI applications will recover high gross margins or high absolute gross profit by combining falling inference costs with workflow pricing power, model routing and operational efficiency. The weak businesses will be exposed as low-value compute resellers.

 

Counter-Thesis

 

AI may structurally reduce application-software margins because capability competition continually increases required compute. Every efficiency gain can be reinvested into longer context, stronger models, more agent steps and richer modalities. If users expect continuous improvement while subscription prices remain anchored to historic SaaS levels, the industry may settle at permanently lower gross margins and require lower valuation multiples.

 

Falsification Criteria

 

The Solten & Co. thesis would weaken if several of the following occur:

 

  • model prices stop declining materially despite continued hardware and algorithmic progress;
  • leading AI applications fail to expand gross margins after multiple years of scale;
  • usage-based and credit pricing produce significant customer resistance and churn;
  • outcome-priced AI cannot sustain pricing because model capabilities commoditize too quickly;
  • agentic workloads consume increasing compute without proportionate customer willingness to pay;
  • model providers capture most application value through vertical integration;
  • public markets consistently assign AI software materially lower valuation multiples even to businesses with strong retention and improving margins.

 

19. What to Watch Next

 

For investors and operators, the next phase should be monitored through concrete unit-economic signals rather than user-count headlines.

 

Gross margin by product or workload. Consolidated company margin can hide whether premium AI features are profitable.

 

Inference cost as a percentage of revenue. This reveals whether model cost is scaling efficiently.

 

Cost per successful workflow. Especially important for agents, where failures and retries can amplify consumption.

 

Model mix. What percentage of workloads require frontier models versus cheaper specialized models?

 

Cache hit rate and routing efficiency. These are increasingly financial metrics.

 

Paid usage architecture. Watch movement from unlimited subscriptions toward credits, metering and hybrid pricing.

 

Heavy-user economics. Averages can hide a loss-making tail.

 

Outcome pricing adoption. Evidence that vendors can price against labor or business results would support margin expansion.

 

Gross-margin trajectory of Anthropic and OpenAI. As frontier providers mature toward public-market scrutiny, these figures will become important benchmarks for the model layer.

 

AI gross-margin disclosures from public SaaS companies. ServiceNow, Snowflake, Salesforce, Cloudflare and others will increasingly reveal whether AI features improve or dilute established software economics.

 

20. Sources & Evidence

 

Primary/regulatory sources

 

Salesforce — FY2026 Form 10-K, subscription/support revenue and cost of revenue

https://www.sec.gov/Archives/edgar/data/1108524/000110852426000060/crm-20260131.htm

 

ServiceNow — FY2025 Form 10-K / earnings materials, subscription gross margin

https://www.sec.gov/Archives/edgar/data/1373715/000137371526000007/now-20251231.htm

 

Datadog — FY2025 Form 10-K, gross margin and cloud infrastructure costs

https://www.sec.gov/Archives/edgar/data/1561550/000162828026008819/ddog-20251231.htm

 

Snowflake — FY2026 Form 10-K, product gross margin and AI-inference infrastructure costs

https://www.sec.gov/Archives/edgar/data/1640147/000164014726000008/snow-20260131.htm

 

OpenAI — GPT-5.6 pricing and July 30, 2026 price-performance update

https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/

https://openai.com/api/pricing/

 

Anthropic — Claude list prices, effective May 27, 2026

https://www-cdn.anthropic.com/files/4zrzovbb/website/3684c2faafb97418665782cea0001f439f74b1d2.pdf

 

Anthropic — Claude Sonnet 5 pricing

https://www.anthropic.com/research/claude-sonnet-5

 

GitHub — Copilot plans and AI Credits pricing

https://github.com/features/copilot/plans

 

Microsoft Azure Research — Agentic Coding in the Wild: Characterizing GitHub Copilot at Production Scale, 2026

https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/ghcp_traces-6.pdf

 

Stanford HAI — AI Index Report 2025, inference cost trends

https://hai.stanford.edu/assets/files/hai_ai-index-report-2025_chapter1_final.pdf

 

Investor / analytical frameworks

 

Battery Ventures — State of AI Report 2025

https://www.battery.com/wp-content/uploads/2026/01/Battery-State-of-AI-Report-2025.pdf

 

High-quality financial reporting

 

Reuters — OpenAI compute spending and reported adjusted gross margin, Feb. 20, 2026

https://www.reuters.com/technology/openai-sees-compute-spend-around-600-billion-by-2030-cnbc-reports-2026-02-20/

 

Reuters Breakingviews — OpenAI margin comparison with Salesforce, Oct. 15, 2025

https://www.reuters.com/commentary/breakingviews/how-infer-method-to-openais-madness-2025-10-15/

 

Reuters Breakingviews — Anthropic gross-margin estimate, Aug. 18, 2026

https://www.reuters.com/commentary/breakingviews/anthropics-deceleration-has-some-ipo-upside-2026-08-18/

 

Reuters — Anthropic revenue run rate tops $65B, Aug. 17, 2026

https://www.reuters.com/technology/anthropic-revenue-run-rate-tops-65-billion-source-says-2026-08-17/

 

Methodological Note

 

This report deliberately rejects three shortcuts common in AI-business commentary.

 

First, it does not treat “AI gross margin” as one universal range. Economics differ by layer and workload.

 

Second, it does not treat free-to-paid conversion as a standalone measure of product-market fit.

 

Third, it does not use old anecdotal compute-cost estimates as evidence of current economics where better contemporary data exist.

 

Comparative gross-margin figures should be interpreted directionally because accounting definitions and business mix differ. Private-company figures for OpenAI and Anthropic are reported or estimated, not audited public filings.

 

About Solten & Co.

 

Solten & Co. is an independent research and analysis firm focused on the AI economy, with deeper research emphasis on AI infrastructure, software AI, Physical AI, robotics, and autonomous systems. We study the technologies, companies, markets, transactions, and capital structures shaping the next phase of AI.

 

Need an independent perspective?

 

Solten & Co. provides independent research and analytical support for investors and decision-makers evaluating companies, markets, and investment opportunities across the AI economy.

 

Discuss a research question · Discuss an investment opportunity · Request independent analysis

 

soltenco.com

 

OpenAI, Anthropic & the New Compute Power Structure

Why the frontier-model race is becoming a contest over capital, cloud distribution, custom silicon, and infrastructure optionality.

Solten & Co. Research Report
Updated: August 20, 2026
Research universe: AI Infrastructure / Software AI / Capital & Deals

Research Passport

  • Publication type: Flagship Research Report
  • Publication version: 1.0
  • Published/updated: August 20, 2026
  • Evidence cut-off: August 20, 2026
  • Research universe: AI Infrastructure / Software AI / Capital & Deals
  • Estimated reading time: ~25 minutes
  • Original research exhibits: 3
  • Primary and high-quality sources: 17+

Intended audience: investors, family offices, venture and growth funds, AI infrastructure operators, strategy leaders and other decision-makers evaluating frontier AI economics. 

Key Findings

  1. Frontier AI is becoming an industrial system, not merely a software-model competition. Capital, power, data centers, silicon, cloud infrastructure, frontier models and distribution are now tightly coupled.

 

  1. The old “OpenAI is locked into Azure” narrative is no longer an adequate description of the market. Microsoft remains central, but OpenAI has deliberately built infrastructure and capital optionality across AWS, NVIDIA, Oracle, CoreWeave and other partners while renegotiating exclusivity.

 

  1. Anthropic’s multi-provider strategy has become materially larger than the roughly $50 billion compute picture often repeated in older analysis. By 2026, its disclosed relationships span AWS Trainium, Google/Broadcom TPUs, Microsoft Azure/NVIDIA capacity, SpaceX GPU infrastructure and additional infrastructure programs.

 

  1. Strategic AI transactions increasingly combine equity, compute purchase commitments, cloud distribution, IP/model access and commercial participation. Headline “investment” values can therefore be misleading unless the economic layers are separated.

 

  1. Long-term compute commitments deserve to be analyzed as quasi-fixed strategic obligations. The central risks are not only model performance but also demand, utilization, silicon obsolescence, falling compute prices, capital dependence, and counterparty concentration.

Table of Contents

  1. The Source Thesis — What the Sources Get Right
  2. OpenAI–Microsoft: The Original Strategic Flywheel
  3. Revenue Sharing: Real, Material — but Frequently Misdescribed
  4. The Biggest Update: OpenAI Is No Longer Simply “Locked into Azure”
  5. OpenAI’s AWS Pivot: From Azure Dependency to Infrastructure Portfolio
  6. Anthropic: Multi-Cloud as Strategy, Not Temporary Compromise
  7. Anthropic’s 2026 Scale-Up Makes the Original Numbers Obsolete
  8. Consumer vs. Enterprise: Useful Distinction, Weak Precision
  9. The Economics: The Real Constraint Is Not “Model Quality” but Cost of Intelligence
  10. The New Strategic Map: Reciprocal Dependence, Not Simple Cloud Control
  11. Capital Is Becoming Part of the Compute Contract
  12. Compute Commitments Are Emerging as a Form of Strategic Debt
  13. What the 2026 Market Says About Anthropic vs. OpenAI
  14. What to Watch Next — Investor Monitoring Framework
  15. Solten & Co. View

Scope & Methodology

Research question. This report examines how the strategic and economic relationships surrounding OpenAI and Anthropic have changed the competitive structure of frontier AI, with particular attention to capital, cloud distribution, compute commitments, silicon strategy, contractual optionality, and infrastructure dependence.

Scope. The report focuses on publicly disclosed or credibly reported relationships involving OpenAI, Anthropic, Microsoft, Amazon/AWS, Google, NVIDIA and selected infrastructure providers. It is not intended as a comprehensive valuation of either OpenAI or Anthropic, nor as an investment recommendation.

Evidence hierarchy. Priority is given to company disclosures, partner announcements, and other primary evidence. High-quality financial reporting is used where material commercial terms are not publicly disclosed. The report avoids treating media-reported terms as audited company facts.

Evidence classes used in this report:

DISCLOSED FACT — directly supported by a primary or authoritative source.

REPORTED TERM — reported by a credible secondary source but not independently disclosed by the relevant counterparty.

ESTIMATE — a quantitative or qualitative assessment based on incomplete public information.

SOLTEN & CO. INTERPRETATION — our synthesis, inference, or analytical conclusion.

Limitations. Large AI infrastructure agreements are unusually difficult to compare. A dollar investment, a multi-year compute purchase commitment, an “up to” gigawatt capacity announcement, and realized installed/utilized capacity are different economic objects. This report therefore avoids combining unlike commitments into a single headline total unless the units and assumptions are explicitly comparable.

What Changed

The sources captured an important structural shift toward infrastructure control, but several of its strongest claims have already aged materially.

OpenAI is no longer well described as a single-cloud, Azure-locked company. Its Microsoft relationship remains strategically important, yet exclusivity has been relaxed and major AWS, NVIDIA, and other infrastructure relationships now create meaningful supplier optionality.

Anthropic’s compute footprint has expanded far beyond the earlier multi-cloud figures frequently cited in 2025-era analysis. The company’s 2026 agreements make infrastructure diversification itself a core strategic capability.

The competitive framing has also changed. The relevant question is no longer simply whether hyperscalers “control” AI labs. The evidence points to reciprocal dependence: model labs need capital and compute, while hyperscalers and silicon vendors need frontier labs as anchor tenants, distribution engines, and validation customers for custom infrastructure.

Finally, the economic object called an “AI investment” is increasingly composite. Equity capital, compute commitments, distribution rights, model/IP access, and commercial revenue participation may sit inside the same strategic relationship. That makes transaction decomposition essential for serious investment analysis.

Executive Summary

The useful insight in the sources is that the frontier-model race can no longer be understood primarily as a benchmark contest. Compute, cloud distribution, capital structure, custom silicon, enterprise access, and contractual flexibility have become strategic variables in their own right.

But the market has moved materially since many of the arrangements described in the sources were formed. The most important update is that OpenAI is no longer accurately described as being structurally locked into Microsoft Azure in the old sense. Microsoft remains a major shareholder and a central strategic partner, but the relationship was progressively loosened in 2025 and 2026. OpenAI now has major compute and distribution relationships with AWS and NVIDIA, can serve products through other cloud providers under the amended Microsoft agreement, and has committed to large-scale multi-provider infrastructure. Microsoft’s OpenAI IP license is now non-exclusive, while OpenAI’s revenue share to Microsoft continues through 2030 subject to a cap.

Anthropic, meanwhile, has gone even further in constructing a deliberately diversified infrastructure strategy. AWS remains its primary cloud and training partner, but Anthropic also uses Google TPUs and NVIDIA GPUs, has Claude available across AWS Bedrock, Google Vertex AI and Microsoft Azure Foundry, and has accumulated multi-gigawatt capacity agreements across providers. In 2026, it also added SpaceX GPU capacity and expanded AWS and Google commitments dramatically.

This changes the central analytical conclusion. The emerging structure is not simply “hyperscalers own the AI labs.” It is better understood as a dense network of reciprocal dependencies: labs need capital, power, and compute; hyperscalers need frontier models to drive cloud demand, custom-silicon adoption, and enterprise AI distribution; chip vendors need anchor customers; investors need exposure to the model layer; and model companies increasingly seek infrastructure optionality to avoid strategic dependence on any single supplier.

The long-term advantage may therefore accrue not to one layer alone, but to actors that control scarce infrastructure while preserving bargaining power across the stack.

1. The Source Thesis — What the Source Gets Right

The source argues that the AI industry has entered a “middle game” in which model quality remains important but infrastructure control, compute access, and cloud distribution increasingly determine who can operate at frontier scale.

That framing is directionally correct.

Training and serving frontier models has become a capital-intensive industrial activity. The relevant inputs are no longer only algorithms and training data. They include data-center capacity, power, accelerators, networking, memory, inference infrastructure, custom silicon, cloud procurement and long-term financing. The frontier-model companies are therefore becoming unusually intertwined with hyperscalers, semiconductor vendors and infrastructure financiers.

OpenAI itself described the 2026 scaling problem in three words: “compute, distribution, and capital.” In February 2026, when announcing $110 billion of new investment, the company said leadership in the next phase would be defined by who could scale infrastructure fast enough to meet demand and convert that capacity into products used at global scale.

This is an important shift in how the sector should be analyzed. A useful company model can no longer stop at product quality, model benchmarks, or subscription growth. It must also ask:

  • Who finances the company’s infrastructure?
  • Which cloud providers distribute the models?
  • What silicon does the company depend on?
  • How much power and capacity has it contracted?
  • Which agreements are exclusive?
  • Which agreements create minimum-purchase or long-term compute obligations?
  • Who receives revenue shares?
  • Where can the company switch providers, and at what economic or technical cost?
  • What does the infrastructure partner receive beyond direct cloud revenue — equity appreciation, custom-silicon validation, enterprise distribution or strategic leverage?

That is the stronger framework behind the source.

2. OpenAI–Microsoft: The Original Strategic Flywheel

Microsoft’s relationship with OpenAI began as a strategic shortcut into frontier AI. Microsoft invested $1 billion in OpenAI in 2019 and subsequently expanded the relationship through additional capital, cloud capacity, IP rights, and distribution arrangements.

The logic was powerful on both sides.

OpenAI received access to a hyperscaler capable of financing and deploying enormous compute clusters. Microsoft received privileged access to frontier-model IP, a differentiated Azure AI offering, and the ability to incorporate OpenAI technology across products such as Copilot and Microsoft 365.

For several years, this created a reinforcing system:

Microsoft capital → OpenAI compute demand → Azure revenue → OpenAI model improvement → Microsoft product differentiation → enterprise Azure demand.

The structure also made Microsoft simultaneously investor, infrastructure provider, distributor, and commercial beneficiary.

In October 2025, OpenAI completed a major recapitalization. Microsoft’s investment in OpenAI Group PBC was valued at approximately $135 billion, representing roughly 27% of the company on an as-converted diluted basis. OpenAI’s nonprofit parent — the OpenAI Foundation — held 26%, while employees and other investors held the remaining 47%.

This part of the source is substantially correct: Microsoft ended up with an economic stake in OpenAI worth about $135 billion, although the precise percentage is better stated as roughly 27%, not a loose 26–30% range.

3. Revenue Sharing: Real, Material — but Frequently Misdescribed

The Microsoft–OpenAI relationship has included revenue-sharing arrangements flowing in both directions.

Historically, reporting indicated that OpenAI agreed to share approximately 20% of revenue with Microsoft through 2030. In 2025, Reuters reported that OpenAI planned to reduce Microsoft’s share over time, but the companies subsequently confirmed that the revenue-sharing arrangement remained in place.

The structure changed again in April 2026.

Microsoft stated that it would no longer pay a revenue share to OpenAI. Revenue-share payments from OpenAI to Microsoft would continue through 2030 at the same percentage, but subject to an overall cap. Reuters subsequently reported, citing The Information, that the total future revenue-sharing obligation had been capped at approximately $38 billion.

This is materially different from the older picture in which reciprocal revenue sharing was treated as a relatively stable permanent mechanism.

The investment implication is important. Microsoft still has several distinct ways to benefit economically from OpenAI:

  1. Equity appreciation through its roughly 27% ownership position.
  2. Revenue-share payments from OpenAI through 2030, subject to the agreed cap.
  3. Azure consumption and broader infrastructure revenue.
  4. Product differentiation and enterprise distribution through Microsoft’s own AI products.
  5. Strategic spillovers into Microsoft’s cloud and developer ecosystem.

The source’s claim that Microsoft captured $865 million through revenue sharing in the first nine months of 2025 should be treated as an externally reported figure rather than a primary-source fact. We did not find a first-party Microsoft or OpenAI disclosure confirming that exact number. It may be useful context, but it should not be presented as audited public financial disclosure.

4. The Biggest Update: OpenAI Is No Longer Simply “Locked into Azure”

This is where the source narrative is now most outdated.

In early 2025, Microsoft still described the OpenAI API as exclusive to Azure and retained a right of first refusal on new compute capacity. The October 2025 agreement preserved Azure API exclusivity and Microsoft’s exclusive IP rights until AGI, while also allowing OpenAI greater freedom to build additional compute elsewhere.

By February 2026, OpenAI and Microsoft jointly clarified that Azure remained the exclusive cloud provider for stateless OpenAI APIs, even while OpenAI pursued additional compute relationships.

Then, in April 2026, the relationship changed more fundamentally.

Microsoft announced an amended agreement under which:

  • Microsoft remains OpenAI’s primary cloud partner.
  • OpenAI products generally ship first on Azure, unless Microsoft cannot or chooses not to support the required capabilities.
  • OpenAI can serve its products to customers through any cloud provider.
  • Microsoft’s license to OpenAI models and products through 2032 became non-exclusive.
  • Microsoft stopped paying revenue share to OpenAI.
  • OpenAI’s revenue-share payments to Microsoft continue through 2030, subject to a cap.

This is a major strategic shift.

The original Microsoft–OpenAI partnership was based on deep bilateral dependence. The revised structure increasingly resembles a major strategic partnership inside a broader multi-cloud and multi-capital network.

OpenAI’s subsequent AWS relationship makes this concrete.

5. OpenAI’s AWS Pivot: From Azure Dependency to Infrastructure Portfolio

In November 2025, OpenAI and AWS announced a $38 billion multi-year agreement under which OpenAI would use AWS infrastructure containing hundreds of thousands of NVIDIA GPUs.

In February 2026, that relationship expanded dramatically.

Amazon committed to invest $50 billion in OpenAI. OpenAI and AWS expanded their infrastructure agreement by another $100 billion over eight years. OpenAI committed to consume approximately 2 gigawatts of Trainium capacity, spanning Trainium3 and Trainium4, beginning to ramp in 2027.

AWS also became the exclusive third-party cloud distribution provider for OpenAI Frontier, while OpenAI and Amazon agreed to co-develop a stateful runtime environment in Amazon Bedrock and customized models for Amazon applications.

By June 2026, OpenAI frontier models and Codex were generally available on AWS.

OpenAI also announced 3 GW of dedicated NVIDIA inference capacity and 2 GW of training capacity on Vera Rubin systems, in addition to infrastructure already running across Microsoft, Oracle Cloud Infrastructure and CoreWeave.

The strategic implication is clear: OpenAI is deliberately creating infrastructure optionality.

This does not make Microsoft unimportant. Microsoft remains the primary cloud partner, major shareholder, and revenue-share recipient. But OpenAI now has multiple meaningful infrastructure and capital relationships, reducing the old single-provider concentration risk and increasing its bargaining flexibility.

6. Anthropic: Multi-Cloud as Strategy, Not Temporary Compromise

Anthropic’s infrastructure model has historically been more diversified than OpenAI’s.

AWS became Anthropic’s primary cloud provider for mission-critical workloads in 2023. In November 2024, Amazon added another $4 billion investment, bringing its total at that time to $8 billion, while AWS became Anthropic’s primary cloud and training partner.

Anthropic simultaneously deepened its Google relationship. In October 2025, it announced plans to use up to one million Google TPUs, with more than one gigawatt of capacity expected online in 2026. Anthropic explicitly described its compute strategy as diversified across Google TPUs, AWS Trainium, and NVIDIA GPUs.

In November 2025, Anthropic expanded into Microsoft Azure as well. Microsoft committed to invest up to $5 billion in Anthropic and NVIDIA up to $10 billion, while Anthropic committed to purchase $30 billion of Azure compute capacity. Claude became available in Microsoft Foundry, giving Anthropic distribution across AWS Bedrock, Google Vertex AI and Microsoft Azure.

This three-cloud distribution position was strategically unusual and important.

7. Anthropic’s 2026 Scale-Up Makes the Original Numbers Obsolete

The source cites Anthropic’s compute commitments at roughly $50 billion. That figure is no longer a useful description of current exposure.

In April 2026, Anthropic announced a new agreement with Google and Broadcom for multiple gigawatts of next-generation TPU capacity, expected to begin coming online in 2027. Anthropic said this was its most significant compute commitment to date.

Later that month, Anthropic expanded its AWS relationship to secure up to 5 GW of new capacity and committed more than $100 billion over ten years to AWS technologies. It said it was already using more than one million Trainium2 chips and that Amazon remained its primary training and cloud provider.

Amazon simultaneously invested another $5 billion, with the possibility of up to $20 billion more in future investment.

Anthropic also added more than 300 MW of NVIDIA GPU capacity through SpaceX’s Colossus infrastructure.

By mid-2026, Anthropic’s infrastructure strategy included:

  • AWS Trainium — primary cloud/training relationship, up to 5 GW of new capacity.
  • Google/Broadcom TPUs — multi-gigawatt next-generation capacity.
  • Microsoft Azure/NVIDIA — $30 billion Azure capacity relationship.
  • SpaceX — more than 300 MW of NVIDIA GPU capacity.
  • Fluidstack — part of a broader $50 billion U.S. infrastructure program.

The strategic pattern is not merely “multi-cloud.” It is active procurement diversification across cloud vendors, silicon architectures, and physical infrastructure providers.

8. Consumer vs. Enterprise: Useful Distinction, Weak Precision

The source contrasts OpenAI as consumer-driven and Anthropic as enterprise-driven. This distinction is broadly useful, but the exact percentages cited in the source should be treated cautiously.

For OpenAI, Reuters reported in late 2025 that roughly 30% of revenue came from enterprise customers, implying that the majority still came from consumer-oriented products, particularly ChatGPT subscriptions. This supports the directional claim that OpenAI was more consumer-weighted than Anthropic.

For Anthropic, multiple sources confirm unusually strong enterprise adoption. In 2025, Reuters described Anthropic’s revenue acceleration as being driven by business demand, particularly coding. Anthropic reported more than 300,000 business customers in October 2025 and more than 500 customers spending over $1 million annualized by early 2026; that figure exceeded 1,000 by April 2026.

However, the claim that “85% of Anthropic revenue comes from B2B API calls” is not supported by the first-party sources reviewed here. Anthropic is clearly enterprise-heavy, but its revenue mix now includes subscriptions, Claude Code, direct enterprise contracts, API activity and hyperscaler distribution. A single percentage may be both unverifiable and quickly obsolete.

The deeper point is more important than the percentages:

OpenAI built a massive direct consumer distribution engine first, then pushed aggressively into enterprise.

Anthropic built stronger early positioning in enterprise and developer workflows, especially coding, while using broad hyperscaler distribution to reach customers inside existing cloud environments.

By 2026, both companies are converging toward enterprise distribution, making the old consumer-versus-enterprise binary less clean.

9. The Economics: The Real Constraint Is Not “Model Quality” but Cost of Intelligence

The source says OpenAI spends roughly $2 for every $1 of revenue. That framing is too simplistic for serious analysis.

OpenAI’s financial profile is undeniably capital-intensive, but reported figures must distinguish cash spending, compute costs, operating losses, and non-cash restructuring charges.

The Financial Times reported that OpenAI generated approximately $13 billion of revenue in 2025 while spending $34 billion. It reported a $39 billion net loss, but roughly $30 billion of that was a non-cash charge related to the prior investor structure. Excluding that charge and other non-cash items, operational losses were reported at around $8 billion.

This means the simple statement “OpenAI spends $2 for every $1 earned” may capture the scale of gross cash requirements at some point, but it is not a reliable representation of ongoing operating unit economics.

The more interesting economic question is how quickly inference costs decline relative to usage growth.

Frontier AI has a structural tension:

Better models → more demand → more inference → more infrastructure spending.

If price per unit of intelligence falls faster than compute efficiency improves, revenue growth may not translate proportionally into margin expansion.

This is why custom silicon and infrastructure bargaining power matter so much.

AWS wants Trainium adoption because custom silicon can reduce dependence on NVIDIA and capture more of the economics inside AWS.

Google wants TPU scale for the same reason.

Microsoft wants OpenAI-driven Azure demand and deeper software integration.

NVIDIA wants long-duration demand visibility from the frontier labs.

The model labs want enough supplier diversity to prevent any single infrastructure provider from capturing too much of their margin.

10. The New Strategic Map: Reciprocal Dependence, Not Simple Cloud Control

The source’s broader takeaway — that hyperscalers may be the true winners — is plausible but incomplete.

Hyperscalers clearly occupy a privileged position because they control data-center footprints, power procurement, networking, cloud distribution, enterprise relationships, and increasingly custom silicon.

But frontier labs also possess leverage.

A leading model company can move enormous volumes of compute procurement, validate a new chip architecture, attract cloud customers, strengthen an enterprise AI platform, and create equity gains for strategic investors.

This creates reciprocal dependence.

Consider AWS and Anthropic.

Anthropic needs AWS infrastructure. But AWS also uses Anthropic as the flagship proof point for Trainium. Project Rainier and more than one million Trainium2 chips give Amazon a reference customer at a scale few others can provide. Anthropic therefore helps AWS validate a strategic attempt to reduce dependence on NVIDIA.

Likewise, OpenAI’s 2 GW Trainium commitment is strategically important to AWS’s custom-silicon business.

The model companies are not merely buyers. They are anchor tenants for an emerging AI industrial infrastructure.

11. Capital Is Becoming Part of the Compute Contract

A striking feature of the current market is the increasing overlap between investor and supplier roles.

Microsoft is both a major OpenAI shareholder and infrastructure partner.

Amazon is both a major Anthropic shareholder and Anthropic’s primary cloud provider. It is now also a $50 billion OpenAI investor and major OpenAI compute supplier.

Google is an Anthropic investor and TPU/cloud partner.

NVIDIA invests in both model companies while also supplying the accelerators on which much of the industry depends.

This creates a new analytical problem for investors: headline “investment” announcements cannot be evaluated independently from procurement commitments, cloud contracts, revenue-sharing agreements and supplier incentives.

A strategic investment may effectively subsidize future infrastructure consumption. A compute commitment may in turn secure capital, distribution, or silicon priority.

Therefore, future Solten & Co. deal analysis should separate at least five economic layers in AI transactions:

  1. Equity investment.
  2. Compute purchase commitments.
  3. Cloud distribution rights.
  4. IP/model licensing rights.
  5. Revenue-share or commercial participation rights.

Without separating these layers, reported transaction values can be misleading.

12. Compute Commitments Are Emerging as a Form of Strategic Debt

Long-term compute commitments are not debt in the legal accounting sense, but economically they can behave like quasi-fixed obligations.

A lab that contracts tens or hundreds of billions of dollars of future capacity is making a bet on continued demand growth, model economics, and capital availability.

This creates several risks:

Demand risk — future AI usage may grow more slowly than contracted capacity.

Price risk — compute prices may fall faster than expected, making old commitments expensive relative to market alternatives.

Technology risk — a contracted silicon architecture may become less competitive.

Capital risk — the company may need continuous financing to fund capacity before operating cash flow catches up.

Utilization risk — infrastructure economics deteriorate sharply if expensive capacity is underused.

Counterparty risk — hyperscalers and infrastructure providers become increasingly exposed to the financial health of a small group of frontier labs.

This is one of the most important areas for future investment research because the market often celebrates giant compute commitments as evidence of confidence while under-analyzing their downside asymmetry.

13. What the 2026 Market Says About Anthropic vs. OpenAI

The competitive picture changed dramatically in 2026.

Anthropic reported run-rate revenue above $30 billion in April 2026, up from approximately $9 billion at the end of 2025. Reuters reported in August 2026 that Anthropic’s annualized run rate had exceeded $65 billion by the end of July.

This is far beyond the scale implied in the source.

OpenAI remains enormous, with unmatched consumer awareness and major enterprise ambitions, but recent reporting suggests stronger competitive pressure from Anthropic in coding and enterprise workloads.

The investment lesson is not that Anthropic has definitively “won.” It is that infrastructure strategy and distribution architecture can materially affect commercial outcomes.

Anthropic’s ability to be present inside all three major cloud ecosystems reduced customer procurement friction and gave it multiple infrastructure paths.

OpenAI, initially more concentrated around Microsoft, has spent 2025–2026 building similar optionality through AWS, NVIDIA, Oracle, CoreWeave and Stargate.

The two firms are therefore converging toward a common strategic requirement: no frontier lab wants to depend on a single source of capital, compute, or distribution.

14. What to Watch Next — Investor Monitoring Framework

For investors analyzing frontier AI, cloud infrastructure or adjacent companies, the following metrics may now be more informative than benchmark leadership alone.

Infrastructure concentration

What percentage of training and inference depends on each provider?

Committed capacity

How much future compute, power, and data-center capacity is contractually committed?

Compute economics

What is the effective cost per training run, per inference token or per unit of delivered intelligence?

Silicon mix

How exposed is the company to NVIDIA versus Trainium, TPU or other accelerators?

Distribution breadth

Can the company sell through AWS, Azure, Google Cloud, and directly?

Revenue concentration

How dependent is growth on consumers, coding tools, API use or a small number of enterprise customers?

Capital dependency

How much external funding is required before free cash flow becomes plausible?

Strategic investor overlap

Are suppliers also shareholders? Do those relationships distort apparent pricing or economics?

Contract flexibility

Can the company shift workloads across providers when technology or economics change?

Infrastructure utilization

Are contracted gigawatts translating into monetized demand?

15. Solten & Co. View

The frontier AI market is evolving from a software race into an industrial system.

That system has at least six tightly coupled layers:

Capital → Power/Data Centers → Silicon → Cloud Infrastructure → Frontier Models → Applications/Distribution.

The key strategic question is no longer only who has the best model.

It is who can secure sufficient capital and infrastructure without surrendering too much economics or strategic flexibility to the companies supplying that infrastructure.

OpenAI’s 2025–2026 evolution is a case study in reducing dependency. Microsoft remains central, but OpenAI has steadily expanded into AWS, NVIDIA, and other infrastructure partners while renegotiating exclusivity.

Anthropic is a case study in diversified infrastructure from an earlier stage. AWS remains primary, but Anthropic has systematically maintained access to multiple clouds and silicon architectures.

The hyperscalers are likely to capture enormous value because the frontier-model boom drives cloud demand, custom-silicon adoption and enterprise AI distribution. But the idea that they will automatically own the economics is too simple.

The more likely equilibrium is a small number of frontier labs and infrastructure giants locked in reciprocal dependence, each attempting to diversify enough to preserve bargaining power.

For investors, the most underappreciated layer may be the contracts connecting them.

Those contracts — compute commitments, distribution rights, strategic investments, revenue shares and silicon partnerships — increasingly determine which companies have flexibility, which carry hidden obligations, and where economic value ultimately accrues.

Fact-Check: Selected Claims from the Source

Claim: Microsoft owns approximately 26–30% of OpenAI, worth about $135B.

Status: Substantially verified. Microsoft disclosed roughly 27%, valued at approximately $135B after the October 2025 recapitalization.

Claim: Microsoft receives roughly 20% of OpenAI revenue through 2030.

Status: Historically supported by reporting; the 2026 amended agreement kept the same percentage through 2030 but added a total cap. Reuters later reported the cap at approximately $38B.

Claim: OpenAI receives roughly 20% of Azure OpenAI/Bing AI revenue.

Status: Reciprocal revenue sharing existed historically, but this is now outdated. Microsoft said in April 2026 that it would no longer pay a revenue share to OpenAI.

Claim: OpenAI committed to purchase $250B of Azure services.

Status: Verified as an October 2025 incremental Azure commitment. However, the broader relationship has since changed materially, and OpenAI has added very large AWS and NVIDIA commitments.

Claim: Microsoft has exclusive API distribution rights until AGI.

Status: Outdated. This described the earlier structure. The April 2026 amendment materially relaxed exclusivity and made Microsoft’s IP license non-exclusive, while OpenAI gained the ability to serve products through other clouds.

Claim: Anthropic has approximately $50B in compute commitments across providers.

Status: Outdated. Anthropic’s 2026 commitments expanded far beyond this number, including more than $100B committed to AWS alone over ten years, multi-gigawatt Google capacity, $30B of Azure capacity and additional GPU infrastructure.

Claim: Anthropic uses up to 1M Google TPUs and around 1 GW of Google capacity.

Status: Verified as the October 2025 announced plan. The Google/Broadcom relationship expanded again in April 2026 into multiple gigawatts of next-generation TPU capacity.

Claim: Amazon invested $8B in Anthropic and AWS is its primary cloud provider.

Status: Verified historically. Amazon’s total investment has since increased, and AWS remains Anthropic’s primary cloud and training provider.

Claim: Anthropic is available through AWS Bedrock, Google Vertex AI and Azure Foundry.

Status: Verified. Anthropic describes Claude as the only frontier AI model available across all three major cloud platforms.

Claim: OpenAI is ~73% consumer revenue and Anthropic ~85% B2B API revenue.

Status: Directionally plausible but not sufficiently supported at those exact percentages by primary evidence reviewed. Reuters reported roughly 30% enterprise revenue for OpenAI in late 2025 and consistently described Anthropic growth as enterprise-led. Exact percentages should not be used without a dated underlying source.

Claim: OpenAI spends roughly $2 for every $1 of revenue.

Status: Oversimplified. OpenAI is highly cash-intensive, but reported losses include major non-cash items and changing compute economics. Use specific period financial data instead of a permanent ratio.

Primary and High-Quality Sources

Microsoft — The next chapter of the Microsoft–OpenAI partnership, Oct. 28, 2025

https://blogs.microsoft.com/blog/2025/10/28/the-next-chapter-of-the-microsoft-openai-partnership/

OpenAI — Our Structure

https://openai.com/our-structure/

OpenAI/Microsoft — Joint Statement, Feb. 27, 2026

https://openai.com/index/continuing-microsoft-partnership/

Microsoft — The next phase of the Microsoft–OpenAI partnership, Apr. 27, 2026

https://blogs.microsoft.com/blog/2026/04/27/the-next-phase-of-the-microsoft-openai-partnership/

Reuters — OpenAI/Microsoft revenue-share cap report, May 12, 2026

https://www.reuters.com/technology/openai-cap-microsoft-revenue-sharing-38-billion-information-reports-2026-05-12/

OpenAI — AWS and OpenAI multi-year strategic partnership, Nov. 3, 2025

https://openai.com/index/aws-and-openai-partnership/

OpenAI — OpenAI and Amazon strategic partnership, Feb. 27, 2026

https://openai.com/index/amazon-partnership/

OpenAI — Scaling AI for everyone, Feb. 27, 2026

https://openai.com/index/scaling-ai-for-everyone/

OpenAI — Frontier models and Codex available on AWS, Jun. 1, 2026

https://openai.com/index/openai-frontier-models-and-codex-are-now-available-on-aws/

Anthropic — Powering the next generation of AI development with AWS, Nov. 22, 2024

https://www.anthropic.com/news/anthropic-amazon-trainium

Anthropic — Expanding our use of Google Cloud TPUs and Services, Oct. 23, 2025

https://www.anthropic.com/news/expanding-our-use-of-google-cloud-tpus-and-services

Anthropic — Microsoft, NVIDIA and Anthropic strategic partnerships, Nov. 2025

https://www.anthropic.com/news/microsoft-nvidia-anthropic-announce-strategic-partnerships

Anthropic — Google/Broadcom next-generation compute expansion, Apr. 6, 2026

https://www.anthropic.com/news/google-broadcom-partnership-compute

Anthropic — Amazon compute expansion, Apr. 20, 2026

https://www.anthropic.com/news/anthropic-amazon-compute

Anthropic — Higher limits and SpaceX compute partnership, 2026

https://www.anthropic.com/news/higher-limits-spacex

Reuters — Anthropic annualized revenue reached $3B on business demand, May 30, 2025

https://www.reuters.com/business/anthropic-hits-3-billion-annualized-revenue-business-demand-ai-2025-05-30/

Reuters — Anthropic revenue run rate tops $65B, Aug. 17, 2026

https://www.reuters.com/technology/anthropic-revenue-run-rate-tops-65-billion-source-says-2026-08-17/

Financial Times — OpenAI spending hit $34B in 2025

https://www.ft.com/content/e15b0d7e-ff6b-4f16-ba7a-4068feddb828

Evidence & Publication Note

Publication-ready research report. The analysis separates disclosed facts, externally reported terms, and Solten & Co. interpretation, and includes original research exhibits covering selected compute capacity, partnership structure, and the evolution toward multi-provider infrastructure portfolios.

Research Exhibits

Exhibit 1 — Selected Announced Frontier-Lab Compute Capacity

The figures below capture disclosed capacity announcements, not directly comparable installed capacity. Timing, silicon, workload type, and “up to” language differ materially across agreements.

Exhibit 2 — AI Strategic Partnerships Are Multi-Layer Transactions

The same counterparty can simultaneously be an investor, compute supplier, distributor, model-access partner and commercial beneficiary. This is why headline investment values alone are poor representations of the underlying economics.

Exhibit 3 — From Bilateral Cloud Partnerships to Multi-Provider Infrastructure Portfolios

The chronology shows the strategic shift from relatively concentrated bilateral relationships toward overlapping networks of capital, compute, and distribution.

Methodological Note

These exhibits are based on disclosed company announcements and high-quality reporting available through August 20, 2026. They intentionally avoid converting unlike commitments into a single headline value. Gigawatts describe capacity; dollars describe investments or contractual purchase obligations; neither is equivalent to realized utilization or economic value. Where terms are described as “up to,” the chart retains that qualification in the underlying analysis.

About Solten & Co.

Solten & Co. is an independent research and analysis firm focused on the AI economy, with deeper research emphasis on AI infrastructure, Physical AI, robotics, and autonomous systems. We study the technologies, companies, markets, transactions, and capital structures shaping the next phase of AI.

Need an independent perspective?

Solten & Co. provides independent research and analytical support for investors and decision-makers evaluating companies, markets, and investment opportunities across the AI economy.

Discuss a research question · Discuss an investment opportunity · Request independent analysis.