The New SaaS Math

How Generative AI Is Rewriting Software Margins, Pricing and Operating Leverage
Solten & Co. Research Report
Published / updated: August 20, 2026
Research universe: Software AI / AI Infrastructure / Economics & Benchmarks

Executive Summary

 

The most common financial shorthand about generative-AI software is that traditional SaaS enjoys 70–90% gross margins while AI-native products fall to 30–60% because every user action consumes tokens, GPUs, or model-API capacity. The direction of that claim is useful; the conclusion is too simple.

 

Generative AI does change the financial architecture of software. It converts part of software delivery from a largely fixed-cost infrastructure problem into a workload-sensitive production cost. A conventional SaaS user may log in more often without materially changing the vendor’s cost to serve. An AI user can create a very different economic outcome: a longer context window, a more capable model, an autonomous agent loop, repeated tool calls or an image/video workload can multiply compute consumption while the customer continues paying the same monthly fee.

 

That makes usage itself a financial variable. It also explains why AI pricing is rapidly moving away from the old idea of an unlimited seat. Credits, metered usage, model tiers, rate limits, and hybrid seat-plus-consumption structures are not merely product packaging. They are mechanisms for allocating inference-cost risk between vendor and customer.

 

But lower gross margins today do not prove that AI software is structurally inferior to SaaS. The cost of intelligence is falling extraordinarily quickly. Stanford’s AI Index found that the cost of querying a model with GPT-3.5-equivalent benchmark performance fell from about $20 per million tokens in November 2022 to $0.07 by October 2024 — more than a 280-fold decline. OpenAI cut the API price of GPT-5.6 Luna by 80% in July 2026 and Terra by 20%, illustrating that model price compression continues. Routing, caching, quantization, speculative decoding, smaller specialized models, and owned infrastructure can all lower unit cost further.

 

The complication is that demand is evolving just as quickly. Agentic software is not simply “chat with more tokens.” Microsoft Azure Research analyzed 13 million GitHub Copilot coding-agent sessions from June 2026 and found 761 million LLM calls and 775 million tool invocations. User archetypes differed by roughly 50× in token consumption; tool failures generated retry loops that materially amplified compute. As software moves from answering questions to performing multi-step work, a single customer task can become dozens of inference events.

 

This produces the central economic tension of AI software: cost per unit of intelligence is falling, but the number of units consumed per customer outcome may rise even faster.

 

Traditional SaaS therefore remains an important benchmark, but not a universal destination. Salesforce generated a fiscal-2026 total gross margin of approximately 77.7%; ServiceNow’s 2025 subscription gross margin was 80%; Datadog’s 2025 gross margin was 80%. Yet even cloud-native, consumption-heavy software already shows lower structural margins: Snowflake’s fiscal-2026 product gross margin was 72%, with the company explicitly citing third-party cloud costs, including AI inference, as a major cost-of-revenue driver. Software margins were never determined only by whether the product was “software”; they reflect where compute, data, support, and infrastructure costs sit in the value chain.

 

At the frontier-model layer, the contrast is sharper. Reuters reported that OpenAI’s adjusted gross margin fell to approximately 33% in 2025 from 40% in 2024 as inference costs increased fourfold. Reuters Breakingviews cited a PitchBook estimate of approximately 44% gross margin for Anthropic in 2026. Those figures are not directly comparable to public SaaS accounting and should not be generalized across all AI applications, but they demonstrate the economic burden of operating frontier models at scale.

 

The most important conclusion is therefore not “AI has lower margins than SaaS.” It is this:

 

AI changes software from a business in which usage is usually desirable into one in which usage must be economically designed.

 

The quality of an AI business increasingly depends on four variables operating together: the value created per unit of inference cost; the vendor’s ability to price that value; the variance and growth of usage; and the rate at which model and infrastructure costs decline. A 55% gross-margin AI product that replaces $1,000 of human labor with $50 of compute may be economically stronger than an 85% gross-margin SaaS tool that saves a customer $20. Conversely, an AI wrapper selling $20 subscriptions while exposing itself to $30 of unbounded monthly inference cost is not a software business with temporarily bad margins; it is a structurally mispriced compute reseller.

 

For investors, this means SaaS multiples should not be applied mechanically to AI revenue. The right question is whether the company has a credible path from raw intelligence consumption to durable gross profit. For founders, it means product design, model architecture and pricing architecture have become inseparable from financial architecture.

 

Key Findings

 

  1. Generative AI does not eliminate SaaS economics; it makes them workload-dependent. The key change is not “software becomes expensive to serve,” but that cost-to-serve varies materially with user behavior, model choice, and task complexity.

 

  1. The widely repeated 30–60% gross-margin range for AI should be treated as a category estimate, not a law. Battery Ventures’ 2025 State of AI framework estimated 0–30% gross margins for some AI application businesses and 30–60% for model inference, but actual company outcomes vary dramatically by workload, pricing, and vertical value.

 

  1. Traditional SaaS itself spans a broad margin range. Public benchmarks show Salesforce near 78% total gross margin, ServiceNow at 80% subscription gross margin, Datadog at 80%, and Snowflake at 72% product gross margin. Consumption intensity already mattered before generative AI.

 

  1. Frontier-model economics remain materially below classic SaaS benchmarks. OpenAI’s reported adjusted gross margin fell to roughly 33% in 2025; Anthropic’s 2026 margin has been estimated around 44%. These are model-provider economics, not universal AI-software margins.

 

  1. Falling model prices create a powerful margin-expansion opportunity, but cost deflation does not automatically become profit. Competition may pass the savings to customers, and agentic products can increase the amount of inference consumed per task.

 

  1. Agentic AI makes the “unlimited seat” structurally dangerous for high-variance workloads. Microsoft’s production-scale GitHub Copilot traces show highly heterogeneous usage and dozens of LLM/tool events per agent session. Heavy users can have radically different costs under the same seat price.

 

  1. The market is already responding through pricing redesign. GitHub Copilot now combines subscription tiers with AI credits; Microsoft increasingly meters agent workloads through Copilot Credits. Credits are not merely monetization tactics — they cap vendor exposure to the long tail of usage.

 

  1. The most important economic metric for AI applications is not gross margin in isolation. It is value created per inference dollar, combined with pricing power and retention. High-value labor substitution can support excellent economics even at gross margins below historic SaaS norms.

 

  1. “Would the product be useful without AI?” is a weak viability test. Many genuinely valuable AI-native products would not exist without AI. The stronger test is whether the workflow creates durable customer value after fully loaded inference, infrastructure, support and deployment costs.

 

  1. AI-native software will likely split into multiple economic classes rather than converge on one margin benchmark: low-margin intelligence resellers, high-margin workflow software, consumption platforms, outcome-priced automation and integrated model/infrastructure businesses.

 

Table of Contents

 

  1. The SaaS Benchmark — What Actually Created the 80% Gross-Margin Ideal
  2. AI Changes COGS, Not the Definition of Software
  3. The Margin Evidence: What We Know and What We Do Not
  4. Why the 30–60% “AI Margin” Rule Is Too Crude
  5. The Unlimited-Usage Trap
  6. Agentic AI Makes Usage Variance a First-Class Financial Problem
  7. Model-Cost Deflation: The Most Important Counterforce
  8. Why Cheaper Intelligence May Not Improve Margins
  9. Pricing Architecture Is Now Part of Systems Architecture
  10. Consumer Freemium: Why Conversion Alone Does Not Diagnose PMF
  11. The Wrapper Question — When Is an AI Product Just Compute Resale?
  12. Vertical AI and the Economics of Labor Substitution
  13. The New AI Software Economic Taxonomy
  14. Valuation: When Does AI Software Deserve a SaaS Multiple?
  15. Risks to Founders and Investors
  16. Industry-Level Effects
  17. Solten & Co. Framework: Value per Inference Dollar
  18. Thesis, Counter-Thesis and Falsification
  19. What to Watch Next
  20. Sources & Evidence

 

Scope & Methodology

 

Research question. This report examines how generative and agentic AI alter the gross-margin structure, pricing design, operating leverage, and valuation logic of software businesses.

 

Scope. The report compares public SaaS and cloud-software benchmarks with reported frontier-model economics, current model API pricing, production-scale agent workload data, and emerging software pricing models. It is not an accounting standard or a forecast for every AI company.

 

Evidence hierarchy. Priority is given to SEC filings, company pricing pages, company disclosures, and academic or production-scale technical research. Reuters and other high-quality financial reporting are used where private-company financial information is not publicly filed. Investor frameworks are treated as estimates and analytical lenses, not audited facts.

 

Evidence classes. DISCLOSED FACT refers to primary or authoritative evidence. REPORTED TERM refers to credible reporting not independently disclosed in full. ESTIMATE refers to calculations or third-party analytical estimates. SOLTEN & CO. INTERPRETATION refers to our synthesis.

 

Critical limitation. Gross margins across the companies cited are not perfectly comparable. Salesforce and Datadog report consolidated gross margin; ServiceNow separately reports subscription gross margin; Snowflake reports product gross margin; OpenAI and Anthropic figures are private-company estimates or reported adjusted measures. The comparison is therefore directional and structural, not an accounting league table.

 

What Changed

 

The popular 2023–2025 framing of AI economics focused on the obvious problem: every generation costs money. By 2026, the more important issue is no longer simply token cost. Three things have changed.

 

First, inference has become substantially cheaper per unit of capability. Model providers now actively segment price-performance tiers and push customers toward smaller models, caching, and batch processing. Raw model cost is becoming an optimization surface, not a fixed tax.

 

Second, software is becoming agentic. A single user instruction may trigger long contexts, repeated reasoning, tool execution, file reads, searches, retries, and subagents. The economic unit is moving from “request” toward “completed workflow.”

 

Third, pricing is adapting. GitHub’s current Copilot structure explicitly combines fixed monthly plans with AI credits whose consumption varies by model and task complexity. That is a structural break from the unlimited-seat intuition inherited from SaaS.

 

The emerging question is therefore not whether AI margins are low today. It is who can convert rapidly falling intelligence costs into durable gross profit before competition passes those savings to customers or expanding usage consumes them.

 

1. The SaaS Benchmark — What Actually Created the 80% Gross-Margin Ideal

 

The mythology of SaaS begins with a true observation: once software has been built, the marginal cost of delivering another copy is small relative to the selling price. Cloud delivery introduced hosting, support, and third-party infrastructure costs, but the economics remained attractive because the vendor could serve many customers from a common code base.

 

This is visible in current public-company filings. Salesforce reported fiscal-2026 revenue of $41.5 billion and gross profit of $32.3 billion, implying total gross margin of about 77.7%. Its subscription and support business was even stronger: $39.4 billion of revenue against $6.8 billion of related cost, or roughly 82.7% gross margin.

 

ServiceNow reported 80% subscription gross margin for full-year 2025. Datadog reported 80% consolidated gross margin. These economics create powerful operating leverage because incremental revenue can finance sales, R&D, and administration while leaving a substantial contribution pool.

 

But the important historical detail is that “software” has never guaranteed 80% margins. Snowflake’s fiscal-2026 product gross margin was 72%. The company’s cost of product revenue rose by $268 million, driven primarily by $248 million of additional third-party cloud infrastructure expense, including costs related to AI inference. Consumption-heavy software already carried a more visible infrastructure burden.

 

The real SaaS advantage is therefore not zero marginal cost. It is predictable marginal cost that is low relative to price.

 

That distinction becomes central in AI.

 

Exhibit 1 — Software Gross Margins Are Becoming More Workload-Sensitive

2. AI Changes COGS, Not the Definition of Software

 

When a SaaS customer uses a CRM dashboard ten times instead of five, the vendor rarely experiences a proportionate increase in direct cost. In AI-native software, use can be directly connected to compute.

 

The relevant cost stack may include model API charges, owned GPU depreciation, cloud inference, retrieval infrastructure, embeddings, vector databases, tool execution, external search, code sandboxes, moderation, storage, observability, and human fallback. Some of these costs are small; some can dominate.

 

More importantly, the cost is not merely per user. It depends on what the user does.

 

A short classification task sent to a small model may cost fractions of a cent. A long-context reasoning task using a frontier model can cost orders of magnitude more. Video generation, voice, and image workloads add different cost curves. Agents can call models repeatedly before returning one visible answer.

 

The result is a financial system in which engagement can improve retention while simultaneously compressing gross margin.

 

This is the opposite of the classic SaaS instinct that “more product usage is always good.” For AI, more usage is good only when the economics of that usage are controlled.

 

3. The Margin Evidence: What We Know and What We Do Not

 

The public evidence supports a meaningful margin gap between mature SaaS and frontier-model providers, but not a universal 30–60% rule for AI applications.

 

Reuters reported in February 2026 that OpenAI’s adjusted gross margin fell to 33% in 2025 from 40% in 2024 after inference costs increased fourfold. OpenAI’s planned compute spending through 2030 is extraordinarily large, illustrating that frontier-model economics remain infrastructure intensive.

 

For Anthropic, Reuters Breakingviews cited a PitchBook estimate of approximately 44% gross margin in August 2026. Anthropic’s revenue run rate had exceeded $65 billion by the end of July, but that growth continues to require massive infrastructure expansion. The margin estimate is useful but should remain explicitly labeled as a third-party estimate rather than an Anthropic disclosure.

 

Battery Ventures’ State of AI 2025 report provided a broader category framework: it estimated gross margins of roughly 0–30% for some AI application businesses and 30–60% for model inference, compared with 80%+ for traditional SaaS applications. The report also argued that margins should improve as token costs fall, routing improves, and pricing moves toward value or outcomes.

 

The important analytical correction is that these are different layers of the stack.

 

A frontier-model company owning training and inference infrastructure is not economically equivalent to an AI application buying a low-cost API. An AI coding agent with long autonomous sessions is not equivalent to a document-classification product. An image generator is not equivalent to compliance software that runs a model once per document and charges for a high-value regulated workflow.

 

There is no single “GenAI gross margin.”

 

4. Why the 30–60% “AI Margin” Rule Is Too Crude

 

Gross margin is an outcome of at least six design choices: workload intensity, model choice, price architecture, customer behavior, infrastructure ownership, and workflow value.

 

Consider two companies with identical $30 monthly subscriptions.

 

Company A performs a few lightweight classifications and spends $2 per active user on inference. Company B provides an agent that reads a repository, reasons over files, executes tools, and retries failures, generating $18 of monthly model and infrastructure cost for a heavy user. They are both “AI software.” Their economics are fundamentally different.

 

Exhibit 2 — $30 Subscription: How Usage-Driven Inference Cost Changes Gross Margin

The illustrative sensitivity model in Exhibit 2 assumes $3 per user of non-inference COGS. At $2 of inference cost, gross margin is about 83%. At $8, it falls to about 63%. At $20, gross margin is roughly 23%.

 

The difference is not category. It is unit economics.

 

This matters because founders and investors can misdiagnose a product by applying a category average. A 45% gross-margin business may be temporarily inefficient and rapidly improving. An 80% gross-margin product may achieve that number only because usage is weak. The direction of margin must be connected to usage, retention and customer value.

 

5. The Unlimited-Usage Trap

 

Fixed subscriptions were ideal for SaaS because the vendor could sell predictability without exposing itself to extreme marginal cost. AI weakens that assumption.

 

The risk is a heavy-user tail. A customer paying $20 or $30 per month can consume far more than that in frontier-model compute if the product allows long contexts, premium models or autonomous agent execution without limits.

 

This is why historical anecdotes such as the widely repeated estimate that ChatGPT once cost hundreds of thousands of dollars per day to operate are less useful than current pricing behavior. Whatever the precise 2023 number, the strategic response across the industry is visible: quotas, credits, model multipliers, rate limits and premium tiers.

 

The market is learning to price variance.

 

GitHub Copilot’s current individual plans show this clearly. The $10 Pro plan includes a defined monthly pool of AI credits; Pro+, Max and enterprise tiers provide progressively larger allowances. GitHub explicitly states that chat, agent mode, code review, cloud agents and CLI interactions consume credits, with cost depending on the model and complexity of the task. Code completion remains unlimited because its cost profile is different.

 

That is not a cosmetic billing change. It is a decomposition of the product into economically distinct workloads.

 

6. Agentic AI Makes Usage Variance a First-Class Financial Problem

 

The economics become more difficult as AI moves from assistant to agent.

 

Microsoft Azure Research’s 2026 production study of GitHub Copilot coding agents is unusually useful because it describes real workload behavior rather than synthetic benchmarks. The dataset covered roughly 13 million sessions from more than 3.2 million users, with 761 million LLM calls, 95 trillion tokens and 775 million tool invocations.

 

That is roughly 58 LLM calls and 60 tool invocations per sampled session on average. The study also found a roughly 50× token range across five user archetypes. Tool failures occurred in about 9% of turns and could drive approximately four times the compute through retry loops. Model switching severely reduced cache reuse.

 

Exhibit 3 — Agentic Software Turns One User Task Into Many Compute Events

This changes product finance in three ways.

 

First, users are not equal-cost seats. A fixed $30 agent license may contain users whose workload differs by an order of magnitude or more.

 

Second, reliability becomes a margin variable. A failed tool call is not merely a UX problem; it can trigger another model loop and another compute bill.

 

Third, systems engineering directly affects gross margin. Cache retention, model routing, context compaction, scheduler design and tool orchestration move from backend optimization to financial strategy.

 

The future CFO of an AI software company will need to understand inference architecture in a way the CFO of a classic SaaS company usually did not.

 

7. Model-Cost Deflation: The Most Important Counterforce

 

The bearish margin story becomes incomplete if it ignores the falling cost of intelligence.

 

Stanford’s 2025 AI Index estimated that the cost of querying a model with GPT-3.5-equivalent MMLU performance fell from $20 per million tokens in November 2022 to $0.07 by October 2024. Depending on workload, the report observed annual inference-price declines ranging from roughly 9× to 900×.

 

The trend has continued through product segmentation and efficiency. OpenAI’s July 30, 2026 update cut GPT-5.6 Luna’s API price by 80% and Terra’s by 20%. Current pricing spans a huge range: GPT-5.6 Luna at $0.20 per million input tokens and $1.20 output; Terra at $2 and $12; Sol at $5 and $30. Anthropic’s current pricing similarly ranges across model tiers, with Sonnet 5 offered at $2 input / $10 output per million tokens through August 2026 before moving to standard pricing.

 

Exhibit 4 — Inference Cost Deflation Is Real — But It Does Not Guarantee Margin Expansion

This creates a path to margin expansion unavailable in many physical businesses. If a workflow generates constant customer value while its inference cost falls 80%, gross profit can expand rapidly without raising price.

 

But there are two catches.

 

The first is competition. If every vendor gains access to cheaper models, customers may capture the savings through lower prices.

 

The second is induced demand. When intelligence becomes cheaper, products may use much more of it.

 

8. Why Cheaper Intelligence May Not Improve Margins

 

The natural assumption is that cheaper tokens mean higher gross margins. That is only true if usage does not expand faster than unit cost falls.

 

AI may exhibit a form of Jevons effect: efficiency lowers the cost of intelligence, which encourages products to consume more intelligence. A support assistant becomes an autonomous resolution agent. A coding autocomplete becomes a coding agent. A document summarizer becomes a full diligence workflow. A chatbot becomes a multi-agent research system.

 

The economic unit therefore changes.

 

Suppose model cost per token falls by 70%, but the product moves from one model call per task to ten calls because autonomy improves customer value. Total inference expense per completed workflow can still rise.

 

This means investors should track cost per customer outcome, not cost per token alone.

 

The relevant metric is not “How cheap is the model?” but “How much intelligence must the product buy to produce one unit of billable value?”

 

9. Pricing Architecture Is Now Part of Systems Architecture

 

Traditional SaaS pricing primarily captured willingness to pay. AI pricing must also manage cost variance.

 

Four structures are emerging.

 

Fixed subscription. The vendor accepts usage risk. This can work when tasks are low-cost, predictable or heavily cached, but it is dangerous when the usage distribution is long-tailed.

 

Subscription plus credits. The customer buys predictable access while the vendor caps extreme consumption. GitHub Copilot is a current example.

 

Pure usage-based pricing. Cost risk passes more directly to the customer. This is natural for APIs and infrastructure but can weaken budget predictability and make software feel like a utility.

 

Seat plus usage. Enterprise customers pay a platform or governance fee plus metered intelligence consumption. This may become the dominant compromise for high-value agentic software.

 

Outcome-based pricing. The vendor charges for completed work, savings or business results rather than tokens. This can create exceptional economics when customer value is large relative to compute, but it shifts execution and measurement risk to the vendor.

 

Exhibit 5 — AI Pricing Is Becoming a Risk-Allocation Mechanism

This is one reason “SaaS versus AI” is the wrong framing. AI is forcing software companies to rediscover pricing as a form of risk management.

 

10. Consumer Freemium: Why Conversion Alone Does Not Diagnose PMF

 

A common argument is that if only a small percentage of hundreds of millions of free AI users convert to paid plans, the business has weak product-market fit. That conclusion is analytically unsound without more information.

 

Freemium businesses deliberately maximize free reach. Low percentage conversion can coexist with tens of millions of paying customers, strong retention, and enormous absolute revenue. Free users may also create distribution, brand, data, or enterprise lead generation.

 

The right questions are cohort-specific: What percentage of high-intent users convert? What is paid-user retention? What is ARPU? What does a free user cost? How much free usage is subsidized by enterprise revenue or strategic distribution? Does the free tier improve acquisition efficiency?

 

For AI, one additional metric becomes essential: contribution margin by user cohort.

 

A low-engagement free user may cost almost nothing. A sophisticated unpaid user running long agent loops may be expensive. “Free MAU” is therefore not one economic category.

 

11. The Wrapper Question — When Is an AI Product Just Compute Resale?

 

“GPT wrapper” is often used as an insult rather than an analytical category. The relevant issue is not whether a company calls an external model. Many valuable software businesses depend on infrastructure they do not own.

 

The economic question is whether the product adds a defensible layer between raw model output and customer value.

 

A weak wrapper has little proprietary workflow, low switching cost, minimal data advantage and prices close to model consumption. Its gross margin is exposed both to upstream model pricing and downstream competition. If the model provider adds the feature directly, the application can disappear.

 

A strong AI application may also use third-party models, but it owns workflow integration, proprietary data context, distribution, compliance, user experience, evaluation systems, business process orchestration or outcome accountability. The model is an input, not the product.

 

The distinction is value capture.

 

If a company buys $1 of model inference and sells $1.30 of undifferentiated output, it is a low-margin reseller. If it buys $10 of inference to automate $500 of legal, accounting or engineering work, the model cost can be economically trivial even if accounting gross margin is only 60% during early scaling.

 

12. Vertical AI and the Economics of Labor Substitution

 

This is why vertical AI may produce some of the strongest businesses despite relatively heavy compute.

 

A legal drafting agent, insurance claims system, healthcare coding workflow, accounting close product or cybersecurity investigation agent can be priced against labor, delay, error rates and regulatory risk rather than against a generic software seat.

 

The economic denominator changes from “software budget” to “cost of work.”

 

A traditional SaaS product might charge $50 per user to make an employee 10% more productive. An AI agent that performs 40% of the employee’s workflow can potentially charge hundreds or thousands of dollars while consuming tens of dollars of compute.

 

This does not guarantee a good business. Vertical products may require implementation, human review, compliance, support and domain-specific data work. But the higher value pool gives the vendor room to absorb inference cost.

 

The strongest AI companies may therefore look less like software licenses and more like digitally delivered labor with software-like scalability.

 

13. The New AI Software Economic Taxonomy

 

Solten & Co. sees at least five emerging economic classes.

 

  1. Intelligence resellers. Products with limited differentiation that mainly package third-party models. Their margins are vulnerable to model commoditization and feature bundling.

 

  1. Workflow software with AI COGS. Products where AI is one variable cost inside a sticky workflow. These may converge toward traditional SaaS-like margins as models become cheaper.

 

  1. Agentic software. Products that perform multi-step work and carry high usage variance. Their economics require credits, usage controls or outcome pricing.

 

  1. AI infrastructure and platforms. Model APIs, inference clouds and data platforms. These are structurally more capital and compute intensive, with margins closer to cloud infrastructure than classic application SaaS.

 

  1. Outcome-priced digital labor. Products priced against completed work or avoided labor cost. These may have lower accounting gross margins than SaaS but much stronger absolute gross profit per customer.

 

The mistake is valuing all five as one category.

 

14. Valuation: When Does AI Software Deserve a SaaS Multiple?

 

High SaaS multiples historically reflected a specific combination: recurring revenue, high gross margin, strong retention, low capital intensity and predictable operating leverage.

 

AI software should deserve comparable valuation only when it demonstrates comparable economic durability, not because revenue is called ARR.

 

Investors should ask:

 

  • Is revenue recurring because customers are contractually committed, or because consumption happened to be high this quarter?
  • Does higher usage expand gross profit or compress it?
  • Is inference cost falling faster than price?
  • Can the company route across models without materially degrading quality?
  • Does the product have enough workflow value to pass compute cost through to customers?
  • Are gross margins expanding with scale?
  • Is customer retention strong after usage limits and pricing changes?
  • How much capital must be invested in infrastructure to support growth?
  • Is the company exposed to a model provider that can vertically integrate into the application?

 

A company with 50% gross margin, 150% net revenue retention and rapidly falling cost-to-serve may deserve more confidence than a company reporting 80% gross margin because customers barely use the product.

 

AI valuation should therefore incorporate margin trajectory and usage economics, not static margin alone.

 

15. Risks to Founders and Investors

 

Pricing risk. A fixed plan can become uneconomic as users discover higher-value or more compute-intensive workflows.

 

Provider risk. Upstream model vendors can change API prices, rate limits, model availability or product scope.

 

Commoditization risk. Falling model prices may reduce costs but also lower barriers to entry and compress application pricing.

 

Usage-growth risk. Greater engagement can raise COGS faster than revenue when pricing is not consumption-aware.

 

Quality-cost risk. Customers may prefer premium models, forcing the vendor to choose between margin and perceived product quality.

 

Agent-loop risk. Tool failures, retries and longer reasoning can create invisible compute leakage.

 

Infrastructure lock-in. Moving between models or clouds can destroy caching, require re-evaluation and create technical switching costs.

 

Capital risk. Companies that own model or inference infrastructure can require extraordinary financing before reaching positive cash flow.

 

Measurement risk. AI companies can present ARR or bookings without disclosing the cost required to produce that revenue. Revenue growth alone can hide deteriorating contribution economics.

 

16. Industry-Level Effects

 

The margin problem will reshape the software stack beyond individual companies.

 

First, model routing becomes financially valuable. Software that automatically chooses the cheapest model capable of completing a task can capture gross-margin expansion without visible product change.

 

Second, caching and context management become economic moats. The ability to reuse context or avoid repeated inference can create meaningful cost advantage.

 

Third, pricing complexity will rise. Buyers will face more credits, consumption units, outcome fees and blended contracts, making procurement and cost forecasting harder.

 

Fourth, AI will blur software and services. Outcome-priced agents may compete directly with labor budgets, business-process outsourcing and professional services rather than only software vendors.

 

Fifth, infrastructure ownership will become a strategic choice. At sufficient scale, companies may move from public APIs to reserved capacity, custom models or owned inference in order to improve margins.

 

Sixth, the gross-margin premium historically attached to software may fragment. Public markets may eventually distinguish high-margin workflow AI from infrastructure-heavy intelligence businesses in the same way they distinguish application SaaS from cloud data platforms today.

 

17. Solten & Co. Framework: Value per Inference Dollar

 

We propose a simple primary lens for AI application economics:

 

Value per Inference Dollar = Customer Economic Value Created / Fully Loaded AI Delivery Cost

 

The numerator should include measurable value: labor displaced, time saved, revenue generated, errors avoided, risk reduced or cycle time compressed.

 

The denominator should include more than API tokens: model calls, retrieval, embeddings, tool execution, hosting, human review, retries, support and any infrastructure directly attributable to the AI workflow.

 

This ratio should then be considered alongside three additional dimensions.

 

Pricing capture. What percentage of created value can the vendor monetize?

 

Usage variance. How widely does cost differ across customers performing ostensibly the same product action?

 

Cost trajectory. Is the fully loaded cost per successful workflow declining over time?

 

An attractive AI business has high customer value, meaningful pricing capture, controllable usage variance and a declining cost curve.

 

This framework is more useful than asking whether a product could exist without AI.

 

18. Thesis, Counter-Thesis and Falsification

 

Solten & Co. Thesis

 

Generative AI will not permanently destroy software gross margins. It will split software into new economic categories and force vendors to price intelligence explicitly. The strongest AI applications will recover high gross margins or high absolute gross profit by combining falling inference costs with workflow pricing power, model routing and operational efficiency. The weak businesses will be exposed as low-value compute resellers.

 

Counter-Thesis

 

AI may structurally reduce application-software margins because capability competition continually increases required compute. Every efficiency gain can be reinvested into longer context, stronger models, more agent steps and richer modalities. If users expect continuous improvement while subscription prices remain anchored to historic SaaS levels, the industry may settle at permanently lower gross margins and require lower valuation multiples.

 

Falsification Criteria

 

The Solten & Co. thesis would weaken if several of the following occur:

 

  • model prices stop declining materially despite continued hardware and algorithmic progress;
  • leading AI applications fail to expand gross margins after multiple years of scale;
  • usage-based and credit pricing produce significant customer resistance and churn;
  • outcome-priced AI cannot sustain pricing because model capabilities commoditize too quickly;
  • agentic workloads consume increasing compute without proportionate customer willingness to pay;
  • model providers capture most application value through vertical integration;
  • public markets consistently assign AI software materially lower valuation multiples even to businesses with strong retention and improving margins.

 

19. What to Watch Next

 

For investors and operators, the next phase should be monitored through concrete unit-economic signals rather than user-count headlines.

 

Gross margin by product or workload. Consolidated company margin can hide whether premium AI features are profitable.

 

Inference cost as a percentage of revenue. This reveals whether model cost is scaling efficiently.

 

Cost per successful workflow. Especially important for agents, where failures and retries can amplify consumption.

 

Model mix. What percentage of workloads require frontier models versus cheaper specialized models?

 

Cache hit rate and routing efficiency. These are increasingly financial metrics.

 

Paid usage architecture. Watch movement from unlimited subscriptions toward credits, metering and hybrid pricing.

 

Heavy-user economics. Averages can hide a loss-making tail.

 

Outcome pricing adoption. Evidence that vendors can price against labor or business results would support margin expansion.

 

Gross-margin trajectory of Anthropic and OpenAI. As frontier providers mature toward public-market scrutiny, these figures will become important benchmarks for the model layer.

 

AI gross-margin disclosures from public SaaS companies. ServiceNow, Snowflake, Salesforce, Cloudflare and others will increasingly reveal whether AI features improve or dilute established software economics.

 

20. Sources & Evidence

 

Primary/regulatory sources

 

Salesforce — FY2026 Form 10-K, subscription/support revenue and cost of revenue

https://www.sec.gov/Archives/edgar/data/1108524/000110852426000060/crm-20260131.htm

 

ServiceNow — FY2025 Form 10-K / earnings materials, subscription gross margin

https://www.sec.gov/Archives/edgar/data/1373715/000137371526000007/now-20251231.htm

 

Datadog — FY2025 Form 10-K, gross margin and cloud infrastructure costs

https://www.sec.gov/Archives/edgar/data/1561550/000162828026008819/ddog-20251231.htm

 

Snowflake — FY2026 Form 10-K, product gross margin and AI-inference infrastructure costs

https://www.sec.gov/Archives/edgar/data/1640147/000164014726000008/snow-20260131.htm

 

OpenAI — GPT-5.6 pricing and July 30, 2026 price-performance update

https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/

https://openai.com/api/pricing/

 

Anthropic — Claude list prices, effective May 27, 2026

https://www-cdn.anthropic.com/files/4zrzovbb/website/3684c2faafb97418665782cea0001f439f74b1d2.pdf

 

Anthropic — Claude Sonnet 5 pricing

https://www.anthropic.com/research/claude-sonnet-5

 

GitHub — Copilot plans and AI Credits pricing

https://github.com/features/copilot/plans

 

Microsoft Azure Research — Agentic Coding in the Wild: Characterizing GitHub Copilot at Production Scale, 2026

https://www.microsoft.com/en-us/research/wp-content/uploads/2026/08/ghcp_traces-6.pdf

 

Stanford HAI — AI Index Report 2025, inference cost trends

https://hai.stanford.edu/assets/files/hai_ai-index-report-2025_chapter1_final.pdf

 

Investor / analytical frameworks

 

Battery Ventures — State of AI Report 2025

https://www.battery.com/wp-content/uploads/2026/01/Battery-State-of-AI-Report-2025.pdf

 

High-quality financial reporting

 

Reuters — OpenAI compute spending and reported adjusted gross margin, Feb. 20, 2026

https://www.reuters.com/technology/openai-sees-compute-spend-around-600-billion-by-2030-cnbc-reports-2026-02-20/

 

Reuters Breakingviews — OpenAI margin comparison with Salesforce, Oct. 15, 2025

https://www.reuters.com/commentary/breakingviews/how-infer-method-to-openais-madness-2025-10-15/

 

Reuters Breakingviews — Anthropic gross-margin estimate, Aug. 18, 2026

https://www.reuters.com/commentary/breakingviews/anthropics-deceleration-has-some-ipo-upside-2026-08-18/

 

Reuters — Anthropic revenue run rate tops $65B, Aug. 17, 2026

https://www.reuters.com/technology/anthropic-revenue-run-rate-tops-65-billion-source-says-2026-08-17/

 

Methodological Note

 

This report deliberately rejects three shortcuts common in AI-business commentary.

 

First, it does not treat “AI gross margin” as one universal range. Economics differ by layer and workload.

 

Second, it does not treat free-to-paid conversion as a standalone measure of product-market fit.

 

Third, it does not use old anecdotal compute-cost estimates as evidence of current economics where better contemporary data exist.

 

Comparative gross-margin figures should be interpreted directionally because accounting definitions and business mix differ. Private-company figures for OpenAI and Anthropic are reported or estimated, not audited public filings.

 

About Solten & Co.

 

Solten & Co. is an independent research and analysis firm focused on the AI economy, with deeper research emphasis on AI infrastructure, software AI, Physical AI, robotics, and autonomous systems. We study the technologies, companies, markets, transactions, and capital structures shaping the next phase of AI.

 

Need an independent perspective?

 

Solten & Co. provides independent research and analytical support for investors and decision-makers evaluating companies, markets, and investment opportunities across the AI economy.

 

Discuss a research question · Discuss an investment opportunity · Request independent analysis

 

soltenco.com