For every $100 in revenue generated by AI model companies, approximately $35-40 flows to the three major cloud providers—AWS, Azure, and GCP—in the form of inference compute fees. From this revenue stream, cloud vendors extract $10-20 in operating profit, translating to an operating margin of roughly 35%-45%.
This is the central finding from Barclays' unit economics research report on the AI industry, published on August 28.
Meanwhile, paid inference margins at AI labs have surged from the mid-teens percentage range in 2025 to 50%-65% or higher in 2026, with adjusted gross margins improving by 30-50 percentage points year-over-year. The momentum is driven by enterprise-grade clients and agentic workflows becoming "must-buy" products in the marketplace.
Barclays analyst Ross Sandler believes current actual margins may exceed reported estimates, though intensifying frontier competition and expanding compute supply are expected to drive gradual normalization over time.
One Industry, Two Financial Profiles
Barclays constructed two hypothetical frontier lab models to dissect the margin differentials. "Lab A" derives approximately 70% of revenue from API and 30% from subscriptions, while "Lab B" is the inverse—80% subscriptions and merely 20% API.
API inherently carries higher inference margins than subscriptions. Combined with differences in training cost allocation and partner revenue-sharing arrangements, the two models exhibit a 17-percentage-point gap in adjusted gross margins—Lab A sits at approximately 55%, while Lab B trails at around 38%.
Revenue recognition methodologies further amplify this distortion. Lab A recognizes indirect API revenue on a gross basis, whereas Lab B applies net basis treatment, sometimes declining to recognize indirect API revenue from strategic partner operations altogether. Barclays draws a parallel to Uber and Lyft: the same core business can produce starkly divergent reported figures purely due to accounting treatment. As AI labs begin disclosing GAAP financials, investors conducting cross-company comparisons must dismantle these accounting discrepancies.
Inference Margins Soaring
Breaking down by product line—
Subscription offerings (such as Claude Code and Codex) carry estimated inference margins around 70%, ironically the lowest among the three product categories. The rationale: AI labs are willing to subsidize token costs to retain users. Subscriptions typically operate on monthly fees with usage caps, and the increased frequency of cap resets recently points to a combination of retention pressures and model efficiency improvements.
Direct API represents the earliest business model for AI labs and delivers the richest margins. Developers at Cursor, Figma, and similar platforms pay based on token consumption. Barclays estimates current API inference margins have surpassed 80%. Improved token efficiency (fewer tokens required to complete identical tasks), nominal API price increases, and infrastructure-level inference optimizations—quantization, speculative decoding techniques, and next-generation compute—all continue unlocking margin expansion. Barclays notes that API inference margins in Q2 2026 stand considerably above charted levels, though a downward regression is anticipated at some future point.
Indirect API delivers an end-user experience identical to direct API, but the billing relationship exists between users and cloud providers. As indirect API's share of revenue grows, divergent revenue recognition approaches across labs will further widen discrepancies in reported financial comparability.
What Cloud Providers Earn
For every $100 of AI lab revenue, Lab A corresponds to $35 in cloud vendor revenue, contributing approximately $11.8 in profit after infrastructure costs—an operating margin near 34%. Lab B, owing to strategic partner revenue-sharing (20% of revenue with cumulative caps), hands cloud providers more—$41 in revenue yielding $19.1 in profit, an operating margin of 47%.
Barclays highlights that revenue-sharing inflates cloud providers' apparent margins; stripping out the sharing arrangement, per-token actual profit remains identical. The sharing arrangement is expected to phase out entirely post-2028.
Agentic subscription products generate additional value for cloud vendors. These stateful runtime offerings often require invoking higher-level software resources like databases, delivering greater value per unit of revenue. In certain scenarios, revenue-sharing arrangements exist between cloud providers and AI labs.
Transition from Training-Led to Inference-Led
Barclays estimates total AI lab revenue will climb from $7 billion in 2024 to $137 billion by 2026, reaching $690 billion by 2028. Year-end ARR figures paint a more aggressive trajectory: approximately $200 billion by end-2026 and $782 billion by end-2028.
Currently, training expenditures still absorb roughly 48% of AI lab revenue, meaning nearly every dollar of lab revenue corresponds to nearly a dollar of cloud vendor income. However, training cost share is declining rapidly—from 96% in 2024 to an expected 35% by 2027 and 30% by 2028. As inference profits progressively outpace training expenditures, AI lab overall profitability will continue improving.
Cloud vendors' AI revenue as a proportion of AI lab revenue is also shrinking: from 153% in 2024 to 90% by 2026, with a projected 73% by 2028. Barclays expects AWS, Azure, and GCP to maintain their respective shares of AI lab compute spending over the next two years, but starting in 2028, secured AI infrastructure projects will come online and become the preferred choice for AI labs—potentially causing the three major cloud providers to gradually cede share in both training and inference domains.
Comments