
On the surface, artificial intelligence (AI) model companies appear to be earning enormous profits. In reality, the cloud computing providers behind them may be the ones enjoying the greater benefits.

Barclays' latest research breaks down profit distribution across the AI industry. It found that for every $100 in revenue generated by model companies (note: approximately 674.6 yuan at the current exchange rate), $35 to $40 (approximately 236.1 to 269.8 yuan at the current exchange rate) flows to the three major cloud platforms, Amazon AWS, Microsoft Azure, and Google GCP, in the form of inference computing fees.
Cloud service providers can earn $10 to $20 in operating profit from this revenue (approximately 67.5 to 134.9 yuan at the current exchange rate), corresponding to operating margins of up to 35% to 45%. This conclusion comes from Barclays' research report on AI industry unit economics, released on August 28, and provides a clearer view of where value is flowing throughout the AI industry chain.
The report also shows that paid inference margins at AI labs have risen sharply, from just above 10% in 2025 to 50% to 65% or more in 2026. Adjusted gross margins increased by 30 to 50 percentage points year over year.
Barclays analyst Ross Sandler said the main factors driving the sharp improvement in margins are enterprise customers and so-called “agentic workflows.” These products have become “must-buy” offerings in the market.
He further believes that AI labs' actual margins may already be higher than the report's estimates. However, as competition among frontier models intensifies and computing capacity continues to expand, this metric is expected to gradually decline eventually.
To help investors better understand the financial differences among AI labs, Barclays constructed two hypothetical frontier AI labs for comparison.
About 70% of “Lab A's” revenue comes from APIs and 30% from subscriptions. “Lab B” has the opposite mix: 80% of its revenue depends on subscriptions, while APIs account for only 20%.
Because API businesses naturally have higher inference margins than subscription businesses, and because training costs and partner revenue-sharing arrangements are allocated differently, the two hypothetical labs show a striking 17-percentage-point gap in adjusted gross margin: approximately 55% for Lab A versus only 38% for Lab B.
Differences in revenue recognition further amplify this distortion.
Lab A recognizes indirect API revenue on a gross basis, while Lab B uses net-basis recognition and may even exclude indirect API revenue operated by strategic partners entirely. Barclays likens this to the difference between Uber and Lyft: Their core businesses are broadly similar, but their financial statements can show completely different figures because of different accounting treatments.
The report cautions that as AI labs begin reporting financial statements under generally accepted accounting principles (GAAP), investors must first understand these accounting differences before comparing companies. Otherwise, meaningful comparisons will be difficult.
Margins by product line
A further breakdown of individual product lines shows that subscription products, such as Claude Code and Codex, have estimated inference margins of about 70%, the lowest among the three major product categories.
The reason is that AI labs are willing to absorb part of the token costs in exchange for user retention. These subscriptions are typically billed monthly with usage caps. Notably, usage-cap resets have become significantly more frequent recently, which may reflect both the user-retention pressure facing AI labs and ongoing improvements in model efficiency.
Direct APIs were the earliest business model for AI labs and are also their highest-margin business line. Developer tools such as Cursor and Figma charge based on the number of tokens consumed by users. Barclays estimates that API inference margins currently exceed 80%.
As models become more token-efficient, meaning fewer tokens are needed to complete the same task, and as nominal API prices rise, continued infrastructure improvements for inference services, including quantization, speculative decoding, and next-generation computing, are creating greater room for profit.
The report specifically notes that API inference margins in the second quarter of 2026 were actually far above the level shown in the chart, though they are still expected to return to normal levels at some point in the future.
Indirect APIs provide end users with an experience broadly similar to that of direct APIs, but users establish a direct billing relationship with cloud service providers. As indirect APIs continue to account for a larger share of total AI lab revenue, differences in revenue recognition among labs will further widen the comparability gap between their financial statements.
How much are cloud providers really earning?
A closer look at cloud providers' profit structure shows that for every $100 in revenue generated by an AI lab (approximately 674.6 yuan at the current exchange rate), Lab A corresponds to about $35 in cloud service revenue (approximately 236.1 yuan at the current exchange rate). After infrastructure costs, the cloud provider can earn about $11.80 in profit (approximately 79.6 yuan at the current exchange rate), representing an operating margin of about 34%.
Lab B is even more profitable. Because of a strategic partner revenue-sharing mechanism, which covers 20% of revenue and has a cumulative cap, cloud providers can earn about $41 in revenue for every $100 of AI revenue generated by Lab B (approximately 674.6 yuan at the current exchange rate). Profit reaches $19.10 (approximately 128.8 yuan at the current exchange rate), producing an operating margin of up to 47%.
It is important to note that Lab B's higher margin comes from the strategic partner revenue-sharing mechanism, which covers 20% of revenue and has a cumulative cap.
Barclays emphasized that this revenue-sharing mechanism inflates cloud providers' apparent margins. Excluding this factor, the actual per-token profit earned by the cloud providers serving the two labs is the same. This revenue-sharing arrangement is expected to gradually decline to zero after 2028.
Agent subscription products can also create additional value for cloud providers. These runtime products, which have state-memory capabilities, often need to call higher-level software resources such as databases, giving them greater value per unit of revenue. In some cases, cloud providers and AI labs also have revenue-sharing arrangements, further increasing cloud providers' actual returns.
At the industry level, Barclays expects AI lab revenue to grow from $7 billion in 2024 (approximately 47.221 billion yuan at the current exchange rate, or approximately NT$220 billion) to $137 billion in 2026 (approximately 924.192 billion yuan, or approximately NT$4.3 trillion), and reach $690 billion in 2028 (approximately 4.65 trillion yuan, or approximately NT$21.8 trillion).
On a year-end annualized recurring revenue (ARR) basis, growth is even more aggressive: ARR is expected to reach approximately $200 billion by the end of 2026 (approximately 1.35 trillion yuan, or approximately NT$6.3 trillion), and is more likely to reach $782 billion by the end of 2028 (approximately 5.28 trillion yuan, or approximately NT$24.7 trillion).
Training expenses currently still account for about 48% of AI lab revenue. This means that for every $1 in revenue earned by an AI lab (approximately 6.7 yuan at the current exchange rate), there is nearly $1 in corresponding revenue for cloud providers (approximately 6.7 yuan at the current exchange rate).
However, training costs as a share of revenue are falling rapidly, from 96% in 2024 to an expected 35% in 2027 and then to 30% in 2028.
The significance of this trend is that as profits from inference gradually exceed training expenses, the overall profitability of AI labs will continue to improve, and the industry's focus will gradually shift from “training-driven” to “inference-driven.”
At the same time, cloud providers' AI revenue as a share of AI lab revenue is declining, from 153% in 2024 to 90% in 2026, and is expected to fall further to 73% in 2028.
Barclays expects AWS, Azure, and GCP to maintain roughly their current shares of AI lab computing expenditures over the next two years. By 2028, however, the industry structure may reach a turning point.
By then, AI infrastructure projects financed with asset-backed funding are expected to come online and become more attractive options for AI labs. This means the current “Big Three” cloud providers may gradually lose share in both AI training and inference, while the current profit landscape undergoes another reshuffle.
