AI token costs surge 40% as IPO-bound firms face margin squeeze

Abstract illustration depicting rising AI token costs through ascending geometric shapes and data visualisation elements

AI companies preparing for public offerings face mounting pressure as token pricing—the cost of processing AI requests—has surged approximately 40% over the past six months, according to TechCrunch AI, threatening the unit economics that underpin their valuations.

The cost inflation arrives at a particularly inopportune moment, with several major AI firms including Anthropic and Perplexity reportedly preparing IPO roadshows for late 2026. The pricing pressure stems from increased demand for high-quality inference on frontier models, combined with constrained GPU capacity and rising energy costs at data centres.

“We’re seeing a fundamental tension between scale and profitability,” one unnamed AI executive told TechCrunch AI. “The cost per token is climbing faster than our efficiency improvements can offset.”

The economics are stark. Whilst model providers have achieved roughly 2x efficiency gains through optimisation techniques over the past year, token pricing has increased by 40-50% for premium inference services. This creates a net cost increase of approximately 20-30% for companies that cannot pass expenses directly to customers—a category that includes most consumer-facing AI applications operating on freemium models.

The business impact divides along clear lines. Infrastructure providers—hyperscalers like Microsoft Azure, Google Cloud, and AWS—stand to benefit from sustained high pricing and increased compute demand. GPU manufacturers, particularly Nvidia, maintain pricing power in a supply-constrained market.

Conversely, application-layer companies face margin compression. Consumer AI services that attracted users with free or low-cost tiers now confront unsustainable unit economics. Enterprise-focused firms possess more flexibility to implement price increases, but risk customer resistance in a competitive market where switching costs remain relatively low.

The IPO implications are substantial. Public market investors will scrutinise gross margins and path to profitability with greater intensity than private backers. Companies that raised at elevated valuations during 2024-2025 must now demonstrate that revenue growth can outpace cost inflation—a challenging proposition when your primary input cost is rising 40% semi-annually.

Historical precedent offers cautionary tales. Cloud computing companies in the early 2010s faced similar cost structure challenges, with many failing to achieve profitability despite strong revenue growth. The difference: cloud infrastructure costs declined steadily over time, whilst AI inference costs show no clear downward trajectory in the near term.

Some firms are responding with architectural changes. TechCrunch AI reports increased adoption of model distillation—using smaller, cheaper models for routine tasks whilst reserving expensive frontier models for complex queries. Others are exploring on-premise deployment options for enterprise customers, shifting infrastructure costs to the client.

The pricing pressure also accelerates consolidation dynamics. Well-capitalised firms can absorb cost increases whilst smaller competitors struggle, potentially leading to acquisition opportunities. Microsoft’s deep integration with OpenAI, for instance, provides cost advantages that independent application developers cannot match.

Energy costs compound the challenge. Data centres running AI workloads consume substantially more power than traditional computing infrastructure, and electricity prices in key markets including Northern Virginia and Dublin have increased 15-20% over the past year. These costs flow through to token pricing with minimal buffering.

Market observers should monitor several indicators in coming months. First, watch for pricing announcements from major model providers—any stabilisation or reduction would significantly improve application-layer economics. Second, track IPO prospectus disclosures around gross margins and customer acquisition costs relative to lifetime value.

Third, observe enterprise contract structures. If large customers negotiate fixed-price agreements, application providers absorb cost volatility; if contracts include cost pass-through provisions, customers bear the risk but may reduce consumption.

The token pricing surge represents a maturation point for the AI industry, forcing a reckoning between growth narratives and profitability realities. Companies approaching public markets must demonstrate sustainable unit economics, not merely revenue growth—a considerably higher bar in an environment of cost inflation.