Rippling's AI Bill Hit 40% of R&D Payroll Before It Built Its Own Spend Tracker
The HR software company watched token costs spiral toward millions in three months, then engineered a gateway and dashboard to cut spending by 63% without reducing usage.
When the CFO Showed the Number
In March, Rippling's executive team sat in a meeting and learned they were burning tokens at a pace that would rival their entire engineering payroll. Chief Financial Officer Adam Swiecicki presented figures that stunned the room: AI token expenditure was tracking toward 40% of the research and development headcount budget. Not 4%. Forty percent. The company was spending as much on inference as it paid in total compensation to four out of every ten engineers.
Month-over-month growth stood at 80%. If the trajectory held, Rippling would spend nearly 90% of its R&D payroll equivalent on tokens within twelve months. Chief Product Officer Matt MacInnis recalls the mood as "incredulous." Management immediately launched what he describes as an urgent audit to map where the money was going and whether the company was seeing returns.
The episode captures a pattern that swept through the software industry in early 2026. Companies handed out API keys to frontier models, engineers defaulted to the newest and most expensive endpoints for every task, and finance teams watched line items balloon. At DailyTechWire, we've tracked similar cost shocks at mid-stage SaaS companies across North America and Asia, though few have disclosed numbers as candidly as Rippling.
The Anatomy of Runaway Spend
Rippling's forensic dive uncovered stark concentration. Roughly 10% to 15% of employees accounted for 60% of total token consumption. One engineer alone was responsible for fifty thousand dollars a month in inference calls. The company used a mix of tools: Cursor, OpenAI, and Anthropic. Employees almost universally defaulted to the latest frontier models, regardless of task complexity.
MacInnis points to misaligned incentives. Inference providers have no structural reason to help customers control spend; their revenue scales with usage. They offer limited visibility into consumption patterns and no interoperability to help enterprises compare or switch models efficiently. The result is a vendor environment that encourages runaway expense rather than optimization.
Rippling did not want to shut off AI access. The productivity gains were real, particularly in engineering workflows. The challenge was to preserve utility while eliminating waste. The company started by negotiating hard spending caps with each provider. Then it built two pieces of infrastructure: an AI gateway to route prompts to the most cost-effective model for each task, and a dashboard to track employee-level spend against measurable output.
Model Selection as Cost Architecture
By mid-2026, enterprises had learned a fundamental lesson: relying on a single model family or vendor is both a performance and a financial liability. Rippling ran its own internal benchmarks and discovered that SpaceX's Grok led on general tasks, but that Z.ai's GLM 5.2 delivered nearly identical performance at 85% lower cost. Databricks and other engineering-heavy firms have similarly adopted GLM 5.2 for coding workflows.
The shift reflects a broader maturation in enterprise AI procurement. Companies now maintain portfolios of models spanning proprietary frontier endpoints, open-weight alternatives, and region-specific options, including Chinese labs. The gateway Rippling built sits between employees and these models, analyzing prompt type, context length, and required reasoning depth, then routing to the least expensive model that meets the threshold.
This is not trivial middleware. Effective routing requires understanding task taxonomy, latency tolerance, and output quality thresholds. Rippling's gateway is now a core component of the AI Spend Console product it launched this week. Enterprises that already operate their own gateways can still use the spend-tracking dashboard, but unlocking governance features requires adopting Rippling's routing layer.
Measuring Productivity, Not Just Tokens
The dashboard maps spending to output at the individual, team, and role level. For engineers, that means tokens consumed per day cross-referenced against lines of code committed and pull requests opened. The system also flags anomalies: high spend paired with frequent rework requests in code review suggests an employee is leaning on AI without critical evaluation.
MacInnis describes the original internal dashboards as "leaderboards," a term borrowed from the tokenmaxxing culture that briefly flourished in early 2026. The new version is more diagnostic than competitive. It surfaces patterns, identifies outliers, and provides managers with data to coach employees on effective prompt design and model selection.
Rippling used the tool internally before productizing it. Token spend dropped from 40% of R&D headcount budget to roughly 15%. Usage, however, did not fall. The company consumed 605 billion tokens in the peak month when the CFO issued his warning. In July, usage climbed back to 600 billion tokens, but the cost was 37% of April's bill. The savings came entirely from routing efficiency and model selection.
The company also identified high-performing users and designated them "AI captains," tasked with training peers. Technology alone does not shift behavior; cultural reinforcement matters. Still, adoption beyond engineering remains uneven. Rippling is piloting AI workflows in customer onboarding, where the dashboard will track tokens against accounts onboarded and data reconciliation tasks completed. MacInnis is blunt about the stakes: if the company cannot link token consumption in general and administrative functions to measurable productivity, access may be restricted.
The End of Universal Access
That conditional framing marks a departure from the way enterprises have historically rolled out collaboration tools. Slack, email, and videoconferencing platforms are provisioned universally. AI inference, at least at Rippling, may not follow that model. If a role or function cannot demonstrate productivity gains that justify token costs, access could be curtailed or gated behind manual approval.
This is a significant shift in enterprise software philosophy. It treats inference as a capital expense tied to output, not as a utility. The implications ripple outward: compensation structures may begin to account for AI leverage, performance reviews may incorporate token efficiency, and hiring criteria may weight prompt literacy alongside domain skills.
MacInnis acknowledges the company is still learning. Measuring productivity in non-engineering contexts is harder. What counts as output for a sales team or a legal function? How do you attribute deal closure or contract review speed to AI assist versus human judgment? Rippling is experimenting with proxies, but the methodology is not yet settled.
A Product, Not Just an Internal Tool
AI Spend Console is now available to Rippling's HR platform subscribers, with usage-based pricing for token consumption. It can also be purchased as a standalone product and integrated with other HR systems of record. The gateway and dashboard bundle is designed for mid-market and enterprise customers managing distributed AI tool sprawl.
The product's existence is itself a signal. Rippling is a human capital management platform; AI spend visibility was not part of its original product roadmap. The company built it because the pain was acute and no vendor offered a comparable solution. That gap suggests the market for AI cost management is still immature, and that enterprises are largely building point solutions in-house.
Whether Rippling's approach becomes a category or remains a feature set absorbed into broader observability platforms is unclear. What is clear is that the era of unconstrained inference spending is closing. Finance teams now scrutinize AI line items with the same rigor they apply to cloud infrastructure. Model selection, routing logic, and employee-level accountability are becoming standard components of enterprise AI operations.
The risk, as Rippling's own trajectory illustrates, is that cost control becomes cost restriction. If productivity measurement frameworks are crude or biased, companies may throttle access prematurely, stifling experimentation and learning. The challenge for 2026 and beyond is to build governance systems that eliminate waste without eliminating exploration.


