The Cost of Unchecked Token Consumption
In early 2026, Rippling discovered that its R&D organization was on a trajectory to spend 90% of its headcount budget on AI tokens. An internal audit revealed that 10–15% of employees were responsible for 60% of total AI spend, with some individual engineers burning $50,000 per month. The primary driver was a default reliance on expensive frontier models for every task, regardless of complexity.
Implementing an AI Gateway and Model Routing
To regain control, Rippling developed an AI Spend Console that functions as an AI gateway. Instead of allowing employees to default to the most expensive models, the system routes prompts to the most cost-effective model capable of handling the specific task. By diversifying their model portfolio—incorporating cheaper, high-performance options like GLM 5.2—Rippling achieved a significant efficiency gain: in July, they processed 600 billion tokens (nearly identical to their peak volume in April) at only 37% of the original cost.
Measuring AI ROI and Productivity
Beyond technical routing, Rippling implemented a management framework to ensure AI spend correlates with actual output. The AI Spend Console provides dashboards that track:
- Individual and Team Spend: Mapping consumption to specific roles.
- Productivity Metrics: Correlating token usage with tangible outputs, such as lines of code or pull requests.
- Quality Control: Identifying high-spend engineers whose code frequently requires rework, signaling inefficient AI usage.
To foster better adoption, the company appointed "AI captains" to mentor employees on effective prompting. Rippling emphasizes that if AI usage cannot be linked to measurable productivity in non-engineering functions (such as G&A or customer onboarding), access will be restricted rather than treated as a universal utility.