Rippling conducted a comprehensive test of 15 AI models using real payroll data, running approximately 2,100 scored attempts per model. The evaluation measured pass rate, cost, and the slowest 10% of task completion times to assess performance on actual personnel, payroll, and financial records. The study revealed that the cheapest AI model performed on par with the most expensive one, according to saastr.com.
The tests involved real production tasks such as calculating headcount by department, adjusting salaries based on conditions, onboarding new hires, scheduling terminations, and entering payment amounts from spreadsheets. Each attempt was graded as pass or fail based on Rippling’s production correctness checks, with unfinished attempts counted as failures. Rippling’s President and CPO Matt MacInnis shared these results to provide transparency on AI model performance in a live business environment.
This evaluation is significant because it moves beyond typical AI leaderboards and synthetic benchmarks, focusing instead on practical application in a B2B SaaS context. The findings suggest that companies can achieve high accuracy and efficiency without incurring the highest costs, challenging assumptions about AI pricing and performance. Rippling’s approach offers a rare data-driven perspective on AI model selection for payroll and HR automation.
The study narrowed the choice to three main models, with Opus 4.6 recommended when accuracy is paramount. Rippling’s detailed scoring and cost analysis provide a valuable reference for businesses seeking to optimize AI deployment in payroll systems, as reported by saastr.com.