Bottleneck Labs tested GPT 5.6 Sol by creating an autonomous agent named Saul to run a real business for 24 hours. Saul started with $350 in working capital and access to a dedicated Mac mini and unlimited tokens. After a full day of continuous operation, Saul ended with a balance of $250.50, losing $99.50, while increasing users from 61 to 66 but generating no new revenue, according to bottlenecklabs.com.
The experiment involved Saul making 1,129 tool calls, including 908 shell commands, and processing 320.7 million prompt tokens. The setup gave Saul full control over business assets and working capital to simulate real-world operations. Despite the agent's nonstop activity and access to resources, it failed to generate revenue and lost money over the 24-hour period, as detailed by Bottleneck Labs.
This test highlights the current limitations of autonomous AI agents in managing real businesses profitably. While GPT 5.6 Sol demonstrated operational capabilities, it struggled with financial decision-making and customer acquisition. The results provide a benchmark for AI-driven business automation, contrasting with human-led startups that typically focus on revenue growth and cost control from day one.
Bottleneck Labs published the full experiment details on their website, including charts of Saul's net worth and operational metrics, offering valuable data for future research on autonomous AI business agents.