AI economics · Technology
Smaller AI systems compete on the cost of being checked
Demo journalism: A fictional software trial puts verification expense, rather than benchmark prestige, at the centre of purchasing decisions.
By Dena Pellwick · 5 min read

In an invented evaluation by software co-operative Merrowstack, a compact language model handles 82% of a narrowly defined invoice-classification workload without escalation. A larger system reaches 89%, but costs more to operate. Neither figure measures general intelligence or reliability outside the test: the documents use a fixed template, and every uncertain classification is routed to a human reviewer.
Once review time enters the model, the apparent price advantage becomes less straightforward. The smaller system generates 180 referrals per 1,000 documents, compared with 110 for its counterpart. At an assumed $2 per review, that gap can outweigh savings in computing. The purchasing decision therefore depends on the workflow surrounding the model, not simply the price charged for each automated response.