Decagon · a16z
Split the stack by task, not by ideology
Decagon runs 90 percent of its agents on open-source models, and the reasoning is not ideological. They evaluate every model on three axes at once: cost, intelligence and latency. A smaller open model, fine-tuned on a specific business task, frequently beats a larger general-purpose frontier model on that task while also being faster and cheaper. That is the rare case where all three axes move the right way together.
What they did not do is pick a side. Open-ended work like trend analysis still goes to frontier models, because that is where breadth actually matters. The split is by task, and it gets re-decided as models change.
Steal it
List your five highest-volume AI calls. Route the narrow, repetitive ones to a fine-tuned small model and leave the open-ended ones on the frontier API.