Finding
GPT-5.6 Luna matched Sol on our hardest question class
Published
In our per-class evals for a production analytics workload, GPT-5.6 Luna matched GPT-5.6 Sol's accuracy on the hardest question class. It did so at a fraction of the price and about half the latency.
We ran per-class evals on a production analytics workload and found a useful crossover. GPT-5.6 Luna matched GPT-5.6 Sol's accuracy on the hardest question class, at a fraction of the price and about half the latency.
Luna sits in the cheap tier of our prices table. Sol is the flagship. The live table carries their current list rates, so we leave the drifting numbers there.
Model capability is a profile that distillation shrinks unevenly. A cheap tier can hold frontier accuracy for one task class while giving up capability elsewhere.
Aggregate leaderboards average over someone else's task mix, which hides this crossover. Because pricing follows the aggregate, per-class measurement exposes capability that is cheap for the work it can handle.
Method note
How we measured it
We built a gold set around the workload's question classes.
We ran each candidate model through a custom harness and used an independently verified grader.
For each class, we recorded accuracy. We also measured latency and calculated cost from current list rates.