In our per-class evals for a production analytics workload, GPT-5.6 Luna matched GPT-5.6 Sol's accuracy on the hardest question class. It did so at a fraction of the price and about half the latency.
We publish when a measurement says something worth repeating, so the feed is quiet by design. Subscribe by RSS, or ask for them by email and a person will add you to the list.