Today2h ago
Analytics
Frontier AI Models: Benchmark Breakdown

Frontier AI Models: Benchmark Breakdown

The frontier got reshuffled in two days. Anthropic shipped Claude Fable 5.1, Meta shipped Muse Spark 1.3. We ran both against Kimi K3 and GPT-5.6 Sol across 9 LiveBench benchmarks.


Fable 5.1 took the crown almost everywhere. But the top spot has a price tag: ~$1.21 per task, 5.5x more than Muse Spark at $0.22. Meta lost on power but won on cost and instruction following. Kimi K3 holds a top reasoning score for pennies, proof open models are catching up fast.