Skip to main content

One post tagged with "benchmarks"

View All Tags

Every major AI lab shipped an agent model this week

ยท 8 min read
Mangat Rai
Creator, Few-Shot Academy

For the last two years, open-weight model releases have mostly been a China story: DeepSeek, Alibaba's Qwen team, Moonshot AI's Kimi, Zhipu's GLM, MiniMax, all shipping frontier-class open weights on a cadence Western labs haven't matched. This week Meta broke that pattern, twice, in one release. It shipped Muse Glimmer, a 30-billion-parameter model with Apache 2.0 weights on Hugging Face, alongside Muse Spark 1.2, a hosted sibling built for long coding sessions. The interesting part isn't that Meta released open weights again. It's that, by Meta's own benchmarks, Glimmer beats two other well-known open-weight models roughly its own size at the kind of multi-step, tool-using tasks that used to separate the labs with the biggest budgets from everyone else.