Meta's Muse Glimmer 30B Targets Local Agentic Coding
A 30-billion-parameter multimodal model built to run locally, released under Apache 2.0 with an eye on agentic coding workflows.
The company's newest flagship targets reasoning and coding while keeping a permissive open-source license.
A 30-billion-parameter multimodal model built to run locally, released under Apache 2.0 with an eye on agentic coding workflows.
Alibaba's Qwen team pushes its largest sparse model yet, activating 95 billion parameters per token from a 2.4-trillion-parameter pool.
Alibaba's latest Qwen release pairs image understanding with text under a permissive Apache-2.0 license.
A new open vision-language model family uses gated cross-attention to enable streaming, low-latency multimodal exchanges.
The Chinese lab pushes a higher-capability checkpoint of its V4 line to Hugging Face under a permissive MIT license.
The 3-billion-parameter vision-language model is tuned for faster multimodal work on edge hardware.
The open-weights text-to-speech model targets low-latency, deployable voice agents across multiple languages.
The company releases open weights for a low-latency, multilingual text-to-speech model aimed at real-time conversational systems.
A small multilingual vision-language model built for instruction following arrives under a research-only license.
Momentum
Benchmarks
| # | Model | Avg. |
|---|---|---|
| 1 | MaziyarPanahi/calme-3.2-instruct-78b | 52.1 |
| 2 | MaziyarPanahi/calme-3.1-instruct-78b | 51.3 |
| 3 | dfurman/CalmeRys-78B-Orpo-v0.1 | 51.2 |
| 4 | MaziyarPanahi/calme-2.4-rys-78b | 50.8 |
| 5 | huihui-ai/Qwen2.5-72B-Instruct-abliterated | 48.1 |
| 6 | Qwen/Qwen2.5-72B-Instruct Qwen · Alibaba | 48.0 |
Lodestones releases an open image generator built on the Krea 2 lineage, aimed squarely at ComfyUI workflows.
The company's newest flagship targets reasoning and coding while keeping a permissive open-source license.

The company's new diffusion model handles text-to-video and image-to-video, with support for joint audio-video generation.
A new open vision-language model family uses gated cross-attention to enable streaming, low-latency multimodal exchanges.
An experimental open-weight model swaps token-by-token decoding for discrete diffusion, generating text in parallel blocks.