Additional benchmark results for mid-size open-source LLMs released on the Medmarks v1.0 leaderboard.
The creators of the Medmarks v1.0 open benchmark suite and leaderboard for LLM medical capabilities have published new results for recent mid-size open-source models, including Gemma, Qwen, Muse, and Nemotron. Their evaluation found that the Gemma 4 31B model currently leads this specific size class in medical tasks.
The continuous release of specific benchmark results highlights the rapid progression of open-source models in specialized domains like medicine.
- –Gemma 4 31B demonstrates strong performance, reinforcing its position in domain-specific tasks.
- –Continuous updates to benchmarks like Medmarks are essential as the open-source landscape evolves quickly with new models like Qwen, Muse, and Nemotron.
- –Open-source mid-size models are becoming increasingly viable for complex, specialized medical use cases, reducing the reliance on massive, closed-source models.
DISCOVERED
1h ago
2026-09-21
PUBLISHED
1h ago
2026-09-21
RELEVANCE
AUTHOR
SophontAI
