Liquid AI launches Pipette benchmark suite
Liquid AI and Artificial Analysis launched Pipette, an open-source benchmark suite for evaluating AI models on real devices. It measures model quality alongside prefill speed, generation speed, latency, memory use, model size, quantization, runtime, and hardware.
Pipette addresses a major blind spot in AI benchmarking: cloud scores rarely predict how models behave on phones, laptops, or other local hardware.
- –Public results help developers choose practical model-and-device combinations
- –Deterministic task-specific scorers avoid relying entirely on opaque LLM judges
- –Tracking quantization, runtime, context length, and memory exposes real deployment tradeoffs
- –The leaderboard could become valuable infrastructure for local AI, though coverage and reproducibility will depend on broader device participation
- –Early results should be treated as configuration-specific, since thermals, software stacks, and hardware differences can materially change performance
DISCOVERED
1d ago
2026-08-25
PUBLISHED
1d ago
2026-08-24
RELEVANCE
AUTHOR
AGTPinsights