Fermion Research launches Neutrino-1 8B
Fermion Research has introduced Neutrino-1 8B, an 8.19 billion parameter decoder-only transformer based on Qwen3-8B that utilizes a proprietary coded ternary-family weight format. At a download size of 2.56 GB (3.88 GB on disk), the model packs its 252 transformer linear layers at one-eighth the footprint of fp16 and decodes weights directly inside matrix kernels. Designed as a unified artifact, a single container serves datacenter GPUs, Apple Silicon MacBooks, and desktop CPUs without conversion while achieving a 72.1 MMLU benchmark score and up to 763 tokens per second via speculative decoding on an NVIDIA H100.
Advanced ternary-coded quantization is solving the memory bandwidth bottleneck for edge AI without discarding model intelligence.
- –Extreme footprint reduction: Compresses an 8B parameter model into under 4 GB on disk, keeping weights bit-packed at rest and during kernel execution.
- –Cross-platform versatility: Eliminates the fragmentation of hardware-specific quantized formats by using a single artifact for GPUs, MacBooks, and CPUs.
- –High memory efficiency: Lowers serving overhead significantly, enabling large context windows alongside weights on consumer-grade hardware.
DISCOVERED
1h ago
2026-07-28
PUBLISHED
5h ago
2026-07-28
RELEVANCE
AUTHOR
handfuloflight