AIR-LLM Broadcasts Weights for Memory-Free Edge Inference
AIR-LLM broadcasts LLM weights from a central radio, then uses RF mixers on edge devices to perform matrix-vector operations without storing the model locally. Evaluations report a 4% perplexity gap on LLaMA-3.1-8B and up to 157.7× lower energy than FP16 inference.
AIR-LLM attacks the memory-movement bottleneck with an unusually radical hardware-software split, but its headline gains remain simulation- and testbed-backed rather than proof of phone-ready deployment.
- –One broadcast can serve many users, making the architecture increasingly attractive as edge-device density rises.
- –RF-domain GEMV keeps prompts and computation at the edge while relocating weight storage to the network.
- –Channel quality, path loss, calibration, thermal noise, and mixer distortion materially affect accuracy.
- –The paper evaluates ray-traced Paris and Munich channels plus a measured RF mixer, not a complete over-the-air commercial system.
- –For developers, this points toward cellular infrastructure as an inference substrate—not merely a connectivity layer.
DISCOVERED
2h ago
2026-10-11
PUBLISHED
2h ago
2026-10-11
RELEVANCE
AUTHOR
Better Stack