AMD acquires AI chip startup Taalas
AMD has acquired a Toronto-based startup that bakes Large Language Model (LLM) weights directly onto silicon instead of relying on high-bandwidth memory (HBM). Early test chips demonstrated running Llama 3.1 8B at around 17,000 tokens per second—roughly 48 times faster than standard NVIDIA GPUs. However, because the model logic is physically hardwired, updating the model requires a complete chip re-spin.
Etching weights onto silicon trades traditional software adaptability for staggering inference performance gains, marking a bold shift toward hyper-specialized AI hardware.
- –Delivering 17k tokens/sec unlocks entirely new real-time voice and agentic application paradigms.
- –The severe drawback of hardware lock-in makes this approach viable primarily for frozen, high-volume foundational models.
- –AMD's acquisition underlines a growing push to bypass traditional memory-bandwidth bottlenecks through unconventional chip design.
DISCOVERED
1h ago
2026-08-08
PUBLISHED
2h ago
2026-08-08
RELEVANCE
AUTHOR
bsormagec