AMD buys a company that tattoos AI models onto silicon — hope you like that model forever
AMD has acquired Toronto-based AI chip startup Taalas, whose approach bakes model weights directly into silicon rather than storing them in HBM. Early benchmarks showed Taalas' first test chip serving Llama 3.1 8B at 16,960 tokens per second — roughly 48x faster than Nvidia GPUs. The tradeoff is significant: any model change beyond a LoRA adapter requires a chip re-spin, though only two metal layers need changing rather than a full redesign. AMD plans to pair Taalas-based accelerators with its Instinct GPU racks, using GPUs for prompt processing and Taalas chips for token generation.