The Reticle Wall: Why AMD's Taalas Bet Works for 8 Billion Parameters and Stops There
AMD's acquisition of Taalas, announced Thursday, is the kind of deal that sounds like science fiction until you read the spec sheet. Taalas doesn't build chips that run AI models — it builds chips that are the model. Weights, dataflow, and matrix multiplication pathways are etched directly into transistors, eliminating the memory bottleneck that dominates GPU inference. The company's first test chip, HC1, runs Meta's Llama 3.1 8B at roughly 17,000 tokens per second — a rate Taalas claimed in February was 73 times…