Industry
AMD Partners with Cerebras to Bet on Low-Latency AI Inference: 5x Improvement in Tokens per Second per Watt
At the Advancing AI 2026 event, AMD announced a partnership with Cerebras to jointly develop a solution for ultra-low latency AI inference. The solution integrates AMD's Helios rack with Cerebras' Wafer-Scale Engine, handling the prompt and token generation stages respectively, and is expected to increase tokens per second per watt by 5 times.
Read the original (opens in a new tab)
News stream data aggregated by AI HOT