EN Submit a tool
Industry

AMD Partners with Cerebras to Bet on Low-Latency AI Inference: 5x Improvement in Tokens per Second per Watt

Published: Source: IT Home (RSS)

ShareXFacebookTelegramWhatsApp

At the Advancing AI 2026 event, AMD announced a partnership with Cerebras to jointly develop a solution for ultra-low latency AI inference. The solution integrates AMD's Helios rack with Cerebras' Wafer-Scale Engine, handling the prompt and token generation stages respectively, and is expected to increase tokens per second per watt by 5 times.

Read the original (opens in a new tab)

News stream data aggregated by AI HOT

Related newsLatest in this category