WHY THIS MATTERS
AMD announced ROCm 10 on August 27, 2026, the latest version of its open software platform for GPU and artificial intelligence computing, marking ten years since the stack's debut as the main rival to Nvidia's CUDA. According to reports from StorageReview, VideoCardz, HPCwire, WCCFTech and AMD's own website, the release introduces ROCm.AI, a toolkit described as AI native that combines the agentic ROCm Hyperloom system, AMD Skills and the ROCm CLI. AMD states that a system configured with ROCm.AI delivers, on average, 3.3 times more inference performance and 2.4 times more training performance compared with ROCm 7, in tests run on a platform with eight Instinct MI355X GPUs using models such as GLM-5, Kimi K2.5 and DeepSeek-R1-0528. The move is widely read as another chapter in AMD's fight against Nvidia for control of AI infrastructure software.In this article
Ten Years of ROCm, One New Release
AMD unveiled ROCm 10 on August 27, 2026, a date that coincides with the tenth anniversary of ROCm, its open software stack for GPU computing and AI workloads. Built as an alternative to CUDA, Nvidia's proprietary environment that still dominates the market for AI training and inference software, ROCm has spent a decade trying to close that gap. Coverage from outlets including StorageReview, VideoCardz, HPCwire and WCCFTech, along with AMD's own official announcement, frames ROCm 10 as the platform's most ambitious release to date, with a stated focus on developer productivity and performance gains across inference and training.
Inside ROCm.AI: Hyperloom, AMD Skills and ROCm CLI
At the center of the announcement is ROCm.AI, a toolkit AMD describes as AI native, built from three components. ROCm Hyperloom is an autonomous agentic system designed to optimize inference workloads end to end, acting on both host side code and GPU kernels directly. It is joined by AMD Skills and the ROCm CLI, which round out the kit with additional automation and control features for developers and operators running AI applications on AMD hardware. According to the company, the goal is to let optimization decisions once handled manually by engineers be guided continuously by AI instead.
The Numbers: 3.3x Inference, 2.4x Training
AMD says a system configured with ROCm.AI delivers, on average, 3.3 times more inference performance and 2.4 times more training performance than ROCm 7 on the same hardware, thanks to AI guided kernel optimization, memory management and scheduling. The tests were measured on a platform with eight Instinct MI355X GPUs, using models such as GLM-5, Kimi K2.5 and DeepSeek-R1-0528. It is worth stressing that these figures are specific to the configurations and optimizations AMD tested: they should not be read as a blanket claim that ROCm 10 is 3.3 times faster than ROCm 7 across every workload or hardware setup.
Another Round in the Fight Against Nvidia
The launch of ROCm 10 and ROCm.AI is being read across the industry as part of AMD's strategy to narrow the lead Nvidia holds thanks to CUDA, still the dominant standard for AI software development on GPUs. By investing in an AI native toolkit that automates optimizations once done by hand, AMD is trying to lower the barrier for developers accustomed to the Nvidia ecosystem and make its Instinct hardware more competitive on performance per dollar spent on AI infrastructure.
Why It Matters for Brazilian Agencies and SMBs
For Brazilian agencies and SMBs running AI agents in marketing and customer service, the AMD versus Nvidia rivalry is rarely an abstract story: GPU infrastructure is the main cost behind every inference call that keeps a production agent running. A more competitive AMD, backed by gains like those claimed for ROCm 10, tends to push AI cloud pricing down over the medium term as cloud providers diversify suppliers and use that competition as leverage in negotiations. For anyone paying for agent infrastructure at scale, every real efficiency gain on the hardware and software side is, ultimately, a possible cut in operating costs.