Back to the notes

ROCm still fights for territory AMD already won

Someone built a tool to make AMD GPUs run local LLMs without the usual driver nightmares. The benchmarks show Vulkan beating HIP by 20 percent.

Processor Technology SOL 20 Computer
Swtpc6800 en:User:Swtpc6800 Michael Holley / Wikimedia Commons. Resized and converted to WebP. Public domain

AMD cards are cheap right now. A used RX 7900 XT costs half what a 4090 does. The problem is not the silicon, it is ROCm, AMD’s answer to CUDA. Installing it on consumer hardware is a multi-hour ritual of kernel patches and prayer. ROCmFix is a script that automates the whole mess. It patches drivers, sets up the runtime, and runs a benchmark suite called InferBench to compare Vulkan and HIP backends for running local LLMs. The interesting bit is not that it works. The interesting bit is the performance gap. InferBench shows Vulkan outrunning HIP by around 20 percent on the same hardware for inference tasks. That is not a rounding error. HIP is AMD’s CUDA clone, the official path for ML work on their cards. Vulkan is the cross-platform graphics API that happens to expose compute shaders. Vulkan was not built for transformers, yet here it is, faster than the thing AMD spent years designing for exactly this workload. The gap suggests HIP is carrying overhead somewhere in the stack. Maybe it is the abstraction layers. Maybe it is compiler maturity. Either way, if you are running Llama 3 on a 7900 XTX at home, you want the Vulkan backend, not the one AMD tells you to use. This matters because Nvidia’s moat is not just CUDA. It is that CUDA works out of the box and you do not spend a Saturday debugging kernel modules. AMD has the hardware. They do not have the install experience. Tools like ROCmFix close that gap, but the fact that Vulkan beats their own ML runtime is a sign the software stack still needs work. I would try this on a second-hand 7900 XT if I had one lying around. The 20 percent speed gain is real, and the setup script means you are not manually editing boot parameters at 2am.


Source: ROCmFix and InferBench – AMD Local-LLM Setup and Vulkan vs. Hip Benchmarking

Back to all notes

Behind the notes

Vikrant
Sharma.

Artificial Intelligence Engineer intern at Voxon Photonics in Adelaide. Studying a Master of Information and Communications Technology at UniSC, with a focus on data, machine learning and security.

Meet the person behind the work