“Ramakrishnan Sivakumar at AMD authored the original vLLM backend, and Jeremy Fowers led the refactor.” One vLLM Stack, Two Architectures: Prototype on RDNA, Scale to an MI300X for the Price of Lunch →
External coverage, in one place: press, community write-ups, and launches involving the work, quoted as published.
“Ramakrishnan Sivakumar at AMD authored the original vLLM backend, and Jeremy Fowers led the refactor.” One vLLM Stack, Two Architectures: Prototype on RDNA, Scale to an MI300X for the Price of Lunch →
“The most significant change with Lemonade 11.5 is the completion of the Lemonade Router that can be used for automatically routing queries to relevant models based on defined policies.” Lemonade 11.5 Local AI Server Released With Completed Lemonade Router →
“Use your stack on whichever hardware you happen to be using at the time.” I switched my local AI setup to AMD's Lemonade after Nvidia support landed, and solved my local AI portability problem →
“How I Fixed vLLM on Strix Halo and Got 3x Better Batch Throughput with Qwen3.5.” Stop Crashing and Start Cooking with vLLM on AMD and Lemonade Server →
“This is an interesting change for a local AI project backed by AMD, since Lemonade was primarily built around Ryzen AI NPUs, Radeon GPUs and x86 CPUs.” AMD-backed Lemonade local AI server adds NVIDIA CUDA support →
“Lemonade, the local AI server solution developed by AMD that is designed to work across their CPUs, GPUs, and NPUs, is out with a new version today that also adds NVIDIA CUDA support.” AMD's Lemonade SDK For Local AI Adds NVIDIA CUDA Support →
“The Lemonade SDK for ‘refreshingly fast local AI’ that is largely developed by AMD engineers as an open-source project continues advancing quite rapidly for serving optimized LLMs on GPUs and NPUs.” AMD's Lemonade SDK For AI Promotes macOS To GA Status, ROCm 7.13 Integrated →
“Lemonade, created by AMD, is a server application plus GUI for running local AI models, similar to projects like LM Studio.” First look: Lemonade serves up local AI with limitations →
572 points · 111 comments on Hacker News. Lemonade by AMD: a fast and open source local LLM server using GPU and NPU →
“AMD decided to find the best way to show off Strix Halo's capabilities. They chose to throw eight local LLMs into a group chat and ask the one question that may ruin family holiday spirit.” AMD demos Strix Halo running 8 AI models to argue over “Is a hot dog a sandwich?” →
“AMD Partnered with Hugging Face to provide a guide on how to accelerate our end-to-end Tiny Agents application using AMD Neural Processing Unit (NPU) and integrated GPU (iGPU).” Hugging Face MCP Course: Lemonade Server unit →
2,326 models published under the ONNX Model Zoo organization on Hugging Face - the corpus rebuilt on TurnkeyML and relaunched at ONNX Community Meetup 2023. ONNX Model Zoo organization on Hugging Face →