Seattle, WA · On-device AI

Ramakrishnan
Sivakumar

Principal ML Software Engineer at AMD. I make large language models run fast on hardware people already own, as co-creator and maintainer of Lemonade, the open-source local AI server.

Current focus

Co-creator & maintainer

Lemonade

Open-source local AI server: LLMs, speech, and image generation on your own hardware, with NPU and GPU acceleration.

Building

Tiered intelligence

Getting every request to the right tier of compute - NPU, GPU, or cloud - through policy-driven model routing: rule, classifier, semantic, and LLM-as-router.

Contributor

Upstream open source

Ongoing contributions across the local-AI stack: vLLM backends, llama.cpp enablement, and the ONNX ecosystem.

Experience

  1. 2014 – 16

    MS, Computer Engineering

    UNC Charlotte
  2. 2016 – 21

    Deep Learning Software Engineer

    Operating System Software Engineer

    Intel
  3. 2021 – 23

    Senior ML Software Engineer

    Groq
  4. 2023 – now

    Principal ML Software Engineer

    Staff ML Software Engineer

    AMD