Current focus
Lemonade
Open-source local AI server: LLMs, speech, and image generation on your own hardware, with NPU and GPU acceleration.
Tiered intelligence
Getting every request to the right tier of compute - NPU, GPU, or cloud - through policy-driven model routing: rule, classifier, semantic, and LLM-as-router.
Upstream open source
Ongoing contributions across the local-AI stack: vLLM backends, llama.cpp enablement, and the ONNX ecosystem.
Experience
-
MS, Computer Engineering
UNC Charlotte -
Deep Learning Software Engineer
Operating System Software Engineer
Intel -
Senior ML Software Engineer
Groq -
Principal ML Software Engineer
Staff ML Software Engineer
AMD