Lemonade
About
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
Details
- License
- Apache-2.0
Explore
- Free, private local AI server (no cloud costs)
- Supports chat, coding, speech, and image generation
- Compatible with OpenAI, Anthropic, and Ollama API standards
- Runs on CPU, GPU, and NPU across multiple platforms
- Built-in Model Manager for browsing and downloading models
- Available as installable server or embeddable binary
1. Install: Windows · Linux · macOS · Docker · Source
2. Get Models: Browse and download with the Model Manager
3. Generate: Try models with the built-in interfaces for chat, image gen, speech gen, and more
4. Mobile: Take your lemonade to go: iOS · Android · Source
5. Connect: Use Lemonade with your favorite apps:
<!-- MARKETPLACE_START -->
<p align="center">
<a href="https://lemonade-server.ai/docs/server/apps/claude-code/" title="Claude Code">
</a> <a href="https://quickthoughts.ca/posts/firefox-chatback-lemonade-sdk/" title="Firefox Chatbot">
</a> <a href="https://lemonade-server.ai/docs/server/apps/anythingLLM/" title="AnythingLLM">
</a> <a href="https://marketplace.dify.ai/plugins/langgenius/lemonade" title="Dify">
</a> <a href="https://github.com/amd/gaia?tab=readme-ov-file#getting-started-guide" title="GAIA">
</a> <a href="https://admcpr.com/local-github-copilot-with-lemonade-server-on-windows" title="GitHub Copilot">
</a> <a href="https://github.com/lemonade-sdk/infinity-arcade" title="Infinity Arcade">
</a> <a href="https://n8n.io/integrations/lemonade-model/" title="n8n">
</a> <a href="https://lemonade-server.ai/docs/server/apps/open-webui/" title="Open WebUI">
</a> <a href="https://lemonade-server.ai/docs/server/apps/open-hands/" title="OpenHands">
</a>
</p>
<p align="center"><em>Want your app featured here? <a href="https://github.com/lemonade-sdk/marketplace">Just submit a marketplace PR!</a></em></p>
<!-- MARKETPLACE_END -->
Lemonade supports multiple inference engines for LLM, speech, TTS, and image generation, and each has its own backend and hardware requirements.
<!-- BEGIN GENERATED: backends-matrix -->
<table>
<thead>
<tr>
<th>Modality</th>
<th>Engine</th>
<th>Backend</th>
<th>Device</th>
<th>OS</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="9"><strong>Text generation</strong></td>
<td rowspan="6"><code>llamacpp</code></td>
<td><code>system</code></td>
<td><code>x86_64</code>/ARM64 CPU, GPU</td>
<td>Linux</td>
</tr>
<tr>
<td><code>metal</code></td>
<td>Apple Silicon GPU</td>
<td>macOS</td>
</tr>
<tr>
<td><code>cuda</code></td>
<td>NVIDIA GPUs (Turing or newer)</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>vulkan</code></td>
<td><code>x86_64</code> CPU, AMD iGPU, AMD dGPU; ARM64 CPU/GPU (Linux)</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>rocm</code></td>
<td>Supported AMD ROCm iGPU/dGPU families</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>cpu</code></td>
<td><code>x86_64</code> CPU; ARM64 CPU (Linux)</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td rowspan="1"><code>flm</code></td>
<td><code>npu</code></td>
<td>XDNA2 NPU</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td rowspan="1"><code>ryzenai-llm</code></td>
<td><code>npu</code></td>
<td>XDNA2 NPU</td>
<td>Windows</td>
</tr>
<tr>
<td rowspan="1"><code>vllm</code> (experimental)</td>
<td><code>rocm</code></td>
<td>Strix Halo iGPU (gfx1151)</td>
<td>Linux</td>
</tr>
<tr>
<td rowspan="6"><strong>Speech-to-text</strong></td>
<td rowspan="5"><code>whispercpp</code></td>
<td><code>npu</code></td>
<td>XDNA2 NPU</td>
<td>Windows</td>
</tr>
<tr>
<td><code>rocm</code></td>
<td>Supported AMD ROCm iGPU/dGPU families</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>vulkan</code></td>
<td><code>x86_64</code> CPU</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>cpu</code></td>
<td><code>x86_64</code> CPU</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>metal</code></td>
<td>Apple Silicon GPU</td>
<td>macOS</td>
</tr>
<tr>
<td rowspan="1"><code>moonshine</code></td>
<td><code>cpu</code></td>
<td><code>x86_64</code>/<code>arm64</code> CPU</td>
<td>Windows, Linux, macOS</td>
</tr>
<tr>
<td rowspan="2"><strong>Text-to-speech</strong></td>
<td rowspan="2"><code>kokoro</code></td>
<td><code>cpu</code></td>
<td><code>x86_64</code> CPU</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>metal</code></td>
<td>Apple Silicon GPU</td>
<td>macOS</td>
</tr>
<tr>
<td rowspan="5"><strong>Image generation</strong></td>
<td rowspan="5"><code>sd-cpp</code></td>
<td><code>rocm</code></td>
<td>Supported AMD ROCm iGPU/dGPU families</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>cuda</code></td>
<td>NVIDIA GPUs (Turing or newer)</td>
<td>Linux</td>
</tr>
<tr>
<td><code>vulkan</code></td>
<td>Vulkan-capable GPUs</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>cpu</code></td>
<td><code>x86_64</code> CPU</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>metal</code></td>
<td>Apple Silicon GPU</td>
<td>macOS</td>
</tr>
</tbody>
</table>
<!-- END GENERATED: backends-matrix -->
To check exactly which recipes/backends are supported on your own machine, run:
lemonade backends
<details>
<summary><small><i> See supported AMD ROCm platforms</i></small></summary>
<br>
<table>
<thead>
<tr>
<th>Architecture</th>
<th>Platform Support</th>
<th>GPU Models</th>
</tr>
</thead>
<tbody>
<tr>
<td><b>gfx1151</b> (STX Halo)</td>
<td>Windows, Ubuntu</td>
<td>Ryzen AI MAX+ Pro 395</td>
</tr>
<tr>
<td><b>gfx120X</b> (RDNA4)</td>
<td>Windows, Ubuntu</td>
<td>Radeon AI PRO R9700, RX 9070 XT/GRE/9070, RX 9060 XT</td>
</tr>
<tr>
<td><b>gfx110X</b> (RDNA3)</td>
<td>Windows, Ubuntu</td>
<td>Radeon PRO W7900/W7800/W7700/V710, RX 7900 XTX/XT/GRE, RX 7800 XT, RX 7700 XT</td>
</tr>
</tbody>
</table>
</details>
<details>
<summary><small><i>** See supported NVIDIA CUDA platforms</i></small></summary>
<br>
<table>
<thead>
<tr>
<th>Compute Capability</th>
<th>Architecture</th>
<th>GPU Models</th>
</tr>
</thead>
<tbody>
<tr>
<td><b>sm_75</b></td>
<td>Turing</td>
<td>RTX 20-series, GTX 16-series, T4</td>
</tr>
<tr>
<td><b>sm_80</b> / <b>sm_86</b></td>
<td>Ampere</td>
<td>RTX 30-series, A100, A40</td>
</tr>
<tr>
<td><b>sm_89</b></td>
<td>Ada Lovelace</td>
<td>RTX 40-series, L40, L4</td>
</tr>
<tr>
<td><b>sm_90</b></td>
<td>Hopper</td>
<td>H100, H200</td>
</tr>
<tr>
<td><b>sm_100</b> / <b>sm_120</b></td>
<td>Blackwell</td>
<td>RTX 50-series, B100, B200</td>
</tr>
</tbody>
</table>
</details>
Initialize the client to use Lemonade Server
client = OpenAI( base_url="http://localhost:13305/api/v1", api_key="lemonade" # required but unused )Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.



