Lemonade

by lemonade-sdk

4.8k 1.2k downloads Not rated yet Apache-2.0

About

Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk

Details

License
Apache-2.0

Explore

- Free, private local AI server (no cloud costs)
- Supports chat, coding, speech, and image generation
- Compatible with OpenAI, Anthropic, and Ollama API standards
- Runs on CPU, GPU, and NPU across multiple platforms
- Built-in Model Manager for browsing and downloading models
- Available as installable server or embeddable binary

1. Install: Windows · Linux · macOS · Docker · Source
2. Get Models: Browse and download with the Model Manager
3. Generate: Try models with the built-in interfaces for chat, image gen, speech gen, and more
4. Mobile: Take your lemonade to go: iOS · Android · Source
5. Connect: Use Lemonade with your favorite apps:

<!-- MARKETPLACE_START -->
<p align="center">
<a href="https://lemonade-server.ai/docs/server/apps/claude-code/" title="Claude Code">Claude Code</a>&nbsp;&nbsp;<a href="https://quickthoughts.ca/posts/firefox-chatback-lemonade-sdk/" title="Firefox Chatbot">Firefox Chatbot</a>&nbsp;&nbsp;<a href="https://lemonade-server.ai/docs/server/apps/anythingLLM/" title="AnythingLLM">AnythingLLM</a>&nbsp;&nbsp;<a href="https://marketplace.dify.ai/plugins/langgenius/lemonade" title="Dify">Dify</a>&nbsp;&nbsp;<a href="https://github.com/amd/gaia?tab=readme-ov-file#getting-started-guide" title="GAIA">GAIA</a>&nbsp;&nbsp;<a href="https://admcpr.com/local-github-copilot-with-lemonade-server-on-windows" title="GitHub Copilot">GitHub Copilot</a>&nbsp;&nbsp;<a href="https://github.com/lemonade-sdk/infinity-arcade" title="Infinity Arcade">Infinity Arcade</a>&nbsp;&nbsp;<a href="https://n8n.io/integrations/lemonade-model/" title="n8n">n8n</a>&nbsp;&nbsp;<a href="https://lemonade-server.ai/docs/server/apps/open-webui/" title="Open WebUI">Open WebUI</a>&nbsp;&nbsp;<a href="https://lemonade-server.ai/docs/server/apps/open-hands/" title="OpenHands">OpenHands</a>
</p>

<p align="center"><em>Want your app featured here? <a href="https://github.com/lemonade-sdk/marketplace">Just submit a marketplace PR!</a></em></p>
<!-- MARKETPLACE_END -->

Lemonade supports multiple inference engines for LLM, speech, TTS, and image generation, and each has its own backend and hardware requirements.

<!-- BEGIN GENERATED: backends-matrix -->
<table>
<thead>
<tr>
<th>Modality</th>
<th>Engine</th>
<th>Backend</th>
<th>Device</th>
<th>OS</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="9"><strong>Text generation</strong></td>
<td rowspan="6"><code>llamacpp</code></td>
<td><code>system</code></td>
<td><code>x86_64</code>/ARM64 CPU, GPU</td>
<td>Linux</td>
</tr>
<tr>
<td><code>metal</code></td>
<td>Apple Silicon GPU</td>
<td>macOS</td>
</tr>
<tr>
<td><code>cuda</code></td>
<td>NVIDIA GPUs (Turing or newer)</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>vulkan</code></td>
<td><code>x86_64</code> CPU, AMD iGPU, AMD dGPU; ARM64 CPU/GPU (Linux)</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>rocm</code></td>
<td>Supported AMD ROCm iGPU/dGPU families</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>cpu</code></td>
<td><code>x86_64</code> CPU; ARM64 CPU (Linux)</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td rowspan="1"><code>flm</code></td>
<td><code>npu</code></td>
<td>XDNA2 NPU</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td rowspan="1"><code>ryzenai-llm</code></td>
<td><code>npu</code></td>
<td>XDNA2 NPU</td>
<td>Windows</td>
</tr>
<tr>
<td rowspan="1"><code>vllm</code> (experimental)</td>
<td><code>rocm</code></td>
<td>Strix Halo iGPU (gfx1151)</td>
<td>Linux</td>
</tr>
<tr>
<td rowspan="6"><strong>Speech-to-text</strong></td>
<td rowspan="5"><code>whispercpp</code></td>
<td><code>npu</code></td>
<td>XDNA2 NPU</td>
<td>Windows</td>
</tr>
<tr>
<td><code>rocm</code></td>
<td>Supported AMD ROCm iGPU/dGPU families
</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>vulkan</code></td>
<td><code>x86_64</code> CPU</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>cpu</code></td>
<td><code>x86_64</code> CPU</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>metal</code></td>
<td>Apple Silicon GPU</td>
<td>macOS</td>
</tr>
<tr>
<td rowspan="1"><code>moonshine</code></td>
<td><code>cpu</code></td>
<td><code>x86_64</code>/<code>arm64</code> CPU</td>
<td>Windows, Linux, macOS</td>
</tr>
<tr>
<td rowspan="2"><strong>Text-to-speech</strong></td>
<td rowspan="2"><code>kokoro</code></td>
<td><code>cpu</code></td>
<td><code>x86_64</code> CPU</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>metal</code></td>
<td>Apple Silicon GPU</td>
<td>macOS</td>
</tr>
<tr>
<td rowspan="5"><strong>Image generation</strong></td>
<td rowspan="5"><code>sd-cpp</code></td>
<td><code>rocm</code></td>
<td>Supported AMD ROCm iGPU/dGPU families</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>cuda</code></td>
<td>NVIDIA GPUs (Turing or newer)
</td>
<td>Linux</td>
</tr>
<tr>
<td><code>vulkan</code></td>
<td>Vulkan-capable GPUs</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>cpu</code></td>
<td><code>x86_64</code> CPU</td>
<td>Windows, Linux</td>
</tr>
<tr>
<td><code>metal</code></td>
<td>Apple Silicon GPU</td>
<td>macOS</td>
</tr>
</tbody>
</table>
<!-- END GENERATED: backends-matrix -->

To check exactly which recipes/backends are supported on your own machine, run:

lemonade backends

<details>
<summary><small><i>
See supported AMD ROCm platforms</i></small></summary>

<br>

<table>
<thead>
<tr>
<th>Architecture</th>
<th>Platform Support</th>
<th>GPU Models</th>
</tr>
</thead>
<tbody>
<tr>
<td><b>gfx1151</b> (STX Halo)</td>
<td>Windows, Ubuntu</td>
<td>Ryzen AI MAX+ Pro 395</td>
</tr>
<tr>
<td><b>gfx120X</b> (RDNA4)</td>
<td>Windows, Ubuntu</td>
<td>Radeon AI PRO R9700, RX 9070 XT/GRE/9070, RX 9060 XT</td>
</tr>
<tr>
<td><b>gfx110X</b> (RDNA3)</td>
<td>Windows, Ubuntu</td>
<td>Radeon PRO W7900/W7800/W7700/V710, RX 7900 XTX/XT/GRE, RX 7800 XT, RX 7700 XT</td>
</tr>
</tbody>
</table>
</details>

<details>
<summary><small><i>** See supported NVIDIA CUDA platforms</i></small></summary>

<br>

<table>
<thead>
<tr>
<th>Compute Capability</th>
<th>Architecture</th>
<th>GPU Models</th>
</tr>
</thead>
<tbody>
<tr>
<td><b>sm_75</b></td>
<td>Turing</td>
<td>RTX 20-series, GTX 16-series, T4</td>
</tr>
<tr>
<td><b>sm_80</b> / <b>sm_86</b></td>
<td>Ampere</td>
<td>RTX 30-series, A100, A40</td>
</tr>
<tr>
<td><b>sm_89</b></td>
<td>Ada Lovelace</td>
<td>RTX 40-series, L40, L4</td>
</tr>
<tr>
<td><b>sm_90</b></td>
<td>Hopper</td>
<td>H100, H200</td>
</tr>
<tr>
<td><b>sm_100</b> / <b>sm_120</b></td>
<td>Blackwell</td>
<td>RTX 50-series, B100, B200</td>
</tr>
</tbody>
</table>
</details>

Initialize the client to use Lemonade Server

client = OpenAI( base_url="http://localhost:13305/api/v1", api_key="lemonade" # required but unused )
No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.

Videos about Lemonade

Relevant YouTube tutorials, setups, and demos