chat
A local LLM running on your GPU via WebGPU. No API key, no server — the conversation never leaves this page.
system prompt
Powered by WebLLM (MLC) — the model runs in a WebGPU context in your tab. The first load
downloads and compiles the weights (cached afterwards); generation and your history stay
entirely on-device.