0-suite /
ai0

chat

A local LLM running on your GPU via WebGPU. No API key, no server — the conversation never leaves this page.

Powered by WebLLM (MLC) — the model runs in a WebGPU context in your tab. The first load downloads and compiles the weights (cached afterwards); generation and your history stay entirely on-device.