ocr
Florence-2 runs on-device (WebGPU, or WebAssembly as a fallback). The image never leaves your device.
result
Florence-2 is a compact vision-language model; here it runs entirely in the browser through
Transformers.js. OCR pulls selectable text out of screenshots and photos; captioning describes
a scene. The first run downloads the model once and caches it.