1. Why You Would Want Offline At All
Cloud models are fast and capable, but there are three critical scenarios where offline local execution wins:
With a browser-local model selected and loaded, inference runs in the browser instead of a hosted model provider. Model/runtime downloads and other app features can still use network services. Avoid entering highly sensitive data unless the full data path is appropriate for it.
After the app, runtime, and model files are available in the browser, local generation may continue without a connection. Browser cache eviction, app reloads, provider calls, and other features may require network access.
Local inference does not use a hosted provider's token quota, but storage, device memory, speed, electricity, model license terms, and browser storage limits still apply. Browser caches can be cleared or evicted.
Browser-local models use your device resources and may have smaller capacity than some hosted models. They can be useful for short tasks when local inference fits your hardware. If you need a larger model or more context, select a connected provider and review its data and pricing terms.
2. WebLLM — Your GPU Runs the Model
What it is: WebLLM uses WebGPU where the browser and hardware support it to run model computations on the device. The selected model files are downloaded, cached by the browser, and loaded into available memory. A browser may later clear cached files, so do not assume the download is permanent.
- • Supported Browsers: Chrome 113+, Edge 113+, Brave with WebGPU enabled. Safari does not support WebGPU yet.
- • Apple Silicon Macs: M1/M2/M3/M4 Macs run WebLLM exceptionally well due to high-bandwidth Unified Memory architecture.
- • Windows PCs: Dedicated GPU with 4GB+ VRAM (NVIDIA GTX 1060 or better) and 8GB–16GB system RAM.
3. The WebLLM Model Catalog — What to Pick
In Model Organizer -> WebLLM provider, Accelerated Logic AI displays curated prebuilt models:
Ultra-fast download, runs on low-end hardware. Great for quick text transformations, routing, and short responses.
Optimal balance between memory footprint, inference speed, and coding accuracy. Fits comfortably on 8GB RAM systems.
Highest quality offline code generation. Requires 12GB–16GB RAM and dedicated GPU VRAM.
Supports image input offline. Upload UI screenshots and prompt "Clone this layout in React".
4. Download, Cache, and Load — Understanding Model States
Each WebLLM model card transitions through 4 distinct visual lifecycle states:
Fetches model shards into browser storage. Real-time percentage progress displays in bottom bar.
Weights stored on local disk (Cache API). Click Load to initialize weights into GPU memory.
Model is active in RAM/VRAM and ready for instant streaming generation.
Removes local weight shards to immediately free up browser disk space.
5. Check the local inference data path
- Equip and Load a WebLLM model (e.g. Llama-3.2-3B).
- Send a test prompt in chat:
Hello in 5 languages. - For a controlled check, keep the same loaded page and model selected, then inspect the browser Network panel during a prompt.
- Confirm the prompt is handled by the local runtime and that no hosted model request is made for this generation. Other background requests may still occur.
6. Storage Management & Deleting Weights
WebLLM model shards are cached in browser Cache Storage. To inspect or manage disk usage:
Checking Storage:
Open Chrome DevTools -> Application tab -> Storage. It shows exact storage used (e.g., 3.8 GB).
Freeing Storage:
In Model Organizer -> WebLLM, click Delete Cache on any cached model card to delete its downloaded weights instantly.
7. Performance Tips for Local WebLLM
Close memory-heavy browser tabs (Figma, YouTube) when loading 7B/8B models to prevent WebGL context loss.
Keep context length under 4096 tokens for local execution by unpinning unused attachments.
8. Gemini Nano — Chrome on-device model where available
Some Chrome builds expose a built-in on-device language model API. Availability, prerequisites, and names of experimental flags can change by Chrome version and device; the workspace shows this provider only when its capability check succeeds.
- Use a supported Chrome build on compatible hardware and check Chrome's current built-in AI documentation for the required setup.
- Open the Model Organizer. The Gemini Nano entry is available only when the browser reports support.
- If the provider is unavailable, use a supported WebLLM, Transformers.js, or local Ollama model instead.
9. WebLLM vs Gemini Nano vs Ollama Comparison
| Feature | Gemini Nano | WebLLM (WebGPU) | Ollama Local |
|---|---|---|---|
| Setup Needed | Supported Chrome build and device | 1-Click Browser Download | Desktop Install + CORS |
| Download Size | Managed by Chrome | 1GB - 4GB in Cache | 2GB - 10GB Local Disk |
| Context Window | ~4096 tokens | 4096 tokens | 8k - 128k tokens |
| Vision Support | No | Yes (Phi-3.5 Vision) | Yes (LLaVA / MiniCPM) |
10. Offline Troubleshooting
Ensure you are using Chrome 113+ or Edge 113+ on desktop. Check chrome://gpu to verify hardware acceleration.
Out of GPU memory. Switch from 8B model to a lighter 1B or 3B model (e.g. Llama-3.2-3B).
Verify both flags in chrome://flags are Enabled and check chrome://components status.