GA Release

The Multi-Agent WebGPU Engine & Interactions API are now generally available.

Part 4: Browser-Local Models — WebLLM & Gemini Nano

Set up supported on-device model runtimes, understand their download requirements, and check what stays local during inference.

Category: Getting Started • Read Time: 20 min read • Updated: August 2026

1. Why You Would Want Offline At All

Cloud models are fast and capable, but there are three critical scenarios where offline local execution wins:

🔒 Local inference after setup

With a browser-local model selected and loaded, inference runs in the browser instead of a hosted model provider. Model/runtime downloads and other app features can still use network services. Avoid entering highly sensitive data unless the full data path is appropriate for it.

✈️ Some work can continue offline

After the app, runtime, and model files are available in the browser, local generation may continue without a connection. Browser cache eviction, app reloads, provider calls, and other features may require network access.

No hosted-model token charge

Local inference does not use a hosted provider's token quota, but storage, device memory, speed, electricity, model license terms, and browser storage limits still apply. Browser caches can be cleared or evicted.

⚖️ The Trade-off:

Browser-local models use your device resources and may have smaller capacity than some hosted models. They can be useful for short tasks when local inference fits your hardware. If you need a larger model or more context, select a connected provider and review its data and pricing terms.

2. WebLLM — Your GPU Runs the Model

What it is: WebLLM uses WebGPU where the browser and hardware support it to run model computations on the device. The selected model files are downloaded, cached by the browser, and loaded into available memory. A browser may later clear cached files, so do not assume the download is permanent.

Hardware & Browser Requirements:
  • • Supported Browsers: Chrome 113+, Edge 113+, Brave with WebGPU enabled. Safari does not support WebGPU yet.
  • • Apple Silicon Macs: M1/M2/M3/M4 Macs run WebLLM exceptionally well due to high-bandwidth Unified Memory architecture.
  • • Windows PCs: Dedicated GPU with 4GB+ VRAM (NVIDIA GTX 1060 or better) and 8GB–16GB system RAM.

3. The WebLLM Model Catalog — What to Pick

In Model Organizer -> WebLLM provider, Accelerated Logic AI displays curated prebuilt models:

🚀 Llama-3.2-1B (~800MB Download)

Ultra-fast download, runs on low-end hardware. Great for quick text transformations, routing, and short responses.

⭐ Llama-3.2-3B Instruct (~2GB Download) [Recommended Starter]

Optimal balance between memory footprint, inference speed, and coding accuracy. Fits comfortably on 8GB RAM systems.

🧠 Llama-3.1-8B or Qwen2.5-Coder-7B (~4GB Download)

Highest quality offline code generation. Requires 12GB–16GB RAM and dedicated GPU VRAM.

👁️ Phi-3.5-Vision-Instruct

Supports image input offline. Upload UI screenshots and prompt "Clone this layout in React".

4. Download, Cache, and Load — Understanding Model States

Each WebLLM model card transitions through 4 distinct visual lifecycle states:

1. Download

Fetches model shards into browser storage. Real-time percentage progress displays in bottom bar.

2. Cached

Weights stored on local disk (Cache API). Click Load to initialize weights into GPU memory.

3. Loaded

Model is active in RAM/VRAM and ready for instant streaming generation.

4. Delete Cache

Removes local weight shards to immediately free up browser disk space.

5. Check the local inference data path

  1. Equip and Load a WebLLM model (e.g. Llama-3.2-3B).
  2. Send a test prompt in chat: Hello in 5 languages.
  3. For a controlled check, keep the same loaded page and model selected, then inspect the browser Network panel during a prompt.
  4. Confirm the prompt is handled by the local runtime and that no hosted model request is made for this generation. Other background requests may still occur.

6. Storage Management & Deleting Weights

WebLLM model shards are cached in browser Cache Storage. To inspect or manage disk usage:

Checking Storage:

Open Chrome DevTools -> Application tab -> Storage. It shows exact storage used (e.g., 3.8 GB).

Freeing Storage:

In Model Organizer -> WebLLM, click Delete Cache on any cached model card to delete its downloaded weights instantly.

7. Performance Tips for Local WebLLM

Memory Headroom

Close memory-heavy browser tabs (Figma, YouTube) when loading 7B/8B models to prevent WebGL context loss.

Context Management

Keep context length under 4096 tokens for local execution by unpinning unused attachments.

8. Gemini Nano — Chrome on-device model where available

Some Chrome builds expose a built-in on-device language model API. Availability, prerequisites, and names of experimental flags can change by Chrome version and device; the workspace shows this provider only when its capability check succeeds.

How to Enable Chrome Gemini Nano:
  1. Use a supported Chrome build on compatible hardware and check Chrome's current built-in AI documentation for the required setup.
  2. Open the Model Organizer. The Gemini Nano entry is available only when the browser reports support.
  3. If the provider is unavailable, use a supported WebLLM, Transformers.js, or local Ollama model instead.

9. WebLLM vs Gemini Nano vs Ollama Comparison

Feature Gemini Nano WebLLM (WebGPU) Ollama Local
Setup Needed Supported Chrome build and device 1-Click Browser Download Desktop Install + CORS
Download Size Managed by Chrome 1GB - 4GB in Cache 2GB - 10GB Local Disk
Context Window ~4096 tokens 4096 tokens 8k - 128k tokens
Vision Support No Yes (Phi-3.5 Vision) Yes (LLaVA / MiniCPM)

10. Offline Troubleshooting

Q:
"WebGPU not supported"

Ensure you are using Chrome 113+ or Edge 113+ on desktop. Check chrome://gpu to verify hardware acceleration.

Q:
Tab crash during model loading

Out of GPU memory. Switch from 8B model to a lighter 1B or 3B model (e.g. Llama-3.2-3B).

Q:
Gemini Nano provider card not showing Active

Verify both flags in chrome://flags are Enabled and check chrome://components status.