GA Release

The Multi-Agent WebGPU Engine & Interactions API are now generally available.

Part 21: Browser Tools — Local Voice Studio, Decks, Transcribe & More

A guided tour of the browser-based tools for Kokoro voiceovers, decks, transcription, 3D scenes, writing, and project organization.

Category: Tooling & Deployment • Read Time: 24 min read • Updated: September 2026

1. One Engine, Many Tools

Accelerated Logic is more than a chat window. Alongside the multi-agent workspace it ships browser-based studios for voiceovers, transcription, presentations, 3D, writing, and project organization. Public guide pages explain what each tool does and where its data goes.

Most AI-assisted creative tools reuse models you have equipped in the Model Organizer. Local Voice Studio is a separate Kokoro speech model, and Transcribe uses Whisper; both run in the browser after their model files download.

Tool Route What it produces Runs on
Logic Studio/logicstudioA Google-Drive-style hub that launches the other tools and organises your decks, transcripts, and notesBrowser
Decks (VIBE)/vibePitch decks, posters, stories, cards & invitations as editable HTML artboardsEquipped models
Transcribe/transcribeSpeech-to-text with timestamps, translation, and SRT/VTT/TXT/JSON exportWhisper on WebGPU/WASM
Local Voice Studio/local-ttsGenerate, preview, and download WAV voiceovers; insert generated audio into a video timelineKokoro 82M in the browser
PolyForge/polyforge3D models with GLTF / OBJ / STL exportEquipped models + WebGL
Humanizer/humanizerRewrites machine-sounding prose; 60+ local rules plus optional AI passLocal rules (+ model)

2. Logic Studio — The Hub

Logic Studio (/logicstudio) is the launcher and file manager for the creative tools. It presents a familiar cloud-drive style grid with folders, recent items, and category filters (Decks, Transcribe, PolyForge, Humanizer, Notes).

What you can do here:

  • Open any tool in an embedded, sandboxed preview frame without leaving the hub — or pop it into a standalone browser tab with the fullscreen/open buttons.
  • Create and organise notes with a rich editor (bold, italic, lists), rename/duplicate/delete items, and search everything with the command box.
  • Manage models in place: because the embedded tools share the chat engine, you can open the Model Organizer from inside the hub and pick which model drives the tools.

Note: the embedded tool frames run with a restricted sandbox and talk to the parent only through validated messages, so a tool page cannot read the chat shell or your stored keys.

3. AI Decks (VIBE) — From a Feeling to an Artboard

Decks (/vibe) turns a description of a mood, an audience, and a purpose into a real HTML artboard rather than a flat image. Supported formats include pitch decks, posters, landing pages, social stories, greeting cards, and invitations.

Why HTML, not an image?

Text stays selectable and editable, layout reflows on small screens, and you can change a colour or headline in the source without regenerating. Image models routinely misspell and distort text; emitting markup produces a true document.

Direct source editing

Every artboard exposes its own HTML. Tweak spacing, swap a palette, or delete a slide, then present full-screen. The first generation is rarely final — editing the part you like beats starting over.

Model choice matters: VIBE uses whatever you have equipped in the Model Organizer. Stronger instruction-following models produce better-structured, multi-slide layouts; small local models work but tend toward simpler compositions.

4. Transcribe — On-Device Speech to Text

Transcribe (/transcribe) runs an OpenAI Whisper model entirely in your browser using WebGPU when available and falling back to a WebAssembly (WASM) backend otherwise. Audio and video are processed on your device — nothing is uploaded for transcription.

Workflow:

  1. Drop in an audio/video file (or record). On first run the chosen Whisper model downloads and is cached.
  2. Transcription produces word/segment timestamps that map the transcript back to the media timeline.
  3. Optionally translate the transcript using your equipped chat model.
  4. Export as .srt, .vtt, .txt, or .json.

If WebGPU is unavailable, the first run may take longer while the WASM backend initialises; Safari and older Firefox use this fallback automatically.

5. Local Voice Studio — Kokoro Text to Speech

Local Voice Studio (/tools/tts.html) creates English WAV speech with Kokoro 82M through Transformers.js. The first use downloads about 100 MB of model files and voice data. Speech generation runs in your browser; the tool does not upload your script to an AI API.

Ways to use it:

  • Paste a script into Local Voice Studio, choose an American or British English voice, preview the result, and download a WAV.
  • Choose Voiceover from the chat + menu, or request a voiceover in your prompt. Once the assistant finishes, a local audio card appears on its reply with playback and WAV download controls.
  • In Decks Studio, generate speech and insert it at the video timeline playhead as an editable audio clip.

The script supports up to 12,000 characters. The browser downloads model files from Hugging Face; generated audio is 24 kHz mono WAV.

6. PolyForge — 3D Modeling in the Browser

PolyForge AI is a WebGL 3D modeling studio launched from Logic Studio. Describe or sculpt a model, preview it in an interactive 3D viewport, and export it in standard interchange formats.

.gltf / .glbWeb & game engines
.objUniversal mesh
.stl3D printing

Like the other tools, PolyForge streams AI assistance through your equipped workspace model via the shared in-page bridge, so your key configuration and provider choices carry over automatically.

7. Humanizer — Make Machine Text Read Naturally

Humanizer (/humanizer) targets the habits that make writing look generated: stacked tricolons, hedging preambles, the "it's not just X, it's Y" frame, uniform sentence length, and vocabulary clustering around words like delve, leverage, robust, seamless.

Two layers:

  • Local rule pass (free, instant): more than 60 deterministic, transparent transformations run entirely in your browser with no network request — the same input always produces the same output.
  • Optional AI rewrite: sends the passage to a model you configured (costs a request against your own key) for a deeper rewrite; you can run the rules after the AI to clean up.

A live readability score and clickable list of flagged phrases shows exactly what changed. No tool can guarantee passing a given AI detector — detectors disagree and false-positive on non-native English — but the rules reliably flatten the mechanical rhythm.

8. Where Tool Files Are Stored

Tool artifacts stay on your device. Decks, transcripts, notes, and 3D projects are saved to the browser's local storage / IndexedDB on your machine; export buttons let you download standard files (SRT, GLTF, HTML) to your own disk. Connect Google Drive from Logic Studio or the workspace if you want a cloud mirror of that work (see Part 17).