1. Three Different Audio Features
Dictation, quick read-aloud, and generated voiceovers use different browser APIs and have different privacy and output behavior:
Dictate prompts hands-free using browser Web Speech Recognition. Transcribes spoken words directly into the composer input field.
Listen to AI responses rendered via native browser Speech Synthesis voices with animated audio equalizer indicators.
Generate downloadable WAV speech on your device with Kokoro 82M, then use it in chat or the Decks video timeline.
2. Speech-to-Text Voice Input
Click the Microphone icon in the composer when the input field is empty (or to append text).
Voice Dictation Workflow:
- Click the Microphone icon to open the full-screen Voice Interface.
- Grant browser microphone permissions when prompted.
- Speak naturally; live transcripts update in real time.
- Click Stop or pause briefly to complete dictation and insert text into the composer.
3. Text-to-Speech Read Aloud
Hover over any assistant message and click Read Aloud (Volume icon).
- Cleaned Audio Output: System tags (
<thinking>,<image_gen>) and code block syntax markers are stripped before synthesis. - Visual Equalizer: Animated soundwave bars indicate active speech synthesis. Click Stop Aloud anytime to silence playback.
- Native OS Voices: Uses default system voices (e.g. Samantha/Alex on macOS, David/Zira on Windows).
4. Create a Voiceover from Chat
In chat, open the + menu and choose Voiceover, or ask for a voiceover in your own prompt. After the assistant finishes its reply, Kokoro creates a downloadable WAV card locally in your browser. Use Local Voice Studio when you want to paste and synthesize a script without asking the chat assistant. The first use downloads about 100 MB of model files; later sessions can reuse the browser cache.
- Choose an American or British English voice from the voiceover card.
- Preview, download, or regenerate the WAV in your browser.
- For a standalone script, open Local Voice Studio; for a video, insert speech at the playhead in Decks Studio.
The script is synthesized on your device. Model weights are downloaded from Hugging Face; the voice generator does not send the script to an AI API.
5. Full-Screen Voice Interface Modal
The full-screen Voice Interface modal features an animated microphone orb with pulsating radial waves, displaying live interim speech hypotheses for extended hands-free prompt composition.
6. Sandboxed Iframe Security & Audio Fallbacks
When Accelerated Logic AI runs inside a restricted preview iframe without explicit audio policy permissions, browsers block window.speechSynthesis.
Resolution: Click "Open in New Tab" at the top right of the workspace preview to run in a top-level browser context where Web Speech APIs operate unimpeded.
7. Hands-Free Developer Workflows
Dictate core feature requirements, allow multi-agent teams to generate application code, and listen to architectural explanations via Read Aloud while inspecting the live Web Sandbox preview pane.
8. Privacy & Data Handling
Processed by standard browser speech services (e.g. Google Cloud Speech in Chrome).
Read Aloud uses browser speech voices. Chat voiceover cards and Local Voice Studio use Kokoro model files downloaded to the browser, then generate audio on-device. Chat prompts and replies sent through a connected cloud model are governed by that selected provider.
9. Voice Troubleshooting
Microphone Denied? Click the address bar lock icon → Site Settings → Microphone → Allow.
No Sound on Read Aloud? Ensure system audio is unmuted and verify that the current tab is not muted in the browser tab bar.