GA Release

The Multi-Agent WebGPU Engine & Interactions API are now generally available.

Part 24: Keyboard Shortcuts, Agent-Directive Tags & Glossary

Every keyboard shortcut and slash command, the XML-style tags the model uses to create skills/memories and call tools, and a plain-language glossary of the whole platform.

Category: Reference • Read Time: 18 min read • Updated: August 2026

1. Keyboard & Mouse Shortcuts

ActionShortcutWhere
Send messageEnterChat composer (Shift+Enter = new line)
Quick open / file searchCtrl / ⌘ + PGitHub Workspace command palette
Cancel / closeEscModals, menus, dialogs
Pause / resume generationPause / Break keyActive agent run
Navigate suggestion list↑ / ↓ then EnterComposer autocomplete
Slash commands/fix · /comments · /optimizeCode workspaces

2. AI Directive Tags

When you ask the assistant to remember something or create a skill, it can emit a structured proposal tag in its reply. The client validates creation proposals and shows a Keep or Discard card above the composer; creation still requires the matching setting and your explicit Keep action. Deletions only run if the corresponding permission setting is enabled.

<create_skill> — proposes a reusable persona/instruction set; saved only after Keep.

<create_memory> — proposes a long-term fact; saved only after Keep and skipped if a duplicate exists.

<delete_skill> / <delete_memory> — remove entries (gated; content_sub matches containing memories).

<mcp_tool_call> — invokes a built-in or MCP tool as JSON; subject to the permission level and confirmation gate.

<image_gen prompt="…"> — requests image generation; a nudge appears if the model mentions an image but omits the tag.

3. Multi-Agent Architectures (Quick Map)

Single AgentOne equipped model handles the whole request.
Standard (Sequential)Orchestrator → workers → finalizer pipeline; requires the fixed orch/finalizer roles.
Planning (Planner–Critic)A planner drafts steps and a critic reviews before execution.
Self-DebateOpposing-view agents challenge an answer to surface flaws.
Document AnalyzerTuned for ingested-file reasoning.
Auto RouterPicks the architecture and model tier from the prompt (unit-tested heuristics).

4. Glossary

WebGPU / WASM: browser APIs for GPU-accelerated and CPU execution; they let local models and transcription run without a server.

WebLLM / Transformers.js: engines that run open-weights LLMs in the browser.

Gemini Nano: Google's small on-device model available in supported Chrome builds.

Puter.js: a free gateway that provides inference without an API key.

Ollama: a local server that runs models on your machine; requires CORS permission (OLLAMA_ORIGINS) to be called from the browser.

MCP (Model Context Protocol): a standard for connecting external tools/servers; external servers are sandboxed to permission-gated calls.

Skill: a reusable, named instruction/persona you (or the model) can save and attach.

Memory: a persisted long-term fact that shapes future responses; semantically matched to relevant prompts.

Thinking traces: the visible per-agent reasoning log streamed during a run.

Infinity Mode: loops a request through critic self-correction up to a configured attempt count.

Prompt queue: messages entered while a run is active are buffered and processed in order with backoff.

SecretStore: the internal helper that centralizes credential reads/writes and keeps keys out of cloud sync.