Repository files navigation

VoiceCanvas logo

VoiceCanvas

Speak a diagram into shape.

中文 · Product docs · Prototype check · Quick start · Contributing

StatusLicenseNodepnpmTypeScript

VoiceCanvas workbench

VoiceCanvas is a voice-first diagram workbench. It turns natural speech into validated graph patches, applies those patches to a Mermaid-backed canvas, and keeps every change reversible through patch history.

The project is currently an early engineering prototype. It is useful for exploring the interaction model, running the local demo, and building toward a real open-source diagram editor where the main workflow is continuous voice editing.

Why This Exists

Most diagram tools make people stop thinking and start operating the UI. Most AI diagram generators handle the first draft better than the tenth edit. VoiceCanvas focuses on the work that happens after the first shape appears: renaming nodes, adding branches, changing flows, confirming ambiguous targets, and undoing a bad patch without rebuilding the whole graph.

The long-term bet is simple: diagrams should change at the speed of a conversation, while the underlying graph remains structured, validated, and exportable.

Highlights

AreaStatusWhat it does
Voice-first command flowWorking prototypeAccepts text segments plus selectable OpenAI Realtime or Gemini Live voice input.
Patch-based editingWorking prototypeConverts commands into atomic graph operations.
Validation layerWorking prototypeChecks patch drafts before they mutate the canvas.
Mermaid rendererWorking prototypeRenders the first diagram surface with Mermaid.
Low-confidence confirmationWorking prototypeShows target candidates before applying ambiguous edits.
History and undoWorking prototypeStores applied patches and restores the previous canvas state.
Realtime voice providersOptionalSupports OpenAI Realtime through the API proxy and Gemini Live through short-lived browser tokens.
External patch compilerOptionalUses an OpenAI-compatible model endpoint when configured.
Local mock compilerBuilt inRuns the demo without model credentials.
JSON exportWorking prototypeExports the current structured canvas document.

Architecture

flowchart LR
A[Mic or text input] --> B[Voice segment]
B --> C[Patch compiler]
C --> D[Validator]
D --> E[Patch engine]
E --> F[Canvas document]
F --> G[Mermaid renderer]
E --> H[Patch history]
H --> I[Undo and export]
Loading

The model is treated as a patch planner. It can propose a draft, but the canvas only changes after the draft passes validation and the patch engine applies it. That split keeps graph state inspectable, makes undo reliable, and keeps voice recognition separate from diagram mutation.

Repository Layout

apps/
web/ React + Vite workbench
api/ Hono API, workspace state, realtime session proxy
packages/
core/ CanvasDoc model, patch engine, validator, Mermaid export
ai/ OpenAI-compatible model patch compiler adapter
eval/ Acceptance cases and metric helpers
docs/
prd/ Product, interaction, roadmap, and system design docs
skills/
voicecanvas-dev-debug-acceptance/

Quick Start

Requirements

  • Node.js 24+
  • pnpm 10+

Run the local workbench

pnpm install
cp .env.example .env
pnpm dev

Open the web app:

http://localhost:5173

The API server runs on:

http://localhost:8787

The Vite dev server proxies /api traffic to the API server. You can run the project without external credentials; empty model settings use the built-in mock patch compiler.

Configuration

Create .env from .env.example and fill only the providers you want to use.

Realtime Voice

OPENAI_API_KEY=
OPENAI_REALTIME_MODEL=gpt-realtime-2
GEMINI_API_KEY=
GEMINI_LIVE_MODEL=gemini-3.1-flash-live-preview

External Patch Compiler

PATCH_COMPILER_API_KEY=
PATCH_COMPILER_BASE_URL=
PATCH_COMPILER_MODEL=
PATCH_COMPILER_PROVIDER=

When the patch compiler variables are empty, VoiceCanvas uses the local mock compiler. That makes the demo easy to run in forks, CI, and offline experiments.

Scripts

CommandDescription
pnpm devStart the web and API apps together.
pnpm dev:webStart only the Vite app.
pnpm dev:apiStart only the Hono API server.
pnpm testRun unit tests across the workspace.
pnpm lintRun lint checks.
pnpm buildBuild apps and type-check packages.
pnpm test:e2eRun Playwright smoke tests.
pnpm check:prototypeRun the full Prototype check suite.
pnpm check:alphaRun the Alpha check suite and local eval report.
pnpm check:alpha:realtimeRun the realtime provider API tests.

API Surface

MethodPathPurpose
GET/healthAPI health check.
GET/api/canvasRead the current workspace snapshot.
POST/api/dev/resetReset the in-memory prototype workspace.
POST/api/commands/text-segmentProcess a text segment as a voice command.
POST/api/patch/compileCompile a patch draft without applying it.
POST/api/patch/applyApply a provided patch draft.
POST/api/patch/confirmConfirm a low-confidence candidate.
POST/api/patch/undoRestore the previous patch state.
GET/api/realtime/providerRead realtime voice provider settings.
POST/api/realtime/openai/sessionProxy WebRTC session offers to OpenAI Realtime.
POST/api/realtime/gemini/tokenIssue a short-lived Gemini Live browser token.
GET/api/export/jsonExport the current CanvasDoc.

Development Notes

  • Source file names use kebab-case.
  • React component exports use PascalCase.
  • React hooks live in apps/web/src/hooks and use use-*.ts.
  • Tests use *.test.ts or *.spec.ts.
  • Generated build and test output should stay out of source control.
  • Workspace packages are private packages for now, while the repository itself is MIT licensed.

Roadmap

StageFocus
PrototypeOne-shot diagram creation, simple local edits, undo, candidate confirmation, JSON export.
AlphaMore reliable continuous editing, selected-object voice commands, local layout improvements, image export.
BetaFlowchart and mind map coverage, stronger evaluation cases, shareable outputs, real user feedback loops.
LaterTeam workspaces, meeting mode, permissions, provider marketplace, richer diagram types.

See docs/prd for the longer product and technical plan.

Contributing

Issues and small focused pull requests are welcome. Before sending a change, run the checks that match your edit:

pnpm test
pnpm lint
pnpm build
pnpm test:e2e

Good first areas:

  • Improve acceptance cases in packages/eval.
  • Add more mock compiler commands in packages/core.
  • Tighten API tests around patch confirmation and undo.
  • Improve workbench interactions in apps/web.
  • Expand provider adapters in packages/ai.

License

MIT. See LICENSE.

About

Voice-first diagram canvas with validated graph patches.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

VoiceCanvas logo

VoiceCanvas

Speak a diagram into shape.

中文 · Product docs · Prototype check · Quick start · Contributing

StatusLicenseNodepnpmTypeScript

VoiceCanvas workbench

VoiceCanvas is a voice-first diagram workbench. It turns natural speech into validated graph patches, applies those patches to a Mermaid-backed canvas, and keeps every change reversible through patch history.

The project is currently an early engineering prototype. It is useful for exploring the interaction model, running the local demo, and building toward a real open-source diagram editor where the main workflow is continuous voice editing.

Why This Exists

Most diagram tools make people stop thinking and start operating the UI. Most AI diagram generators handle the first draft better than the tenth edit. VoiceCanvas focuses on the work that happens after the first shape appears: renaming nodes, adding branches, changing flows, confirming ambiguous targets, and undoing a bad patch without rebuilding the whole graph.

The long-term bet is simple: diagrams should change at the speed of a conversation, while the underlying graph remains structured, validated, and exportable.

Highlights

AreaStatusWhat it does
Voice-first command flowWorking prototypeAccepts text segments plus selectable OpenAI Realtime or Gemini Live voice input.
Patch-based editingWorking prototypeConverts commands into atomic graph operations.
Validation layerWorking prototypeChecks patch drafts before they mutate the canvas.
Mermaid rendererWorking prototypeRenders the first diagram surface with Mermaid.
Low-confidence confirmationWorking prototypeShows target candidates before applying ambiguous edits.
History and undoWorking prototypeStores applied patches and restores the previous canvas state.
Realtime voice providersOptionalSupports OpenAI Realtime through the API proxy and Gemini Live through short-lived browser tokens.
External patch compilerOptionalUses an OpenAI-compatible model endpoint when configured.
Local mock compilerBuilt inRuns the demo without model credentials.
JSON exportWorking prototypeExports the current structured canvas document.

Architecture

flowchart LR
A[Mic or text input] --> B[Voice segment]
B --> C[Patch compiler]
C --> D[Validator]
D --> E[Patch engine]
E --> F[Canvas document]
F --> G[Mermaid renderer]
E --> H[Patch history]
H --> I[Undo and export]
Loading

The model is treated as a patch planner. It can propose a draft, but the canvas only changes after the draft passes validation and the patch engine applies it. That split keeps graph state inspectable, makes undo reliable, and keeps voice recognition separate from diagram mutation.

Repository Layout

apps/
web/ React + Vite workbench
api/ Hono API, workspace state, realtime session proxy
packages/
core/ CanvasDoc model, patch engine, validator, Mermaid export
ai/ OpenAI-compatible model patch compiler adapter
eval/ Acceptance cases and metric helpers
docs/
prd/ Product, interaction, roadmap, and system design docs
skills/
voicecanvas-dev-debug-acceptance/

Quick Start

Requirements

  • Node.js 24+
  • pnpm 10+

Run the local workbench

pnpm install
cp .env.example .env
pnpm dev

Open the web app:

http://localhost:5173

The API server runs on:

http://localhost:8787

The Vite dev server proxies /api traffic to the API server. You can run the project without external credentials; empty model settings use the built-in mock patch compiler.

Configuration

Create .env from .env.example and fill only the providers you want to use.

Realtime Voice

OPENAI_API_KEY=
OPENAI_REALTIME_MODEL=gpt-realtime-2
GEMINI_API_KEY=
GEMINI_LIVE_MODEL=gemini-3.1-flash-live-preview

External Patch Compiler

PATCH_COMPILER_API_KEY=
PATCH_COMPILER_BASE_URL=
PATCH_COMPILER_MODEL=
PATCH_COMPILER_PROVIDER=

When the patch compiler variables are empty, VoiceCanvas uses the local mock compiler. That makes the demo easy to run in forks, CI, and offline experiments.

Scripts

CommandDescription
pnpm devStart the web and API apps together.
pnpm dev:webStart only the Vite app.
pnpm dev:apiStart only the Hono API server.
pnpm testRun unit tests across the workspace.
pnpm lintRun lint checks.
pnpm buildBuild apps and type-check packages.
pnpm test:e2eRun Playwright smoke tests.
pnpm check:prototypeRun the full Prototype check suite.
pnpm check:alphaRun the Alpha check suite and local eval report.
pnpm check:alpha:realtimeRun the realtime provider API tests.

API Surface

MethodPathPurpose
GET/healthAPI health check.
GET/api/canvasRead the current workspace snapshot.
POST/api/dev/resetReset the in-memory prototype workspace.
POST/api/commands/text-segmentProcess a text segment as a voice command.
POST/api/patch/compileCompile a patch draft without applying it.
POST/api/patch/applyApply a provided patch draft.
POST/api/patch/confirmConfirm a low-confidence candidate.
POST/api/patch/undoRestore the previous patch state.
GET/api/realtime/providerRead realtime voice provider settings.
POST/api/realtime/openai/sessionProxy WebRTC session offers to OpenAI Realtime.
POST/api/realtime/gemini/tokenIssue a short-lived Gemini Live browser token.
GET/api/export/jsonExport the current CanvasDoc.

Development Notes

  • Source file names use kebab-case.
  • React component exports use PascalCase.
  • React hooks live in apps/web/src/hooks and use use-*.ts.
  • Tests use *.test.ts or *.spec.ts.
  • Generated build and test output should stay out of source control.
  • Workspace packages are private packages for now, while the repository itself is MIT licensed.

Roadmap

StageFocus
PrototypeOne-shot diagram creation, simple local edits, undo, candidate confirmation, JSON export.
AlphaMore reliable continuous editing, selected-object voice commands, local layout improvements, image export.
BetaFlowchart and mind map coverage, stronger evaluation cases, shareable outputs, real user feedback loops.
LaterTeam workspaces, meeting mode, permissions, provider marketplace, richer diagram types.

See docs/prd for the longer product and technical plan.

Contributing

Issues and small focused pull requests are welcome. Before sending a change, run the checks that match your edit:

pnpm test
pnpm lint
pnpm build
pnpm test:e2e

Good first areas:

  • Improve acceptance cases in packages/eval.
  • Add more mock compiler commands in packages/core.
  • Tighten API tests around patch confirmation and undo.
  • Improve workbench interactions in apps/web.
  • Expand provider adapters in packages/ai.

License

MIT. See LICENSE.

About

Voice-first diagram canvas with validated graph patches.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

VoiceCanvas logo

VoiceCanvas

Speak a diagram into shape.

中文 · Product docs · Prototype check · Quick start · Contributing

StatusLicenseNodepnpmTypeScript

VoiceCanvas workbench

VoiceCanvas is a voice-first diagram workbench. It turns natural speech into validated graph patches, applies those patches to a Mermaid-backed canvas, and keeps every change reversible through patch history.

The project is currently an early engineering prototype. It is useful for exploring the interaction model, running the local demo, and building toward a real open-source diagram editor where the main workflow is continuous voice editing.

Why This Exists

Most diagram tools make people stop thinking and start operating the UI. Most AI diagram generators handle the first draft better than the tenth edit. VoiceCanvas focuses on the work that happens after the first shape appears: renaming nodes, adding branches, changing flows, confirming ambiguous targets, and undoing a bad patch without rebuilding the whole graph.

The long-term bet is simple: diagrams should change at the speed of a conversation, while the underlying graph remains structured, validated, and exportable.

Highlights

AreaStatusWhat it does
Voice-first command flowWorking prototypeAccepts text segments plus selectable OpenAI Realtime or Gemini Live voice input.
Patch-based editingWorking prototypeConverts commands into atomic graph operations.
Validation layerWorking prototypeChecks patch drafts before they mutate the canvas.
Mermaid rendererWorking prototypeRenders the first diagram surface with Mermaid.
Low-confidence confirmationWorking prototypeShows target candidates before applying ambiguous edits.
History and undoWorking prototypeStores applied patches and restores the previous canvas state.
Realtime voice providersOptionalSupports OpenAI Realtime through the API proxy and Gemini Live through short-lived browser tokens.
External patch compilerOptionalUses an OpenAI-compatible model endpoint when configured.
Local mock compilerBuilt inRuns the demo without model credentials.
JSON exportWorking prototypeExports the current structured canvas document.

Architecture

flowchart LR
A[Mic or text input] --> B[Voice segment]
B --> C[Patch compiler]
C --> D[Validator]
D --> E[Patch engine]
E --> F[Canvas document]
F --> G[Mermaid renderer]
E --> H[Patch history]
H --> I[Undo and export]
Loading

The model is treated as a patch planner. It can propose a draft, but the canvas only changes after the draft passes validation and the patch engine applies it. That split keeps graph state inspectable, makes undo reliable, and keeps voice recognition separate from diagram mutation.

Repository Layout

apps/
web/ React + Vite workbench
api/ Hono API, workspace state, realtime session proxy
packages/
core/ CanvasDoc model, patch engine, validator, Mermaid export
ai/ OpenAI-compatible model patch compiler adapter
eval/ Acceptance cases and metric helpers
docs/
prd/ Product, interaction, roadmap, and system design docs
skills/
voicecanvas-dev-debug-acceptance/

Quick Start

Requirements

  • Node.js 24+
  • pnpm 10+

Run the local workbench

pnpm install
cp .env.example .env
pnpm dev

Open the web app:

http://localhost:5173

The API server runs on:

http://localhost:8787

The Vite dev server proxies /api traffic to the API server. You can run the project without external credentials; empty model settings use the built-in mock patch compiler.

Configuration

Create .env from .env.example and fill only the providers you want to use.

Realtime Voice

OPENAI_API_KEY=
OPENAI_REALTIME_MODEL=gpt-realtime-2
GEMINI_API_KEY=
GEMINI_LIVE_MODEL=gemini-3.1-flash-live-preview

External Patch Compiler

PATCH_COMPILER_API_KEY=
PATCH_COMPILER_BASE_URL=
PATCH_COMPILER_MODEL=
PATCH_COMPILER_PROVIDER=

When the patch compiler variables are empty, VoiceCanvas uses the local mock compiler. That makes the demo easy to run in forks, CI, and offline experiments.

Scripts

CommandDescription
pnpm devStart the web and API apps together.
pnpm dev:webStart only the Vite app.
pnpm dev:apiStart only the Hono API server.
pnpm testRun unit tests across the workspace.
pnpm lintRun lint checks.
pnpm buildBuild apps and type-check packages.
pnpm test:e2eRun Playwright smoke tests.
pnpm check:prototypeRun the full Prototype check suite.
pnpm check:alphaRun the Alpha check suite and local eval report.
pnpm check:alpha:realtimeRun the realtime provider API tests.

API Surface

MethodPathPurpose
GET/healthAPI health check.
GET/api/canvasRead the current workspace snapshot.
POST/api/dev/resetReset the in-memory prototype workspace.
POST/api/commands/text-segmentProcess a text segment as a voice command.
POST/api/patch/compileCompile a patch draft without applying it.
POST/api/patch/applyApply a provided patch draft.
POST/api/patch/confirmConfirm a low-confidence candidate.
POST/api/patch/undoRestore the previous patch state.
GET/api/realtime/providerRead realtime voice provider settings.
POST/api/realtime/openai/sessionProxy WebRTC session offers to OpenAI Realtime.
POST/api/realtime/gemini/tokenIssue a short-lived Gemini Live browser token.
GET/api/export/jsonExport the current CanvasDoc.

Development Notes

  • Source file names use kebab-case.
  • React component exports use PascalCase.
  • React hooks live in apps/web/src/hooks and use use-*.ts.
  • Tests use *.test.ts or *.spec.ts.
  • Generated build and test output should stay out of source control.
  • Workspace packages are private packages for now, while the repository itself is MIT licensed.

Roadmap

StageFocus
PrototypeOne-shot diagram creation, simple local edits, undo, candidate confirmation, JSON export.
AlphaMore reliable continuous editing, selected-object voice commands, local layout improvements, image export.
BetaFlowchart and mind map coverage, stronger evaluation cases, shareable outputs, real user feedback loops.
LaterTeam workspaces, meeting mode, permissions, provider marketplace, richer diagram types.

See docs/prd for the longer product and technical plan.

Contributing

Issues and small focused pull requests are welcome. Before sending a change, run the checks that match your edit:

pnpm test
pnpm lint
pnpm build
pnpm test:e2e

Good first areas:

  • Improve acceptance cases in packages/eval.
  • Add more mock compiler commands in packages/core.
  • Tighten API tests around patch confirmation and undo.
  • Improve workbench interactions in apps/web.
  • Expand provider adapters in packages/ai.

License

MIT. See LICENSE.

About

Voice-first diagram canvas with validated graph patches.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

VoiceCanvas logo

VoiceCanvas

Speak a diagram into shape.

中文 · Product docs · Prototype check · Quick start · Contributing

StatusLicenseNodepnpmTypeScript

VoiceCanvas workbench

VoiceCanvas is a voice-first diagram workbench. It turns natural speech into validated graph patches, applies those patches to a Mermaid-backed canvas, and keeps every change reversible through patch history.

The project is currently an early engineering prototype. It is useful for exploring the interaction model, running the local demo, and building toward a real open-source diagram editor where the main workflow is continuous voice editing.

Why This Exists

Most diagram tools make people stop thinking and start operating the UI. Most AI diagram generators handle the first draft better than the tenth edit. VoiceCanvas focuses on the work that happens after the first shape appears: renaming nodes, adding branches, changing flows, confirming ambiguous targets, and undoing a bad patch without rebuilding the whole graph.

The long-term bet is simple: diagrams should change at the speed of a conversation, while the underlying graph remains structured, validated, and exportable.

Highlights

AreaStatusWhat it does
Voice-first command flowWorking prototypeAccepts text segments plus selectable OpenAI Realtime or Gemini Live voice input.
Patch-based editingWorking prototypeConverts commands into atomic graph operations.
Validation layerWorking prototypeChecks patch drafts before they mutate the canvas.
Mermaid rendererWorking prototypeRenders the first diagram surface with Mermaid.
Low-confidence confirmationWorking prototypeShows target candidates before applying ambiguous edits.
History and undoWorking prototypeStores applied patches and restores the previous canvas state.
Realtime voice providersOptionalSupports OpenAI Realtime through the API proxy and Gemini Live through short-lived browser tokens.
External patch compilerOptionalUses an OpenAI-compatible model endpoint when configured.
Local mock compilerBuilt inRuns the demo without model credentials.
JSON exportWorking prototypeExports the current structured canvas document.

Architecture

flowchart LR
A[Mic or text input] --> B[Voice segment]
B --> C[Patch compiler]
C --> D[Validator]
D --> E[Patch engine]
E --> F[Canvas document]
F --> G[Mermaid renderer]
E --> H[Patch history]
H --> I[Undo and export]
Loading

The model is treated as a patch planner. It can propose a draft, but the canvas only changes after the draft passes validation and the patch engine applies it. That split keeps graph state inspectable, makes undo reliable, and keeps voice recognition separate from diagram mutation.

Repository Layout

apps/
web/ React + Vite workbench
api/ Hono API, workspace state, realtime session proxy
packages/
core/ CanvasDoc model, patch engine, validator, Mermaid export
ai/ OpenAI-compatible model patch compiler adapter
eval/ Acceptance cases and metric helpers
docs/
prd/ Product, interaction, roadmap, and system design docs
skills/
voicecanvas-dev-debug-acceptance/

Quick Start

Requirements

  • Node.js 24+
  • pnpm 10+

Run the local workbench

pnpm install
cp .env.example .env
pnpm dev

Open the web app:

http://localhost:5173

The API server runs on:

http://localhost:8787

The Vite dev server proxies /api traffic to the API server. You can run the project without external credentials; empty model settings use the built-in mock patch compiler.

Configuration

Create .env from .env.example and fill only the providers you want to use.

Realtime Voice

OPENAI_API_KEY=
OPENAI_REALTIME_MODEL=gpt-realtime-2
GEMINI_API_KEY=
GEMINI_LIVE_MODEL=gemini-3.1-flash-live-preview

External Patch Compiler

PATCH_COMPILER_API_KEY=
PATCH_COMPILER_BASE_URL=
PATCH_COMPILER_MODEL=
PATCH_COMPILER_PROVIDER=

When the patch compiler variables are empty, VoiceCanvas uses the local mock compiler. That makes the demo easy to run in forks, CI, and offline experiments.

Scripts

CommandDescription
pnpm devStart the web and API apps together.
pnpm dev:webStart only the Vite app.
pnpm dev:apiStart only the Hono API server.
pnpm testRun unit tests across the workspace.
pnpm lintRun lint checks.
pnpm buildBuild apps and type-check packages.
pnpm test:e2eRun Playwright smoke tests.
pnpm check:prototypeRun the full Prototype check suite.
pnpm check:alphaRun the Alpha check suite and local eval report.
pnpm check:alpha:realtimeRun the realtime provider API tests.

API Surface

MethodPathPurpose
GET/healthAPI health check.
GET/api/canvasRead the current workspace snapshot.
POST/api/dev/resetReset the in-memory prototype workspace.
POST/api/commands/text-segmentProcess a text segment as a voice command.
POST/api/patch/compileCompile a patch draft without applying it.
POST/api/patch/applyApply a provided patch draft.
POST/api/patch/confirmConfirm a low-confidence candidate.
POST/api/patch/undoRestore the previous patch state.
GET/api/realtime/providerRead realtime voice provider settings.
POST/api/realtime/openai/sessionProxy WebRTC session offers to OpenAI Realtime.
POST/api/realtime/gemini/tokenIssue a short-lived Gemini Live browser token.
GET/api/export/jsonExport the current CanvasDoc.

Development Notes

  • Source file names use kebab-case.
  • React component exports use PascalCase.
  • React hooks live in apps/web/src/hooks and use use-*.ts.
  • Tests use *.test.ts or *.spec.ts.
  • Generated build and test output should stay out of source control.
  • Workspace packages are private packages for now, while the repository itself is MIT licensed.

Roadmap

StageFocus
PrototypeOne-shot diagram creation, simple local edits, undo, candidate confirmation, JSON export.
AlphaMore reliable continuous editing, selected-object voice commands, local layout improvements, image export.
BetaFlowchart and mind map coverage, stronger evaluation cases, shareable outputs, real user feedback loops.
LaterTeam workspaces, meeting mode, permissions, provider marketplace, richer diagram types.

See docs/prd for the longer product and technical plan.

Contributing

Issues and small focused pull requests are welcome. Before sending a change, run the checks that match your edit:

pnpm test
pnpm lint
pnpm build
pnpm test:e2e

Good first areas:

  • Improve acceptance cases in packages/eval.
  • Add more mock compiler commands in packages/core.
  • Tighten API tests around patch confirmation and undo.
  • Improve workbench interactions in apps/web.
  • Expand provider adapters in packages/ai.

License

MIT. See LICENSE.

About

Voice-first diagram canvas with validated graph patches.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

VoiceCanvas logo

VoiceCanvas

Speak a diagram into shape.

中文 · Product docs · Prototype check · Quick start · Contributing

StatusLicenseNodepnpmTypeScript

VoiceCanvas workbench

VoiceCanvas is a voice-first diagram workbench. It turns natural speech into validated graph patches, applies those patches to a Mermaid-backed canvas, and keeps every change reversible through patch history.

The project is currently an early engineering prototype. It is useful for exploring the interaction model, running the local demo, and building toward a real open-source diagram editor where the main workflow is continuous voice editing.

Why This Exists

Most diagram tools make people stop thinking and start operating the UI. Most AI diagram generators handle the first draft better than the tenth edit. VoiceCanvas focuses on the work that happens after the first shape appears: renaming nodes, adding branches, changing flows, confirming ambiguous targets, and undoing a bad patch without rebuilding the whole graph.

The long-term bet is simple: diagrams should change at the speed of a conversation, while the underlying graph remains structured, validated, and exportable.

Highlights

AreaStatusWhat it does
Voice-first command flowWorking prototypeAccepts text segments plus selectable OpenAI Realtime or Gemini Live voice input.
Patch-based editingWorking prototypeConverts commands into atomic graph operations.
Validation layerWorking prototypeChecks patch drafts before they mutate the canvas.
Mermaid rendererWorking prototypeRenders the first diagram surface with Mermaid.
Low-confidence confirmationWorking prototypeShows target candidates before applying ambiguous edits.
History and undoWorking prototypeStores applied patches and restores the previous canvas state.
Realtime voice providersOptionalSupports OpenAI Realtime through the API proxy and Gemini Live through short-lived browser tokens.
External patch compilerOptionalUses an OpenAI-compatible model endpoint when configured.
Local mock compilerBuilt inRuns the demo without model credentials.
JSON exportWorking prototypeExports the current structured canvas document.

Architecture

flowchart LR
A[Mic or text input] --> B[Voice segment]
B --> C[Patch compiler]
C --> D[Validator]
D --> E[Patch engine]
E --> F[Canvas document]
F --> G[Mermaid renderer]
E --> H[Patch history]
H --> I[Undo and export]
Loading

The model is treated as a patch planner. It can propose a draft, but the canvas only changes after the draft passes validation and the patch engine applies it. That split keeps graph state inspectable, makes undo reliable, and keeps voice recognition separate from diagram mutation.

Repository Layout

apps/
web/ React + Vite workbench
api/ Hono API, workspace state, realtime session proxy
packages/
core/ CanvasDoc model, patch engine, validator, Mermaid export
ai/ OpenAI-compatible model patch compiler adapter
eval/ Acceptance cases and metric helpers
docs/
prd/ Product, interaction, roadmap, and system design docs
skills/
voicecanvas-dev-debug-acceptance/

Quick Start

Requirements

  • Node.js 24+
  • pnpm 10+

Run the local workbench

pnpm install
cp .env.example .env
pnpm dev

Open the web app:

http://localhost:5173

The API server runs on:

http://localhost:8787

The Vite dev server proxies /api traffic to the API server. You can run the project without external credentials; empty model settings use the built-in mock patch compiler.

Configuration

Create .env from .env.example and fill only the providers you want to use.

Realtime Voice

OPENAI_API_KEY=
OPENAI_REALTIME_MODEL=gpt-realtime-2
GEMINI_API_KEY=
GEMINI_LIVE_MODEL=gemini-3.1-flash-live-preview

External Patch Compiler

PATCH_COMPILER_API_KEY=
PATCH_COMPILER_BASE_URL=
PATCH_COMPILER_MODEL=
PATCH_COMPILER_PROVIDER=

When the patch compiler variables are empty, VoiceCanvas uses the local mock compiler. That makes the demo easy to run in forks, CI, and offline experiments.

Scripts

CommandDescription
pnpm devStart the web and API apps together.
pnpm dev:webStart only the Vite app.
pnpm dev:apiStart only the Hono API server.
pnpm testRun unit tests across the workspace.
pnpm lintRun lint checks.
pnpm buildBuild apps and type-check packages.
pnpm test:e2eRun Playwright smoke tests.
pnpm check:prototypeRun the full Prototype check suite.
pnpm check:alphaRun the Alpha check suite and local eval report.
pnpm check:alpha:realtimeRun the realtime provider API tests.

API Surface

MethodPathPurpose
GET/healthAPI health check.
GET/api/canvasRead the current workspace snapshot.
POST/api/dev/resetReset the in-memory prototype workspace.
POST/api/commands/text-segmentProcess a text segment as a voice command.
POST/api/patch/compileCompile a patch draft without applying it.
POST/api/patch/applyApply a provided patch draft.
POST/api/patch/confirmConfirm a low-confidence candidate.
POST/api/patch/undoRestore the previous patch state.
GET/api/realtime/providerRead realtime voice provider settings.
POST/api/realtime/openai/sessionProxy WebRTC session offers to OpenAI Realtime.
POST/api/realtime/gemini/tokenIssue a short-lived Gemini Live browser token.
GET/api/export/jsonExport the current CanvasDoc.

Development Notes

  • Source file names use kebab-case.
  • React component exports use PascalCase.
  • React hooks live in apps/web/src/hooks and use use-*.ts.
  • Tests use *.test.ts or *.spec.ts.
  • Generated build and test output should stay out of source control.
  • Workspace packages are private packages for now, while the repository itself is MIT licensed.

Roadmap

StageFocus
PrototypeOne-shot diagram creation, simple local edits, undo, candidate confirmation, JSON export.
AlphaMore reliable continuous editing, selected-object voice commands, local layout improvements, image export.
BetaFlowchart and mind map coverage, stronger evaluation cases, shareable outputs, real user feedback loops.
LaterTeam workspaces, meeting mode, permissions, provider marketplace, richer diagram types.

See docs/prd for the longer product and technical plan.

Contributing

Issues and small focused pull requests are welcome. Before sending a change, run the checks that match your edit:

pnpm test
pnpm lint
pnpm build
pnpm test:e2e

Good first areas:

  • Improve acceptance cases in packages/eval.
  • Add more mock compiler commands in packages/core.
  • Tighten API tests around patch confirmation and undo.
  • Improve workbench interactions in apps/web.
  • Expand provider adapters in packages/ai.

License

MIT. See LICENSE.

About

Voice-first diagram canvas with validated graph patches.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

VoiceCanvas logo

VoiceCanvas

Speak a diagram into shape.

中文 · Product docs · Prototype check · Quick start · Contributing

StatusLicenseNodepnpmTypeScript

VoiceCanvas workbench

VoiceCanvas is a voice-first diagram workbench. It turns natural speech into validated graph patches, applies those patches to a Mermaid-backed canvas, and keeps every change reversible through patch history.

The project is currently an early engineering prototype. It is useful for exploring the interaction model, running the local demo, and building toward a real open-source diagram editor where the main workflow is continuous voice editing.

Why This Exists

Most diagram tools make people stop thinking and start operating the UI. Most AI diagram generators handle the first draft better than the tenth edit. VoiceCanvas focuses on the work that happens after the first shape appears: renaming nodes, adding branches, changing flows, confirming ambiguous targets, and undoing a bad patch without rebuilding the whole graph.

The long-term bet is simple: diagrams should change at the speed of a conversation, while the underlying graph remains structured, validated, and exportable.

Highlights

AreaStatusWhat it does
Voice-first command flowWorking prototypeAccepts text segments plus selectable OpenAI Realtime or Gemini Live voice input.
Patch-based editingWorking prototypeConverts commands into atomic graph operations.
Validation layerWorking prototypeChecks patch drafts before they mutate the canvas.
Mermaid rendererWorking prototypeRenders the first diagram surface with Mermaid.
Low-confidence confirmationWorking prototypeShows target candidates before applying ambiguous edits.
History and undoWorking prototypeStores applied patches and restores the previous canvas state.
Realtime voice providersOptionalSupports OpenAI Realtime through the API proxy and Gemini Live through short-lived browser tokens.
External patch compilerOptionalUses an OpenAI-compatible model endpoint when configured.
Local mock compilerBuilt inRuns the demo without model credentials.
JSON exportWorking prototypeExports the current structured canvas document.

Architecture

flowchart LR
A[Mic or text input] --> B[Voice segment]
B --> C[Patch compiler]
C --> D[Validator]
D --> E[Patch engine]
E --> F[Canvas document]
F --> G[Mermaid renderer]
E --> H[Patch history]
H --> I[Undo and export]
Loading

The model is treated as a patch planner. It can propose a draft, but the canvas only changes after the draft passes validation and the patch engine applies it. That split keeps graph state inspectable, makes undo reliable, and keeps voice recognition separate from diagram mutation.

Repository Layout

apps/
web/ React + Vite workbench
api/ Hono API, workspace state, realtime session proxy
packages/
core/ CanvasDoc model, patch engine, validator, Mermaid export
ai/ OpenAI-compatible model patch compiler adapter
eval/ Acceptance cases and metric helpers
docs/
prd/ Product, interaction, roadmap, and system design docs
skills/
voicecanvas-dev-debug-acceptance/

Quick Start

Requirements

  • Node.js 24+
  • pnpm 10+

Run the local workbench

pnpm install
cp .env.example .env
pnpm dev

Open the web app:

http://localhost:5173

The API server runs on:

http://localhost:8787

The Vite dev server proxies /api traffic to the API server. You can run the project without external credentials; empty model settings use the built-in mock patch compiler.

Configuration

Create .env from .env.example and fill only the providers you want to use.

Realtime Voice

OPENAI_API_KEY=
OPENAI_REALTIME_MODEL=gpt-realtime-2
GEMINI_API_KEY=
GEMINI_LIVE_MODEL=gemini-3.1-flash-live-preview

External Patch Compiler

PATCH_COMPILER_API_KEY=
PATCH_COMPILER_BASE_URL=
PATCH_COMPILER_MODEL=
PATCH_COMPILER_PROVIDER=

When the patch compiler variables are empty, VoiceCanvas uses the local mock compiler. That makes the demo easy to run in forks, CI, and offline experiments.

Scripts

CommandDescription
pnpm devStart the web and API apps together.
pnpm dev:webStart only the Vite app.
pnpm dev:apiStart only the Hono API server.
pnpm testRun unit tests across the workspace.
pnpm lintRun lint checks.
pnpm buildBuild apps and type-check packages.
pnpm test:e2eRun Playwright smoke tests.
pnpm check:prototypeRun the full Prototype check suite.
pnpm check:alphaRun the Alpha check suite and local eval report.
pnpm check:alpha:realtimeRun the realtime provider API tests.

API Surface

MethodPathPurpose
GET/healthAPI health check.
GET/api/canvasRead the current workspace snapshot.
POST/api/dev/resetReset the in-memory prototype workspace.
POST/api/commands/text-segmentProcess a text segment as a voice command.
POST/api/patch/compileCompile a patch draft without applying it.
POST/api/patch/applyApply a provided patch draft.
POST/api/patch/confirmConfirm a low-confidence candidate.
POST/api/patch/undoRestore the previous patch state.
GET/api/realtime/providerRead realtime voice provider settings.
POST/api/realtime/openai/sessionProxy WebRTC session offers to OpenAI Realtime.
POST/api/realtime/gemini/tokenIssue a short-lived Gemini Live browser token.
GET/api/export/jsonExport the current CanvasDoc.

Development Notes

  • Source file names use kebab-case.
  • React component exports use PascalCase.
  • React hooks live in apps/web/src/hooks and use use-*.ts.
  • Tests use *.test.ts or *.spec.ts.
  • Generated build and test output should stay out of source control.
  • Workspace packages are private packages for now, while the repository itself is MIT licensed.

Roadmap

StageFocus
PrototypeOne-shot diagram creation, simple local edits, undo, candidate confirmation, JSON export.
AlphaMore reliable continuous editing, selected-object voice commands, local layout improvements, image export.
BetaFlowchart and mind map coverage, stronger evaluation cases, shareable outputs, real user feedback loops.
LaterTeam workspaces, meeting mode, permissions, provider marketplace, richer diagram types.

See docs/prd for the longer product and technical plan.

Contributing

Issues and small focused pull requests are welcome. Before sending a change, run the checks that match your edit:

pnpm test
pnpm lint
pnpm build
pnpm test:e2e

Good first areas:

  • Improve acceptance cases in packages/eval.
  • Add more mock compiler commands in packages/core.
  • Tighten API tests around patch confirmation and undo.
  • Improve workbench interactions in apps/web.
  • Expand provider adapters in packages/ai.

License

MIT. See LICENSE.

About

Voice-first diagram canvas with validated graph patches.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

VoiceCanvas logo

VoiceCanvas

Speak a diagram into shape.

中文 · Product docs · Prototype check · Quick start · Contributing

StatusLicenseNodepnpmTypeScript

VoiceCanvas workbench

VoiceCanvas is a voice-first diagram workbench. It turns natural speech into validated graph patches, applies those patches to a Mermaid-backed canvas, and keeps every change reversible through patch history.

The project is currently an early engineering prototype. It is useful for exploring the interaction model, running the local demo, and building toward a real open-source diagram editor where the main workflow is continuous voice editing.

Why This Exists

Most diagram tools make people stop thinking and start operating the UI. Most AI diagram generators handle the first draft better than the tenth edit. VoiceCanvas focuses on the work that happens after the first shape appears: renaming nodes, adding branches, changing flows, confirming ambiguous targets, and undoing a bad patch without rebuilding the whole graph.

The long-term bet is simple: diagrams should change at the speed of a conversation, while the underlying graph remains structured, validated, and exportable.

Highlights

AreaStatusWhat it does
Voice-first command flowWorking prototypeAccepts text segments plus selectable OpenAI Realtime or Gemini Live voice input.
Patch-based editingWorking prototypeConverts commands into atomic graph operations.
Validation layerWorking prototypeChecks patch drafts before they mutate the canvas.
Mermaid rendererWorking prototypeRenders the first diagram surface with Mermaid.
Low-confidence confirmationWorking prototypeShows target candidates before applying ambiguous edits.
History and undoWorking prototypeStores applied patches and restores the previous canvas state.
Realtime voice providersOptionalSupports OpenAI Realtime through the API proxy and Gemini Live through short-lived browser tokens.
External patch compilerOptionalUses an OpenAI-compatible model endpoint when configured.
Local mock compilerBuilt inRuns the demo without model credentials.
JSON exportWorking prototypeExports the current structured canvas document.

Architecture

flowchart LR
A[Mic or text input] --> B[Voice segment]
B --> C[Patch compiler]
C --> D[Validator]
D --> E[Patch engine]
E --> F[Canvas document]
F --> G[Mermaid renderer]
E --> H[Patch history]
H --> I[Undo and export]
Loading

The model is treated as a patch planner. It can propose a draft, but the canvas only changes after the draft passes validation and the patch engine applies it. That split keeps graph state inspectable, makes undo reliable, and keeps voice recognition separate from diagram mutation.

Repository Layout

apps/
web/ React + Vite workbench
api/ Hono API, workspace state, realtime session proxy
packages/
core/ CanvasDoc model, patch engine, validator, Mermaid export
ai/ OpenAI-compatible model patch compiler adapter
eval/ Acceptance cases and metric helpers
docs/
prd/ Product, interaction, roadmap, and system design docs
skills/
voicecanvas-dev-debug-acceptance/

Quick Start

Requirements

  • Node.js 24+
  • pnpm 10+

Run the local workbench

pnpm install
cp .env.example .env
pnpm dev

Open the web app:

http://localhost:5173

The API server runs on:

http://localhost:8787

The Vite dev server proxies /api traffic to the API server. You can run the project without external credentials; empty model settings use the built-in mock patch compiler.

Configuration

Create .env from .env.example and fill only the providers you want to use.

Realtime Voice

OPENAI_API_KEY=
OPENAI_REALTIME_MODEL=gpt-realtime-2
GEMINI_API_KEY=
GEMINI_LIVE_MODEL=gemini-3.1-flash-live-preview

External Patch Compiler

PATCH_COMPILER_API_KEY=
PATCH_COMPILER_BASE_URL=
PATCH_COMPILER_MODEL=
PATCH_COMPILER_PROVIDER=

When the patch compiler variables are empty, VoiceCanvas uses the local mock compiler. That makes the demo easy to run in forks, CI, and offline experiments.

Scripts

CommandDescription
pnpm devStart the web and API apps together.
pnpm dev:webStart only the Vite app.
pnpm dev:apiStart only the Hono API server.
pnpm testRun unit tests across the workspace.
pnpm lintRun lint checks.
pnpm buildBuild apps and type-check packages.
pnpm test:e2eRun Playwright smoke tests.
pnpm check:prototypeRun the full Prototype check suite.
pnpm check:alphaRun the Alpha check suite and local eval report.
pnpm check:alpha:realtimeRun the realtime provider API tests.

API Surface

MethodPathPurpose
GET/healthAPI health check.
GET/api/canvasRead the current workspace snapshot.
POST/api/dev/resetReset the in-memory prototype workspace.
POST/api/commands/text-segmentProcess a text segment as a voice command.
POST/api/patch/compileCompile a patch draft without applying it.
POST/api/patch/applyApply a provided patch draft.
POST/api/patch/confirmConfirm a low-confidence candidate.
POST/api/patch/undoRestore the previous patch state.
GET/api/realtime/providerRead realtime voice provider settings.
POST/api/realtime/openai/sessionProxy WebRTC session offers to OpenAI Realtime.
POST/api/realtime/gemini/tokenIssue a short-lived Gemini Live browser token.
GET/api/export/jsonExport the current CanvasDoc.

Development Notes

  • Source file names use kebab-case.
  • React component exports use PascalCase.
  • React hooks live in apps/web/src/hooks and use use-*.ts.
  • Tests use *.test.ts or *.spec.ts.
  • Generated build and test output should stay out of source control.
  • Workspace packages are private packages for now, while the repository itself is MIT licensed.

Roadmap

StageFocus
PrototypeOne-shot diagram creation, simple local edits, undo, candidate confirmation, JSON export.
AlphaMore reliable continuous editing, selected-object voice commands, local layout improvements, image export.
BetaFlowchart and mind map coverage, stronger evaluation cases, shareable outputs, real user feedback loops.
LaterTeam workspaces, meeting mode, permissions, provider marketplace, richer diagram types.

See docs/prd for the longer product and technical plan.

Contributing

Issues and small focused pull requests are welcome. Before sending a change, run the checks that match your edit:

pnpm test
pnpm lint
pnpm build
pnpm test:e2e

Good first areas:

  • Improve acceptance cases in packages/eval.
  • Add more mock compiler commands in packages/core.
  • Tighten API tests around patch confirmation and undo.
  • Improve workbench interactions in apps/web.
  • Expand provider adapters in packages/ai.

License

MIT. See LICENSE.

About

Voice-first diagram canvas with validated graph patches.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

VoiceCanvas logo

VoiceCanvas

Speak a diagram into shape.

中文 · Product docs · Prototype check · Quick start · Contributing

StatusLicenseNodepnpmTypeScript

VoiceCanvas workbench

VoiceCanvas is a voice-first diagram workbench. It turns natural speech into validated graph patches, applies those patches to a Mermaid-backed canvas, and keeps every change reversible through patch history.

The project is currently an early engineering prototype. It is useful for exploring the interaction model, running the local demo, and building toward a real open-source diagram editor where the main workflow is continuous voice editing.

Why This Exists

Most diagram tools make people stop thinking and start operating the UI. Most AI diagram generators handle the first draft better than the tenth edit. VoiceCanvas focuses on the work that happens after the first shape appears: renaming nodes, adding branches, changing flows, confirming ambiguous targets, and undoing a bad patch without rebuilding the whole graph.

The long-term bet is simple: diagrams should change at the speed of a conversation, while the underlying graph remains structured, validated, and exportable.

Highlights

AreaStatusWhat it does
Voice-first command flowWorking prototypeAccepts text segments plus selectable OpenAI Realtime or Gemini Live voice input.
Patch-based editingWorking prototypeConverts commands into atomic graph operations.
Validation layerWorking prototypeChecks patch drafts before they mutate the canvas.
Mermaid rendererWorking prototypeRenders the first diagram surface with Mermaid.
Low-confidence confirmationWorking prototypeShows target candidates before applying ambiguous edits.
History and undoWorking prototypeStores applied patches and restores the previous canvas state.
Realtime voice providersOptionalSupports OpenAI Realtime through the API proxy and Gemini Live through short-lived browser tokens.
External patch compilerOptionalUses an OpenAI-compatible model endpoint when configured.
Local mock compilerBuilt inRuns the demo without model credentials.
JSON exportWorking prototypeExports the current structured canvas document.

Architecture

flowchart LR
A[Mic or text input] --> B[Voice segment]
B --> C[Patch compiler]
C --> D[Validator]
D --> E[Patch engine]
E --> F[Canvas document]
F --> G[Mermaid renderer]
E --> H[Patch history]
H --> I[Undo and export]
Loading

The model is treated as a patch planner. It can propose a draft, but the canvas only changes after the draft passes validation and the patch engine applies it. That split keeps graph state inspectable, makes undo reliable, and keeps voice recognition separate from diagram mutation.

Repository Layout

apps/
web/ React + Vite workbench
api/ Hono API, workspace state, realtime session proxy
packages/
core/ CanvasDoc model, patch engine, validator, Mermaid export
ai/ OpenAI-compatible model patch compiler adapter
eval/ Acceptance cases and metric helpers
docs/
prd/ Product, interaction, roadmap, and system design docs
skills/
voicecanvas-dev-debug-acceptance/

Quick Start

Requirements

  • Node.js 24+
  • pnpm 10+

Run the local workbench

pnpm install
cp .env.example .env
pnpm dev

Open the web app:

http://localhost:5173

The API server runs on:

http://localhost:8787

The Vite dev server proxies /api traffic to the API server. You can run the project without external credentials; empty model settings use the built-in mock patch compiler.

Configuration

Create .env from .env.example and fill only the providers you want to use.

Realtime Voice

OPENAI_API_KEY=
OPENAI_REALTIME_MODEL=gpt-realtime-2
GEMINI_API_KEY=
GEMINI_LIVE_MODEL=gemini-3.1-flash-live-preview

External Patch Compiler

PATCH_COMPILER_API_KEY=
PATCH_COMPILER_BASE_URL=
PATCH_COMPILER_MODEL=
PATCH_COMPILER_PROVIDER=

When the patch compiler variables are empty, VoiceCanvas uses the local mock compiler. That makes the demo easy to run in forks, CI, and offline experiments.

Scripts

CommandDescription
pnpm devStart the web and API apps together.
pnpm dev:webStart only the Vite app.
pnpm dev:apiStart only the Hono API server.
pnpm testRun unit tests across the workspace.
pnpm lintRun lint checks.
pnpm buildBuild apps and type-check packages.
pnpm test:e2eRun Playwright smoke tests.
pnpm check:prototypeRun the full Prototype check suite.
pnpm check:alphaRun the Alpha check suite and local eval report.
pnpm check:alpha:realtimeRun the realtime provider API tests.

API Surface

MethodPathPurpose
GET/healthAPI health check.
GET/api/canvasRead the current workspace snapshot.
POST/api/dev/resetReset the in-memory prototype workspace.
POST/api/commands/text-segmentProcess a text segment as a voice command.
POST/api/patch/compileCompile a patch draft without applying it.
POST/api/patch/applyApply a provided patch draft.
POST/api/patch/confirmConfirm a low-confidence candidate.
POST/api/patch/undoRestore the previous patch state.
GET/api/realtime/providerRead realtime voice provider settings.
POST/api/realtime/openai/sessionProxy WebRTC session offers to OpenAI Realtime.
POST/api/realtime/gemini/tokenIssue a short-lived Gemini Live browser token.
GET/api/export/jsonExport the current CanvasDoc.

Development Notes

  • Source file names use kebab-case.
  • React component exports use PascalCase.
  • React hooks live in apps/web/src/hooks and use use-*.ts.
  • Tests use *.test.ts or *.spec.ts.
  • Generated build and test output should stay out of source control.
  • Workspace packages are private packages for now, while the repository itself is MIT licensed.

Roadmap

StageFocus
PrototypeOne-shot diagram creation, simple local edits, undo, candidate confirmation, JSON export.
AlphaMore reliable continuous editing, selected-object voice commands, local layout improvements, image export.
BetaFlowchart and mind map coverage, stronger evaluation cases, shareable outputs, real user feedback loops.
LaterTeam workspaces, meeting mode, permissions, provider marketplace, richer diagram types.

See docs/prd for the longer product and technical plan.

Contributing

Issues and small focused pull requests are welcome. Before sending a change, run the checks that match your edit:

pnpm test
pnpm lint
pnpm build
pnpm test:e2e

Good first areas:

  • Improve acceptance cases in packages/eval.
  • Add more mock compiler commands in packages/core.
  • Tighten API tests around patch confirmation and undo.
  • Improve workbench interactions in apps/web.
  • Expand provider adapters in packages/ai.

License

MIT. See LICENSE.

About

Voice-first diagram canvas with validated graph patches.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages