Skip to content

Flow

A voice-driven, AI-native OS prototype: you speak, and Claude answers aloud while assembling the on-screen interface in real time. Built solo in 24 hours at the Anthropic × Y Combinator hackathon.

TL;DR

My partDesigned a single streamed-JSON contract, { speech, canvas }: Claude returns one object, the client parses it and mutates the canvas store directly, so each spoken sentence maps to one canvas action.

ROLE
Team of 4
CONTEXT
Anthropic × Y Combinator Hackathon · 42AI
STATUS
Done
RESULT
A runnable prototype: speaking a request makes Claude answer aloud and assemble the matching interface on a spatial canvas, built end to end in the 24-hour hackathon window.
LAST UPDATED
26 Sep 2026
voicepush-to-talkWhisperlocal · Web WorkerClaudeVercel AI SDK{ speech, canvas }one JSON streamElevenLabsneural TTScanvaswidgets fly in
Fig. — how it works
1

What I built

Team of 4

  1. Designed a single streamed-JSON contract, { speech, canvas }: Claude returns one object, the client parses it and mutates the canvas store directly, so each spoken sentence maps to one canvas action.
  2. Built the spatial canvas and widget registry — camera pan and zoom, a layout manager and around thirteen widget types (quiz, lesson, mail compose, task list).
  3. Wired the voice stack: local Whisper speech-to-text in a Web Worker, and ElevenLabs text-to-speech with a native speechSynthesis fallback.
  4. Added a two-agent demo layer on Claude Haiku (an intent router and a progress tracker) and a Gmail integration over MCP with a mock-inbox fallback.
2

Key choices

One streamed JSON contract instead of tool calling
Claude returns a single { speech, canvas } object the client parses directly, keeping speech and UI in lockstep with less latency than round-tripping tool calls.
Two small Haiku agents for a robust demo
An intent router and a progress tracker let the scripted scenario run in any order, regardless of the exact wording.
Browser-side API access, demo-only
The Anthropic key ships to the client, deliberately accepted for a local hackathon demo and documented as not production-safe.
3

Results

A runnable prototype: speaking a request makes Claude answer aloud and assemble the matching interface on a spatial canvas, built end to end in the 24-hour hackathon window.

4

Limits

  • Local demo only: the API key is exposed to the browser and there is no backend or auth, so it is not deployable as is.
  • Chrome or Edge only, and it needs a microphone; premium voice needs an ElevenLabs key.
  • The content is scripted around one seeded scenario; the Gmail path falls back to a mock inbox without OAuth.
  • No tests or CI: it is a proof of concept, not maintained software.

Questions about this project? → Email me