Flow
A voice-driven, AI-native OS prototype: you speak, and Claude answers aloud while assembling the on-screen interface in real time. Built solo in 24 hours at the Anthropic × Y Combinator hackathon.
TL;DR
My partDesigned a single streamed-JSON contract, { speech, canvas }: Claude returns one object, the client parses it and mutates the canvas store directly, so each spoken sentence maps to one canvas action.
- ROLE
- Team of 4
- CONTEXT
- Anthropic × Y Combinator Hackathon · 42AI
- STATUS
- Done
- STACK
- TypeScriptReactViteTailwind CSSVercel AI SDKAnthropic API (Claude Sonnet + Haiku)Whisper (transformers.js)ElevenLabsZustandFramer MotionZodKaTeXClaude Code
- RESULT
- A runnable prototype: speaking a request makes Claude answer aloud and assemble the matching interface on a spatial canvas, built end to end in the 24-hour hackathon window.
- LAST UPDATED
- 26 Sep 2026
1
What I built
Team of 4
- Designed a single streamed-JSON contract, { speech, canvas }: Claude returns one object, the client parses it and mutates the canvas store directly, so each spoken sentence maps to one canvas action.
- Built the spatial canvas and widget registry — camera pan and zoom, a layout manager and around thirteen widget types (quiz, lesson, mail compose, task list).
- Wired the voice stack: local Whisper speech-to-text in a Web Worker, and ElevenLabs text-to-speech with a native speechSynthesis fallback.
- Added a two-agent demo layer on Claude Haiku (an intent router and a progress tracker) and a Gmail integration over MCP with a mock-inbox fallback.
2
Key choices
- One streamed JSON contract instead of tool calling
- Claude returns a single { speech, canvas } object the client parses directly, keeping speech and UI in lockstep with less latency than round-tripping tool calls.
- Two small Haiku agents for a robust demo
- An intent router and a progress tracker let the scripted scenario run in any order, regardless of the exact wording.
- Browser-side API access, demo-only
- The Anthropic key ships to the client, deliberately accepted for a local hackathon demo and documented as not production-safe.
3
Results
A runnable prototype: speaking a request makes Claude answer aloud and assemble the matching interface on a spatial canvas, built end to end in the 24-hour hackathon window.
4
Limits
- Local demo only: the API key is exposed to the browser and there is no backend or auth, so it is not deployable as is.
- Chrome or Edge only, and it needs a microphone; premium voice needs an ElevenLabs key.
- The content is scripted around one seeded scenario; the Gmail path falls back to a mock inbox without OAuth.
- No tests or CI: it is a proof of concept, not maintained software.
Questions about this project? → Email me