Skip to main content
Use this guide when you want users to hold a button, ask your app a question out loud, and hear the answer spoken back.

What you’re building

The finished button in Chrome. Hold it, ask, release, and the answer plays back in about ten seconds; the wait while it thinks is shortened in this clip. The follow-up says only “it”, and the agent knows it means the sky.

You: Why is the sky blue?Agent: The sky looks blue because sunlight is made of many colors, and when it passes through the atmosphere, air molecules scatter the short blue wavelengths much more than the longer red ones. That scattered blue light reaches our eyes from every direction, so the whole sky glows blue. This effect is called Rayleigh scattering.You: Why does it turn orange at sunset?Agent: At sunset, sunlight has to travel through much more atmosphere before reaching your eyes, because the sun sits low on the horizon. Along that long path, most of the blue light gets scattered away, leaving the reds and oranges to dominate what you see.
The button records the microphone while it is held down. When it is released, the browser posts the recording to one route on your server, which makes three OpenRouter calls in a row: elevenlabs/scribe-v2 turns the speech into text, a chat model writes a short spoken answer, and elevenlabs/eleven-v4-turbo reads it out. The route returns the transcript, the answer, and the MP3, and the browser plays it. The API key stays on the server, and the browser keeps the conversation so follow-up questions work. By the end, you will have:
  1. A /api/talk route that takes a recording and returns a spoken answer.
  2. A hold-to-talk button that records, sends, and plays the reply.
  3. A conversation that remembers earlier turns.
Building with a coding agent? Paste the URL of this page into Claude Code, Codex, or Cursor and ask it to add this push-to-talk button to your app.

Before you start

You need:
  • An OpenRouter API key available as OPENROUTER_API_KEY
  • Bun to run the server
  • A browser with a microphone
This guide uses elevenlabs/scribe-v2 for transcription, anthropic/claude-haiku-5.5 for answers, and elevenlabs/eleven-v4-turbo for speech. Any chat model in the catalog works for the middle step. Speech-to-Text and Text-to-Speech list the other audio models and every request parameter. See the Scribe v2 and Eleven v4 Turbo model pages for pricing.

Step 1: Add the server route

Create server.ts. handleTalk reads the recording and the conversation so far from a form post, then transcribes, answers, and speaks. The system prompt asks for short plain sentences, because everything the model writes is read aloud, and for one opening audio tag such as [warmly], which Eleven v4 Turbo performs instead of saying. Browsers record WebM, except Safari, which records MP4 audio, so the route picks the format from the file’s type. If any of the three calls fails, the route returns the OpenRouter error with a 502 so the browser can show why, and a malformed history field starts a fresh conversation instead of failing the turn.
server.ts
Start it:
Then, in a second terminal, post any short voice memo to the route, here question-1.m4a, to check the server half on its own:
speech is the spoken answer as a base64 MP3.

Step 2: Add the button

The page is one button and a list for the conversation. Put the three files next to server.ts; Bun bundles index.html and the files it links when the server starts.
index.html
talk.js does the work. Pressing the button asks for the microphone and starts a MediaRecorder. Releasing it, or dragging off it, stops the recorder, and the recording goes to /api/talk with the conversation so far. The reply comes back as base64, so it plays straight from a data: URL. The audio tag is stripped before the answer is shown, because the listener hears it as delivery rather than words. A press shorter than 400 ms is ignored, so an accidental tap does not send an empty recording. Any failure, whether the network drops, OpenRouter rejects a request, or the browser blocks playback, puts the button back to Hold to talk and logs the reason to the browser console as Talk failed:.
talk.js
The styles give each state its own look, so users can tell when the button is listening. touch-action: none stops a long press on a phone from scrolling the page or opening a menu.
talk.css

Step 3: Try it

Open the page, allow the microphone, and hold the button while you ask a question. Release it and the button shows Thinking… and then Speaking… while the answer plays. Each turn takes about ten seconds from release to the first word of the reply, and speech is the longest stage, so the system prompt’s two-or-three-sentence limit is what keeps it short.

Use it in your own app

handleTalk takes a standard Request and returns a Response, so it works in any fetch-style server on Bun or Node, which provide the process.env and Buffer it uses. Copy server.ts into your app without the index.html import and the Bun.serve block, and mount handleTalk at /api/talk, for example as export const POST = handleTalk in a Next.js app/api/talk/route.ts. Then copy the button, talk.js, and talk.css into the page that needs it.

Troubleshooting

Microphone errors show in the conversation list. Errors from OpenRouter show in the browser console after Talk failed:, and in the /api/talk response body.

Check your work

Hold the button and ask “Why is the sky blue?”, then ask “Why does it turn orange at sunset?”. Both questions should appear in the list as you said them, and the second answer should be about the sky even though the question never names it. The answers should play in the george voice, with no bracketed tag on screen or in the audio.