Skip to main content
Build a minimal voice-agent loop in Next.js: the browser captures speech, a server route decides what to say, and Rime speaks the reply. This starter uses browser speech recognition and a stub response function; replace both with production components. Keep RIME_API_KEY on the server. Call Rime over HTTPS with fetch; no Rime-specific npm package is required. The optional WebSocket bridge adds the only new dependency, ws. Architecture: the browser never talks to Rime directly. Your API key stays server-side, and the browser calls your own /api/tts route:

Prerequisites

  • A Rime API key from the API Tokens page. Put it in .env.local as RIME_API_KEY=...
  • A Next.js 14+ project using the App Router (npx create-next-app@latest defaults work)

1. Server: the TTS proxy route

Create app/api/tts/route.ts (or src/app/api/tts/route.ts if your project uses src/):
app/api/tts/route.ts
Verify it works before touching the frontend:
test.mp3 should be playable audio. If you get a 401, check that RIME_API_KEY is set in .env.local and the dev server was restarted after adding it.

2. Client: mic in, Rime audio out

Create app/voice-agent/page.tsx. It uses the browser’s built-in SpeechRecognition for input (Chrome/Edge/Safari), a stub respond() function as the agent brain, and your /api/tts route for the voice:
app/voice-agent/page.tsx
Open http://localhost:3000/voice-agent, select Speak, say something, and the agent answers in Rime’s astra voice.

3. Stream over WebSockets when latency matters

The HTTP route above is the right default for most apps. Switch to WebSockets for incremental synthesis, where audio starts while the rest of the sentence is still being generated, or for barge-in handling with word-level timestamps. Two things make the WebSocket setup different:
  1. Rime’s WebSocket endpoints authenticate with an Authorization header, which browser WebSockets cannot send. The bridge must live on your server, where it also keeps the API key off the client.
  2. Next.js route handlers can’t hold WebSocket connections, so you need a small custom server.
Install ws, then create server.mjs in the project root:
server.mjs
Point your dev/start scripts at it:
package.json
On the client, connect to your bridge and play chunks as they arrive:
The Coda WebSocket reference defines the full /ws3 schema, including chunk, timestamps, done, and error events plus the flush, clear, and eos operations. The WebSocket overview covers segmentation and interruption patterns.

Production building blocks

WebSocket API overview

Endpoints, word-level timestamps, context IDs, and interruption handling.

Voices

Swap astra for any Coda voice. Coda covers eight languages, and each voice serves one of them.

Streaming formats

Choose between Opus, MP3, WAV, PCM, and μ-law for your latency budget.

LiveKit & Pipecat

Use the ready-made Rime plugins when a framework should own transport and orchestration.