import asyncio
import websockets
class RimeClient:
def __init__(self, speaker, api_key):
self.url = f"wss://users-ws.rime.ai/ws?speaker={speaker}&modelId=mistv2&audioFormat=mp3"
self.auth_headers = {
"Authorization": f"Bearer {api_key}"
}
self.audio_data = b''
async def send_tokens(self, websocket, message):
for token in message:
await websocket.send(token)
async def handle_audio(self, websocket):
while True:
try:
audio = await websocket.recv()
except websockets.exceptions.ConnectionClosedOK:
break
self.audio_data += audio
async def run(self, message):
async with websockets.connect(self.url, additional_headers=self.auth_headers) as websocket:
await asyncio.gather(
self.send_tokens(websocket, message),
self.handle_audio(websocket),
)
def save_audio(self, file_path):
with open(file_path, 'wb') as f:
f.write(self.audio_data)
message = [
"This ",
"is ",
"a ",
"test ",
"<CLEAR>",
"This ",
"is ",
"a ",
"sentence, ",
"that ",
"will ",
"produce ",
"audio ",
"across ",
"two ",
"messages.",
"<EOS>",
]
client = RimeClient("cove", api_key="xxx")
asyncio.run(client.run(message))
client.save_audio("output.mp3")
Websocket
Websockets
Mist v2 plain-text WebSocket (/ws): send text, receive raw audio bytes.
import asyncio
import websockets
class RimeClient:
def __init__(self, speaker, api_key):
self.url = f"wss://users-ws.rime.ai/ws?speaker={speaker}&modelId=mistv2&audioFormat=mp3"
self.auth_headers = {
"Authorization": f"Bearer {api_key}"
}
self.audio_data = b''
async def send_tokens(self, websocket, message):
for token in message:
await websocket.send(token)
async def handle_audio(self, websocket):
while True:
try:
audio = await websocket.recv()
except websockets.exceptions.ConnectionClosedOK:
break
self.audio_data += audio
async def run(self, message):
async with websockets.connect(self.url, additional_headers=self.auth_headers) as websocket:
await asyncio.gather(
self.send_tokens(websocket, message),
self.handle_audio(websocket),
)
def save_audio(self, file_path):
with open(file_path, 'wb') as f:
f.write(self.audio_data)
message = [
"This ",
"is ",
"a ",
"test ",
"<CLEAR>",
"This ",
"is ",
"a ",
"sentence, ",
"that ",
"will ",
"produce ",
"audio ",
"across ",
"two ",
"messages.",
"<EOS>",
]
client = RimeClient("cove", api_key="xxx")
asyncio.run(client.run(message))
client.save_audio("output.mp3")
The Rime API authenticates every request with a bearer token in the
Discards buffered text that hasn’t been synthesized yet. Send it when the user interrupts. It doesn’t cancel a synthesis already in progress, so stop playback in your client and drop any audio that still arrives.
Synthesizes whatever text is buffered, without waiting for a sentence boundary, and sends the audio.
Synthesizes whatever text is buffered, sends the audio, and then the server closes the connection.
Authorization header: Authorization: Bearer YOUR_API_KEY. See API authentication for how to create a key.
Overview
This WebSocket accepts plain text and returns raw audio bytes in the format you set withaudioFormat. You pass every synthesis parameter as a query parameter when you open the connection, and those settings apply for the life of the connection.
The API buffers input until it detects a sentence boundary, typically at ., ?, or !, so synthesis starts once it has enough text for natural prosody. The wait is most noticeable on the first sentence. After that, later text usually arrives while earlier audio is still playing, and the buffering becomes largely invisible. To change when synthesis starts, set segment.
Messages
Send
Send bare text, not JSON.This will be converted to audio via websockets
Receive
Your client receives raw audio bytes in theaudioFormat set at connection time. Append the messages in order to rebuild the audio.
<FF>^@^@^@9LAME3.100^AP^@^@^@^@^@^@^@^@^T<A0>$^D>"^@^@<A0>^@^@<A8><C0><BA><9D>G^N^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@
^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@^@
^@^@^@^@^@^@^@^@^@^@
Commands
Send these strings as text messages to control the buffer.<CLEAR>
Discards buffered text that hasn’t been synthesized yet. Send it when the user interrupts. It doesn’t cancel a synthesis already in progress, so stop playback in your client and drop any audio that still arrives.
<FLUSH>
Synthesizes whatever text is buffered, without waiting for a sentence boundary, and sends the audio.
<EOS>
Synthesizes whatever text is buffered, sends the audio, and then the server closes the connection.
Variable parameters
string
required
Must be a voice from the Rime voice catalog.
string
Set to
mistv2.string
One of
mp3, mulaw, or pcmstring
default:"eng"
If provided, the language must match the language spoken by the selected speaker. Verify the pairing in the Rime voice catalog.
bool
default:"false"
When set to true, adds pauses between words enclosed in angle brackets. The number inside the brackets specifies the pause duration in milliseconds.
Example: “Hi. <200> I’d love to have a conversation with you.” adds a 200ms pause between the first and second sentences.
Example: “Hi. <200> I’d love to have a conversation with you.” adds a 200ms pause between the first and second sentences.
bool
default:"false"
When set to true, you can specify the phonemes for a word enclosed in curly brackets.
Example: “{h’El.o} World” pronounces “Hello” with the phonemes you supplied. Learn more about custom pronunciation.
Example: “{h’El.o} World” pronounces “Hello” with the phonemes you supplied. Learn more about custom pronunciation.
string
Comma-separated list of speed values applied to words in square brackets. Values < 1.0 speed up speech, > 1.0 slow it down.
Example: “This is [slow] and [fast]”, use “3, 0.5” to make “slow” slower and “fast” faster.
int
The value, if provided, must be between 4000 and 44100. Default: 22050
float
default:"1.0"
Adjusts the speed of speech. Lower than 1.0 is faster and higher than 1.0 is slower.Note: this is the legacy Mist v2 convention. Coda and Mist v3 invert it, so for those models higher than 1.0 is faster.
bool
default:"false"
mist/mistv2 only. Skips text normalization before synthesis. This reduces latency, but digits and abbreviations reach the model unexpanded and may be mispronounced.
string
default:"bySentence"
Controls how text is segmented for synthesis. Available options:
- “immediate” - Synthesizes text immediately without waiting for complete sentences
- “never” - Never segments the text, waits for explicit flush or EOS
- “bySentence” (default) - Waits for complete sentences before synthesis
immediate=true in query params is equivalent to segment=immediate. If a null value is provided, it will default to “bySentence”.import asyncio
import websockets
class RimeClient:
def __init__(self, speaker, api_key):
self.url = f"wss://users-ws.rime.ai/ws?speaker={speaker}&modelId=mistv2&audioFormat=mp3"
self.auth_headers = {
"Authorization": f"Bearer {api_key}"
}
self.audio_data = b''
async def send_tokens(self, websocket, message):
for token in message:
await websocket.send(token)
async def handle_audio(self, websocket):
while True:
try:
audio = await websocket.recv()
except websockets.exceptions.ConnectionClosedOK:
break
self.audio_data += audio
async def run(self, message):
async with websockets.connect(self.url, additional_headers=self.auth_headers) as websocket:
await asyncio.gather(
self.send_tokens(websocket, message),
self.handle_audio(websocket),
)
def save_audio(self, file_path):
with open(file_path, 'wb') as f:
f.write(self.audio_data)
message = [
"This ",
"is ",
"a ",
"test ",
"<CLEAR>",
"This ",
"is ",
"a ",
"sentence, ",
"that ",
"will ",
"produce ",
"audio ",
"across ",
"two ",
"messages.",
"<EOS>",
]
client = RimeClient("cove", api_key="xxx")
asyncio.run(client.run(message))
client.save_audio("output.mp3")

