Fields
| Field | Type | Required | Description | Example |
|---|---|---|---|---|
Input | components.SpeechInput | ✅ | Text to synthesize, or a list of turns for multi-speaker input. Each turn has its own text, voice, and instructions. Multi-speaker input is currently supported by Gemini TTS models only. | Hello world |
InputReferences | []components.SpeechInputReference | ➖ | Reference content for stateless voice cloning or voice design. Audio mode: one to three input_audio parts, each optionally paired with a text part carrying its transcript (a single clip accepts its transcript before or after it; with multiple clips each transcript immediately follows its clip); only routed to endpoints that support voice cloning (and multiple references when more than one part is sent). Image mode: exactly one image_url part; only routed to endpoints that support image references. The two modes cannot be mixed. An empty array is treated as no reference. | [ { “input_audio”: { “data”: “data:audio/wav;base64,UklGRuQXDABXQVZF…” }, “type”: “input_audio” }, { “text”: “I used to rule the world.”, “type”: “text” } ] |
Instructions | *string | ➖ | Delivery instructions for the whole request, such as tone, pacing, or emotion. Supported by OpenAI gpt-4o-mini-tts and Gemini TTS models. Ignored by other providers. | Speak in a warm and friendly tone. |
Model | string | ✅ | TTS model identifier | mistralai/voxtral-mini-tts-2603 |
Provider | *components.SpeechRequestProvider | ➖ | Provider configuration: data policy routing preferences (zdr, data_collection) and provider-specific passthrough options | |
ResponseFormat | *components.SpeechRequestResponseFormat | ➖ | Audio output format | pcm |
SessionID | *string | ➖ | A unique identifier for grouping related requests (e.g., a conversation or agent workflow). Used for observability grouping in Broadcast and private logging; never sent to the provider. If provided in both the request body and the x-session-id header, the body value takes precedence. Maximum of 256 characters. | session-1234 |
Speed | *float64 | ➖ | Playback speed multiplier. Honored by models that support it (e.g. OpenAI TTS). Other providers either ignore it or return a 400 for a non-default value when the model has no speed control. | 1 |
Trace | *components.TraceConfig | ➖ | Metadata for observability and tracing. Known keys (trace_id, trace_name, span_name, generation_name, parent_span_id) have special handling. Additional keys are passed through as custom metadata to configured broadcast destinations. | { “trace_id”: “trace-abc123”, “trace_name”: “my-app-trace” } |
User | *string | ➖ | A unique identifier representing your end-user. Forwarded to Broadcast and private logging as the end-user id; never sent to the provider. | user-1234 |
Voice | *string | ➖ | Voice identifier (provider-specific). | en_paul_neutral |