Fields
| Field | Type | Required | Description | Example |
|---|---|---|---|---|
Channel | *int64 | ➖ | Zero-based audio channel index for the word, present when the provider transcribes channels separately | 0 |
Confidence | *float64 | ➖ | Provider confidence for the word from 0 to 1, present when the provider returns per-word confidence | 0.98 |
End | float64 | ✅ | Word end time in seconds | 0.4 |
Speaker | *int64 | ➖ | Speaker index for the word, present when the provider returns diarization data | 0 |
SpeakerLabel | *string | ➖ | Provider speaker label for the word, present when the provider labels speakers with a string | speaker_0 |
Start | float64 | ✅ | Word start time in seconds | 0 |
Type | *components.STTWordType | ➖ | Kind of entry; omitted or “word” for spoken words, “audio_event” for non-speech sounds the provider tags with timestamps | word |
Word | string | ✅ | The transcribed word, or the event tag such as “(laughter)” when type is audio_event | Hello |