Skip to content

Realtime speech with word timestamps

SmallestAITtsLiveRealtimeClient streams Lightning audio over a WebSocket so playback can start before synthesis finishes. Audio arrives as base64 chunk frames, interleaved with word_timestamp frames when wordTimestamps is on, and the stream ends with a complete frame.

Word timestamps are a WebSocket-only feature — the synchronous HTTP and SSE routes accept the flag but ignore it. They are available on English and Hindi voices.

This example assumes using SmallestAI; is in scope and apiKey contains your SmallestAI API key.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
using var client = new SmallestAIClient(apiKey);

var (voiceId, _) = await PickContrastingVoicesAsync(client, TestContext.CancellationToken);

await realtime.ConnectAsync(cancellationToken: TestContext.CancellationToken);

await realtime.SendTtsSynthesizeAsync(
    voiceId: voiceId,
    text: "Streaming this sentence so playback can start before synthesis finishes.",
    wordTimestamps: true,
    cancellationToken: TestContext.CancellationToken);

var audioBytes = 0;
var words = new List<TtsLiveEventData>();
var completed = false;

await foreach (var @event in realtime.ReceiveUpdatesAsync(TestContext.CancellationToken))
{
    switch (@event.Status)
    {
        case TtsLiveEventStatus.Chunk:
            audioBytes += @event.GetAudioBytes()?.Length ?? 0;
            break;

        case TtsLiveEventStatus.WordTimestamp when @event.Data is { } data:
            words.Add(data);
            break;

        case TtsLiveEventStatus.Complete:
            completed = true;
            break;

        default:
            break;
    }

    if (completed)
    {
        break;
    }
}