Skip to content

Transcribe with speaker diarization

Pulse returns the transcript, word-level speaker labels and timestamps from a single pre-recorded call. Pass diarize: true for speaker attribution, and add wordTimestamps: true when you want that attribution per word rather than only per utterance.

Speaker ids are zero-indexed and stable within one response, not across requests. Use pulse rather than pulse-pro — Pulse Pro is English-only and does not diarize.

This example assumes using SmallestAI; is in scope and apiKey contains your SmallestAI API key.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
using var client = new SmallestAIClient(apiKey);

var audio = await BuildTwoSpeakerWavAsync(client, TestContext.CancellationToken);

var result = await client.SpeechToText.TranscribeAsync(
    model: WavesV1SttPostParametersModel.Pulse,
    language: WavesV1SttPostParametersLanguage.En,
    request: audio,
    wordTimestamps: WavesV1SttPostParametersWordTimestamps.True,
    diarize: WavesV1SttPostParametersDiarize.True,
    cancellationToken: TestContext.CancellationToken);

var transcription = result.TranscriptionResponse;

// Every word carries its own timing plus the speaker it was attributed to.

// Two voices went in, so at least two speakers should come back.
var speakers = transcription.Utterances?
    .Where(utterance => utterance.Speaker != null)
    .Select(utterance => utterance.Speaker!.Value)
    .Distinct()
    .ToList() ?? [];