deepgram-rust-text-to-speech
Use ao implementar conversão de texto em fala da Deepgram no SDK Rust, incluindo seleção do modelo Aura, flags de recurso de fala, manipulação de arquivo de saída ou fluxo de bytes, e…
npx skills add https://github.com/deepgram/deepgram-rust-sdk --skill deepgram-rust-text-to-speechUsing Deepgram Text-to-Speech (Rust SDK)
Use this skill when generating audio from text with the Rust SDK's Speak surface.
When to use this product
- Converting text into audio files with
speak_to_file(...). - Streaming TTS bytes with
speak_to_stream(...). - Selecting Aura voices and output encodings with
speak::options::Options. - Synthesizing with Flux TTS (
/v2/speak) in batch withflux_speak_to_file(...)orflux_speak_to_stream(...), or turn by turn over the WebSocket withflux_request(options).handle().
Authentication
For a TTS-only install:
[dependencies]
deepgram = { version = "0.10.1", default-features = false, features = ["speak"] }
tokio = { version = "1", features = ["full"] }
futures = "0.3"
# Only add `bytes = "1"` if you need to name `bytes::Bytes` in your own signatures.
# The code below relies on type inference and does not import bytes directly.
let dg = deepgram::Deepgram::new(std::env::var("DEEPGRAM_API_KEY")?)?;
- API keys use
Authorization: Token <api_key>. - Aura text-to-speech (
/v1/speak) is REST only in this crate: a saved file or a stream of bytes. Flux TTS (/v2/speak) has both transports in thespeak::fluxmodule:Speak::flux_speak_to_file/flux_speak_to_streamfor batch andSpeak::flux_request(options).handle()for the streaming WebSocket (FluxSpeakHandle:speak,flush,interrupt,configure_speed,close,receive).
Quick start
Quick start: save audio to a file
use std::{path::Path, time::Instant};
use deepgram::{
speak::options::{Container, Encoding, Model, Options},
Deepgram,
};
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
let api_key = std::env::var("DEEPGRAM_API_KEY")?;
let dg = Deepgram::new(&api_key)?;
let options = Options::builder()
.model(Model::AuraAsteriaEn)
.encoding(Encoding::Linear16)
.sample_rate(16000)
.container(Container::Wav)
.build();
let start = Instant::now();
dg.text_to_speech()
.speak_to_file("Hello from Rust.", &options, Path::new("output.wav"))
.await?;
println!("Time to download audio: {:.2?}", start.elapsed());
Ok(())
}
Quick start: stream response bytes
use deepgram::{
speak::options::{Container, Encoding, Model, Options},
Deepgram,
};
use futures::stream::StreamExt;
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
let api_key = std::env::var("DEEPGRAM_API_KEY")?;
let dg = Deepgram::new(&api_key)?;
let options = Options::builder()
.model(Model::AuraAsteriaEn)
.encoding(Encoding::Linear16)
.sample_rate(16000)
.container(Container::Wav)
.build();
let mut stream = dg
.text_to_speech()
.speak_to_stream("Hello from Rust.", &options)
.await?;
while let Some(chunk) = stream.next().await {
println!("received {} bytes", chunk.len());
}
Ok(())
}
Quick start: Flux TTS batch (REST)
use std::path::Path;
use deepgram::{
speak::flux::options::{Model, Options},
Deepgram,
};
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
let api_key = std::env::var("DEEPGRAM_API_KEY")?;
let dg = Deepgram::new(&api_key)?;
// model is required; encoding defaults to mp3 on the batch transport.
let options = Options::builder(Model::FluxHaleyEn).build();
dg.text_to_speech()
.flux_speak_to_file("Your appointment is confirmed.", &options, Path::new("flux-tts-batch.mp3"))
.await?;
Ok(())
}
Quick start: Flux TTS streaming (WebSocket)
use deepgram::{
speak::flux::{
options::{Encoding, Model, Options},
response::FluxSpeakResponse,
},
Deepgram,
};
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
let api_key = std::env::var("DEEPGRAM_API_KEY")?;
let dg = Deepgram::new(&api_key)?;
let options = Options::builder(Model::FluxHaleyEn)
.encoding(Encoding::Linear16)
.sample_rate(24_000)
.build();
// Connect to the /v2/speak WebSocket and get the handle.
let speak = dg.text_to_speech();
let mut handle = speak.flux_request(options).handle().await?;
// Stream the turn's text in, end the turn with flush, then close.
handle.speak("Hello! ").await?;
handle.speak("This audio was synthesized over the Flux TTS websocket.").await?;
handle.flush().await?;
handle.close().await?;
// Binary audio frames and JSON control events arrive on receive().
let mut saw_session_metadata = false;
while let Some(response) = handle.receive().await {
match response? {
FluxSpeakResponse::Audio(_bytes) => { /* play or buffer the audio */ }
FluxSpeakResponse::SpeechMetadata(_) => { /* end-of-turn signal */ }
FluxSpeakResponse::FatalError { code, description, .. } => {
return Err(format!("fatal Flux TTS error {code}: {description}").into());
}
FluxSpeakResponse::SessionMetadata { .. } => {
saw_session_metadata = true;
break;
}
_ => {}
}
}
if !saw_session_metadata {
return Err("session ended before its terminal SessionMetadata event".into());
}
Ok(())
}
handle.interrupt(playback_offset_ms) cancels the active turn on barge-in and the server answers with SpeechInterrupted; handle.configure_speed(speed) changes the speaking rate mid-session and the server answers with ConfigureSuccess or ConfigureFailure. See examples/speak/flux/websocket_synthesize.rs for the full event loop, including saving the audio as a WAV file.
Key parameters
- Entrypoints:
Deepgram::text_to_speech(),Speak::speak_to_file(...),Speak::speak_to_stream(...). - TTS
Optionsbuilder fields:model,encoding,sample_rate,container,bit_rate. - Model enum lives in
deepgram::speak::options::Modeland includes voices such asAuraAsteriaEn,AuraLunaEn,AuraOrionEn, plusCustomId(String). speak_to_stream(...)returnsimpl Stream<Item = bytes::Bytes>.- Flux TTS entrypoints:
Speak::flux_speak_to_file(text, &options, path),Speak::flux_speak_to_stream(text, &options),Speak::flux_request(options).handle()returningFluxSpeakHandle. - Flux TTS
Options::builder(Model)methods:encoding,sample_rate,speed,expressivity,mip_opt_out,tag, and the REST-onlycontainer,bit_rate,callback,callback_method,priority_low.priority_low()serializes aspriority=low. Models areflux-{voice}-{language}, for exampleModel::FluxHaleyEn; the WebSocket rejects REST-only options and the compressed encodings (mp3,opus,flac,aac) withDeepgramError::InvalidOptions.
API reference (layered)
- In-repo
README.mdsrc/speak/rest.rssrc/speak/options.rssrc/speak/flux/(rest.rs,websocket.rs,options.rs,response.rs)examples/speak/rest/text_to_speech_to_file.rsexamples/speak/rest/text_to_speech_to_stream.rsexamples/speak/flux/batch_synthesize.rsexamples/speak/flux/websocket_synthesize.rs
- OpenAPI
- Raw spec:
https://developers.deepgram.com/openapi.yaml - Endpoint reference:
https://developers.deepgram.com/reference/text-to-speech/speak-request
- Raw spec:
- AsyncAPI
- Rust SDK support: the Flux TTS
/v2/speakWebSocket viaSpeak::flux_request(options).handle(); the Aura/v1/speakWebSocket is not implemented in this crate - Raw spec:
https://developers.deepgram.com/asyncapi.yaml
- Rust SDK support: the Flux TTS
- Context7
/llmstxt/developers_deepgram_llms_txt
- Product docs
https://developers.deepgram.com/docs/text-to-speechhttps://developers.deepgram.com/docs/tts-rest
Gotchas
- Aura TTS is REST only in this crate. The only TTS WebSocket in
src/speak/is the Flux TTS client inspeak::flux::websocket; Aura voices are rejected on/v2/speak. - Pick encoding/container pairs deliberately. For raw output use
Container::None; for.wavoutput useContainer::Wav. speak_to_stream(...)still uses the REST endpoint. It streams HTTP response bytes fromPOST /v1/speak; the WebSocket surface isSpeak::flux_requeston/v2/speak.- Use API keys with
Token. Do not send API keys asBearer.
Example files in this repo
examples/speak/rest/text_to_speech_to_file.rsexamples/speak/rest/text_to_speech_to_stream.rsexamples/speak/flux/batch_synthesize.rsexamples/speak/flux/websocket_synthesize.rs
Central product skills
For cross-language Deepgram product knowledge — the consolidated API reference, documentation finder, focused runnable recipes, third-party integration examples, and MCP setup — install the central skills:
npx skills add deepgram/skills
This SDK ships language-idiomatic code skills; deepgram/skills ships cross-language product knowledge (see api, docs, recipes, examples, starters, setup-mcp).