Skip to main content

audio

client.audio is the Speech surface: turn recorded audio into text, and turn text into spoken audio. It exposes the two general capability lanes that are not chat — asr (speech-to-text) and tts (text-to-speech) — as two methods:

Two facts set this namespace apart from the rest of the SDK:

  • Its own routes, not chat.completions. Speech does not go through chat.completions.createaudio.transcriptions(...) and audio.speech(...) hit their own dedicated routes directly. You never pick a voice model, a GPU, or a quantization; Pareta resolves the serving model behind the lane, exactly as model="auto" does for chat.
  • Metered per minute of audio. Both lanes are metered against your org balance by audio duration — input length for transcription, output length for synthesis — not by tokens. An empty balance raises InsufficientCreditsError (402). Top-up is browser-only; the SDK exposes neither balance nor payment methods.

Both calls go through the client's transport, so auth, retries, and typed error mapping apply exactly as they do everywhere else.

All examples use the synchronous Pareta client. Every method has an async twin with the same signature on AsyncPareta; see Async.

from pareta import Pareta

pa = Pareta.from_env() # reads PARETA_API_KEY (and optional PARETA_BASE_URL)
import { Pareta } from "pareta";

const pa = Pareta.fromEnv();
const t = await pa.audio.transcriptions("clip.wav", { language: "en" });
t.text; // the transcript
const s = await pa.audio.speech("hello there");
await s.save("out.wav"); // decoded bytes (Node)

audio.transcriptions

def transcriptions(self, audio, *, language: str | None = None) -> Transcription

Route: POST /v1/audio/transcriptions

Speech-to-text (the asr lane). Hands your audio to Pareta, returns the transcript plus the detected language and the metered duration.

  • audio (required): the clip to transcribe, in any of three forms — a path (str / os.PathLike) to an audio file, raw audio bytes, or an already base64-encoded string. A path or bytes are read and base64-encoded for you; a string that does not name a file is assumed to be base64 and passed through. An empty base64 string raises ValueError; anything that is not a path, bytes, or string raises TypeError.
  • language (optional): an ISO language hint (e.g. "en", "es"). Omit it to auto-detect across the supported languages.

Metered per minute of input audio.

result = pa.audio.transcriptions("meeting-clip.wav")

print(result.text) # the transcript
print(result.language) # detected (or the hint you passed)
print(result.duration_s) # input length that was metered (per minute)

audio= is flexible about where the bytes come from — a path, an in-memory buffer, or a pre-encoded string all work — and language is a hint, not a requirement:

# Raw bytes (e.g. from a recorder or an upload), with a language hint.
with open("call.ogg", "rb") as f:
result = pa.audio.transcriptions(f.read(), language="en")

# A Transcription stringifies to its transcript.
print(str(result)) # same as result.text (empty string if None)

Returns a Transcription.


audio.speech

def speech(self, text: str, *, voice: str | None = None) -> Speech

Route: POST /v1/audio/speech

Text-to-speech (the tts lane). Synthesizes spoken audio from text and returns a Speech whose .audio is the decoded bytes — call .save(path) to write a file.

  • text (required): the text to speak. Empty or whitespace-only text raises ValueError before any request goes out.
  • voice (optional): a voice id. Omit it for the default (Kokoro) voice.

Metered per minute of output audio.

speech = pa.audio.speech("Pareta turns text into speech in one call.")
speech.save("out.wav")

print(speech.format) # container/codec, e.g. "wav"
print(speech.sample_rate) # Hz
print(speech.duration_s) # output length that was metered (per minute)

.save() returns the same Speech, so you can chain it; pick a voice with voice=:

audio_bytes = pa.audio.speech(
"Your contract has been processed.",
voice="af_heart",
).save("notice.wav").audio # write the file and keep the bytes

Returns a Speech.


Async

Every method above has an async twin on AsyncPareta with an identical signature; the methods are coroutines. Speech.save(...) is a local file write, not a network call, so it is the same on both clients.

import asyncio
from pareta import AsyncPareta

async def main():
async with AsyncPareta.from_env() as pa:
result = await pa.audio.transcriptions("clip.wav", language="en")
print(result.text, result.duration_s)

speech = await pa.audio.speech("Hello from Pareta.")
speech.save("hello.wav")
print(speech.format, speech.sample_rate)

asyncio.run(main())

Response objects

Every object keeps the raw server JSON: call .to_dict() for lossless access to anything not yet surfaced as a typed field, and index it dict-style (result["..."]) as an escape hatch.

Transcription

From audio.transcriptions. Stringifies to .text (or "" when absent).

FieldTypeNotes
textstr | NoneThe transcript
languagestr | NoneDetected language, or the hint you passed
duration_sfloat | NoneInput audio length that was metered (per minute)

Speech

From audio.speech.

FieldTypeNotes
audiobytesThe synthesized audio, base64-decoded to raw bytes (b"" if empty)
audio_base64str | NoneThe raw base64 payload as returned by the server
sample_rateint | NoneSample rate in Hz
duration_sfloat | NoneOutput audio length that was metered (per minute)
formatstr | NoneContainer/codec of the returned audio, e.g. "wav"

Speech also has one method:

MethodReturnsNotes
save(path)SpeechWrite the decoded .audio bytes to path (str / os.PathLike); returns self for chaining

See also

  • tasks: how speech data is scored when you benchmark it (WER for transcripts).
  • chat: the OpenAI-compatible model="auto" inference surface for the chat-style capabilities, metered the same way.
  • Errors and metering: InsufficientCreditsError, the per-minute metering, and the full exception hierarchy.