Skip to main content

Running inference

You run inference through chat.completions.create, which has the same shape as the OpenAI chat completions API. Pass model="auto", a list of messages, and you get a ChatCompletion back. Set stream=True and you get an iterator of token deltas instead.

Pareta is OpenAI-compatible on the wire, so you can run inference with this SDK, with the openai package, or with raw HTTP, whichever fits your stack. This SDK's extra value is the control plane (evals on your own data, auto metrics); for plain inference the two are interchangeable.

A few platform truths that shape this page:

  • There is no model to pick. model is the literal string "auto"; Pareta selects the best-fit specialist behind it. Real open-weight model ids never reach you; the backend resolves them. You never pick a GPU.
  • Inference is metered against your org balance. A successful completion debits your balance — one debit per request, for the model that produces the final response, no matter how many internal model calls Pareta makes to serve it. If the balance is empty, the call raises InsufficientCreditsError (402). Top-up is browser-only; the SDK has no balance or payment surface.

The model ID auto​

The model id for every request is the literal string "auto". Pareta selects the best-fit open-weight specialist for the workload, checks the result against the workload's quality bar — output parsing and schema validation (including your response_format), constrained decoding where the task defines it, and model-graded checks on tasks configured for them — and automatically escalates to a frontier model when the specialist path cannot meet the bar. A task with no deployed specialist that holds the bar is served by a frontier model directly, and a frontier-served response is returned as served. Verification is a quality gate, not a guarantee that an AI output can never be wrong.

One request, one debit, for the model that produces the final response; routing, verification, retries, and non-serving attempts are included, and a request that errors out bills $0. Streaming works the same way — the answer streams token by token, and the SSE stream carries : pareta-progress <stage> comments (planning / executing / answering) you can surface as status.

completion = client.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "…"}],
)

Everything below — setup, streaming, async, errors — is that one call in different shapes.

Setup​

Mint a pareta_sk_ key in the dashboard, export it, and build the client from the environment:

export PARETA_API_KEY=pareta_sk_...

Python

from pareta import Pareta

pa = Pareta.from_env() # reads PARETA_API_KEY (+ optional PARETA_BASE_URL)

TypeScript

import { Pareta } from "pareta";

const pa = Pareta.fromEnv(); // reads PARETA_API_KEY (+ optional PARETA_BASE_URL)

from_env() is the recommended path. You can also pass the key explicitly: Pareta(api_key="pareta_sk_..."). The client is a context manager, so with Pareta.from_env() as pa: cleans up the HTTP connection for you.

A basic completion​

Pass model="auto" and a non-empty messages list in OpenAI format. You get back a ChatCompletion.

Python

from pareta import Pareta

with Pareta.from_env() as pa:
resp = pa.chat.completions.create(
model="auto", # the only model id
messages=[
{"role": "system", "content": "You extract structured fields from documents."},
{"role": "user", "content": "What is the invoice total?\n\nINVOICE\nTotal due: $4,210.00"},
],
)

print(resp.choices[0].message.content)
print(resp.usage.total_tokens, "tokens")

TypeScript

import { Pareta } from "pareta";

const pa = Pareta.fromEnv();
const resp = await pa.chat.completions.create({
model: "auto", // the only model id
messages: [
{ role: "system", content: "You extract structured fields from documents." },
{ role: "user", content: "What is the invoice total?\n\nINVOICE\nTotal due: $4,210.00" },
],
});

console.log(resp.choices[0].message.content);
console.log(resp.usage.totalTokens, "tokens");

model and messages are both required. The SDK raises ValueError before sending if model is falsy or messages is empty, so a malformed call fails fast without burning a request.

The ChatCompletion shape​

create() returns a ChatCompletion. The fields mirror OpenAI:

Python

resp.id # str | None
resp.model # str | None: echoes "auto"
resp.created # int | None: Unix timestamp
resp.choices # list[Choice]
resp.choices[0].index # int | None
resp.choices[0].finish_reason # "stop", "length", ...
resp.choices[0].message.role # "assistant"
resp.choices[0].message.content # str | None: the generated text
resp.usage.prompt_tokens # int | None
resp.usage.completion_tokens # int | None
resp.usage.total_tokens # int | None

TypeScript

resp.id // string | null
resp.model // string | null: echoes "auto"
resp.created // number | null: Unix timestamp
resp.choices // Choice[]
resp.choices[0].index // number | null
resp.choices[0].finishReason // "stop", "length", ...
resp.choices[0].message.role // "assistant"
resp.choices[0].message.content // string | null: the generated text
resp.usage.promptTokens // number | null
resp.usage.completionTokens // number | null
resp.usage.totalTokens // number | null

Every response object keeps the raw server JSON. If a field isn't surfaced as a typed property, reach it with resp.to_dict() or resp["..."]. Nothing the API returns is lost behind the typed layer.

Passthrough parameters​

Any extra keyword you pass goes straight into the request body, so the full OpenAI parameter set is available without the SDK enumerating it:

Python

resp = pa.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "Summarize this contract clause: ..."}],
temperature=0.2,
max_tokens=512,
top_p=0.9,
)

TypeScript

const resp = await pa.chat.completions.create({
model: "auto",
messages: [{ role: "user", content: "Summarize this contract clause: ..." }],
temperature: 0.2,
max_tokens: 512,
top_p: 0.9,
});

temperature, max_tokens, top_p, stop, seed, and friends all pass through unchanged.

max_tokens (or max_completion_tokens) is a hard cap on the answer, and a lower value is always honored exactly. Without one, Pareta sizes the budget to the request, up to 4,096 tokens for an open-ended text answer. If you expect a longer answer, such as a long report or a full-file rewrite, send a larger max_tokens: it raises the budget to as much as 8,192 tokens on Pareta-hosted models (less when the input itself nearly fills the model's context). Structured outputs (response_format) get a larger budget of their own. An answer that runs out of budget ends with finish_reason: "length", so check it before treating a long answer as complete.

Structured outputs (response_format)​

auto supports OpenAI Structured Outputs: pass a response_format and the response content is a single JSON document that conforms to your schema, however the request was served. Pareta enforces the schema with constrained decoding on its own specialists, validates every answer against your schema before delivery, and escalates automatically when an answer doesn't conform. If no conformant answer can be produced at all, the request fails with an error and you are not billed — nonconformant JSON is never delivered.

Schema instructions written only in the prompt ("reply with JSON matching …") are not enforced by specialists; response_format with a json_schema is the contract, and it is what constrains decoding. Send it on every request whose output your code parses.

Python

resp = pa.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "Classify: 'refund my order'"}],
response_format={
"type": "json_schema",
"json_schema": {
"name": "intent",
"schema": {
"type": "object",
"properties": {
"label": {"type": "string", "enum": ["refund", "billing", "other"]},
"confidence": {"type": "number"},
},
"required": ["label"],
"additionalProperties": False,
},
},
},
)
data = json.loads(resp.choices[0].message.content) # parses: validated against the schema before delivery

TypeScript

const resp = await pa.chat.completions.create({
model: "auto",
messages: [{ role: "user", content: "Classify: 'refund my order'" }],
response_format: {
type: "json_schema",
json_schema: {
name: "intent",
schema: {
type: "object",
properties: {
label: { type: "string", enum: ["refund", "billing", "other"] },
confidence: { type: "number" },
},
required: ["label"],
additionalProperties: false,
},
},
},
});

{"type": "json_object"} (JSON mode) is also supported: the content is guaranteed to parse as JSON, with no schema check. A malformed response_format returns a 400 with a specific message — it is never silently ignored.

Notes:

  • Nested objects, arrays, enum, required, and additionalProperties: false are all enforced. Schemas do not need to satisfy OpenAI's strict-mode restrictions (optional fields are fine).
  • The root schema must declare "type": "object" (same as OpenAI). Local $ref/$defs are supported; external $ref URLs are rejected with a 400. Schemas are capped at 50KB.
  • The format keyword (date-time, email, …) is not enforced — matching OpenAI, which also ignores format in structured outputs. Use pattern or enum when the shape matters.
  • With response_format set, temperature is not applied (structured serving runs at its canonical settings), and streamed responses deliver the JSON as a single buffered chunk after validation rather than token-by-token.
  • Clarifying questions are suppressed: a structured request always answers with schema-conformant JSON, never a prose question.

Tool calling (tools)​

model="auto" supports OpenAI-compatible function calling. Pass a tools array and the response comes back with tool_calls whose arguments are valid JSON for the schema you declared. Pareta serves tool turns on open models trained for tool use, and escalates when the turn needs more capability than they have.

You drive the loop. One request is one turn: Pareta returns either a final message or the tool calls it wants made. You execute them, append the results as role: "tool" messages, and call again.

Python

tools = [{
"type": "function",
"function": {
"name": "get_stock_price",
"description": "Get the current trading price for a ticker symbol.",
"parameters": {
"type": "object",
"properties": {"symbol": {"type": "string"}},
"required": ["symbol"],
},
},
}]

messages = [{"role": "user", "content": "What is NVDA trading at?"}]
resp = pa.chat.completions.create(model="auto", messages=messages, tools=tools)

call = resp.choices[0].message.tool_calls[0]
args = json.loads(call.function.arguments) # parses, guaranteed
result = get_stock_price(**args) # your function

messages += [
{"role": "assistant", "tool_calls": resp.choices[0].message.tool_calls},
{"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)},
]
final = pa.chat.completions.create(model="auto", messages=messages, tools=tools)
print(final.choices[0].message.content)

TypeScript

const tools = [{
type: "function",
function: {
name: "get_stock_price",
description: "Get the current trading price for a ticker symbol.",
parameters: {
type: "object",
properties: { symbol: { type: "string" } },
required: ["symbol"],
},
},
}];

let messages = [{ role: "user", content: "What is NVDA trading at?" }];
const resp = await pa.chat.completions.create({ model: "auto", messages, tools });

const call = resp.choices[0].message.tool_calls[0];
const args = JSON.parse(call.function.arguments);
const result = await getStockPrice(args.symbol);

messages = [
...messages,
{ role: "assistant", tool_calls: resp.choices[0].message.tool_calls },
{ role: "tool", tool_call_id: call.id, content: JSON.stringify(result) },
];
const final = await pa.chat.completions.create({ model: "auto", messages, tools });

tool_choice works as it does on OpenAI: "auto" (the default) lets the model decide, "required" forces some tool call, "none" forbids them, and {"type": "function", "function": {"name": "..."}} forces a specific one. parallel_tool_calls is passed through.

Notes:

  • tools and response_format cannot be combined — the request returns a 400. The two guarantees are enforced by different machinery, and accepting both would silently drop the schema guarantee rather than honor it.
  • stream=True with tools returns a 400 today. Retry without it.
  • n > 1 returns a 400.
  • If the serving model has been idle and is cold, the turn waits for it to start (this can take several minutes) and is answered by the frontier tier if it doesn't come up in time. Use a client timeout of at least 10 minutes, and send an Idempotency-Key header so an automatic client retry never bills the same turn twice: a retry with the same key and body gets the first attempt's response back (waiting for it if it is still running) instead of running the turn again (details). A 503 with a Retry-After header means the serving model could not be reached at all; retry the call.
  • Billing is per turn, at the same rates as any other auto request: a turn answered by Pareta's open models bills at the published per-token rates, and a turn escalated to the frontier tier bills at that provider's list price. Every response carries the usual X-Pareta-Billed (exact debit) and X-Pareta-Frontier-Would-Have-Cost (estimated) headers plus the route diagnostics (X-Pareta-Served-By, X-Pareta-Escalated, X-Pareta-Specialist, X-Pareta-Route-Version), so a multi-turn tool loop shows its cost and its route turn by turn.

Streaming​

Set stream=True and create() returns an iterator of ChatCompletionChunk objects instead of a single ChatCompletion. Each chunk carries a delta (not a message); the incremental text is at chunk.choices[0].delta.content.

Python

with Pareta.from_env() as pa:
stream = pa.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "Summarize this support ticket thread in two sentences for the agent handoff note:\n\n..."}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="", flush=True)
print()

TypeScript

const pa = Pareta.fromEnv();
const stream = pa.chat.completions.create({
model: "auto",
messages: [{ role: "user", content: "Summarize this support ticket thread in two sentences for the agent handoff note:\n\n..." }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0].delta.content || "");
}
console.log();

ChatCompletionChunk has the same schema as ChatCompletion; it exists as a distinct type only for hinting. Guard delta.content with or "": the first and last chunks of a stream often carry role or finish metadata with no text.

The stream is data-only SSE and always terminates on a [DONE] sentinel, which the SDK consumes for you, so the iterator simply ends. Note that retries only cover the initial handshake. Once tokens are flowing, a mid-stream drop raises immediately rather than silently resuming.

Async​

AsyncPareta mirrors the sync client. Methods are async def; for streaming you await the call once, then async for over the chunks.

Python

import asyncio
from pareta import AsyncPareta

async def main():
async with AsyncPareta.from_env() as pa:
# Non-streaming
resp = await pa.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "What is the invoice total?"}],
)
print(resp.choices[0].message.content)

# Streaming
stream = await pa.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "Summarize this support ticket thread in two sentences for the agent handoff note:\n\n..."}],
stream=True,
)
async for chunk in stream:
print(chunk.choices[0].delta.content or "", end="", flush=True)
print()

asyncio.run(main())

TypeScript

// There is no AsyncPareta in TypeScript — the one Pareta client is already
// Promise-only. Every I/O method returns a Promise you `await`; streaming
// returns an AsyncIterable you drive with `for await`.
import { Pareta } from "pareta";

const pa = Pareta.fromEnv();

// Non-streaming
const resp = await pa.chat.completions.create({
model: "auto",
messages: [{ role: "user", content: "What is the invoice total?" }],
});
console.log(resp.choices[0].message.content);

// Streaming
const stream = pa.chat.completions.create({
model: "auto",
messages: [{ role: "user", content: "Summarize this support ticket thread in two sentences for the agent handoff note:\n\n..." }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0].delta.content || "");
}
console.log();

Reading the cost of a request​

Every completion tells you what it cost without any SDK: the X-Pareta-Billed response header is the exact debit in micro-USD, and X-Pareta-Frontier-Would-Have-Cost is an estimate of what a single frontier call on the same prompt would have cost (usage tokens × frontier list price). The difference is your estimated per-request savings. Streamed responses deliver the same two numbers as SSE comment lines just before [DONE]. See the HTTP API reference for details.

Handling metering and not-ready errors​

Two error cases are specific to running inference. Both subclass ParetaError, so a single except ParetaError is a fine catch-all; the specific classes let you branch.

Python

from pareta import (
Pareta,
InsufficientCreditsError, # 402: org balance empty
EndpointNotReadyError, # 503: a serving backend is warming / briefly unavailable
)

with Pareta.from_env() as pa:
try:
resp = pa.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "What is the invoice total?\n\nINVOICE\nTotal due: $4,210.00"}],
)
print(resp.choices[0].message.content)
except InsufficientCreditsError:
# Balance hit zero. Top up in the dashboard (billing is browser-only);
# the SDK exposes no balance or payment surface.
print("Out of credit. Top up in the dashboard, then retry.")
except EndpointNotReadyError:
# A serving backend behind auto is warming. The SDK already retried
# the 503 with backoff; wait briefly and retry the call.
print("Backend warming — retry shortly.")

TypeScript

import {
Pareta,
InsufficientCreditsError, // 402: org balance empty
EndpointNotReadyError, // 503: a serving backend is warming / briefly unavailable
} from "pareta";

const pa = Pareta.fromEnv();
try {
const resp = await pa.chat.completions.create({
model: "auto",
messages: [{ role: "user", content: "What is the invoice total?\n\nINVOICE\nTotal due: $4,210.00" }],
});
console.log(resp.choices[0].message.content);
} catch (e) {
if (e instanceof InsufficientCreditsError) {
// Balance hit zero. Top up in the dashboard (billing is browser-only);
// the SDK exposes no balance or payment surface.
console.log("Out of credit. Top up in the dashboard, then retry.");
} else if (e instanceof EndpointNotReadyError) {
// A serving backend behind auto is warming. The SDK already retried
// the 503 with backoff; wait briefly and retry the call.
console.log("Backend warming — retry shortly.");
} else {
throw e;
}
}

Transient failures (429 rate limits, 5xx, connection timeouts) are retried automatically with exponential backoff, max_retries times (default 2). You only see RateLimitError or APITimeoutError after retries are exhausted. See Errors for the full hierarchy.

Using the OpenAI SDK instead​

Because Pareta is one OpenAI-compatible endpoint, you don't need this SDK to call it. Use your Pareta API key, point the OpenAI client at Pareta's base URL, and set the model ID to auto. Note the /v1 suffix the OpenAI client expects:

Python

from openai import OpenAI

client = OpenAI(api_key="pareta_sk_...", base_url="https://api.pareta.ai/v1")

resp = client.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "What is the invoice total?"}],
)
print(resp.choices[0].message.content)

TypeScript

import OpenAI from "openai";

const client = new OpenAI({ apiKey: "pareta_sk_...", baseURL: "https://api.pareta.ai/v1" });

const resp = await client.chat.completions.create({
model: "auto",
messages: [{ role: "user", content: "What is the invoice total?" }],
});
console.log(resp.choices[0].message.content);

Tooling that discovers model ids by listing keeps working too: models.list() (GET /v1/models) returns exactly one entry, "auto" — there is only one model id to call. Field details in the models reference.

Streaming, temperature, max_tokens, response_format, and tools behave as described on this page; the sections above state what is supported and what returns a 400, and the HTTP API reference covers the rest of the route map. Metering still applies: a zero balance returns a 402, which the openai client surfaces as its own status error. Reach for the Pareta SDK when you want typed errors and the control plane: running evals on your own data and reading auto's metrics (core concepts).