Quickstart
Pareta is one OpenAI-compatible endpoint with one model ID, auto, built for
the tasks your application runs again and again. Send a request and Pareta
selects the best-fit open-weight specialist when a deployed specialist holds
the workload's quality bar, checks the result against that bar, and
automatically escalates to a frontier model when the specialist path cannot
meet it; when no deployed specialist holds the bar, a frontier model serves
the request. You are billed once, for the model that
produces the final response, and inference is metered against your org's
balance.
This page gets one production call working, then shows the pieces around it:
streaming, metering, and benchmarking auto against frontier models on your
own data.
The 30-second version
A support-ticket classifier: a fixed label set, one request per ticket, and
response_format with a json_schema as the output contract. Pareta
constrains decoding to the schema and validates the result before delivery,
so label is one of yours — or the call fails with an error and is not billed (see Structured outputs).
Python
import json
from pareta import Pareta
client = Pareta.from_env() # reads PARETA_API_KEY
LABELS = ["billing", "refund", "login", "bug", "feature_request", "other"]
TICKET_FORMAT = { # the output contract
"type": "json_schema",
"json_schema": {
"name": "ticket_label",
"schema": {
"type": "object",
"properties": {"label": {"type": "string", "enum": LABELS}},
"required": ["label"],
"additionalProperties": False,
},
},
}
def classify(ticket: str) -> str:
completion = client.chat.completions.create(
model="auto",
messages=[
{"role": "system", "content": "Classify the customer support ticket into exactly one label."},
{"role": "user", "content": ticket},
],
response_format=TICKET_FORMAT,
)
return json.loads(completion.choices[0].message.content)["label"]
print(classify("I was charged twice for my March invoice.")) # one of LABELS, e.g. "billing"
TypeScript
import { Pareta } from "pareta";
const client = Pareta.fromEnv(); // reads PARETA_API_KEY
const LABELS = ["billing", "refund", "login", "bug", "feature_request", "other"];
const TICKET_FORMAT = { // the output contract
type: "json_schema",
json_schema: {
name: "ticket_label",
schema: {
type: "object",
properties: { label: { type: "string", enum: LABELS } },
required: ["label"],
additionalProperties: false,
},
},
};
async function classify(ticket: string): Promise<string> {
const completion = await client.chat.completions.create({
model: "auto",
messages: [
{ role: "system", content: "Classify the customer support ticket into exactly one label." },
{ role: "user", content: ticket },
],
response_format: TICKET_FORMAT,
});
return JSON.parse(completion.choices[0].message.content ?? "{}").label;
}
console.log(await classify("I was charged twice for my March invoice.")); // one of LABELS, e.g. "billing"
That is the product. Everything below — benchmarking auto against frontier
models on your own data and monitoring spend and estimated savings — exists to
prove and operate that one call.
Schema instructions written only in the prompt are not enforced;
response_format with a json_schema is the contract. See
Structured outputs for what
is validated, JSON mode, and the limits.
- Prove it:
client.evalswith"auto"among the candidates (see Evaluation) — per-contender quality + cost on YOUR data. - Watch it:
client.auto.metrics()— requests, success rate, spend, and estimated savings against frontier list prices. - Compare it:
client.auto.compare_frontier(...)— one prompt against a frontier vendor, metered, for a side-by-side.
The later sections on this page reuse TICKET_FORMAT from the example above.
Install
pip install pareta # or: uv add pareta / poetry add pareta
Authenticate
Mint a pareta_sk_ key in the dashboard (key management is browser-only) and
export it. Pareta.from_env() reads PARETA_API_KEY (and an optional
PARETA_BASE_URL).
export PARETA_API_KEY="pareta_sk_..."
The SDK only ever consumes a key. It never creates, lists, or revokes them, and it never exposes your balance or payment methods. Topping up credit is browser-only.
Find out how your data will be scored
There is no model to pick and nothing to deploy — send any generation job to
the model ID auto. The one lookup you'll ever do is for benchmarking: give
tasks.match a plain-English description of your dataset and it tells you how
an eval will score rows like yours:
Python
from pareta import Pareta
pa = Pareta.from_env() # reads PARETA_API_KEY
m = pa.tasks.match("extract key fields from contracts")
print(m.type) # "task" — a benchmarked task covers this
if m.chosen:
print(m.chosen.task_id) # e.g. "contract-key-fields"
TypeScript
import { Pareta } from "pareta";
const pa = Pareta.fromEnv(); // reads PARETA_API_KEY
const m = await pa.tasks.match("extract key fields from contracts");
console.log(m.matched); // true — a benchmarked task covers this
console.log(m.chosen?.taskId); // e.g. "contract-key-fields"
m.type is one of four verdicts: "task" (a benchmarked task fits),
"capability" (a general lane — chat, coding, vision, … — covers it),
"unsupported" (Pareta does not cover this; a correct answer, not an error),
or "none" (the router was unavailable and the lexical fallback found nothing
confident). Whatever the verdict names, running the job is always the same
call: chat.completions.create(model="auto", ...).
pa.tasks.list() and pa.tasks.retrieve(task_id) return the scoring
definitions eval rows are validated against — see tasks.
Stream the response
Pass stream=True to get an iterator of ChatCompletionChunk. The incremental
text lives on chunk.choices[0].delta.content (it can be None on the first
and last chunks, so guard it). Streaming is for free-text outputs such as the
handoff summary below; a response_format response arrives as one validated
chunk instead.
Python
for chunk in pa.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "Summarize this support ticket thread in two sentences for the agent handoff note:\n\n..."}],
stream=True,
):
print(chunk.choices[0].delta.content or "", end="", flush=True)
print()
TypeScript
for await (const chunk of pa.chat.completions.create({
model: "auto",
messages: [{ role: "user", content: "Summarize this support ticket thread in two sentences for the agent handoff note:\n\n..." }],
stream: true,
})) {
process.stdout.write(chunk.choices[0].delta.content || "");
}
console.log();
Extra OpenAI parameters (temperature, max_tokens, top_p, and so on) pass
straight through as keyword arguments.
Cost and credit
Every successful completion debits your org's balance — one debit per request,
no matter how many internal model calls Pareta makes to serve it. If the
balance is empty, the call raises InsufficientCreditsError (HTTP 402).
Top-up is browser-only.
Python
from pareta import InsufficientCreditsError
try:
resp = pa.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "I was charged twice for my March invoice."}],
response_format=TICKET_FORMAT,
)
except InsufficientCreditsError:
print("Out of credit — top up in the dashboard.")
TypeScript
import { InsufficientCreditsError } from "pareta";
try {
const resp = await pa.chat.completions.create({
model: "auto",
messages: [{ role: "user", content: "I was charged twice for my March invoice." }],
response_format: TICKET_FORMAT,
});
} catch (e) {
if (e instanceof InsufficientCreditsError) {
console.log("Out of credit — top up in the dashboard.");
} else {
throw e;
}
}
Evaluation runs are metered the same way (auto plus frontier compute). An
EvalRun reports its billed total on run.cost, a Decimal in dollars floored
to whole cents (so a sub-cent run reads Decimal("0.00")); the raw value is on
run.cost_micro_usd. See Evals.
Clean up
There is nothing running on your account to stop — the serving fleet behind
auto is Pareta's to operate. Cleanup is just closing the client (or using it
as a context manager).
Python
pa.close()
TypeScript
// No close() in TS: the client owns no connection (it uses the native fetch),
// so there is nothing to release and no context-manager form to wrap it in.
Python
# Context-manager form closes the HTTP client for you.
with Pareta.from_env() as pa:
resp = pa.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "I was charged twice for my March invoice."}],
response_format=TICKET_FORMAT,
)
TypeScript
// No context manager in TS — just construct and use it; nothing to close.
const pa = Pareta.fromEnv();
const resp = await pa.chat.completions.create({
model: "auto",
messages: [{ role: "user", content: "I was charged twice for my March invoice." }],
response_format: TICKET_FORMAT,
});
List what you can call
models.list() returns the OpenAI-compatible model list. It has exactly one
entry — "auto" — which is the point: the id you pass to
chat.completions.create(model=...) is never a decision.
Python
for m in pa.models.list():
print(m.id, m.owned_by)
TypeScript
for (const m of await pa.models.list()) {
console.log(m.id, m.ownedBy);
}
Async
AsyncPareta mirrors the sync client; resource methods are async def and
streams are async iterators.
Python
import asyncio
import json
from pareta import AsyncPareta
async def main():
async with AsyncPareta.from_env() as pa:
resp = await pa.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "I was charged twice for my March invoice."}],
response_format=TICKET_FORMAT,
)
print(json.loads(resp.choices[0].message.content)["label"])
asyncio.run(main())
TypeScript
// There is no AsyncPareta in TypeScript — the single `Pareta` client is already
// async. Every I/O method returns a Promise (await it), and streams are async
// iterables (`for await`).
import { Pareta } from "pareta";
const pa = Pareta.fromEnv();
const resp = await pa.chat.completions.create({
model: "auto",
messages: [{ role: "user", content: "I was charged twice for my March invoice." }],
response_format: TICKET_FORMAT,
});
console.log(JSON.parse(resp.choices[0].message.content ?? "{}").label);
Already using the OpenAI SDK?
You do not need this SDK just to run inference. Use your Pareta API key, point
the openai client at Pareta's base URL, and set the model ID to auto:
Python
from openai import OpenAI
client = OpenAI(api_key="pareta_sk_...", base_url="https://api.pareta.ai/v1")
resp = client.chat.completions.create(
model="auto",
messages=[
{"role": "system", "content": "Classify the customer support ticket into exactly one label."},
{"role": "user", "content": "I was charged twice for my March invoice."},
],
response_format=TICKET_FORMAT,
)
TypeScript
import OpenAI from "openai";
const client = new OpenAI({ apiKey: "pareta_sk_...", baseURL: "https://api.pareta.ai/v1" });
const resp = await client.chat.completions.create({
model: "auto",
messages: [
{ role: "system", content: "Classify the customer support ticket into exactly one label." },
{ role: "user", content: "I was charged twice for my March invoice." },
],
response_format: TICKET_FORMAT,
});
This SDK's unique value is everything around that call: benchmark auto on
your own data (evals) and watch traffic and estimated savings
(auto.metrics()) — from code.
Next steps
- Core concepts — the verified-inference path behind
auto, the quality bar and escalation, how your data is scored, metering, and the eval → production funnel. - Evals — benchmark
autoagainst frontier baselines on your own data. - Errors — the
ParetaErrorhierarchy and retry behavior.