Pareta
Pareta is a verified-inference provider built for repeatable production AI workloads. Through one OpenAI-compatible endpoint and one model ID, "auto", your application can access Pareta’s fleet of open-weight specialists, including several distilled in-house. Pareta checks every result against the workload’s quality bar and automatically escalates to a frontier model whenever the specialist path cannot meet it. You are charged only for the model that produces the final response. For workloads handled by Pareta specialists, serving costs can be 30–140× lower than the comparable frontier model, and you can verify the quality and economics on your own production data.
These docs cover the three things you do with that endpoint, all behind one pareta_sk_ key:
- Run inference with the model ID
auto. OpenAI-compatible and metered. This SDK (Python + TypeScript) and the stockopenaiclient are interchangeable for inference; streaming,response_format, andtoolsare covered in Running inference. Nothing to deploy. - Evaluate it on your own data. Run
autohead-to-head against frontier baselines on your rows, with one scorer for every contender, and read per-contender quality and cost before you move traffic. - Read what each request cost. Every chat-completion response carries
X-Pareta-Billed(the exact debit, in micro-USD) andX-Pareta-Frontier-Would-Have-Cost(an estimate: usage tokens × frontier list price); direct endpoints (/v1/embeddings,/v1/rerank, audio, images) carry no receipt headers and report their cost in the dashboard.
A few platform facts shape the whole API:
- Models and GPUs are hidden. You never pick either. Pareta selects the best-fit specialist per request when a deployed specialist holds the workload's quality bar; otherwise a frontier model serves the request. Hardware is Pareta's problem.
- Frontier (vendor) ids are in the clear. They appear as eval baselines and in
auto.compare_frontier(); every open-weight specialist stays behindauto. - Inference and evals are metered against your org balance. A successful call debits credit once per request, no matter how many internal model calls Pareta makes to serve it: routing, verification, retries, and non-serving attempts are included. An empty balance raises
InsufficientCreditsError(402). An eval run reports its billed total onrun.cost(dollars). Top-up is browser-only; the SDK never touches billing.
Pareta is infrastructure for AI tasks embedded in products and automated workflows, not a destination chatbot or general coding copilot.
Install
Python
pip install pareta # or: uv add pareta / poetry add pareta
TypeScript
npm install pareta # or: pnpm add pareta / yarn add pareta / bun add pareta
Python or TypeScript? Both SDK clients are at full parity. The one design difference: Python ships sync (
Pareta) and async (AsyncPareta) clients; TypeScript has a single Promise-onlyPareta(every method isasync). Code samples throughout these docs show Python and TypeScript side by side.
A first production call
Mint a pareta_sk_ key in the dashboard, export it as PARETA_API_KEY, and send a repeatable task. The example below classifies support tickets into a fixed label set. The output contract is response_format with a json_schema: Pareta constrains decoding to the schema and validates the result before delivery, so label is always one of yours.
Python
import json
from pareta import Pareta
pa = Pareta.from_env() # reads PARETA_API_KEY
LABELS = ["billing", "refund", "login", "bug", "feature_request", "other"]
TICKET_FORMAT = { # the output contract
"type": "json_schema",
"json_schema": {
"name": "ticket_label",
"schema": {
"type": "object",
"properties": {"label": {"type": "string", "enum": LABELS}},
"required": ["label"],
"additionalProperties": False,
},
},
}
def classify(ticket: str) -> str:
resp = pa.chat.completions.create(
model="auto",
messages=[
{"role": "system", "content": "Classify the customer support ticket into exactly one label."},
{"role": "user", "content": ticket},
],
response_format=TICKET_FORMAT,
)
return json.loads(resp.choices[0].message.content)["label"]
print(classify("I was charged twice for my March invoice.")) # one of LABELS, e.g. "billing"
TypeScript
import { Pareta } from "pareta";
const pa = Pareta.fromEnv(); // reads PARETA_API_KEY
const LABELS = ["billing", "refund", "login", "bug", "feature_request", "other"];
const TICKET_FORMAT = { // the output contract
type: "json_schema",
json_schema: {
name: "ticket_label",
schema: {
type: "object",
properties: { label: { type: "string", enum: LABELS } },
required: ["label"],
additionalProperties: false,
},
},
};
async function classify(ticket: string): Promise<string> {
const resp = await pa.chat.completions.create({
model: "auto",
messages: [
{ role: "system", content: "Classify the customer support ticket into exactly one label." },
{ role: "user", content: ticket },
],
response_format: TICKET_FORMAT,
});
return JSON.parse(resp.choices[0].message.content ?? "{}").label;
}
console.log(await classify("I was charged twice for my March invoice.")); // one of LABELS, e.g. "billing"
Schema instructions written only in the prompt ("reply with JSON") are not enforced; response_format with a json_schema is the contract. See Structured outputs for what is validated, JSON mode, and the limits.
Already using the OpenAI SDK?
The same call works from the stock openai client: use your Pareta API key, point the OpenAI client at Pareta's base URL, and set the model ID to auto. TICKET_FORMAT is the schema from the example above.
Python
from openai import OpenAI
client = OpenAI(api_key="pareta_sk_...", base_url="https://api.pareta.ai/v1")
resp = client.chat.completions.create(
model="auto",
messages=[
{"role": "system", "content": "Classify the customer support ticket into exactly one label."},
{"role": "user", "content": "I was charged twice for my March invoice."},
],
response_format=TICKET_FORMAT,
)
print(resp.choices[0].message.content) # {"label": "..."}
TypeScript
import OpenAI from "openai";
const client = new OpenAI({ apiKey: "pareta_sk_...", baseURL: "https://api.pareta.ai/v1" });
const resp = await client.chat.completions.create({
model: "auto",
messages: [
{ role: "system", content: "Classify the customer support ticket into exactly one label." },
{ role: "user", content: "I was charged twice for my March invoice." },
],
response_format: TICKET_FORMAT,
});
console.log(resp.choices[0].message.content); // {"label": "..."}
Two changed strings, api_key and base_url, plus the model ID auto. See Migrating from the OpenAI SDK for the full walkthrough and when the pareta SDK's control plane (evals on your own data, auto metrics) earns the install.
Prove it on your data, then operate it
- Evaluate.
pa.evals.runs.create(...)runsautoand the frontier baselines you name on your own rows, applies one scorer to every contender, and reports quality, cost, and the metered run total. An evaluation set is 5–50 items. See Evaluating on your own data and the worked example. - Monitor.
pa.auto.metrics()rolls up yourautotraffic: requests, success rate, spend, and estimated savings against frontier list prices. See Cost & quality monitoring. - Handle errors. Every failure is a typed
ParetaErrorsubclass; an empty balance is a 402, a cold serving backend is a retried 503. See Errors, retries & timeouts.
Guide
Start-to-finish, in reading order. Every page shows Python and TypeScript. See the guide index.
- Installation & authentication — install
pareta(pip or npm), authenticate with apareta_sk_key, make a first metered call. - Quickstart — a structured production call end to end (support-ticket classification with
response_format), streaming, metering, then benchmarkingautoagainst frontier models on your data. - Core concepts — the verified-inference path behind
auto, the quality bar and escalation, hidden models and hardware, how evals score your data, metering, and the eval → production funnel. - Running inference —
chat.completions.create, structured outputs, tool calling, streaming, passthrough params, cost headers,models.list, and metering errors. - Evaluating on your own data — benchmark
autoagainst frontier baselines withevals.setsandevals.runs: quality/CIs/cost, and the metered run total. - Errors, retries & timeouts — the
ParetaErrorhierarchy, which errors to catch, and the retry policy. - Async & concurrency — Python's
AsyncParetavs TypeScript's Promise-only client, and fanning out concurrent calls. - Configuration — API key, base URL, timeouts, retries, and injecting a custom HTTP client.
Examples
Copy-paste workflows for real jobs, in both languages. See the examples index.
- Benchmark auto on your own data — eval
autoagainst frontier baselines and readrun.cost. - Document extraction (PDF/image) — the blob-task loop: upload documents, benchmark
autoon them, send documents in production. - Streaming chat completions — iterate chat chunks and accumulate text.
- Concurrent calls — fan out inference and eval calls (
asyncio.gather/Promise.all). - Cost & quality monitoring — read what calls cost and watch your
autotraffic withauto.metrics(). - Migrating from the OpenAI SDK — keep using
openaiagainst Pareta, and when to switch topareta.
Developer tools
The tools below are supported integrations for reaching the same endpoint and control plane from a shell or an AI coding agent; they are not the product. All of them use the same pareta_sk_ key.
- The
paretaCLI — the whole control plane as a shell command (pip install "pareta[cli]"), tables or--json. - MCP server — expose the control plane to an AI agent (Claude Code, Codex, Claude Desktop, Cursor) as tools (
pip install "pareta[mcp]"). - The
/paretaskill — aSKILL.mdthat drives the CLI as a slash command in Claude Code and Codex. - Connect OpenClaw to Pareta — point an OpenClaw agent at
/agent/v1as its primary model: open-weight turns with native tool calls, frontier escalation when a turn needs it.
Reference
Field-by-field API docs. Signatures are shown in Python; the TypeScript API mirrors them (camelCase names, options objects, awaited), see any guide page for the TS form. See the reference index.
- Client —
Pareta(and Python'sAsyncPareta):from_env/fromEnv, constructor params, lifecycle, and the resource namespaces. - chat.completions —
chat.completions.create, return types, streaming, and the error surface. - models —
models.list()and theModelfields. - tasks —
list/retrieve/matchand their response models. - evals —
evals.sets,evals.runs, andevals.frontierModels. - audio —
audio.transcriptions(speech-to-text) andaudio.speech(text-to-speech), metered per minute. - Exceptions — the
ParetaErrorhierarchy and status-to-class mapping. - Response types — every response object plus the
.costvs.costMicroUsdmoney convention. - Underlying HTTP API — the
/v1routes the SDK wraps (language-neutral). - Agent API (
/agent/v1) — the wire contract for agent runtimes (OpenClaw): fields, streaming, session pinning, billing, errors.
Support
- Community Slack — pareta.ai/slack: the
#supportchannel is where the Pareta team answers questions. - Email — support@pareta.ai.