Skip to main content

Pareta

Pareta is a verified-inference provider built for repeatable production AI workloads. Through one OpenAI-compatible endpoint and one model ID, "auto", your application can access Pareta’s fleet of open-weight specialists, including several distilled in-house. Pareta checks every result against the workload’s quality bar and automatically escalates to a frontier model whenever the specialist path cannot meet it. You are charged only for the model that produces the final response. For workloads handled by Pareta specialists, serving costs can be 30–140× lower than the comparable frontier model, and you can verify the quality and economics on your own production data.

These docs cover the three things you do with that endpoint, all behind one pareta_sk_ key:

  • Run inference with the model ID auto. OpenAI-compatible and metered. This SDK (Python + TypeScript) and the stock openai client are interchangeable for inference; streaming, response_format, and tools are covered in Running inference. Nothing to deploy.
  • Evaluate it on your own data. Run auto head-to-head against frontier baselines on your rows, with one scorer for every contender, and read per-contender quality and cost before you move traffic.
  • Read what each request cost. Every chat-completion response carries X-Pareta-Billed (the exact debit, in micro-USD) and X-Pareta-Frontier-Would-Have-Cost (an estimate: usage tokens × frontier list price); direct endpoints (/v1/embeddings, /v1/rerank, audio, images) carry no receipt headers and report their cost in the dashboard.

A few platform facts shape the whole API:

  • Models and GPUs are hidden. You never pick either. Pareta selects the best-fit specialist per request when a deployed specialist holds the workload's quality bar; otherwise a frontier model serves the request. Hardware is Pareta's problem.
  • Frontier (vendor) ids are in the clear. They appear as eval baselines and in auto.compare_frontier(); every open-weight specialist stays behind auto.
  • Inference and evals are metered against your org balance. A successful call debits credit once per request, no matter how many internal model calls Pareta makes to serve it: routing, verification, retries, and non-serving attempts are included. An empty balance raises InsufficientCreditsError (402). An eval run reports its billed total on run.cost (dollars). Top-up is browser-only; the SDK never touches billing.

Pareta is infrastructure for AI tasks embedded in products and automated workflows, not a destination chatbot or general coding copilot.

Install​

Python

pip install pareta # or: uv add pareta / poetry add pareta

TypeScript

npm install pareta # or: pnpm add pareta / yarn add pareta / bun add pareta

Python or TypeScript? Both SDK clients are at full parity. The one design difference: Python ships sync (Pareta) and async (AsyncPareta) clients; TypeScript has a single Promise-only Pareta (every method is async). Code samples throughout these docs show Python and TypeScript side by side.

A first production call​

Mint a pareta_sk_ key in the dashboard, export it as PARETA_API_KEY, and send a repeatable task. The example below classifies support tickets into a fixed label set. The output contract is response_format with a json_schema: Pareta constrains decoding to the schema and validates the result before delivery, so label is always one of yours.

Python

import json
from pareta import Pareta

pa = Pareta.from_env() # reads PARETA_API_KEY

LABELS = ["billing", "refund", "login", "bug", "feature_request", "other"]

TICKET_FORMAT = { # the output contract
"type": "json_schema",
"json_schema": {
"name": "ticket_label",
"schema": {
"type": "object",
"properties": {"label": {"type": "string", "enum": LABELS}},
"required": ["label"],
"additionalProperties": False,
},
},
}

def classify(ticket: str) -> str:
resp = pa.chat.completions.create(
model="auto",
messages=[
{"role": "system", "content": "Classify the customer support ticket into exactly one label."},
{"role": "user", "content": ticket},
],
response_format=TICKET_FORMAT,
)
return json.loads(resp.choices[0].message.content)["label"]

print(classify("I was charged twice for my March invoice.")) # one of LABELS, e.g. "billing"

TypeScript

import { Pareta } from "pareta";

const pa = Pareta.fromEnv(); // reads PARETA_API_KEY

const LABELS = ["billing", "refund", "login", "bug", "feature_request", "other"];

const TICKET_FORMAT = { // the output contract
type: "json_schema",
json_schema: {
name: "ticket_label",
schema: {
type: "object",
properties: { label: { type: "string", enum: LABELS } },
required: ["label"],
additionalProperties: false,
},
},
};

async function classify(ticket: string): Promise<string> {
const resp = await pa.chat.completions.create({
model: "auto",
messages: [
{ role: "system", content: "Classify the customer support ticket into exactly one label." },
{ role: "user", content: ticket },
],
response_format: TICKET_FORMAT,
});
return JSON.parse(resp.choices[0].message.content ?? "{}").label;
}

console.log(await classify("I was charged twice for my March invoice.")); // one of LABELS, e.g. "billing"

Schema instructions written only in the prompt ("reply with JSON") are not enforced; response_format with a json_schema is the contract. See Structured outputs for what is validated, JSON mode, and the limits.

Already using the OpenAI SDK?​

The same call works from the stock openai client: use your Pareta API key, point the OpenAI client at Pareta's base URL, and set the model ID to auto. TICKET_FORMAT is the schema from the example above.

Python

from openai import OpenAI

client = OpenAI(api_key="pareta_sk_...", base_url="https://api.pareta.ai/v1")

resp = client.chat.completions.create(
model="auto",
messages=[
{"role": "system", "content": "Classify the customer support ticket into exactly one label."},
{"role": "user", "content": "I was charged twice for my March invoice."},
],
response_format=TICKET_FORMAT,
)
print(resp.choices[0].message.content) # {"label": "..."}

TypeScript

import OpenAI from "openai";

const client = new OpenAI({ apiKey: "pareta_sk_...", baseURL: "https://api.pareta.ai/v1" });

const resp = await client.chat.completions.create({
model: "auto",
messages: [
{ role: "system", content: "Classify the customer support ticket into exactly one label." },
{ role: "user", content: "I was charged twice for my March invoice." },
],
response_format: TICKET_FORMAT,
});
console.log(resp.choices[0].message.content); // {"label": "..."}

Two changed strings, api_key and base_url, plus the model ID auto. See Migrating from the OpenAI SDK for the full walkthrough and when the pareta SDK's control plane (evals on your own data, auto metrics) earns the install.

Prove it on your data, then operate it​

  • Evaluate. pa.evals.runs.create(...) runs auto and the frontier baselines you name on your own rows, applies one scorer to every contender, and reports quality, cost, and the metered run total. An evaluation set is 5–50 items. See Evaluating on your own data and the worked example.
  • Monitor. pa.auto.metrics() rolls up your auto traffic: requests, success rate, spend, and estimated savings against frontier list prices. See Cost & quality monitoring.
  • Handle errors. Every failure is a typed ParetaError subclass; an empty balance is a 402, a cold serving backend is a retried 503. See Errors, retries & timeouts.

Guide​

Start-to-finish, in reading order. Every page shows Python and TypeScript. See the guide index.

  • Installation & authentication — install pareta (pip or npm), authenticate with a pareta_sk_ key, make a first metered call.
  • Quickstart — a structured production call end to end (support-ticket classification with response_format), streaming, metering, then benchmarking auto against frontier models on your data.
  • Core concepts — the verified-inference path behind auto, the quality bar and escalation, hidden models and hardware, how evals score your data, metering, and the eval → production funnel.
  • Running inference — chat.completions.create, structured outputs, tool calling, streaming, passthrough params, cost headers, models.list, and metering errors.
  • Evaluating on your own data — benchmark auto against frontier baselines with evals.sets and evals.runs: quality/CIs/cost, and the metered run total.
  • Errors, retries & timeouts — the ParetaError hierarchy, which errors to catch, and the retry policy.
  • Async & concurrency — Python's AsyncPareta vs TypeScript's Promise-only client, and fanning out concurrent calls.
  • Configuration — API key, base URL, timeouts, retries, and injecting a custom HTTP client.

Examples​

Copy-paste workflows for real jobs, in both languages. See the examples index.

Developer tools​

The tools below are supported integrations for reaching the same endpoint and control plane from a shell or an AI coding agent; they are not the product. All of them use the same pareta_sk_ key.

  • The pareta CLI — the whole control plane as a shell command (pip install "pareta[cli]"), tables or --json.
  • MCP server — expose the control plane to an AI agent (Claude Code, Codex, Claude Desktop, Cursor) as tools (pip install "pareta[mcp]").
  • The /pareta skill — a SKILL.md that drives the CLI as a slash command in Claude Code and Codex.
  • Connect OpenClaw to Pareta — point an OpenClaw agent at /agent/v1 as its primary model: open-weight turns with native tool calls, frontier escalation when a turn needs it.

Reference​

Field-by-field API docs. Signatures are shown in Python; the TypeScript API mirrors them (camelCase names, options objects, awaited), see any guide page for the TS form. See the reference index.

  • Client — Pareta (and Python's AsyncPareta): from_env/fromEnv, constructor params, lifecycle, and the resource namespaces.
  • chat.completions — chat.completions.create, return types, streaming, and the error surface.
  • models — models.list() and the Model fields.
  • tasks — list/retrieve/match and their response models.
  • evals — evals.sets, evals.runs, and evals.frontierModels.
  • audio — audio.transcriptions (speech-to-text) and audio.speech (text-to-speech), metered per minute.
  • Exceptions — the ParetaError hierarchy and status-to-class mapping.
  • Response types — every response object plus the .cost vs .costMicroUsd money convention.
  • Underlying HTTP API — the /v1 routes the SDK wraps (language-neutral).
  • Agent API (/agent/v1) — the wire contract for agent runtimes (OpenClaw): fields, streaming, session pinning, billing, errors.

Support​