Skip to main content

Core concepts

Pareta is one OpenAI-compatible endpoint with one model ID, auto, built for repeatable production workloads. This page covers the handful of ideas the rest of the SDK assumes you understand: the verified-inference path behind auto (how Pareta selects a specialist, checks the result against the workload's quality bar, and escalates), one interface per data shape, tasks (how evals score your data), open-weight vs frontier models, why models and hardware are hidden, how metering works, and the funnel that ties them together (prove auto on your data, ship it, watch the metrics).

Every code block below is runnable as written. They all start from a client:

Python

from pareta import Pareta

pa = Pareta.from_env() # reads PARETA_API_KEY (+ optional PARETA_BASE_URL)

TypeScript

import { Pareta } from "pareta";

const pa = Pareta.fromEnv(); // reads PARETA_API_KEY (+ optional PARETA_BASE_URL)

from_env() is the path you want in almost every case. The explicit form is Pareta(api_key="pareta_sk_...", base_url="https://api.pareta.ai"); arguments are keyword-only. See Authentication for key minting (browser-only) and The client for timeouts, retries, and the async AsyncPareta mirror.

The verified-inference path behind auto​

Every request you send with the model ID auto takes the same path. Pareta selects the best-fit open-weight specialist for the workload (several are distilled in-house), checks the result against the workload's quality bar, and automatically escalates to a frontier model when the specialist path cannot meet it. When no deployed specialist holds the bar for a task, the request is served by a frontier model directly; a frontier-served response is returned as served. One request, one bill, for the model that produces the final response — routing, verification, retries, and non-serving attempts are included.

"Checks" means output parsing and schema validation (including your response_format), constrained decoding where the task defines it, and model-graded checks on tasks where the specialist's measured quality calls for them. Verification is a quality gate, not a guarantee that an AI output can never be wrong.

There is nothing to deploy and no model to pick — "which model?" is the question Pareta answers for you, per request. models.list() reflects that: it returns exactly one entry.

Python

for m in pa.models.list():
print(m.id) # exactly one entry: "auto"

TypeScript

for (const m of await pa.models.list()) {
console.log(m.id); // exactly one entry: "auto"
}

Calling auto is plain chat — see Inference is OpenAI-compatible below. The surfaces around that call live on pa.auto:

  • auto.metrics() — your org's auto traffic, rolled up: requests + success rate (30d), spend, hourly p50/p95/error buckets (7d), daily success cells (30d), and the estimated savings against frontier list prices.
  • auto.compare_frontier(model=..., messages=...) (TypeScript auto.compareFrontier({ model, messages })) — one prompt against a frontier vendor for a side-by-side with "auto". Metered at the vendor's actual token cost; a failed vendor call bills $0. Allowed models: gpt-5.5, gemini-3-5-flash, gemini-3-1-pro, claude-sonnet-4-6.

One interface per data shape​

You never choose a model anywhere on Pareta — you choose the data shape, and each shape has exactly one interface:

Your dataInterfaceBehind it
messages in, text outchat.completions with model="auto"the verified-inference path behind auto
a query + documents to rankreranka Pareta reranking specialist
text to turn into vectorsembeddingsa Pareta open-weight embedding specialist
audio in / audio outaudiothe speech lanes
a prompt to turn into an imageimagesthe image-generation lane

Within every shape, routing, model choice, and escalation are Pareta's job. The separate routes exist because vectors, ranked lists, images, and audio bytes don't fit the chat message contract — not because there is anything to navigate.

Only the chat shape goes through auto's route → verify → escalate loop and carries the X-Pareta-* receipt and diagnostics headers. Embeddings, reranking, images, and audio are direct endpoints: each is served by its own Pareta model at its own per-unit rate, outside that loop, and auto never routes a chat request to them. model="auto" belongs to chat.completions only — on the direct endpoints it has no meaning.

Tasks: how evals score your data​

Internally, the quality bar behind auto comes from a catalog of benchmarked tasks — Pareta has measured open-weight and frontier models against each on real data. You don't navigate that catalog to use Pareta. You meet it in exactly one place: benchmarking on your own data, where a task says how your results are scored — the row shape your dataset must follow and the scorer that grades outputs against your labels.

Every task has a stable id (e.g. "contract-key-fields"), a default_scorer (the function that grades a model's output — field-F1, nDCG@10, WER, judge panel), and a has_blob_input flag (true when the rows carry documents or images, not just text).

Python

for task in pa.tasks.list():
print(task.id, task.default_scorer, "blob" if task.has_blob_input else "text")

# Fetch one task, optionally with sample rows to see its input shape
t = pa.tasks.retrieve("contract-key-fields", examples_n=3)
print(t.id, t.default_scorer, t.has_blob_input)

TypeScript

for (const task of await pa.tasks.list()) {
console.log(task.id, task.defaultScorer, task.hasBlobInput ? "blob" : "text");
}

// Fetch one task, optionally with sample rows to see its input shape
const t = await pa.tasks.retrieve("contract-key-fields", { examplesN: 3 });
console.log(t.id, t.defaultScorer, t.hasBlobInput);

Rather than reading the scorer list, describe your dataset in plain English and tasks.match tells you how it will be scored:

Python

m = pa.tasks.match("vendor invoices with labeled line items and totals")
if m.matched:
print("grade with:", m.chosen.task_id) # the grading contract your eval will resolve to

TypeScript

const m = await pa.tasks.match("vendor invoices with labeled line items and totals");
if (m.matched && m.chosen) {
console.log("grade with:", m.chosen.taskId);
}

match() raises ValueError on an empty query. A no-match answer is a statement about scoring — no benchmarked task fits that description — not about serving: generation work always goes to model="auto". See Tasks for the full matcher surface.

Open-weight vs frontier models​

Two kinds of model stand behind every task:

  • Open-weight models are Pareta's specialists — the models auto selects between, including several distilled in-house. Pareta benchmarks them, serves them, and picks between them — you never call one directly or learn its identity.
  • Frontier models are hosted vendor models (OpenAI, Google, Anthropic, and so on). They play two roles: the model auto escalates to when the specialist path cannot meet the workload's quality bar (or serves directly when no deployed specialist holds the bar), and the baseline you measure auto against in evals. An eval shows, on your rows, whether auto meets the bar the frontier baseline sets, and what each contender costs.

Frontier (vendor) ids appear in the clear — those are public products — in exactly two places: eval baselines and auto.compare_frontier(). To enumerate the frontier roster you can evaluate against (annotated for your eval set's bound grading contract via task=eval_set.task_id), use evals.frontier_models:

Python

for fm in pa.evals.frontier_models(task=eval_set.task_id):
print(fm.id, fm.vendor, "vision" if fm.vision else "text",
"(benchmarked)" if fm.benchmarked else "")

TypeScript

for (const fm of await pa.evals.frontierModels("contract-key-fields")) {
console.log(fm.id, fm.vendor, fm.vision ? "vision" : "text",
fm.benchmarked ? "(benchmarked)" : "");
}

Passing task= annotates each model's benchmarked flag (measured on that task) and filters the roster by capability (for example, only vision-capable models are returned for document tasks). Feed the id values into an eval run's frontier= list.

Models are hidden​

You never pick a model, and open-weight identities never cross the API. The only model id you send is "auto"; the only model ids you read back are "auto" and frontier vendor ids in eval and comparison results (result.model_id).

This is a feature, not an omission. There are no open-model ids to look up, hard-code, or keep current — when Pareta promotes a better specialist behind a task, your requests get it on the next call, with no code change and no migration.

Hardware is hidden​

You never choose a GPU, tensor-parallel degree, quantization scheme, or serving mode. The specialists behind "auto" run on serving classes Pareta resolves from its registry, and capacity — warm pools, autoscaling, cold starts — is Pareta's problem. The one place serving infrastructure surfaces in the SDK is EndpointNotReadyError (503): a serving backend behind auto is warming or briefly unavailable. The SDK retries 503s automatically, so you rarely see it.

Inference is OpenAI-compatible​

Call auto through chat.completions.create with model="auto". The request and response match the OpenAI chat schema, so the official openai client works against the same base URL and key.

Python

resp = pa.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "Extract the contract effective date."}],
temperature=0, # extra OpenAI params pass straight through
)
print(resp.choices[0].message.content)
print(resp.usage.total_tokens)

TypeScript

const resp = await pa.chat.completions.create({
model: "auto",
messages: [{ role: "user", content: "Extract the contract effective date." }],
temperature: 0, // extra OpenAI params pass straight through
});
console.log(resp.choices[0].message.content);
console.log(resp.usage.totalTokens);

Streaming yields ChatCompletionChunk objects; the incremental text is on chunk.choices[0].delta.content:

Python

for chunk in pa.chat.completions.create(model="auto", messages=[...], stream=True):
print(chunk.choices[0].delta.content or "", end="")

TypeScript

for await (const chunk of pa.chat.completions.create({ model: "auto", messages: [{ role: "user", content: "Summarize this support ticket thread in two sentences:\n\n..." }], stream: true })) {
process.stdout.write(chunk.choices[0].delta.content || "");
}

create() raises ValueError up front if model or messages is empty. See Running inference for streaming details and the async iterator form.

Metering and billing​

Both inference and evals are metered against your organization's balance.

  • Inference: a successful chat.completions.create() debits the org balance once per request, for the model that produces the final response — no matter how many internal model calls Pareta makes to serve it. Routing, verification, retries, and non-serving attempts are included.
  • Speech: the pa.audio namespace (pa.audio.transcriptions(...), pa.audio.speech(...), Python-only) is billed per minute of audio — see Audio.
  • Evals: evals.runs.create() debits for the compute it spends: "auto" and any frontier baselines you include. A FAILED run is not charged.
  • Frontier comparisons: auto.compare_frontier() is metered at the vendor's actual token cost; a failed vendor call bills $0.
  • Empty balance: every path raises InsufficientCreditsError (HTTP 402).

Python

from pareta import InsufficientCreditsError

try:
resp = pa.chat.completions.create(model="auto", messages=[{"role": "user", "content": "Extract the contract effective date."}])
except InsufficientCreditsError:
print("Top up the org balance in the dashboard, then retry.")

TypeScript

import { InsufficientCreditsError } from "pareta";

try {
const resp = await pa.chat.completions.create({ model: "auto", messages: [{ role: "user", content: "Extract the contract effective date." }] });
} catch (e) {
if (e instanceof InsufficientCreditsError) {
console.log("Top up the org balance in the dashboard, then retry.");
} else {
throw e;
}
}

Topping up is browser-only. The SDK never exposes the balance, payment methods, or top-up. It only consumes credit and surfaces the 402 when there is none.

Reading cost off an eval run​

An eval run reports what it cost. The SDK follows one money convention (SDK_PLAN §6): the billed total is floored to whole cents so the SDK never overstates a charge, while sub-cent precision stays available in micro-USD.

  • run.cost is a Decimal in dollars, floored to cents. A 5 µUSD run reads Decimal("0.00").
  • run.cost_micro_usd is the raw integer (1_000_000 = $1.00).
  • Per-item unit rates such as result.mean_cost_micro_usd stay in micro-USD. Flooring them to cents would erase the auto-vs-frontier comparison that is the whole point.

Python

print(run.cost) # Decimal("0.42"): billed dollars, floored to cents
print(run.cost_micro_usd) # 420715: raw micro-USD

TypeScript

console.log(run.cost); // "0.42": billed dollars (string), floored to cents
console.log(run.costMicroUsd); // 420715: raw micro-USD

The proof funnel​

The pieces above compose into one path from "I have a job" to "auto is running it in production, cheaper." This is the recommended flow:

eval on YOUR data -> model="auto" in production -> watch the metrics
  1. Eval "auto" against frontier baselines on your own data. Public benchmarks are a starting point; your rows are the deciding vote.
  2. Ship model="auto" — the same call, now carrying production traffic.
  3. Watch auto.metrics() — requests, success rate, spend, estimated savings against frontier list prices.

Python

from pareta import Pareta

pa = Pareta.from_env()

# 1. Evaluate "auto" against frontier baselines on YOUR rows.
# Pass items + prompt to create the eval set inline — Pareta works out
# the scoring from them — or use an existing set id.
run = pa.evals.runs.create(
prompt="extract the effective date from each contract",
items=[
{"input": {"contract_text": "...your contract text..."}, "expected_output": {"effective_date": "2026-01-01"}},
# ...more rows...
],
models=["auto"],
frontier="benchmarked", # the benchmarked frontier baselines
wait=True, # block until the run is terminal
)

# 2. Read results (quality + cost), then ship the same call to production
for r in sorted(run.results, key=lambda r: (r.quality_mean or 0), reverse=True):
print(r.model_id, r.kind, r.quality_mean, r.mean_cost_micro_usd, f"n={r.n_succeeded}")

print("eval cost:", run.cost) # Decimal dollars, floored to cents

resp = pa.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "...your contract text..."}],
)

# 3. Watch it in production
m = pa.auto.metrics()
print(m["requests_30d"], m["success_rate_30d"], m["savings_vs_frontier_micro_usd_30d"])

TypeScript

import { Pareta } from "pareta";

const pa = Pareta.fromEnv();

// 1. Evaluate "auto" against frontier baselines on YOUR rows.
// Pass items + prompt to create the eval set inline — Pareta works out
// the scoring from them — or use an existing set id.
const run = await pa.evals.runs.create({
prompt: "extract the effective date from each contract",
items: [
{ input: { contract_text: "...your contract text..." }, expected_output: { effective_date: "2026-01-01" } },
// ...more rows...
],
models: ["auto"],
frontier: "benchmarked", // the benchmarked frontier baselines
wait: true, // block until the run is terminal
});

// 2. Read results (quality + cost), then ship the same call to production
for (const r of [...run.results].sort((a, b) => (b.qualityMean ?? 0) - (a.qualityMean ?? 0))) {
console.log(r.modelId, r.kind, r.qualityMean, r.meanCostMicroUsd, `n=${r.nSucceeded}`);
}

console.log("eval cost:", run.cost); // dollar string, floored to cents

const resp = await pa.chat.completions.create({
model: "auto",
messages: [{ role: "user", content: "...your contract text..." }],
});

// 3. Watch it in production
const m = await pa.auto.metrics();
console.log(m.requests_30d, m.success_rate_30d, m.savings_vs_frontier_micro_usd_30d);

A few notes on the eval call:

  • Provide either eval_set=<id> (an existing set) or items=... + prompt=... to create one inline (task= is optional — pass it to pin a specific task). With neither, create() raises ValueError.
  • frontier= accepts None/"none" (no baselines), an explicit list of frontier ids, "all" (every frontier model for the task), or "benchmarked" (only the frontier models measured on this task, vision-filtered for document tasks). Keyword resolution needs to know the task; with eval_set=, the SDK looks the task up for you.
  • wait=True polls until the run reaches "completed" or "failed" (run.is_terminal), then returns the final EvalRun. For document tasks, attach binaries with evals.sets.upload_document(...) before running.

For the full eval API (building sets, attaching documents, inline vs. existing sets, and polling semantics) see Evaluating models. For the catalog and matcher surface in depth, see Tasks.

Errors at a glance​

Every SDK error subclasses ParetaError. The status-mapped subclasses let you branch on what went wrong without inspecting status codes:

ExceptionStatusWhen
AuthenticationError401bad or missing key
InsufficientCreditsError402org out of credit (top up in the dashboard)
PermissionDeniedError403the user lacks permission
NotFoundError404unknown task or run
ConflictError409transient contention (auto-retried)
RateLimitError429throttled (auto-retried)
EndpointNotReadyError503a serving backend behind auto is warming or briefly unavailable (auto-retried)
BadRequestError400/422malformed request
APIConnectionError / APITimeoutErrorn/atransport failure (auto-retried)

Python

import pareta

try:
resp = pa.chat.completions.create(model="auto", messages=[{"role": "user", "content": "Extract the contract effective date."}])
except pareta.EndpointNotReadyError:
print("A backend is warming; retries are exhausted — try again shortly.")
except pareta.InsufficientCreditsError:
print("Out of credit. Top up in the dashboard.")
except pareta.ParetaError as e:
print("request failed:", e)

TypeScript

import { EndpointNotReadyError, InsufficientCreditsError, ParetaError } from "pareta";

try {
const resp = await pa.chat.completions.create({ model: "auto", messages: [{ role: "user", content: "Extract the contract effective date." }] });
} catch (e) {
if (e instanceof EndpointNotReadyError) {
console.log("A backend is warming; retries are exhausted — try again shortly.");
} else if (e instanceof InsufficientCreditsError) {
console.log("Out of credit. Top up in the dashboard.");
} else if (e instanceof ParetaError) {
console.log("request failed:", e);
} else {
throw e;
}
}

See Error handling for the full hierarchy, the request_id attribute for support, and the retry policy.