Extraction: documents and contracts
Pull structured fields out of documents — a scanned invoice, a PDF, a contract — and get JSON back. You'll extract vendor/total/line-items from a real invoice image, then key legal fields from a contract's text.
Extraction is a chat job: image or text in, JSON out, so it all goes through
chat.completions.create(model="auto"), OpenAI-compatible on the wire. One
interface, whatever the document looks like; each request is one debit against
the org balance regardless of internal routing.
Setup
Install the SDK (installation guide) and set
PARETA_API_KEY.
Python
from pareta import Pareta
pa = Pareta.from_env() # reads PARETA_API_KEY (and optional PARETA_BASE_URL)
TypeScript
import { Pareta } from "pareta";
const pa = Pareta.fromEnv(); // reads PARETA_API_KEY (and optional PARETA_BASE_URL)
Visual documents (images and PDFs)
A visual document goes in as OpenAI content parts: one text part stating the
fields you want, one image_url part carrying the pixels as a base64 data URI.
Send the field schema as response_format on every call whose output your code
parses: the prompt describes the task, but schema instructions in the prompt
alone are not enforced — the json_schema is the contract Pareta constrains
decoding to and validates before delivery, so the response parses without
ceremony. This example downloads a real invoice from the public examples
dataset and extracts five fields from it.
Python
import base64, json, urllib.request
INVOICE_FORMAT = { # the output contract
"type": "json_schema",
"json_schema": {
"name": "invoice_fields",
"schema": {
"type": "object",
"properties": {
"vendor": {"type": "string"},
"invoice_no": {"type": "string"},
"date": {"type": "string"},
"total": {"type": "number"},
"line_items": {
"type": "array",
"items": {
"type": "object",
"properties": {
"description": {"type": "string"},
"quantity": {"type": "number"},
"amount": {"type": "number"},
},
"required": ["description", "amount"],
"additionalProperties": False,
},
},
},
"required": ["vendor", "invoice_no", "date", "total", "line_items"],
"additionalProperties": False,
},
},
}
url = ("https://raw.githubusercontent.com/Pareta-AI/example-datasets"
"/main/invoice-extraction/documents/0.jpg")
img_b64 = base64.b64encode(urllib.request.urlopen(url).read()).decode()
resp = pa.chat.completions.create(
model="auto",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "Extract the vendor, invoice number, date, total, "
"and line items from this invoice."},
{"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{img_b64}"}},
],
}],
response_format=INVOICE_FORMAT,
max_tokens=1024,
)
fields = json.loads(resp.choices[0].message.content) # validated against the schema before delivery
TypeScript
const INVOICE_FORMAT = { // the output contract
type: "json_schema",
json_schema: {
name: "invoice_fields",
schema: {
type: "object",
properties: {
vendor: { type: "string" },
invoice_no: { type: "string" },
date: { type: "string" },
total: { type: "number" },
line_items: {
type: "array",
items: {
type: "object",
properties: {
description: { type: "string" },
quantity: { type: "number" },
amount: { type: "number" },
},
required: ["description", "amount"],
additionalProperties: false,
},
},
},
required: ["vendor", "invoice_no", "date", "total", "line_items"],
additionalProperties: false,
},
},
};
const url = "https://raw.githubusercontent.com/Pareta-AI/example-datasets" +
"/main/invoice-extraction/documents/0.jpg";
const imgB64 = Buffer.from(await (await fetch(url)).arrayBuffer()).toString("base64");
const resp = await pa.chat.completions.create({
model: "auto",
messages: [{
role: "user",
content: [
{ type: "text", text: "Extract the vendor, invoice number, date, total, " +
"and line items from this invoice." },
{ type: "image_url", image_url: { url: `data:image/jpeg;base64,${imgB64}` } },
],
}],
response_format: INVOICE_FORMAT,
max_tokens: 1024,
});
const fields = JSON.parse(resp.choices[0].message.content ?? ""); // validated against the schema before delivery
Expected output for that document:
{
"vendor": "Bradley-Andrade",
"invoice_no": "97159829",
"date": "2015-09-18",
"total": 978.12,
"line_items": [ ... ]
}
PDFs
You can hand a PDF into the same image_url slot. When a base64 data URI
declares application/pdf — or its decoded bytes start with the %PDF magic,
so a wrong or missing MIME type is still caught — the platform rasterizes it
server-side before inference: the PDF block is replaced with one PNG image
block per rendered page, up to the first 8 pages (sibling keys like
detail are preserved on each page). For longer documents, split the PDF or
send pre-rasterized page images. A PDF that cannot be rendered returns a 400
with a clear message rather than forwarding raw bytes a vision model would
reject. Plain http(s) image
URLs and URL-encoded (non-base64) data URIs pass through untouched.
Python
pdf_b64 = base64.b64encode(open("contract.pdf", "rb").read()).decode()
{"type": "image_url", "image_url": {"url": f"data:application/pdf;base64,{pdf_b64}"}}
TypeScript
const pdfB64 = (await readFile("contract.pdf")).toString("base64");
({ type: "image_url", image_url: { url: `data:application/pdf;base64,${pdfB64}` } });
Full runnable example: python/extraction/visual_document.py · typescript/extraction/visual-document.ts
Contract fields (text)
When the document is already text, there is nothing special to do: the contract
goes in as ordinary string content on the same chat surface. Declare the fields
in the response_format schema — that is the enforced contract; the prompt only
describes them — and give the model an explicit out (null) for anything the
text does not state, in the schema (["string", "null"]) and in the prompt,
otherwise it will guess. Sample contract text from CUAD (Hendrycks et al.),
CC-BY-4.0.
Python
import json
from pathlib import Path
contract_text = Path("data/sample-contract.txt").read_text()
CONTRACT_FORMAT = { # the output contract
"type": "json_schema",
"json_schema": {
"name": "contract_fields",
"schema": {
"type": "object",
"properties": {
"parties": {"type": "array", "items": {"type": "string"}},
"agreement_date": {"type": ["string", "null"]},
"effective_date": {"type": ["string", "null"]},
"governing_law": {"type": ["string", "null"]},
},
"required": ["parties", "agreement_date", "effective_date", "governing_law"],
"additionalProperties": False,
},
},
}
prompt = ("Extract the parties, agreement date, effective date, and governing law. "
"parties is a list of legal entity names; dates are YYYY-MM-DD; "
"use null for anything the text does not state.")
resp = pa.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": f"{prompt}\n\n---\n\n{contract_text}"}],
response_format=CONTRACT_FORMAT,
max_tokens=512,
)
fields = json.loads(resp.choices[0].message.content) # validated against the schema before delivery
print(fields["parties"], fields["governing_law"])
TypeScript
import { readFile } from "node:fs/promises";
const contractText = await readFile("data/sample-contract.txt", "utf8");
const CONTRACT_FORMAT = { // the output contract
type: "json_schema",
json_schema: {
name: "contract_fields",
schema: {
type: "object",
properties: {
parties: { type: "array", items: { type: "string" } },
agreement_date: { type: ["string", "null"] },
effective_date: { type: ["string", "null"] },
governing_law: { type: ["string", "null"] },
},
required: ["parties", "agreement_date", "effective_date", "governing_law"],
additionalProperties: false,
},
},
};
const prompt = "Extract the parties, agreement date, effective date, and governing law. " +
"parties is a list of legal entity names; dates are YYYY-MM-DD; " +
"use null for anything the text does not state.";
const resp = await pa.chat.completions.create({
model: "auto",
messages: [{ role: "user", content: `${prompt}\n\n---\n\n${contractText}` }],
response_format: CONTRACT_FORMAT,
max_tokens: 512,
});
const fields = JSON.parse(resp.choices[0].message.content ?? ""); // validated against the schema before delivery
console.log(fields.parties, fields.governing_law);
Full runnable example: python/extraction/contract_fields.py · typescript/extraction/contract-fields.ts
See also
- Inference (OpenAI-compatible) — the full chat surface: streaming, extra params, error handling.
- Chat reference —
chat.completions.createrequest and response shapes. - Document extraction end-to-end — the full document workflow, from your own files to production.
- Prove it on your own data: evaluate on your data — benchmark
"auto"against frontier baselines on your own documents before you commit.