Medical coding (ICD-10)
Turn a clinical discharge summary into ICD-10-CM codes with one chat call. The
summary goes to model="auto" as plain text; response_format with a
json_schema pins the output to {"codes": [...]} — a JSON object holding an
array of ICD-10-CM code strings; you parse the object and have structured codes.
Medical coding is a text-in, structured-text-out job, so it rides the standard
OpenAI-compatible chat surface — the one interface for every text workload.
"auto" is the only model id; the completion is metered against your org
balance, one debit per request regardless of internal routing.
Setup
Python
from pareta import Pareta
pa = Pareta.from_env() # reads PARETA_API_KEY (and optional PARETA_BASE_URL)
TypeScript
import { Pareta } from "pareta";
const pa = Pareta.fromEnv(); // reads PARETA_API_KEY (and optional PARETA_BASE_URL)
See installation for keys and environment.
Code a discharge summary
The output contract is response_format: a json_schema declaring
{"codes": [...]} — an object with one array of code strings. Pareta constrains
decoding to that schema on its own specialists and validates every answer
against it before delivery, so the content always parses into exactly that
shape. The prompt describes the task; schema instructions written only in the
prompt are not enforced, so the schema is the contract your code relies on.
With response_format set, temperature is not applied — structured serving
runs at its canonical settings. max_tokens just needs headroom for the
array; 512 covers a typical inpatient stay.
Python
PROMPT = "Assign ICD-10-CM codes for the discharge summary below.\n\n" + DISCHARGE_SUMMARY
resp = pa.chat.completions.create(
model="auto",
response_format={
"type": "json_schema",
"json_schema": {
"name": "icd_codes",
"strict": True,
"schema": {"type": "object", "properties": {"codes": {"type": "array", "items": {"type": "string"}}},
"required": ["codes"], "additionalProperties": False},
},
},
messages=[{"role": "user", "content": PROMPT}],
max_tokens=512,
)
raw = resp.choices[0].message.content or ""
TypeScript
const PROMPT = "Assign ICD-10-CM codes for the discharge summary below.\n\n" + DISCHARGE_SUMMARY;
const resp = await pa.chat.completions.create({
model: "auto",
response_format: {
type: "json_schema",
json_schema: {
name: "icd_codes",
strict: true,
schema: { type: "object", properties: { codes: { type: "array", items: { type: "string" } } },
required: ["codes"], additionalProperties: false },
},
},
messages: [{ role: "user", content: PROMPT }],
max_tokens: 512,
});
const raw = resp.choices[0].message.content ?? "";
Full runnable example: python/icd-coding/icd_coding.py · typescript/icd-coding/icd-coding.ts
Parse the codes
The content is validated against the schema before delivery, so it is always a
bare JSON object — no markdown fence, no prose — and json.loads is the whole
parser. Keep one guard anyway: check that codes really is a list, so a wrong
shape fails loudly at the boundary rather than flowing downstream.
Python
import json
def parse_codes(text: str) -> list[str]:
codes = json.loads(text)["codes"] # schema-conformant JSON, by contract
if not isinstance(codes, list):
raise ValueError(f"expected a list of codes, got {type(codes).__name__}")
return codes
for code in parse_codes(raw):
print(code)
TypeScript
function parseCodes(text: string): string[] {
const codes = JSON.parse(text).codes; // schema-conformant JSON, by contract
if (!Array.isArray(codes)) throw new Error(`expected a list of codes, got ${typeof codes}`);
return codes;
}
for (const code of parseCodes(raw)) {
console.log(code);
}
Full runnable example: python/icd-coding/icd_coding.py · typescript/icd-coding/icd-coding.ts
Nothing to pick
There is no coding model in this example because there is nothing to name:
"auto" recognizes medical-coding traffic and routes it internally to the
right serving path, per request, server-side. Your code stays a plain chat call
with a json_schema contract — the routing is Pareta's job, not a parameter.
Full runnable example: python/icd-coding/icd_coding.py · typescript/icd-coding/icd-coding.ts
See also
- Inference (OpenAI-compatible) — the full chat surface, streaming, extra params.
- Chat reference —
chat.completions.createrequest and response shapes. - Streaming chat — token-by-token output for long generations.
- Evaluating on your data — benchmark
"auto"on your own summaries, metered like inference.