Pricing

Simple, prepaid, and per token. Point your OpenAI-compatible client at Pareta, use model: "auto", and pay for the tokens of the answer — the routing, verification, and orchestration that produced it are on us.

$0.05

per 1M input tokens
Every model: "auto" answer served by Pareta's own models — one rate, no model tiers to pick.

$0.10

per 1M output tokens
Real usage tokens, metered per request. The response reports exactly what it cost.

How billing works

You pay for the answer, not the machinery. Behind one request Pareta may plan, route, verify, and re-check — none of that is billed. The meter charges the answer model's own token consumption, once per request.

When only a frontier model clears the bar, auto says so and serves it: that request is billed at the provider's list price instead of the token rates. You see which happened on every response — the X-Pareta-Billed header carries the exact amount, and the dashboard reconciles to the penny.

Prepaid, with $30 free. New accounts start with a $30 credit and no card. Top up when you're ready; calls on an empty balance return a clean 402, never a surprise invoice.

What that buys, measured

The same battery behind the homepage numbers — production serving path, real metered billing:

Task Quality vs frontier Cost vs frontier
Intent classification113%88× cheaper
Text embedding109%32× cheaper
ICD-10 medical coding106%138× cheaper
Invoice extraction103%142× cheaper
Document reranking101%31× cheaper
Contract field extraction96%20× cheaper

Retrieval models

standalone rates · see Retrieval models
pareta-embed  $0.004 / 1M input tokens
pareta-rerank  $0.025 / 1,000 documents

Images & audio

unit-priced per generation / minute
Image generation, editing, speech-to-text and text-to-speech are metered per unit at the rates shown in the product — same prepaid balance, same receipts.
Start with $30 on us. No card, no sales call — sign up, get a key, and the first several hundred million tokens are free. Questions about volume or dedicated capacity: info@pareta.ai.
Start free →

Source: measured Pareta benchmark runs, July 2026, on the production serving path customers call. Frontier cost = usage tokens × vendor list price; Pareta = actual metered billing. Quality normalized to the strongest frontier model per task. Frontier set: GPT-5.5, Claude Opus 4.7/4.8, Sonnet 4.6, Gemini 3.5 Flash, OpenAI text-embedding-3-large. Token rates apply to answers served by Pareta-hosted models; frontier-served answers bill at provider list price.