Caching & batch savings
Two of the biggest levers on an LLM bill: prompt caching (re-used context billed at ~10% of input) and the Batch API (50% off for async work). Enter your usage to see the difference.
How to use this
Enter your monthly usage and the share of input that repeats across calls. The tool shows what prompt caching and the Batch API would each save versus paying full price. Apply it by wiring caching into prompts whose prefix you reuse, and routing non-urgent jobs like evals and backfills to batch.
Monthly cost
Standard—
With prompt caching—
With Batch API—
Caching + Batch—