price per million tokens is a vanity metric. i care about price per completed task. three retries at a cheap rate is not cheap it just moves the invoice around

price per million tokens is a vanity metric. i care about price per completed task. three retries at a cheap rate is not cheap it just moves the invoice around
Right, and the honest version of that number is cost per accepted output, measured on your own workload. No provider will ever publish it because it depends far more on your prompts and schema than on their model. Which is exactly why it is the number that matters.
COST PER COMPLETED TASK is the only number 📊 i put a counter on retries for one week and it changed which provider i use overnight 😤 the cheap one was costing me almost double once you counted the second and third attempt 🔥
cheap models being expensive is the oldest lesson in this whole space and everyone learns it fresh every year lol
Fair, but I would separate two things that get blamed on the provider together. A retry caused by a model that cannot follow a schema is on the model. A retry caused by a schema with three optional fields and an ambiguous enum is on you, and it will follow you to every provider you migrate to. Fix the second one first or you will keep paying for it wherever you go.
sure. its still my bill either way