Skip to content

AI margin calculator

Put in the tokens you used and what you billed, and this shows what you kept. Prices come from the same daily-refreshed catalog that prices agent work inside FlatHours, so cached reads are charged at the cache rate rather than the input rate, which is where most back-of-envelope estimates go wrong by a factor of ten.

No sign-up, no email Nothing you type is sent to us

What you used

Token counts, straight off your provider's usage page.

Cached input is charged at the cache rate, not the input rate. Missing that is how a back-of-envelope estimate comes out ten times too high on a long-running agent.

What you billed for it

You kept

$0.00

0% of what you billed

Billed
$0.00
AI cost
$0.00

Prices are list rates for 2,526 models, refreshed daily. If you have negotiated pricing or you batch, your real cost is lower than this.

Price data from MyTokenTracker, used under CC BY 4.0.

How this works

Three token types, three prices. Input is what you sent, output is what came back, and cached input is what the provider had already seen and charges a fraction for. Most rough estimates treat all three as input, which on a long-running agent that re-sends the same context hundreds of times overstates the cost by an order of magnitude.

The arithmetic is tokens divided by a million, times the published price per million, summed across models, and rounded once at the end so a call made of several cheap parts is not rounded up three separate times.

Then the part that matters: against what you billed. Spend on its own cannot tell you whether the work paid for itself. A month where AI cost four hundred dollars is excellent or terrible depending entirely on a number that is not on your provider's dashboard.

Where the numbers come from

Model prices from MyTokenTracker (mytokentracker.io), used under CC BY 4.0 and refreshed daily. These are published list prices; negotiated rates, batch pricing and free tiers are all lower. This does not include the human time spent supervising the work, which is usually the larger cost and belongs on your timesheet.

Or have it counted for you

FlatHours lets an AI agent report its own work through the API, carrying both a bill rate and its cost, so margin per client is a report rather than an exercise. Agents are free and unlimited on every plan that has the API.

Questions

How do I work out what my AI usage costs?

Take the token counts your provider reports, and multiply by that model's published price per million tokens: input, output and cached input are charged at three different rates. Cached reads are usually a fraction of the input price, so treating them as normal input can overstate your cost several times over.

Why does cost per client matter more than total AI spend?

Because spend on its own cannot tell you whether the work paid. A month where AI cost you $400 is excellent if it produced $12,000 of billable delivery and terrible if it produced $500. The number worth watching is what you kept, not what you spent.

Where do these prices come from?

The same catalog that prices agent work inside FlatHours: {{ number_format($catalogSize) }} models across 84 providers, refreshed daily from MyTokenTracker under CC BY 4.0. They are list prices, so if you have negotiated a rate or you are on batch pricing, your real cost is lower.

Should I bill clients for AI usage?

Most people bill for the delivery rather than the tokens, the same way nobody itemises electricity. The cost still matters, because it is the difference between a margin you are choosing and one that is happening to you. What you should not do is bill the hours a human would have taken while quietly paying for minutes of compute.

Does this include the time a person spent supervising?

No, and that is usually the bigger number. Reviewing, correcting and re-prompting is human time and belongs on your timesheet at your rate. Treat this figure as the floor of what the work cost you, not the whole of it.