You’re talking about the cost to the consumer, where I was talking about the energy cost of an individual unit of work by a given model, but they’re both relevant. The underlying energy cost of generating a token at a given level of capability has been falling as hardware and inference become more efficient. Those efficiency gains can feed through into lower token prices for users, but also a more capable model can often complete the same task with fewer tokens, fewer retries and less prompting.
So even if the headline price of a new frontier model looks similar or higher, the actual cost, both in energy and money, of getting a given piece of work done can still fall substantially. It’ll keep doing so too, its still in its infancy.
You’re talking about the cost to the consumer, where I was talking about the energy cost of an individual unit of work by a given model, but they’re both relevant. The underlying energy cost of generating a token at a given level of capability has been falling as hardware and inference become more efficient. Those efficiency gains can feed through into lower token prices for users, but also a more capable model can often complete the same task with fewer tokens, fewer retries and less prompting.
So even if the headline price of a new frontier model looks similar or higher, the actual cost, both in energy and money, of getting a given piece of work done can still fall substantially. It’ll keep doing so too, its still in its infancy.