Before AI, Technology costs were easy to control. Executives via the finance department would control costs through the hiring process and procurement process. Cost were controlled by limiting the number of developers. In actual fact, most organisations do not control costs, they control the day rate or the amount paid per day (week, month or year). Each year (or quarter), the executives would decide how much they want to spend on technology for the rest of the year and adjust head count accordingly. The main benefit of this approach is that the maximum cost for a time period (e.g. a month or a year) was controlled.
Very few organisations have effective control of the costs of investments. There are too many sources of uncertainty and variability. Finance can control the maximum cost of a team per month but they cannot control the number of months that the team will be required to deliver an investment. A widely used industry heuristic (but not in finance) is that the cost of the investment will always cost twice what you think it will cost. Todd Little’s research showed the actual versus estimate follows a log normal (Wiebel) distribution with a mean of x2 and in ten percent of cases x4.
This approach of controlling head count meant that technology projects never unexpectedly spend more than budget in a month or year, even if the cost of the investments themselves resemble a bus crash in slow motion. The source of bankruptcy is not going to be unexpected cost overspend for a period, even though an investment may eventually drag an organisation down in a tragically visible manner.
AI token usage will lead to a loss of control due to unplanned spending. Unlike headcount, maximum cost of token usage cannot be controlled. A developer (AI or human) fires off an agentic job to generate code (or some other task). For the past few months, the cost of an agentic job has been $100, however for some reason this task is different and the cost is $50,000. As with human developers, the actuals versus estimates of jobs follows a distribution which has a long tail.
The obvious answer is to provide a limited number of tokens for each developer. Now the developers keep hitting the limit half way though jobs and both time and tokens are lost… which drive down speed and efficiency. The manager handing out extra tokens soon gets fed up with the approval process and automates it. If you think you can control costs using spending limits, think how comfortable you are going to feel about telling your boss the following…
- “We sent all the developers home because we ran out of tokens.”
- “The developers wanted the day off so they created an agent to burn through all the tokens generating images of cats doing cool things.”
- “We hit the monthly token limit for this project, so we are switching to a project with tokens left.”
Although controlling token usage will be a nightmare, a bigger nightmare will be the transparency that token usage will bring. Investments will be assigned a token budget. It is well known that development teams often need to switch developer time between budgets…. A few days of support that gets booked as time spent on the strategic initiative. Development teams get away with it because there is no way to track what a developer is actually working on. Token usage will be tracked against the investment being worked on. At the end of the month, the business investor will look at the tokens charged to their investment and demand to know what has been done. “Jazz hands explanation” that the Agent was just doing some thinking or working on some technical debt wont be a justification.
The solution? I do not know. I’d be keen to hear from practitioners who have cracked this one.
It is clear that Technology Departments will need to be better at providing transparency into the money they spend and which investments or operational activity it gets spent on. ( I do have a solution for this. đ )
Leave a Reply