The speed with which AI generates code is changing the way we understand software development. But measuring productivity by counting lines, tasks or prompts can lead us to the same mistake as always, confusing activity with results. The measure that matters is much simpler: it is how much product we get for the resources we use.
There is a trap in the current conversation about Artificial Intelligence Agents, and it is that we are confusing capacity with productivity.
Just because a machine can program does not mean that a company will produce more software. Likewise, being able to write thousands of lines of code in seconds does not mean that it will be cheaper; even the fact that an Agent is capable of working for hours without a break, testing thousands of alternatives or solving problems that would require enormous human efforts does not answer the question that should concern us.
The question is much simpler, how much useful software do we get for the money we are spending? Everything else is secondary.
Because if we are not able to answer this, deciding between developers, suppliers, AI Agents or hybrid teams seems a lot like gambling. And managing budgets based on bets has never been a good methodology.
A recent study offered an interesting clue and was that incorporating AI into the work of programmers allowed the volume of code produced to increase by more than 50%.
But that data alone does not prove that a company is more productive. We can write twice as much code and not deliver twice as much product, we can even write a lot more and end up spending more to review, test and maintain it.
That is precisely the problem, the metric should not be how much code AI generates, but how much useful software we get for the cost necessary to produce it.
A company has clients, budgets, deadlines and products to deliver and that is why it should care much less if one model outperforms another by two points in a test and much more about how much product it achieves with the resources it assigns.
In software, this difference is fundamental. We should not measure the productivity of an Agent by the number of lines of code it generates. Nor that of a developer by the number of tasks he completes. Nor that of a supplier because of the number of people it places on a project.
The output a company buys is not lines of code, hours of work, or tokens. The result is software that works and adds value. And this is where Agents change the rules of the game.
Until now, when we needed to develop software, the decision was relatively simple, do it with our team or hire a provider. Now there is a third possibility, using AI Agents. And we will probably soon stop thinking of these alternatives as watertight compartments. The most interesting option may be a fourth, hybrid teams in which people and Agents work together.
The temptation, then, is to assume that AI will automatically win because it looks cheaper. And we already know that he has no salary, no vacations, no schedule. It can work continuously and generate code at a speed that no human developer can match. Now, there is a small problem and that is that the code is not a product.
An Agent can generate an impressive amount of code and then force us to dedicate enormous resources to reviewing, integrating, testing, fixing and maintaining it. At the same time, a developer can write relatively little code and deliver functionality that generates a lot of value and, for this reason, it is not useful to compare human hours with computing hours. You have to compare results.
If a human team produces a certain amount of useful software for a certain cost and a group of Agents produces more for that same cost, we will have an answer. If the opposite happens, too. And if a team made up of people using Agents manages to produce much more than either of them alone, we will have another answer, we no longer need to discuss it for years because we can measure it.
In fact, the arrival of Agents gives us an extraordinary opportunity to do something that the sector has been trying to do for a long time without much success and that is to truly measure the productivity of software development.
We can measure how much product our internal team achieves, how much a supplier achieves, how much Agents produce working on their own and, furthermore, how much a human team produces when it incorporates AI into its way of working. Then you just have to put those results in relation to the cost. The discussion changes completely.
We no longer need to speculate about whether programmers will disappear, whether providers will no longer be needed, or whether Agents will end up doing all the work. We can observe what happens in each case. We may even encounter a paradox and that is that an extraordinarily powerful machine is too expensive for certain problems and that the human being, apparently less sophisticated, remains economically unbeatable, however, the exact opposite can also happen.
The important thing is that now we have the possibility to verify it and that is why we should be careful about replacing one obsession with another. For years we measured development by hours, people or lines of code. Now we run the risk of starting to measure it through prompts, tokens or the number of Agents deployed. It would be the same error with a different technology because at the end of the chain there are no tokens, no prompts, nor Agents, what there is is product. And that should be the metric that matters.
The arrival of Agents gives us an extraordinary opportunity to do something that the sector has been trying for a long time: measure the productivity of software development
Every time we incorporate Artificial Intelligence into development we should be able to answer two very simple questions: Are we producing more useful software than before? How much is it costing us to get it?
Only then can we decide if we need five developers, five Agents, or two developers working with twenty Agents. We may discover that AI can replace entire teams, or that it turns certain human teams into extraordinarily productive organizations, or we may even find that for some jobs humans are still surprisingly cheap. We do not have to choose the answer by intuition, we just have to measure it because the true revolution in Artificial Intelligence does not consist of having more AI, it consists of achieving more product with the same resources.
Until we can prove that this is happening, no matter how spectacular the demo is, we will still have exactly the same thing, an impressive demo and a bill behind it.
Julián Gómez Bejarano, Chief Digital Officer LedaMC
