Arnau Mustieles Arola, Executive Manager of the Systems RUN Services Competence Center at Stratesys

An artificial intelligence demonstration can be prepared in a short time and surprise any management committee. Even so, the challenge is not in showing what AI is capable of doing in a controlled environment, but in ensuring that this capacity works in a stable, safe and useful way in real operations. It is precisely in the distance between demonstration and reliability in operation that it is decided whether AI brings value or only expectations.

And the potential is real: Generative AI can summarize documents, classify requests, locate information, compose responses or generate code in seconds. The difficulty appears when capabilities still in development are presented as if they were already ready for any organization.

AI is much easier to demonstrate than to convert into an operational capability. A controlled case is enough to obtain in seconds a result that previously required time and specialized knowledge.

This ease amplifies expectations generated in a context in which manufacturers project an almost miraculous future with their models, networks are filled with success stories, courses to “develop AI” in three days and prompts presented as universal solutions, while service providers present ongoing developments as a final product.

And that leads customers to expect that their suppliers already have capabilities ready to improve their productivity and reduce costs. The expectation is understandable, but it should be managed rigorously: between a promising test and an industrialized solution there is usually a relevant part of the effort, cost and risk of the project.

To better understand this distance, it is worth separating two dimensions that, when talking about artificial intelligence, are frequently confused: the technological maturity of a tool and the depth with which it is integrated into our daily work.

The first dimension is maturity, which measures technological reliability across three different situations. At the first level is what is technically possible: the model manages to solve a task, but only under very specific conditions. Then what has been consistently demonstrated appears: the tool has already been tested with real cases, it offers stable results and, above all, we are clear about its limits. Finally, there is what is ready for production: the solution can operate in the day-to-day life of the company with security and control, knowing how to manage imperfect data, exceptions and their impact on the business.

The second dimension is depth, which defines the degree of autonomy and scope of AI within our tasks. In its most basic form, the general-purpose co-pilot appears, which helps to write, summarize or search for information, while the user provides the context and always reviews the final result. It can also be applied to a specific activity, such as sorting through a request bin or drafting a response. A higher level of integration allows you to connect AI with systems and data, always based on permissions, to reduce manual searches and operational steps. Finally, an agent with the ability to act no longer only suggests or searches for information, but rather coordinates different tools and executes actions on its own within a workflow.

These two dimensions can intersect in very different ways. We can have a solution with very deep integration into our systems, but that is still technologically very immature. On the contrary, a simple co-pilot with limited functions may be perfectly ripe for production.

Nor should we be obsessed with always achieving maximum integration. For many activities, using a basic co-pilot well provides great value almost immediately. It is a path with less complexity and lower risk, since it works as an intelligent support for a human who, supposedly, is the one who provides the definitive criteria.

Whatever the depth chosen, an AI solution is still software: it needs data, security, testing, operation and maintenance. But LLMs require additional assessments.

The first challenge is veracity: the model can generate a plausible and well-written answer without being really supported by the available information. This forces you to control which sources you use, if they are updated and if the answer is really supported by them.

The second is to evaluate quality when there is no single correct answer. In these cases, it is not enough to pass a closed set of tests: acceptance criteria must be defined, representative cases evaluated and behavior observed during operation, which may change when the model, sources or context are modified.

The requirements will depend on the use case, its autonomy and the associated risk. In a co-pilot, the focus is primarily on protecting information, establishing usage criteria, and ensuring user review. When AI is applied to structured tasks, it is necessary to clearly define what it receives, what it should return, and how it is evaluated. If it is integrated with other systems, permit management, traceability and control of the information used are added. And if it can perform actions, you also have to decide what it can do on its own, what requires approval, and how to reverse incorrect action.

For example, in an application maintenance service, an outdated knowledge base can inadvertently deteriorate responses. An incorrect recommendation can also delay finding the solution or cause another problem in production.

All this forces us to move from the potential to the real case. In the face of constantly evolving artificial intelligence, our role as consumers must be strategic. The key is to choose the appropriate level of depth for each business need, understanding that greater integration and autonomy usually imply more implementation, operation and control effort.

AI is not a universal solution. In some tasks it offers capabilities that surpass people; In others, it extends your work as a co-pilot and, in many, it simply allows you to do the same thing faster or more comfortably. The real challenge is to evaluate the return on each use case and anticipate technological dependency, including a possible increase in supplier costs.

By Arnau Mustieles Arola, Executive Manager of the Systems RUN Services Competence Center at Stratesys