Last year, I attended a workshop that my primary-school-aged son was taking part in. The children were divided into small groups. A ceramic object was placed in front of them, and their task was to recreate it with their own hands.
They had received a short introduction beforehand. They were shown how to use the material, the basic principles of shaping it, and the object they were expected to make. But the training was not deep enough to teach the craft itself. The properties of clay, balance of form, proportion, drying, and how to recognise and correct mistakes were naturally left at a surface level.
When the activity began, the children got to work with real enthusiasm. They talked, laughed, and played while trying to shape the clay in front of them. Every child participated, added their own interpretation, and enjoyed the satisfaction of making something tangible.
By the end of the workshop, the tables were full of ceramic pieces. But many bore little resemblance to the original model. The proportions were off. Some parts were missing. Some forms had collapsed. Others had become something entirely different from what had been intended.
And yet, the children were happy.
They had made something. They had used their hands, put in effort, and ended with a visible result. They were proud of it. Most did not notice how far the result had drifted from the original example.
That experience made me think about how organizations are using AI.
Producing with AI is not the same as producing the right result
AI now enables people to write content, build software prototypes, analyse information, prepare client presentations, draft strategy documents, and propose technical solutions in fields where they may have little prior experience.
That accessibility is valuable. AI lowers the threshold for creating. It provides speed, confidence, and a useful first draft. CB Insights notes that low-code and no-code tools are now allowing teams without deep AI expertise to build and deploy AI agents.
But a crucial distinction is often missed:
Producing an output with AI does not mean that the output is correct, reliable, fit for purpose, or usable in real operating conditions.
Someone can generate a polished report. Yet its data may be incomplete, its assumptions may be wrong, and its recommendations may not fit the company’s customers, operating model, technical debt, or regulatory obligations.
A team can create a software prototype that appears to work in a few hours. Yet that software may be insecure, impossible to scale, costly to maintain, non-compliant, or unable to survive real production traffic.
A leader can ask AI for a strategic plan and receive compelling headings, tables, and action items. Yet the plan may not reflect the company’s market, customers, financial realities, operational constraints, or priorities.
Just as the children felt successful because they had produced a ceramic object, teams can feel successful because an AI system has generated an output. But success is not simply creating something. Success means solving the right problem, at the right quality, with measurable value, in a way that holds up in the real world.
The data supports the concern
This is more than an observation or a metaphor. In CB Insights’ Q4 2025 survey of 59 executives, 80% said AI-agent adoption was a priority, but 40% could not track — or did not know — the return on investment from those agents.
In other words, enterprises are deploying AI agents faster than they can measure them.
Moreover, the metrics organizations use often favour easy-to-count efficiency indicators over whether the system is creating meaningful business outcomes. In the same CB Insights research, productivity gains were tracked by 63% of respondents, cost savings by 58%, and time saved by 58%. Only 25% tracked revenue impact.
That is the organizational equivalent of the ceramics workshop: there is an object on the table, the child is pleased, and the activity looks complete — but no one has systematically checked the object against the model, the intended function, or an explicit quality standard.
The real issue is not the tool. It is missing context.
An AI model rarely has the full context required to do an organization’s work well. It does not automatically know a company’s history, customers, workflows, frontline realities, technical debt, legal obligations, data quality, team capabilities, or hidden operational risks.
CB Insights makes this point directly: generic agents fail in complex enterprise environments without relevant context. Its Q4 2025 research also identifies integration challenges with existing systems and internal expertise gaps as the leading barriers to implementing AI agents.
Consider a maritime organization asking AI to design a ship-to-shore data-synchronization architecture. The response may look modern and persuasive: cloud services, APIs, event streams, message queues, and an analytics layer.
But the proposal remains incomplete until key operating questions have been answered:
- How intermittent, expensive, or latency-constrained is connectivity at sea?
- What are the real bandwidth, packet-loss, and delay characteristics of satellite communication?
- Which protocols do legacy operational systems actually support?
- What must continue working when the vessel is offline?
- Which data is personal, commercially sensitive, or regulated?
- What are the safety, maintenance, and compliance consequences of a synchronization failure?
- Which measures will determine whether the system is successful?
Without these answers, an AI-generated architecture can resemble an attractive ceramic piece: impressive on a table, but unable to perform the function it was meant to serve.
The most dangerous failure is silent failure
One of the hardest aspects of AI output is that it does not always announce when it is wrong. Fluent prose, apparently working code, and persuasive slides can create the appearance of a correct analysis, a secure system, or a deployable strategy.
CB Insights calls this silent failure: agents can fail silently in complex environments without mechanisms to monitor, test, and evaluate their behaviour. That is why observability and evaluation became the most active generative-AI market tracked by CB Insights by deal count across 91 markets.
The main barriers cited by organizations make the risk concrete:
| Deployment barrier | Share of respondents | What it means in practice |
|---|---|---|
| Reliability and security | 47% | Output accuracy and data privacy are not assured |
| Implementation and integration | 41% | Connecting AI to legacy systems and real enterprise data is difficult |
| Talent and change management | 35% | Technical deployment alone is insufficient; expertise and adoption discipline are needed |
This is why the issue is bigger than writing better prompts. It is about designing the full system: domain knowledge, data infrastructure, integration, validation, and accountability.
What creates real value
Choosing a capable model is necessary, but it is not sufficient. Organizations need to build several conditions together.
| Required condition | What happens without it | Practical response |
|---|---|---|
| Domain expertise | The wrong problem may be solved and errors may go unnoticed | Include subject-matter experts in design, testing, and approval |
| Trusted data and context | Output becomes superficial, misleading, or irrelevant | Provide current, verified enterprise knowledge with defined access rules |
| System integration | The agent produces recommendations but has little workflow impact | Connect data sources, workflows, and permissions securely |
| Clear objective | AI produces many things, but not the right thing | Define the decision, operational goal, and intended outcome first |
| Success metrics | “It looks good” becomes the definition of success | Measure quality, accuracy, error rate, speed, cost, risk, and business impact |
| Observability and evaluation | Silent failures move into production | Establish test sets, pre-release validation, continuous monitoring, and feedback loops |
| Human oversight | Confident errors reach customers or decision-makers | Design risk-based human approval and exception handling |
Human oversight is not a last-minute control added because AI is weak. It is part of the system design. CB Insights similarly identifies a human-in-the-loop model as a way to mitigate reliability and security risk through oversight and co-creation.
Experimentation is not production
There is an important nuance in the ceramics story. The children’s outputs were not failures if the purpose of the workshop was to have fun, understand the material, spark curiosity, and gain confidence in making things. By that standard, the workshop succeeded.
The same distinction applies to AI. Imperfect first outputs can be entirely appropriate — and often useful — for ideation, drafting, exploration, creating research questions, and rapid prototyping. In these settings, AI’s value is not that it always produces a finished answer; it is that it accelerates thought and experimentation.
But the standard changes for a client proposal, production software, an investment decision, a regulated process, a safety-critical operation, or corporate strategy. In those situations, “an output was generated” is not a meaningful success criterion.
The real question is:
Can this output do the intended job in real conditions — reliably, audibly, and with measurable results?
The optimistic conclusion: the workshop is maturing
This is not an argument against AI. It is an argument for using AI in a way that can earn trust and create durable value.
As the gaps become more visible, a new production layer is emerging to address them: observability tools that monitor agent behaviour, memory and knowledge-access systems that preserve enterprise context, evaluation and testing capabilities, and cost-to-outcome tools that connect AI spending to real business results.
In other words, the workshop is no longer only handing out clay. It is slowly acquiring the molds, measuring tools, kiln monitors, and experienced instructors it was missing.
AI puts the clay in our hands. But creating the right form still requires domain knowledge, context, trusted data, disciplined iteration, measurement, and human responsibility.
The biggest risk in the AI era is not that people will be unable to create. It is that they will be able to create too much, too quickly, and mistake visible output for real value.
Not every polished sentence is sound analysis. Not every working demo is secure software. Not every attractive presentation is executable strategy. And not every ceramic shape is a faithful recreation of the model in front of it.
Note: The industry data cited in this article is drawn from CB Insights research published between 2024 and 2026 on AI-agent adoption, ROI, observability, enterprise context, and deployment barriers.