Why Sovereign Infrastructure Is Emerging as the Next Competitive Advantage
How Can Organizations Ensure AI Agents Move from Demo to Real-World Deployment? - MIT Sloan Management Review Middle East How Can Organizations Ensure AI Agents Move from Demo to Real-World Deployment? - MIT Sloan Management Review Middle East

How Can Organizations Ensure AI Agents Move from Demo to Real-World Deployment?

Many AI agent pilots show strong potential in demos, yet only a fraction are deployed at scale. What keeps them from reaching production?

Topics

  • [Image: Nomita Samaiyar/MITSMR Middle East]

    Key Takeaways

    01

    A working demo doesn’t mean it will work at scale. Pilots run in clean, controlled conditions. Real companies are messy: different teams, inconsistent data, and systems that don’t talk to each other. The gap lies here.

    02

    You can’t blame the bot or the vendor. Every AI agent needs a real person responsible for its actions.

    03

    Speeding up an old process doesn’t automatically mean savings. The real payoff comes from redesigning the workflow around the AI.

     

    In 2024, Swedish fintech company Klarna launched its AI assistant powered by OpenAI. In the first month, the AI handled 2.3 million conversations,  accounting for two-thirds of Klarna’s customer service chats. It did the work of 700 full-time employees and reduced repeat inquiries by 25%.

    Co-founder and CEO Sebastian Siemiatkowski called it a “breakthrough in customer interaction,” adding that everyone should “test, test, test, and explore.” During its Q3 2025 earnings call, Klarna said its AI agent was doing the work of 853 full-time human agents and had saved them $60 million.

    However, the optimism was short-lived. In 2025, Siemiatkowski admitted that the company may have cut costs too aggressively with AI, which led to lower service quality and customer dissatisfaction. “We probably overindexed a little bit on that, and then in the last six months we have been trying to course correct,” he said. The company began rehiring human agents. 

    The model delivered on what it ought to. However, without humans in the loop, cracks began to surface. 

    Klarna’s AI agent isn’t the only one to fall short. This points to a larger pattern. AI pilots cannot succeed on their own; they need a supportive environment to scale. Ibrahim AlJallaf, Chief Operating Officer, Inception42 (a G42 company), says, “The pattern we see most consistently isn’t the agent’s reasoning failing; it’s the environment around it not being ready.”

    Good Demo Does Not Mean It’s Production Ready

    A pilot excels in a controlled setting because the inputs are clean, the systems are available on demand, and only one team drives the interaction. Meanwhile, production is an altogether different environment, with systems updating at different times and incomplete or messy data across teams. 

    In a controlled environment, AI agents covered a limited range of systems, processes, and organizational units. “Once deployed at scale, the agent encounters the full complexity of the enterprise, including subsidiaries with different operating models, inconsistent processes, uneven data quality, and numerous local exceptions. As a result, production outcomes may be better, worse, or simply different from those observed during the pilot,” says Antonio Rizzi, VP, Solution Consulting – EMEA South, ServiceNow.

    The result? A pilot that appears finished and well-built but is far from production-ready status. “A demo proves the model works. Production proves the organization has the integration, the governance, and the ownership structure to operate around it,” says AlJallaf. 

    Joe Dunleavy, Regional CTO and Global Head of Dava.X AI Group at Endava, recommends knowing what you are building for before you actually do: “If you don’t know the problem, it’s very hard to come up with a solution to it.”

    A few key fault lines to watch out for include integration debt, governance, procurement, bad data, change management, unclear ownership, and return on investment (ROI).

    When an agent makes a mistake, the incident remains the organization’s responsibility—not the agent’s.

    — Antonio Rizzi, VP, Solution Consulting – EMEA South, ServiceNow

    Data is at the Center of AI

    There are two main facts about data in AI. First, data serves as the core of any AI system. Second, it does not have to be uniform across all departments and areas. Rizzi agrees with this. He’s seen a pilot do well in one department, but falter when scaled. “The initial team had clean, structured data, while other functions had inconsistent data and unclear ownership,” he shares.

    Analysts at Forrester note that poor data quality emerged as the single biggest factor limiting the effectiveness of companies’ adoption and scaling of generative AI. Poor-quality data can result in inaccurate predictions, biased outcomes, and confidence errors. Dunleavy calls a weak foundation a “very normal problem.” 

    “A weak foundation of data structure, architecture, and well-understood data will mean that your pilot will only go so far, and then it will not scale,” he adds.

    An MIT study revealed that only 5% of AI projects deliver measurable financial gains — 95% fail due to poor data quality.

    Research Context

    • This article is based on interviews with: 

    Ibrahim AlJallaf, Chief Operating Officer, Inception42 (a G42 company)

    Antonio Rizzi, VP, Solution Consulting – EMEA South, ServiceNow

    Joe Dunleavy, Regional CTO and Global Head of Dava.X AI Group at Endava

    • MIT State of AI in Business 2025: 95% of corporate generative AI pilots failed to deliver measurable financial impact. 
    • Forrester: Poor data quality emerged as the single biggest factor limiting the effectiveness of companies’ adoption and scaling of generative AI.
    • Gartner Forecast: Worldwide AI spending will grow 47% to touch $2.59 trillion in 2026

     

    Who Absorbs the AI Budget?

    Gartner predicts that worldwide AI spending will grow 47% to touch $2.59 trillion in 2026. So, who takes responsibility for the budget for developing the initiative in-house? It is here that our two experts pose contrasting views. For AlJallaf, the budget sits with IT infrastructure and operations, “that’s where it should stay.” For him, work such as legacy integrations, data pipes, and monitoring rarely falls under an AI or innovation budget. 

    “The infrastructure investment, data pipelines, legacy connectors, and monitoring layers are a standing operational cost that benefits every future deployment, not just one pilot. Framing it that way makes it easier for CFOs to approve, because they’re not funding a single experiment; they’re funding capability that compounds,” he adds. 

    Meanwhile, for Rizzi, it sits with a dedicated AI function led by a Chief AI Officer, where the “boring” work is usually split across budgets. “The AI team may fund foundational data-readiness projects such as data cataloging, ownership and stewardship, quality remediation, metadata and lineage, access controls, and trusted data pipelines. IT is then often expected to fund the architectural work: connecting legacy systems and applications, exposing APIs and MCP tools, implementing monitoring and security, and extending the agentic harness,” he says. 

    He adds that funding works better when the AI office, IT, and the business owner come together for an initiative, linking “those investments to a measurable workflow outcome.”

    A demo proves the model works. Production proves the organization has the integration, the governance, and the ownership structure to operate around it.

    — Ibrahim AlJallaf, Chief Operating Officer, Inception42

    Who Answers When the Agent Makes a Mistake 

    In April, US-based software startup PocketOS was hit critically after an AI agent wiped out its entire production database and backups in seconds. “An AI coding agent, Cursor running Anthropic’s Claude Opus 4.6, deleted our production database and all volume-level backups in a single API call to Railway,” said Jeremy Crane, founder of PocketOS. “It took 9 seconds.” In another case, an AI agent from Replit wiped the production database of startup SaaStr.

    Data loss is not a question of if, but of when. When something goes wrong, who takes accountability? “When an agent makes a mistake, the incident remains the organization’s responsibility—not the agent’s,” clarifies Rizzi, adding that “accountability cannot be delegated to AI.”

    Organizations must establish a governance system that maps every AI system to a handler (business owner, process owner, data owner, and technical owner) to have a clear accountability and escalation roadmap. 

    AlJallaf believes accountability should lie with those who own the process in which the agent is operating, “Not the vendor, and not an abstract ‘AI team.” “If an agent is handling procurement exceptions, the incident sits with procurement operations, the same way it would if a junior analyst made the error,” he says.

    The two outright reject the notion that the AI or the vendor is at fault in any way. 

    On designing for accountability, the Inception42 COO says, “Every agent we put into production has a defined owner before it goes live, someone whose job includes reviewing its output, not just someone who requested the pilot.”

    Prove AI ROI Before Integrating it Into Workflow

    After massive investments in capital and effort, your AI agents have been introduced into the workflow. With a predefined checklist in hand, how does an organization and a leader measure the ROI delivered by the agent over 12 months? Rizzi advises measuring ROI against a baseline established before the agent is deployed and linked to a specific business workflow. “Productivity gains should not automatically be treated as financial savings: the hours released must translate into lower costs, avoided hiring, increased output or capacity that can genuinely be redeployed,” he says.

    AlJallaf prioritizes getting the finance department on board. “The first thing we had to get finance to accept is that if you drop an agent into a workflow designed for people and only measure whether that same workflow got faster, you’ll always underestimate the return.” 

    The advantage of speed over old processes isn’t the goal. The real gain is the possibility of what can happen once the workflow is redesigned, keeping the agent and the human at its center. 

    “We avoid vague productivity claims that finance teams can’t audit,” he says.

    What Leaders Must Do Differently

    Role

    Action Required

    C-Suite 

    Governance above all. Leaders should put a framework in place before any AI pilot launches — one that requires lineage tracking, clear sign-off steps, and ongoing monitoring for every project. 

    Operational Leaders

    Leaders should initially keep high-risk, irreversible, or poorly structured activities out of scope, including actions with broad system permissions, sensitive customer decisions, or processes where data quality and accountability are unclear

    Boards & Governance

    The board should not ask “what can this agent do”; it needs to ask what the workflow actually needs and where an agentic solution genuinely multiplies what the team can do, versus where a human has to stay. 

    MIT Sloan Management Review Middle East invites you to the second edition of AI Research Forum — “Agentic AI: From Mandate to Momentum” taking place on 22 October 2026 in Dubai. 

    To partner, speak, or attend AIRF, click here.

    Topics

    More Like This

    You must to post a comment.

    First time here? : Comment on articles and get access to many more articles.

    ×