Stop Piloting, Start Delegating Your AI


The pilot purgatory most enterprises sit in is a management failure, not a technology failure
Every large company I work with can show me an AI agent that works. It reconciles invoices, drafts the first pass of a supplier negotiation, resolves the routine half of a customer service queue, or writes and tests code faster than the team that reviews it. Then I ask the question that matters, which is how much of the business actually runs through it, and the answer is almost always the same. The agent itself is in a pilot. It has been in a pilot for eleven months. There is a slide about expanding its scope next quarter, and there was a nearly identical slide last quarter.
The conventional explanation for this is that the technology is not ready, that the models still hallucinate, that governance is immature, that the risks of letting software act on the company's behalf are simply too high to move faster. I want to argue the opposite. The agents are, in most cases, ready enough. What is not ready is the management system around them, and pilot purgatory is what it looks like when an organization has bought a capability it has not yet decided to trust.
The numbers describe a delegation problem
The evidence that AI stalling in a perpetual pilot is a management problem rather than a technical one is sitting in the data. Research my firm, Genpact, conducted with HFS Research this summer surveyed more than 2,000 executives across sixteen industries and fourteen functions. The research estimates roughly 18 trillion dollars in enterprise value trapped behind four accumulated debts in data, process, technology, and talent; it additionally finds that 85 percent of leaders acknowledge those debts are actively limiting what they can get from AI. The detail underneath that headline is where the delegation problem becomes visible. Only about a third of enterprise data is ready for AI to use, and more than 40 percent of AI and analytics initiatives are failing on data quality alone. Roughly 40 percent of employee time still disappears into manual and ungoverned processes, which means the workflows agents are being asked to join were never designed to be handed to anyone, human or machine. Core systems average a decade in age, and only about a third of the workforce is ready to operate alongside agents at all. Most tellingly, more than half of the companies surveyed have no funded initiative to address core system aging, and only 6 percent are what the full report calls “proven debt resolvers” with revitalization programs operating at scale. Nearly 13 percent of average function spend is now flowing into AI, and it is flowing into organizations that have not cleared the ground for optimal AI intervention.
Read those figures together and the story is not that agents fail. The story is that agents are being deployed into organizations that never wrote down how their own work gets done, never cleaned the data the work depends on, and never decided who would be accountable for an outcome that a machine produced. A pilot is a very comfortable place to leave that decision unmade.
A pilot is a way of not deciding
Consider what a pilot actually is. It is a controlled environment in which the agent's output is reviewed by a human before it touches anything real, in which the volume is small enough that nobody's job depends on it, and in which failure has no consequences beyond a lessons-learned document. Those are precisely the conditions under which the agent can never demonstrate value, because the value of an agent is that it acts without waiting for a person, at a volume no person could sustain. An agent that is supervised on every action is an expensive intern, and the company is paying for a capability it has structurally forbidden itself from using.
Anyone who has managed people will recognize this pattern; it is exactly how weak delegation looks. A manager who assigns work and then reviews every line of it has not delegated anything; they have simply added a step. Real delegation means defining the outcome, setting the boundaries within which the person is free to act, agreeing what will be measured, and then getting out of the way while retaining accountability for the result. Pilot purgatory is what happens when leaders are willing to give an agent a task but unwilling to give it a mandate, and the reason is rarely that the technology has not earned it. It is that nobody has done the uncomfortable organizational work of deciding what the mandate should be.
That work, deciding how best to operationalize AI agents, is hard in ways that have nothing to do with software. You have to decide which exceptions the agent may resolve alone and which must escalate, which means someone has to inventory the exceptions, which means admitting that the documented process covers perhaps two thirds of what actually happens. You have to give the agent clean, current data, which means funding the unglamorous remediation your data debt has deferred for a decade. You have to attach a named executive to the output, which means someone senior accepts that a machine acting under their authority may make a visible mistake. None of that is a research problem. All of it is a leadership problem, and a pilot lets everyone postpone it indefinitely.
What do the companies that escaped AI pilot purgatory have in common?
The organizations I have watched move from pilot to production did not have better models, and in several cases they had noticeably worse data than peers who remained stuck. What they had was a different sequence. They picked a single end-to-end process with a measurable outcome and a clear owner, they rebuilt that process around what the agent could do rather than inserting the agent into the process as designed for people, and they gave the agent real authority within defined limits from the first week, with a human reviewing outcomes rather than actions. In one consumer business I advised, the pivotal moment was not a technical milestone at all. It was the operations lead agreeing to be measured on the agent's cycle time and error rate as though it were a member of her team, at which point the agent stopped being an experiment and started being a delegated function with a manager.
Three moves follow from this for any leader, or any soon-to-be leader, facing a pilot that has overstayed its welcome. First, set a date by which the pilot either goes to production with real authority or is shut down, because an open-ended pilot is a decision deferred, and deferred decisions accumulate debt of their own. Second, redesign the workflow before you scale the agent, since an agent dropped into a process built for humans will inherit every manual handoff and every undocumented exception that process contains. Third, assign a named owner who is accountable for the agent's outcomes, not its supervision, because a process that nobody is measured on is a process that nobody is running.
Why does this matter for the HBS classes of 2027 and 2028?
Most readers of this newspaper will not spend their careers deciding whether to build agents. They will spend their careers deciding what to delegate to them, and they will inherit organizations that have been pretending to make that decision for years. The single most valuable skill in the first decade of your post-MBA career may turn out to be the one the case method has always tried to teach, which is the ability to define an outcome, set the boundaries, assign the accountability, and then let go.
Technology has already made that possible. What it cannot do is make the decision for you, and the companies still sitting in pilot purgatory are the proof.

Charisma Glassman (HBS Class of GMP 2023 and EMBA Columbia Business School 2026) is Global Head of AI Advisory for Retail, Consumer, and eCommerce at Genpact, where she leads AI transformation programs across the Americas, Europe, Australasia, and Japan.




Comments