The gap between demoing an AI agent and running a team of them inside a regulated enterprise is not a gap of months. It is a gap of discipline. Pilots tolerate ambiguity about who decides what, which system holds the truth, and what happens when the model is wrong. Production does not, particularly in functions like underwriting, sales, and service, where a bad automated decision reaches a customer before anyone reviews it. Vishy Kasinadhuni, a Senior Director of Enterprise Architecture who co-created a production multi-agent orchestration framework powering autonomous workflows, puts the problem bluntly: “Everyone is talking about AI agents. Far fewer have a team of them in production inside a large enterprise with real customers and real risk on the line. That is where the real work begins.” His argument is that the companies struggling with agents are not struggling with models. They are struggling because they built the agents first and went looking for the process afterward.
The Workflow Comes Before The Agents
Most agent projects start from capability. A team sees what a model can do, builds something that does it, then hunts for a place in the business to put it. Vishy inverts the sequence entirely. “Start with the workflow, then design the agents,” he says. “Map the business process end-to-end, where the decisions are made. Where does the work change hands?” That mapping exercise sounds unglamorous next to model selection, and it is exactly why it gets skipped. It is also the only thing that tells you how many agents you need, and what each one is for.
The discipline that follows is scope. “Each agent should own a clear job with defined inputs and outputs,” Vishy says, with the orchestrator deciding “who acts, when, and with what context.” Read that as a warning against the generalist agent, the single assistant pointed at a whole department and expected to figure it out. Agents with fuzzy boundaries fail in ways nobody can diagnose, because there is no contract to check the output against. Agents with defined inputs and outputs fail visibly, at a known handoff, which is the difference between a bug and a mystery. “When the design is solid,” he says, “everything built on top of it runs with a purpose.” The inverse is the more useful observation: when the design is not solid, nothing built on top of it can be fixed, only replaced.
Governance Is Architecture, Not Policy
The standard enterprise reflex is to let the technology ship and let the control function catch up. Security review arrives late, legal arrives later, and the pilot that worked beautifully in a sandbox spends two quarters in remediation. Vishy treats that sequence as a design failure rather than an inconvenience. “Build governance into the foundation,” he says. “Every agent needs an identity, the right entitlements, and policies that define what it can see and do.”
Identity for a non-human actor is a bigger idea than it first appears. An agent that can read a customer record, call an internal API, and trigger a downstream action is an employee in every meaningful operational sense, minus the badge and the training. Without an identity, there is no entitlement model, no audit record, and no way to answer the questions a regulator or a board member will eventually ask: What was this thing permitted to do? And who permitted it? Vishy describes designing a multi-tenant AI gateway and a policy engine from the start, which is the architectural form that answer takes. The payoff is counterintuitive and worth stating plainly, because it cuts against how most organizations experience controls. “When security and control live in the platform, teams move faster with confidence, and leaders can trust every action the system takes.” Governance embedded in the platform removes the per-project negotiation. Governance bolted on afterward reintroduces it every time.
Keep Humans Where Judgment Lives
The marketing language around agentic AI points toward full autonomy, and that framing has quietly set the wrong benchmark for enterprise teams. Vishy draws the division of labor differently. “The strongest frameworks keep humans involved at the moments that matter most,” he says. “Agents handle coordination and routine steps. People handle judgment, exceptions, and final decisions.” That is not a transitional arrangement pending better models. It is a description of what each party is good at, and in underwriting or service escalation, the exceptions are precisely where the money and the liability sit.
Making that split real requires instrumentation, not intention. Vishy specifies clear checkpoints, full audit trails, and strong observability “so everyone can see what happened and why.” The last three words carry the weight. Plenty of systems can tell you what an agent did. Far fewer can reconstruct the context it was handed, the policy it operated under, and the reasoning path it followed, which is what any serious review demands after an incident. Build that in and a human checkpoint becomes a decision point with real information behind it. Leave it out and the checkpoint degrades into a rubber stamp, where a person approves an output they have no practical means of evaluating, and the organization gets the appearance of oversight at the cost of its substance.
The through line in Vishy’s method is that none of the hard parts are about AI. Process mapping, identity and entitlements, audit, and clear ownership of decisions are the same foundations that separate functioning enterprise systems from expensive ones. Agents simply raise the cost of getting them wrong, because the system now acts on its own schedule. “Start with the workflow, make governance part of the architecture, and design for real collaboration between people and AI,” he says. “That’s how multi-agent systems deliver results in production.” The sequence matters as much as the content.
Follow Vishy Kasinadhuni on LinkedIn for more insights on enterprise architecture, multi-agent orchestration, and AI governance at scale.










