For the past two years, the main enterprise AI question was simple: how do we give employees an AI assistant?
That question is already becoming outdated.
The next stage of enterprise AI is not one extremely powerful assistant sitting in front of every worker. It is a network of specialized AI agents, each responsible for a different part of a business process, working together under a shared set of rules.
One agent might gather information. Another might check the numbers. A third might compare the result with company policy. A fourth might prepare an action. A fifth might review the work before anything reaches a customer, regulator, employee, or payment system.
That is a very different model from asking a chatbot a question.
It also fits New York unusually well.
New York businesses run some of the most complicated knowledge workflows in the world. A bank does not simply “analyze a company.” Research, risk, legal, compliance, credit, operations, client service, trading, and reporting may all touch the same decision. A healthcare system does not simply “schedule a patient.” Insurance, provider availability, clinical need, location, records, referrals, and patient preferences can all become part of the workflow.
The same pattern appears in advertising, insurance, real estate, law, media, professional services, technology, and commerce.
A single AI agent can certainly call multiple tools and complete several steps. But as the number of systems, rules, responsibilities, and approval points grows, one general-purpose agent begins to carry too much context and too much authority.
That is where the multi-agent model becomes interesting.
The shift is already visible in New York. BNY now describes its digital employees as multi-agent systems. Citi has launched an internal agent platform called Arc and publicly describes a future wealth workflow in which a team of agents works together. JPMorganChase lists multi-step, multi-agent tasks as a formal AI research area. S&P Global is designing products for autonomous multi-agent systems, while New York-headquartered Kyndryl is building enterprise infrastructure around multi-agent orchestration.
But there is an important qualification.
Multi-agent AI is not automatically better AI. It is more expensive, harder to observe, harder to secure, and easier to over-engineer. Anthropic has reported large performance gains from a multi-agent research system on some complex research tasks, but it also found that multi-agent systems can consume roughly 15 times as many tokens as ordinary chat interactions.
The real opportunity for New York enterprises, therefore, is not to create as many agents as possible.
It is to decide where dividing work between specialized agents creates enough additional business value to justify the extra complexity.
That is the question this article will answer.
The Short Version: One AI Agent Is Starting to Look Like One Employee Trying to Run an Entire Department
The first generation of enterprise generative AI was mostly about answers. Employees asked questions, requested summaries, drafted emails, searched internal information, or generated code.
The second generation gave AI tools.
An agent could search a database, open a file, use software, run code, call an API, or complete a transaction.
Multi-agent AI introduces a third idea: division of labor.
Instead of asking one model to understand everything, remember everything, use every tool, follow every policy, and complete every step, an enterprise can create several agents with narrower responsibilities.
A research agent researches. A risk agent checks risk. A documentation agent creates the record. A policy agent checks the action against company rules. An execution agent performs an approved action. An orchestrator decides which agents should work and in what order.
This is not just a theoretical architecture anymore. OpenAI’s Agents SDK formally supports orchestration using manager agents, specialist agents, handoffs, and agents that call other agents as tools. Microsoft Agent Framework 1.0 similarly supports enterprise multi-agent orchestration. The Agent2Agent protocol, or A2A, moved under Linux Foundation governance and reported support from more than 150 organizations by April 2026.

At the same time, the plumbing around agents is becoming more standardized. The Model Context Protocol, or MCP, is increasingly used to connect agents with enterprise tools and data, while A2A focuses on communication between agents. MCP’s July 2026 specification added substantial scalability and authorization improvements, and its August roadmap explicitly highlighted agent identity and enterprise security as priorities.
That matters because the biggest barrier to multi-agent enterprise software may not ultimately be model intelligence.
It may be coordination.
What Exactly Is a Multi-Agent AI System?
The simplest way to understand multi-agent AI is to compare it with a team.
Imagine a New York private equity firm conducting diligence on an acquisition.
A single AI agent could theoretically receive the instruction:
“Analyze this company and tell me whether we should investigate it further.”
The agent could search documents, review financial statements, study competitors, examine management commentary, identify risks, and write a report.
That sounds convenient. It is also a lot to ask from one system.
A multi-agent approach breaks the assignment into separate responsibilities.
The financial agent examines revenue, margins, debt, cash flow, and working capital. The market agent investigates competitors and industry trends. The customer agent studies reviews and customer concentration. The legal agent searches for litigation and regulatory issues. The synthesis agent receives structured findings from the specialists and prepares the investment committee brief.
None of those agents necessarily needs the same instructions, the same model, the same tools, or the same permissions.
That separation is the point.
Multi-agent does not simply mean “more AI”
A workflow that makes five calls to the same chatbot is not necessarily multi-agent.
A useful multi-agent architecture normally contains some meaningful separation between roles. Different agents may have different objectives, context, tool access, memory, instructions, security permissions, or models.
Anthropic describes multi-agent systems as systems in which several agents work together, with its research architecture using a lead agent that delegates work to specialized subagents running in parallel. OpenAI documentation similarly describes manager-and-specialist architectures and handoffs between agents.
The important word is specialized.
If every agent receives the same prompt, accesses the same data, and performs the same search, an enterprise has probably multiplied cost without creating meaningful specialization.
The basic enterprise architecture
A practical multi-agent workflow may look like this:
| Layer | Job |
| Intake agent | Understand the request and collect missing information |
| Orchestrator | Decide which specialist agents should work |
| Specialist agents | Perform separate pieces of research, analysis, checking, or execution |
| Verification agent | Check evidence, calculations, policy, or completeness |
| Human approval | Review high-impact or uncertain decisions |
| Execution agent | Perform the authorized transaction or system update |
| Audit layer | Record what happened, why it happened, and which systems were touched |
The exact structure will vary. A three-agent system may be enough for one workflow, while another process could require ten specialists.
More agents should never be the goal.
Cleaner separation of work should be the goal.
Why the Single-Agent Model Starts to Break
Single agents remain extremely useful. In many cases, they are exactly what a company should use.
The problem appears when businesses keep adding responsibilities.
First the agent receives access to documents.
Then CRM.
Then email.
Then finance data.
Then customer support.
Then the legal knowledge base.
Then calendar access.
Then the ability to update Salesforce.
Then permission to send messages.
Then payment authority.
Eventually one AI system is expected to understand hundreds of instructions, thousands of possible tools, several security policies, multiple business functions, and a huge amount of context.
The architecture begins to resemble one employee who has simultaneously been asked to become a banker, lawyer, analyst, operations manager, database administrator, security officer, and executive assistant.
That is not specialization. It is overload.
Context becomes a scarce resource
Every AI model has limits on how much useful information it can actively manage.
Even when context windows become extremely large, filling them with every policy, customer record, tool description, conversation, instruction, workflow state, and document can create noise.
Subagents provide a different solution.
Instead of putting everything into one context, the research agent gets research context. The compliance agent gets compliance context. The customer agent receives the customer information it actually needs.
Anthropic has specifically identified context isolation as one reason multi-agent systems can work well. Its research architecture lets subagents explore deeply in separate contexts and return compressed findings to the lead agent.
Tool overload becomes a design problem
A general-purpose enterprise agent may eventually have access to hundreds or thousands of actions.
That introduces confusion.
Which customer database should it use? Which financial system is authoritative? Which tool should update an address? Which application contains the final approved contract? Which action requires permission?
Specialization narrows the choice.
The accounts-receivable agent might have six carefully selected tools. The recruiting agent might have another five. The security agent might not be allowed to access customer communications at all.
A narrower tool surface can make systems easier to test and easier to govern.
Permissions become easier to understand
The moment an agent moves from reading information to changing something, permissions become central.
A research agent may need read access.
A refund agent may need authority to issue refunds up to $100.
A treasury agent might be permitted to prepare a transaction but never release funds.
A compliance agent may need permission to stop a workflow but no permission to execute it.
That separation is much harder when one enormous agent possesses every capability.
NIST has made agent identity and authorization a major focus of its 2026 work, specifically warning that organizations need controls around agents accessing different tools, datasets, and applications.
Multi-agent architecture gives enterprises a natural place to apply those boundaries.
Why This Shift Is Happening Now
Enterprises could have built systems involving multiple models before 2026.
The difference is that several pieces of infrastructure are maturing at roughly the same time.
Models are better at planning. Tool use is becoming more reliable. Agent frameworks now include orchestration features. Protocols are emerging for connecting agents with tools and with other agents. Observability is improving. Identity and governance are moving closer to the center of product design.
The pieces are beginning to look less like separate experiments and more like a software stack.
Agent frameworks are becoming production infrastructure
OpenAI’s current Agents SDK supports specialists, handoffs, agents used as tools, guardrails, tracing, sandbox environments, and programmable orchestration. Microsoft Agent Framework reached version 1.0 in April 2026 and combines ideas from Semantic Kernel and AutoGen, including multi-agent orchestration and support for A2A and MCP.
This changes the enterprise calculation.
Engineering teams no longer have to invent every coordination mechanism themselves.
Agents are getting a language for speaking to other agents
Google introduced A2A in 2025 as a protocol for agent-to-agent communication. The project later moved to the Linux Foundation, and by April 2026 the foundation reported participation from more than 150 organizations as well as production use across several industries.
A2A is particularly important for large enterprises because companies rarely use one vendor for everything.
A Microsoft-built agent may need to cooperate with an internal Python agent. A finance vendor may expose another specialist. A customer-service platform may have its own agent. An acquired business might operate a completely different technology stack.
If every connection requires custom code, multi-agent architecture becomes expensive very quickly.
Common protocols reduce that problem.
MCP is helping standardize access to tools and data
MCP solves a different part of the puzzle.
It gives AI applications a standardized method for accessing tools and context. Its 2026 roadmap has increasingly focused on enterprise issues such as authorization, auditability, scalability, and agent identity.
Think of the distinction this way.
MCP helps an agent communicate with things.
A2A helps an agent communicate with other agents.
Enterprises will still use ordinary APIs, event streams, databases, queues, and custom integrations. These protocols do not replace traditional software engineering.
They simply make agentic systems easier to connect.
Original NYC Tech Journal Research: What Public Evidence Actually Shows in New York
The phrase “multi-agent AI” is now used so frequently that it can become difficult to distinguish a real implementation from a product announcement.
So NYC Tech Journal created a small original public-signal audit.
The objective was not to rank companies or claim that public information reveals everything happening inside them. It was to answer a narrower question:
What concrete signs can we find that large New York organizations are moving from AI assistants toward action-taking, orchestrated, or multi-agent systems?
Our methodology
We reviewed nine large organizations headquartered in New York City or deeply tied to New York’s enterprise economy.
For each organization, we used first-party company sources wherever possible and coded five public signals:
Explicit multi-agent language means the organization specifically discusses multiple agents, agent teams, or multi-agent systems rather than simply using the word “agentic.”
Agent or orchestration platform means the company describes infrastructure designed to create, connect, deploy, or manage agents.
Connected operational workflow means the AI system is described as interacting with real workflows, applications, transactions, infrastructure, or business processes rather than only generating text.
Governance or control signal means the source discusses permissions, auditing, policy controls, identity, bounded authority, or similar operational safeguards.
Quantified scale signal means the company publishes a measurable number showing meaningful AI deployment or platform scale. The number does not have to refer exclusively to multi-agent AI, so this column should be interpreted as evidence of enterprise AI readiness rather than proof of multi-agent deployment.
We only marked a signal when the public source supported it. We did not infer that a company had deployed multi-agent AI simply because its executives discussed the technology.
Table 1: NYC Enterprise Multi-Agent Public Signal Audit
| Organization | Explicit multi-agent signal | Agent/orchestration platform | Connected operational workflow | Governance/control signal | Quantified scale signal |
| BNY | Yes | Yes | Yes | Yes | Yes |
| Citi | Yes | Yes | Yes | Yes | Yes |
| JPMorganChase | Yes, R&D | Not clearly disclosed | Yes, R&D | Yes | Not in reviewed source |
| S&P Global | Yes | Yes | Yes | Yes | Not directly comparable |
| Kyndryl | Yes | Yes | Yes | Yes | Yes |
| American Express | Not explicit | Yes | Yes | Yes | Yes |
| Verizon | Not explicit | Yes | Yes | Not explicit in reviewed evidence | Not comparable |
| NYU Langone Health | Not explicit | Not publicly described as orchestration platform | Yes | Not explicit in reviewed evidence | Yes |
| BlackRock | Not explicit | Copilot rather than multi-agent platform | Mainly decision support | Yes | Not in reviewed source |
Sources include BNY’s annual report and proxy materials, Citi’s Arc announcement, JPMorganChase’s AI research program, S&P Global product and research materials, Kyndryl’s agentic architecture announcements, American Express’s 2026 shareholder letter and ACE platform, Verizon’s agentic network products, NYU Langone’s FY2026 report, and BlackRock’s Aladdin Copilot documentation.
Chart 1: How Common Are the Signals in Our Nine-Organization Sample?
| Public signal | Organizations | Share of sample |
| Connected operational workflow | 8 of 9 | 89% |
| Governance/control evidence | 7 of 9 | 78% |
| Agent/orchestration platform | 6 of 9 | 67% |
| Explicit multi-agent language | 5 of 9 | 56% |
| Some quantified enterprise AI scale evidence | 5 of 9 | 56% |
Visual view
Connected workflow ████████░ 8/9
Governance/control ███████░░ 7/9
Agent platform ██████░░░ 6/9
Explicit multi-agent █████░░░░ 5/9
Quantified AI scale █████░░░░ 5/9
This is not a statistically representative survey of New York companies. It is a deliberately selected enterprise sample designed to study public architecture signals, and the categories involve some judgment despite the conservative coding methodology.
The result is still revealing.
The public conversation has already moved much further than “employees are using chatbots.”
Original Finding #1: Multi-Agent Architecture Is Arriving Faster Than Clear Evidence of Multi-Agent Production
More than half of the organizations in our sample now have public material that explicitly mentions multi-agent systems or agent teams.
That sounds like widespread adoption.
It is not.
The strongest evidence in the sample comes from BNY. Its 2025 annual report describes its digital employees as multi-agent systems, while its 2026 proxy says 99% of employees were on its Eliza platform, with 160 enterprise AI solutions in production and 134 digital employees live.
That is unusually concrete.
Citi is also important. Arc is an internal platform for building and scaling agents, and Citi describes those agents as monitored, auditable, and governed. The company says more than 80% of the 180,000 employees with access to Citi AI tools use them regularly. But Citi’s wealth example describes a future team of AI agents preparing information for bankers, so it should not be presented as proof that this specific multi-agent workflow is already live.
JPMorganChase provides another type of evidence. Its AI research organization explicitly lists work on “multi-step multi-agent tasks,” including computer-use actions and human assistance, while its security organization is already publishing detailed thinking about runtime controls for autonomous agents. That demonstrates serious technical investment, but public R&D does not automatically equal enterprise-wide deployment.
S&P Global and Kyndryl provide strong architecture signals as technology suppliers. S&P Global has designed data infrastructure for environments ranging from controlled pipelines to autonomous multi-agent systems. Kyndryl explicitly markets multi-agent design and orchestration through its Agentic AI Framework and Kyndryl Bridge.
The pattern matters.
Enterprise architecture is moving toward multi-agent faster than public evidence of large-scale multi-agent production.
That is exactly what we should expect during an infrastructure transition.
Original Finding #2: New York’s Economic Structure Makes Agent Orchestration More Important Here Than in Many Markets
New York has another reason to pay close attention.
The city contains an enormous concentration of expensive, complex, information-heavy work.
NYC Comptroller data show that in May 2026 the city had approximately 520,640 financial-activities jobs, 221,000 information jobs, and 806,300 professional and business-services jobs. Together, those three categories contained about 1.548 million jobs out of 4.854 million total nonfarm jobs.
That is approximately 31.9% of all New York City payroll employment.
But employment tells only part of the story.
The Comptroller’s 2025 economic report lists 2024 average wages of $309,863 in financial activities, $215,222 in information, and $150,621 in professional and business services. The report groups these as the city’s high-wage office-using sectors.
Using the Comptroller’s reported employment and average-wage figures, NYC Tech Journal calculated a simple payroll proxy.
Table 2: Economic Surface Area of NYC’s High-Wage Office Sectors
| Sector | 2024 employment | Average wage | Estimated payroll proxy |
| Financial Activities | 483,264 | $309,863 | ~$149.7B |
| Information | 216,299 | $215,222 | ~$46.6B |
| Professional & Business Services | 751,800 | $150,621 | ~$113.2B |
| Combined | 1,451,363 | — | ~$309.5B |
NYC Tech Journal calculation using NYC Comptroller employment and wage data.
Using the same report’s citywide employment and average-wage figures gives a total payroll proxy of roughly $547.7 billion.
That means these three sectors represented about 31.3% of the jobs but approximately 56.5% of the payroll proxy in the 2024 dataset.
Chart 2: NYC High-Wage Office Sectors — Jobs Versus Payroll Proxy
Share of jobs
██████░░░░░░░░░░░░░░ 31.3%
Share of payroll proxy
███████████░░░░░░░░░ 56.5%
The calculation is intentionally simple. Employment multiplied by average annual wage is not the same thing as a formal measure of total compensation or economic output, and sector classifications contain many jobs that will not be suitable for AI automation.
But it shows why enterprise AI matters so much in New York.
A large part of the city’s wage base sits inside organizations where employees coordinate information, approvals, analysis, documents, customers, transactions, and compliance.
These are exactly the environments where workflow orchestration becomes valuable.
Original Finding #3: New York Already Shows Exceptionally High AI Intensity, but Most Enterprise AI Is Still Narrow
The NYC Comptroller’s February 2026 analysis used Anthropic Economic Index data to study work-related Claude use.
It found that New York’s per-capita Claude use was higher than every U.S. state, trailing only Washington, D.C. The analysis also found an extraordinary concentration in computer and mathematical occupations: roughly 43% of New Yorkers’ work-related Claude conversations were associated with that occupational group.
Yet broader company adoption still appears shallow in many organizations.
The Federal Reserve Bank of Atlanta’s 2026 survey of nearly 750 corporate executives found that more than half of surveyed firms had already invested in AI and more than 80% expected to invest during 2026. Large firms were considerably more likely to invest.
The NYC Comptroller, drawing on that survey and other data, noted that 57% of adopting firms had integrated AI into no more than three business functions. It also described current generative AI use as still heavily centered on writing, document analysis, and information search rather than end-to-end process automation.
That gap is critical.
New York has intense AI usage.

New York also has large enterprises with the technical foundations to build agents.
But much of enterprise AI is still local rather than organizational.
Multi-agent systems are one possible bridge from individual productivity to company-level workflow redesign.
BNY Offers One of the Clearest New York Examples of the Shift
BNY deserves special attention because its language has moved beyond vague promises.
Its strategy includes scaling agentic solutions and enabling increasingly autonomous workflows. The company explicitly describes its digital employees as multi-agent systems that work alongside human employees.
That wording tells us something about where enterprise architecture is heading.
The unit of automation is getting larger.
A traditional automation handles one task.
A copilot helps one worker.
A single agent may handle one workflow.
A multi-agent digital employee can potentially coordinate several pieces of a workflow across different systems.
BNY’s finance organization has also explained the likely progression clearly. Near-term value can come from focused agents that retrieve, structure, analyze, and summarize information. Over time, organizations may move toward more end-to-end digital employees operating across multiple systems and process steps under controls and human oversight.
That is likely to be a much more realistic enterprise path than jumping directly from ChatGPT-style assistants to fully autonomous digital departments.
Citi Shows Why the Platform Layer Matters
Citi’s Arc platform offers another lesson.
The difficult part of enterprise agents is not creating one impressive demo.
It is creating a repeatable system for building the next 100 agents.
Arc is designed as a platform through which developers can build and scale agents across the company. Citi says agents will be monitored, auditable, and governed, and gives an example in which a team of agents could prepare information for a wealth banker by gathering portfolio data, analyzing markets, modeling scenarios, and presenting the result before the client meeting.
The architectural insight is bigger than the specific example.
Large companies probably do not want 40 departments independently inventing 40 agent stacks.
They need shared identity.
Shared logging.
Shared policy controls.
Shared model access.
Shared evaluation.
Shared data connectors.
Shared cost monitoring.
And then they need freedom for individual teams to build specialized agents on top of that foundation.
Multi-agent strategy therefore becomes partly a platform strategy.
JPMorganChase Shows Why Security Has to Move to Runtime
The more agents act, the less useful static security becomes on its own.
JPMorganChase describes agentic AI as a structural change because agents do not merely produce outputs; they can take action. The firm’s security writing argues that controls must operate at the point of execution and produce complete, tamper-evident runtime records.
This becomes even more important in a multi-agent system.
Suppose a research agent sends a recommendation to a transaction agent.
Who authorized the first agent?
Is the second agent allowed to trust it?
Which human identity is the transaction being performed on behalf of?
What happens when an agent delegates to another agent with broader permissions?
Can an outside agent impersonate an internal one?
What happens when a malicious instruction hidden inside a document gets passed between three agents?
These are no longer prompt-writing questions.
They are identity, security, software architecture, and governance questions.
The Most Useful Multi-Agent Pattern for Enterprises May Be the Manager-and-Specialist Model
There are many ways to coordinate agents.
One of the easiest enterprise patterns to understand is a manager with specialists.
The manager receives the business request.
It then calls specialist agents for narrow jobs.
The specialists return structured results.
The manager combines those results and decides what should happen next.
OpenAI’s agent documentation describes a closely related “agents as tools” pattern in which a manager remains responsible for the final output while calling specialists for bounded tasks.
This structure has an important governance advantage.
Responsibility remains relatively clear.
Example: investment research
A Wall Street research system might contain:
| Agent | Responsibility |
| Market agent | Industry data and market structure |
| Filing agent | SEC filings and financial statements |
| Earnings agent | Calls, guidance, and management commentary |
| Valuation agent | Multiples and financial model |
| Risk agent | Contradictory evidence and downside cases |
| Citation agent | Source verification |
| Lead agent | Final research package |
The valuation agent does not need permission to send an email.
The earnings agent does not need payment access.
The citation agent should not be able to change the financial model.
Permissions follow responsibilities.
That is a powerful design principle.
Parallel Agents Are Valuable When Work Can Truly Happen in Parallel
A second useful architecture sends several agents to work simultaneously.
This is especially useful for research.
Instead of asking one agent to investigate 50 companies sequentially, a lead agent can divide the companies among specialized workers.
Anthropic reports that its multi-agent research architecture outperformed its single-agent setup by 90.2% on an internal research evaluation involving breadth-heavy tasks. The company also reported that parallelization significantly reduced research time in complex cases.
But this advantage depends on the work.
Ten independent searches can happen in parallel.
Ten steps that each require the result of the previous step cannot.
The business question should therefore come before the architecture question.
Can this workflow be divided into independent pieces?
If the answer is no, adding agents may simply create coordination overhead.
Handoffs Work Better When Responsibility Needs to Move
Another pattern involves handoffs.
An intake agent may begin a customer conversation.
When it identifies a billing issue, it hands control to a billing specialist.
A technical issue goes to a technical agent.
A potential fraud case goes to a security agent.
OpenAI’s Agents SDK treats handoffs as a core orchestration pattern and distinguishes them from manager-based systems. With a handoff, the specialist can become the active agent rather than merely returning a result to the original manager.
This design can work particularly well in customer service, employee support, IT service management, healthcare navigation, and other environments with clear specialist roles.
The biggest mistake is creating too many routes.
If the routing system contains 70 similar agents with unclear responsibilities, the enterprise has recreated an organizational bureaucracy in software.
Multi-Agent Finance Could Become One of New York’s Largest Opportunities
Finance is an obvious testing ground because work already moves through specialized roles.
Consider institutional onboarding.
A client submits information.
One team validates identity.
Another checks sanctions.
Another performs know-your-customer analysis.
Another validates legal documents.
Another creates internal accounts.
Another configures entitlements.
Another reviews exceptions.
Another approves the final relationship.

A multi-agent architecture maps naturally to that workflow.
The key is not replacing every person.
The key is reducing the time humans spend carrying information between systems.
Credit analysis offers another strong example
S&P Global has already launched an agentic Credit Memo Builder designed to aggregate information from different sources and accelerate preparation of credit materials. Its broader architecture work describes agents that can observe, suggest, act with human approval, or act autonomously within bounded policy.
A bank could take that idea further.
A document agent extracts borrower information.
A financial agent calculates ratios.
A sector agent studies market conditions.
A covenant agent reviews the loan agreement.
A risk agent looks for inconsistencies.
A memo agent produces the first draft.
A senior credit officer still makes the credit decision.
That final distinction matters.
Multi-agent AI should not be confused with the removal of human accountability.
Legal Work Is Naturally Multi-Agent Because Review Already Happens in Layers
A corporate legal matter rarely involves only one activity.
Research may lead to drafting.
Drafting leads to factual verification.
Then citation checking.
Then policy review.
Then negotiation analysis.
Then revisions.
Then approval.
Instead of one legal AI agent trying to perform the entire process, firms can separate those functions.
A contract extraction agent should not necessarily be the same agent that judges whether a term complies with company policy. A drafting agent should not be the same system responsible for independently checking the draft.
The second agent creates useful friction.
That pattern can be described as maker-checker AI.
One system creates.
Another system checks.
In industries where errors are expensive, that may be more useful than simply making the first agent larger.
Advertising Could Move From One Creative Agent to an Agent Team
New York’s advertising sector provides another example.
A single marketing agent might create copy.
A multi-agent campaign system could work very differently.
A planning agent identifies the audience. A research agent studies customer behavior. A creative agent produces concepts. A brand agent checks tone. A compliance agent reviews prohibited claims. A media agent determines placement. A measurement agent monitors results. The lead system reallocates work when performance changes.
That is closer to a miniature operating system for a campaign than a chatbot.
The strategic benefit does not come from giving each task a fashionable AI label.
It comes from closing the loop between observe, decide, execute, measure, and adjust.
Multi-Agent Healthcare Will Need Much Tighter Boundaries
Healthcare shows both the promise and the danger of this architecture.
NYU Langone reported that more than 1,200 inpatient and emergency providers were supported by ambient AI documentation tools and that it launched an AI agent on its Find a Doctor service in February 2026. The agent can interpret patient preferences and connect them with current information on clinicians, location, insurance, language, and availability.
A future multi-agent patient-navigation workflow could include separate insurance, scheduling, referral, clinical-routing, and communication agents.
But the permissions should differ sharply.
An appointment agent may be allowed to book available slots.
A clinical agent may summarize information but require clinician review.
An insurance agent may access coverage information but not the complete medical record.
Multi-agent systems make this separation possible.
They also make it necessary.
American Express Shows How Agents Could Start Transacting
American Express is particularly interesting because payments force the agent conversation out of the world of answers.
The company is building Agentic Commerce Experiences, or ACE, to support transactions performed by AI agents. It has also announced protection for certain registered agent purchases and is developing authentication and intent mechanisms around agent transactions.
Its 2026 shareholder letter describes AI plans spanning commerce, travel, dining, expense reporting, marketing, software development, and customer service. It also says AI-assisted development tools had reached more than 11,000 engineering professionals, with agentic coding tools being scaled further.
This illustrates an important future state.
Enterprises will eventually need to coordinate not only their own agents but also outside agents acting on behalf of customers, suppliers, employees, and partners.
That turns multi-agent interoperability into a business issue.
What Happens When Your Customer Is an Agent?
Imagine a hotel.
Today, a human customer visits a website.
Tomorrow, a travel agent run by an AI company may ask the hotel booking agent for availability.
The travel agent compares the result with an airline agent.
A payment agent receives purchase authorization.
A loyalty agent checks benefits.
A corporate travel policy agent confirms the trip is allowed.
Five separate systems may participate in one purchase.
The customer experience is no longer one person clicking through one interface.
It becomes negotiation between software systems.
This is one reason protocols such as A2A matter. Google’s original A2A announcement specifically framed the protocol around agents collaborating across siloed enterprise applications, vendors, and frameworks.
For New York businesses, the strategic question becomes:
Are our systems ready to serve software customers as well as human customers?
The Agent KPI Dashboard Every New York Enterprise Should Build
Traditional AI metrics are not enough.
Counting prompts does not tell you whether an agent system works.
Counting active users does not tell you whether the workflow is economically valuable.
A serious multi-agent deployment needs operational measures.
Table 3: Multi-Agent Enterprise KPI Dashboard
| KPI | What it tells you |
| Workflow completion rate | Whether the full job actually gets finished |
| Human intervention rate | How often the system needs rescue |
| Escalation accuracy | Whether the right problems reach people |
| Agent handoff failure rate | Whether coordination is breaking |
| Tool-call success rate | Whether external systems work reliably |
| Verification failure rate | How often checker agents reject work |
| Cost per completed workflow | Whether the system is economically viable |
| Median completion time | Whether automation actually makes work faster |
| Rework rate | Whether humans must redo completed work |
| Unauthorized-action attempts | Whether permissions are correctly designed |
| Source traceability | Whether outputs can be tied back to evidence |
| Business outcome | Revenue, cost, risk, speed, conversion, resolution, or another real result |
The most important metric is usually not agent accuracy in isolation.
It is successful business outcomes per dollar of system cost.
An agent that scores 95% on a technical benchmark but requires constant human correction may be less useful than a simpler system that reliably resolves a narrow workflow.
Multi-Agent Systems Need an Agent Ledger
Every serious enterprise should be able to answer a basic question:
What did the agents actually do?
That sounds obvious.
It is not.
A multi-agent system can produce dozens or hundreds of intermediate actions. One agent reads a document. Another calls an API. Another sends a message. Another rejects a recommendation. Another requests human approval.
If the final result is wrong, the company needs to reconstruct the path.
That requires an agent ledger.
For each important workflow, store the initiating identity, agents involved, model versions, data sources, tool calls, permissions, handoffs, approvals, actions, failures, costs, and final outcome.
This aligns with the direction of current security work. NIST’s 2026 agent initiatives emphasize identity, authorization, auditing, and the novel security risks created by agent systems. JPMorganChase similarly argues for runtime controls and complete records for autonomous systems.
Without this layer, debugging becomes guesswork.
Every Agent Needs an Identity
Human employees do not normally share one administrator account.
AI agents should not either.
A finance agent should have a finance-agent identity.
A marketing agent should have a marketing-agent identity.
A refund agent should have precisely defined refund permissions.
An agent working for a specific employee may also need to carry information about the human authority under which it is acting.
This becomes even more important when delegation occurs.
Agent A may be allowed to ask Agent B for research.
That does not mean Agent A should automatically inherit Agent B’s ability to move money.
Identity should travel with the action.
Authority should not silently expand during a handoff.
NIST has specifically highlighted identification, authorization, auditing, and non-repudiation as emerging agent-security concerns.
That sounds like security language.
It is also basic operational design.
The Biggest New Security Risk Is the Chain Reaction
Single-agent failure is one thing.
Multi-agent failure can propagate.
A research agent reads a malicious web page containing hidden instructions.
The compromised result is sent to a planning agent.
The planning agent tells an execution agent to perform an action.
The execution agent has legitimate credentials.
A bad instruction has now moved through a trusted internal chain.
OWASP’s agentic security work identifies insecure inter-agent communication, privilege abuse, context injection, tool misuse, rogue agents, and cascading failures among the emerging risks for multi-agent applications.
NIST has also said that commenters to its agent-security RFI broadly agreed that AI agents introduce novel threats and that existing cybersecurity practices will need adaptation.
This means multi-agent architecture should contain blast-radius controls.
An agent should only be able to damage the smallest possible part of the system.
Human Approval Should Be Designed Into the Workflow
One of the weakest enterprise patterns is adding a button labeled “human approval” after the architecture has already been designed.
Approval needs to be meaningful.
The reviewer should know what changed.
The reviewer should see the evidence.
The reviewer should understand uncertainty.
The reviewer should know which actions will happen after approval.
The system should also define which decisions never require approval.
If humans must approve every low-risk step, automation becomes pointless.
If humans approve nothing, the company may create unacceptable risk.
The objective is risk-based autonomy.
A low-risk internal research agent might operate freely.
A customer-message agent could send normal replies but escalate unusual cases.
A payment agent might automatically settle transactions below a strict threshold and require approval above it.
A legal agent may draft but never sign.
Autonomy should be earned by the workflow, not granted because the model sounds intelligent.
Multi-Agent Does Not Always Mean Better
This point deserves emphasis.
A company should not turn every process into a collection of agents.
Anthropic’s own multi-agent research found the architecture most useful for valuable tasks involving heavy parallelization, large amounts of information, and many complex tools. It also found substantial additional token consumption and noted that tasks with many dependencies or a need for shared context can be poor fits.
A single agent is usually better when the workflow is simple
If the task requires four sequential steps and one database, one well-designed agent may be enough.
Adding a manager and three specialists may create additional latency, cost, and failure points.
Deterministic software is better when the rules are deterministic
If a tax calculation has a fixed formula, use the formula.
If a payment threshold can be checked with ordinary code, use ordinary code.
If the system needs to validate whether a date is later than another date, it does not need an AI agent.
The strongest enterprise architecture will mix AI with traditional software.
Agents handle ambiguity.
Code handles certainty.
Workflows are better than agents when the path should never change
A process that must always perform steps A, B, C, and D in exactly that order does not need an AI system deciding what to do next.
A deterministic workflow engine is usually safer.
Agentic reasoning becomes useful when the correct path depends on what the system discovers.
That distinction can save companies enormous amounts of money.
Multi-Agent Economics Will Matter More Than Demo Quality
An impressive demo can hide terrible economics.
Each agent may generate model calls.
Each specialist may search data.
Verification adds additional calls.
Retry logic adds more.
Longer context adds cost.
Parallel work can multiply consumption very quickly.
Anthropic has reported that ordinary agents in its research setting used about four times as many tokens as chat interactions, while multi-agent systems used roughly 15 times as many. Those numbers are not universal pricing rules, but they show why enterprises need economic discipline.
Goldman Sachs Research similarly expects agentic AI to drive enormous growth in token consumption and argues that enterprise adoption requires careful orchestration and properly structured data.

This changes how AI projects should be evaluated.
Do not ask:
“How much does one model call cost?”
Ask:
“How much does one successfully completed business outcome cost?”
A Better Formula for Agent ROI
Consider a customer-support workflow.
The old process costs $18 per resolved case.
A multi-agent system costs $1.70 in model and infrastructure expenses per attempted case.
That looks cheap.
But if 30% of cases require human intervention and another 10% require rework, the economics change.
The correct calculation needs to include:
AI infrastructure.
Human review.
Failed attempts.
Rework.
Vendor software.
Integration costs.
Monitoring.
Security.
Engineering.
And the economic value of faster completion.
The unit of measurement should be the finished workflow.
Table 4: A Simple Agent ROI Model
| Measure | Example |
| AI + infrastructure per attempted workflow | $1.70 |
| Human review allocation | $2.50 |
| Rework allocation | $1.20 |
| Platform/monitoring allocation | $0.80 |
| True automated workflow cost | $6.20 |
| Prior human workflow cost | $18.00 |
| Gross process saving | $11.80 |
The exact numbers will be different for every company.
The method is what matters.
Measure the whole system.
Start With the Workflow, Not the Number of Agents
A business leader should never begin a project by saying:
“We need a five-agent system.”
That is architecture-first thinking.
Start with the process.
Map what actually happens today.
Where does information enter?
Who checks it?
Where does work wait?
Which decisions require judgment?
Which decisions follow fixed rules?
Which applications are involved?
Which mistakes are expensive?
Where must a human remain accountable?
Only then should the team decide how many agents are needed.
In many cases, the answer will be one.
In some cases, it will be three.
Occasionally, the business process will genuinely justify a much larger system.
A Practical 90-Day Multi-Agent Pilot for a New York Enterprise
A sensible pilot should be narrow enough to control but valuable enough to matter.
Do not begin with the CEO’s inbox.
Do not begin with autonomous trading.
Do not begin with unrestricted payments.
Do not begin with a workflow that can create a major legal liability.
Choose a repetitive knowledge process containing several clearly separable steps.
Days 1-15: Map the process
Document the workflow exactly as it operates today.
Record volumes, cycle time, failure rate, human hours, systems, approval points, and current cost.
Without a baseline, the company will later have no credible way to prove ROI.
Days 16-30: Build the single-agent baseline
This step is often skipped.
Do not skip it.
Build the simplest credible version with one agent and deterministic software where appropriate.
Measure its quality.
This gives the multi-agent architecture something to beat.
Days 31-45: Split only the weak areas
Look at where the baseline fails.
Perhaps research quality is poor because the context becomes too large.
Perhaps verification is inconsistent.
Perhaps the agent becomes confused by too many tools.
Create specialist agents only for those problems.
The architecture should grow because evidence demands it.
Days 46-60: Add permissions and observability
Give every agent an identity.
Define accessible data.
Define usable tools.
Define forbidden actions.
Add traces.
Store handoffs.
Record failures.
Set cost limits.
This is the point where a prototype begins becoming enterprise software.
Days 61-75: Run shadow mode
Let the agents complete the work without actually executing high-impact actions.
Compare their result with the human process.
Study disagreement.
Do not hide errors inside averages.
A 95% success rate sounds impressive until the remaining 5% contains the company’s most expensive cases.
Days 76-90: Release bounded autonomy
Allow automation where performance is strong and consequences are limited.
Keep approval where risk remains high.
Then measure the economic outcome against the original baseline.
Only after that should the company expand the system.
Table 5: The Multi-Agent Pilot Scorecard
| Question | Pilot target |
| Does the multi-agent version beat the single-agent baseline? | Yes, measurably |
| Is cycle time lower? | Quantified |
| Is cost per completed workflow lower? | Quantified |
| Are severe errors controlled? | Defined threshold |
| Can every action be traced? | 100% for material actions |
| Does each agent have bounded permissions? | Yes |
| Are human escalation rules explicit? | Yes |
| Can the workflow recover from agent failure? | Tested |
| Can one agent be replaced without rebuilding everything? | Preferably yes |
| Is there a real business owner? | Required |
If a project cannot answer these questions, it probably is not ready for broad deployment.
Multi-Agent Systems Will Change Enterprise Software Design
The long-term change may be larger than adding AI features to existing applications.
Traditional enterprise software is organized around screens.
Open CRM.
Open ERP.
Open email.
Open research platform.
Open ticketing system.
Open HR software.
The user becomes the integration layer.
People carry information between applications.
Agents can invert that relationship.
The user states an objective.
The agent system interacts with the applications.
If that architecture becomes common, enterprise software increasingly becomes infrastructure behind the agent layer.
S&P Global has already described a future in which agents need data that can move across multiple systems and remain traceable to authoritative sources.
That is why data architecture suddenly matters so much.
The Best Enterprise Data Will Be Agent-Ready
Agents do not magically repair bad information.
If customer identifiers differ across six databases, multi-agent systems inherit the problem.
If financial numbers cannot be traced to an authoritative source, agents cannot make them authoritative.
If permissions are inconsistent, agents amplify the inconsistency.
If APIs are undocumented, agents struggle to use them reliably.
Agent-ready data needs several properties.
It needs clear ownership.
It needs consistent identifiers.
It needs useful metadata.
It needs permissions.
It needs lineage.
It needs freshness information.
It needs accessible interfaces.
And it needs a system of record.
This may sound less exciting than autonomous AI.
It is probably more important.
New York Companies Should Build an Agent Control Plane Before They Build an Agent Army
As agents multiply, enterprises will need a common management layer.
Think of it as the control plane for the digital workforce.
It should answer:
Which agents exist?
Who owns them?
What model does each use?
What data can each access?
What tools can each call?
Who can invoke them?
Which agents can communicate?
What is each agent costing?
How often does each fail?
What changed in the newest version?
Can the company immediately disable one?
Kyndryl’s current strategy provides a useful example of this direction. The company describes Kyndryl Bridge and its Agentic AI Framework as infrastructure for deploying, orchestrating, securing, and governing agentic workflows, while separately introducing policy-as-code capabilities for enterprise agent governance.
The principle applies regardless of vendor.
Companies need control before scale.
Multi-Agent AI Will Create a New Form of Organizational Design
There is also a people question.
What does a department look like when humans work with groups of agents instead of using individual copilots?
Managers may eventually manage both.
An analyst might supervise research agents.
An operations manager might own an automated workflow.
A compliance professional may define the policies that agents must obey.
A product manager may manage a catalog of AI capabilities.
A security team may assign agent identities and privileges.
This means AI transformation becomes operating-model transformation.
The company has to decide which work belongs to people, which work belongs to software, and which work should be shared.
Human judgment may become more important, not less
As machines perform more of the collection and coordination work, humans may spend proportionally more time on ambiguous decisions.
Should we trust this customer?
Should we accept this unusual risk?
Should we launch this campaign?
Should we escalate this patient?
Should we make this investment?
Should we sign this agreement?
Those are not simply information-retrieval problems.
They involve accountability.
That will remain difficult to automate responsibly.
The Biggest Mistake Will Be Recreating Corporate Bureaucracy With AI
Humans already understand one major weakness of large organizations.
Too many handoffs.
A document passes through seven teams.
Everyone adds comments.
Nobody owns the result.
Meetings multiply.
Decisions slow down.
A badly designed multi-agent system can reproduce the same problem at machine speed.
Agent A asks Agent B.
Agent B requests information from Agent C.
Agent C sends the request to Agent D.
Agent D rejects it.
Agent B retries.
Agent E reviews the retry.
The company has invented digital bureaucracy.
That is why architecture should remain simple.
Every additional agent should have a reason to exist.
A Good Rule: Add Another Agent Only When Separation Creates Value
There are four especially strong reasons to split one agent into multiple agents.
Context isolation
Different specialists require substantially different information, and combining everything reduces quality.
Parallel work
Independent pieces can run simultaneously and meaningfully reduce completion time.
Specialization
Different tasks require different prompts, models, tools, skills, or domain knowledge.
Separation of authority
Different actions should have different permissions or independent review.
Anthropic has similarly identified context isolation, parallel execution, and specialization as key scenarios in which multi-agent architectures can be useful.
If none of these apply, keep the architecture simpler.
What New York Enterprise Leaders Should Do Now
The opportunity is not to announce a multi-agent strategy.
It is to find one workflow where specialization can beat a well-designed single agent.
Pick a process with measurable pain.
Build the single-agent baseline.
Identify exactly where it fails.
Add specialists only where they solve a real problem.
Give agents narrow identities.
Keep deterministic rules deterministic.
Require traceability.
Measure the finished workflow.
Price the whole system.
And expand only after the economics work.
That sequence sounds conservative.
For enterprise AI, conservative architecture can produce aggressive business results.
Five Predictions for New York’s Multi-Agent Era
The first change will be the rise of internal agent platforms. Large enterprises will increasingly prefer a common layer for identity, governance, monitoring, models, connectors, and agent deployment rather than allowing every business unit to create isolated systems.
The second will be specialization by function. Instead of one “company AI,” businesses will develop finance, legal, sales, security, operations, healthcare, marketing, and engineering agents with tightly controlled responsibilities.
The third will be agent-to-agent commerce. Customers will increasingly arrive through software acting on their behalf, forcing businesses to expose inventory, pricing, bookings, account actions, and transactions in machine-usable ways. American Express’s current agentic-commerce work offers an early example of the infrastructure being built for this world.
The fourth will be a major expansion of agent identity and security infrastructure. NIST’s new AI Agent Standards Initiative, its work on agent authorization, and emerging standards such as OWASP’s Agent Control Standard all point toward a future where agent permissions and runtime controls become first-class enterprise infrastructure.
The fifth will be a move away from counting AI users and toward measuring completed work.
That last shift may be the most important.
The Bigger Story: Enterprise AI Is Starting to Organize Itself Like a Workforce
The single AI agent is not disappearing.
It will remain the right architecture for an enormous number of applications.
But the enterprise problem is changing.
Companies are no longer asking only whether AI can answer questions.
They are asking whether AI can complete work that crosses departments, databases, applications, policies, approvals, and organizational boundaries.
One general-purpose agent can only absorb so much complexity before the design becomes fragile.
Multi-agent systems offer another path.
Divide the work.
Keep context focused.
Give specialists limited tools.
Separate creation from verification.
Separate analysis from execution.
Give every agent an identity.
Record every important action.
Bring humans in where judgment and accountability matter.
That model is particularly relevant to New York because so much of the city’s economic value is created inside complicated knowledge workflows. Our analysis of NYC Comptroller data shows that finance, information, and professional/business services represent less than one-third of city jobs but well over half of our simple payroll proxy. Meanwhile, public disclosures from major New York organizations show the architecture moving steadily from copilots toward action-taking agents, orchestration platforms, and explicit multi-agent systems.
The transition is still early.

That should make business leaders more disciplined, not less interested.
The winners of the next phase of enterprise AI are unlikely to be the companies that deploy the greatest number of agents.
They will be the companies that understand exactly which work should be divided, which work should stay together, where machines should act, where people should decide, and how all of it can operate as one reliable system.
For New York enterprises, that is the real promise of multi-agent AI.



