For the first few years of the generative AI boom, most businesses used AI in roughly the same way. An employee opened a chat window, typed a question, received an answer, and then went back to doing the actual work.
That model was useful. An AI assistant could summarize a document, rewrite an email, explain a spreadsheet, draft marketing copy, or help an analyst think through a problem. But the employee still owned almost every next step.
That is starting to change.
Across New York, a new group of AI companies is building software that does not simply answer questions. These systems can research information, make decisions within set rules, update business systems, create finished work, contact customers, schedule activities, monitor changing conditions, process financial operations, and continue working after the original prompt has been given.
That is the important difference between an AI assistant and an AI agent.
An assistant helps a person perform a task.
An agent can be given part of the task itself.
The difference may sound small, but economically it is enormous. Once software can move from generating information to executing work, companies can start redesigning workflows rather than simply making individual employees slightly faster.
McKinsey’s 2025 global AI survey found that 23% of respondents said their organizations were already scaling an agentic AI system somewhere in the business, while another 39% were experimenting with agents. The same research also showed that enterprise-wide scaling was still limited, which is an important reminder that the agent market remains early.
By August 2026, Deloitte found an even clearer gap between ambition and readiness. Only 5% of organizations in its survey said their business processes were highly prepared for AI agents, while just 15% had scaled orchestrated, cross-functional multi-agent adoption. At the same time, 74% of leaders expected nearly half of their business processes to be redesigned or rebuilt around AI agents within four years.
That gap between what companies want and what their operating systems can currently support may become one of the biggest enterprise technology opportunities of the next several years.
And New York is unusually well positioned to become a major testing ground.
NYCEDC has identified more than 2,000 AI startups in New York City, along with more than 40,000 workers in the metro area with AI skills. Its research also points to more than 25,000 technology startups and over 1,200 active venture capital firms in the broader city ecosystem.
But the more important advantage may be the industries sitting beside those technology companies.
New York has finance, insurance, law, accounting, healthcare, advertising, real estate, media, professional services and large enterprise operations concentrated within a few subway stops of one another. These industries contain thousands of workflows that are expensive, repetitive, information-heavy and difficult to automate with old software.
Those are precisely the workflows modern AI agents are being designed to attack.
The Short Version: Assistants Answer. Agents Execute.
The easiest way to understand the shift is to look at who owns the next step.
Suppose an investment banker asks an AI system to analyze a company.
An assistant might summarize the company’s earnings reports and suggest valuation considerations. The banker then opens Excel, gathers market data, updates a model, builds slides and writes the investment memo.
An agent can potentially be given a broader instruction such as: analyze this company, update the comparable-company model, build the presentation, identify important risks and prepare the materials for review.

The banker still decides whether the work is correct and what recommendation should ultimately be made. But much more of the mechanical workflow has moved from the human to the software.
That changes the economics of AI.
AI assistants improve individual tasks
Most AI assistants begin with a human request and end with a human receiving information.
They are excellent for tasks such as writing, summarizing, brainstorming, explaining, searching and answering questions. They can substantially increase productivity without changing the basic structure of the workflow.
The employee remains the operator.
Copilots work beside the employee
A copilot goes somewhat further.
It may understand company information, suggest actions, generate content inside existing software or help complete several connected steps. However, the human usually remains actively involved.
Think of the copilot as an extremely capable coworker sitting beside you.
AI agents accept delegated work
Agents change the interaction.
Instead of asking:
“What should I do?”
the user can increasingly say:
“Do this.”
The agent may decide which tools to use, gather the required information, perform several steps, create an output and report back when the task is finished.
Autonomous workflows go one step further
The most advanced systems do not necessarily wait for a new prompt.
They can respond to events.
A customer becomes overdue.
A new sales lead enters the CRM.
A referral arrives.
A contract is signed.
An invoice reaches its due date.
A portfolio company publishes earnings.
The software detects the event and begins the appropriate workflow according to rules established by the organization.
That is where AI starts becoming part of the operating model rather than simply another productivity application.
AI Assistants vs Copilots vs Agents vs Autonomous Workflows
The labels used by technology vendors are not always consistent. One company may call a product an agent even when it behaves more like a sophisticated copilot, while another may avoid the agent label despite offering meaningful automation.
The more useful approach is to examine capabilities rather than branding.
| Capability | AI Assistant | AI Copilot | AI Agent | Autonomous Workflow |
| Answers questions | Yes | Yes | Yes | Yes |
| Generates content | Yes | Yes | Yes | Yes |
| Uses company context | Sometimes | Usually | Usually | Usually |
| Performs multiple steps | Limited | Moderate | Strong | Strong |
| Chooses tools | Limited | Sometimes | Yes | Yes |
| Changes external systems | Rare | Limited | Often | Yes |
| Remembers workflow state | Limited | Sometimes | Often | Usually |
| Works after initial prompt | Rare | Limited | Often | Yes |
| Responds automatically to events | No | Rare | Sometimes | Yes |
| Needs human approval every step | Usually | Often | Not necessarily | Usually only for exceptions |
| Best mental model | Adviser | Coworker | Delegate | Digital operating process |
The core difference is therefore not whether the software contains an LLM.
It is how much responsibility the company is prepared to delegate to the software.
NYC Tech Journal Original Research: The New York Agentic Work Index
To understand what this transition actually looks like in New York, NYC Tech Journal conducted an original structured analysis of publicly available product information from 11 AI companies with a meaningful New York presence.
This is not a ranking of which company has the “best AI.”
Instead, the research asks a narrower question:
How far has each public product model moved from answering questions toward executing business work?
How We Built the Dataset
The research was conducted using public information available through September 10, 2026.
A company had to meet three conditions to enter the sample.
First, it needed a clear New York operating presence, headquarters, founding history or major product team.
Second, it needed an enterprise AI product designed around an identifiable business workflow.
Third, enough first-party documentation had to be publicly available to determine what the product can currently do.
Future promises were not counted as current capabilities unless the company said the feature was already available.
The final sample contained:
Clay, Rogo, Regal, EliseAI, Tabs, Tennr, Basis, Norm Ai, Credal, Hebbia and Ramp.
These businesses cover sales, finance, accounting, legal work, compliance, customer communications, housing, healthcare and enterprise knowledge work.
Clay, for example, began in Brooklyn and now operates an NYC office in Chelsea. Hebbia lists its address at 233 Spring Street. Regal describes itself as built in NYC and based in Manhattan. Tennr lists its office at 345 Hudson Street. Norm Ai lists a New York address at 7 World Trade Center. Ramp lists its business address on West 23rd Street.
The Five-Part Scoring Model
Each system received between zero and two points across five dimensions.
| Dimension | 0 Points | 1 Point | 2 Points |
| Workflow scope | Answers or single task | Several connected steps | End-to-end workflow |
| Tool access | No meaningful tools | Mostly retrieval/read access | Can work across systems or use read/write tools |
| Action authority | Recommends | Creates drafts/artifacts | Performs operational actions |
| State and time horizon | Single interaction | Reusable/session workflow | Persistent, scheduled or background work |
| Supervision model | Human drives every step | Selective approval/review | Exception-based supervised autonomy |
The maximum score is 10.
For this analysis, scores of 0–3 represent an assistant model, 4–6 a copilot model, 7–8 an agent model and 9–10 an autonomous-workflow model.
Ambiguous capabilities were scored conservatively.
Importantly, these scores measure publicly documented product architecture and operating capability, not independently verified accuracy. Company product pages naturally present their software in the strongest possible light, so the index should not be treated as an independent performance benchmark.
What We Found
Chart 1: NYC Agentic Work Index
| Company | Main Workflow | Score / 10 | Model |
| Clay | Sales and GTM | 9 | Autonomous workflow |
| Tabs | Billing and revenue | 9 | Autonomous workflow |
| Tennr | Healthcare operations | 9 | Autonomous workflow |
| Ramp | Finance and spend | 9 | Autonomous workflow |
| Rogo | Institutional finance | 8 | Agent |
| Regal | Customer communications | 8 | Agent |
| EliseAI | Housing and healthcare | 8 | Agent |
| Credal | Enterprise agent platform | 8 | Agent |
| Basis | Accounting | 7 | Agent |
| Hebbia | Finance/legal knowledge work | 7 | Agent |
| Norm Ai | Legal and compliance | 6 | Copilot-to-agent boundary |
Average score: 8.0 / 10
Median score: 8.0 / 10
Again, the sample intentionally focuses on companies already building workflow-oriented AI. A finding that most of them score highly does not mean most NYC businesses have reached the same level of automation.
The interesting finding appears when the individual dimensions are separated.
Original Finding #1: Reasoning Is No Longer the Main Bottleneck
Every company in our sample publicly described software capable of handling multi-step workflows.
That was the strongest dimension in the entire dataset.
The weaker dimensions were what happened after the reasoning was completed.
Chart 2: Share of Sample Showing the Strongest Level of Each Capability
| Capability | Companies Scoring 2/2 | Share of Sample |
| Multi-step/end-to-end workflow | 11 of 11 | 100% |
| Cross-system/tool capability | 8 of 11 | 73% |
| Operational action authority | 7 of 11 | 64% |
| Persistent/background state | 5 of 11 | 45% |
| Exception-based autonomy | 2 of 11 | 18% |
This may be the most important finding in the research.
The agent market is no longer primarily trying to prove that language models can reason through several steps.
The harder questions are:
Can the agent access the right systems?
Can it change those systems?
Can it remember what happened yesterday?
Can it continue working without another prompt?
Can the company safely allow the system to act without checking every single decision?
That is why the next wave of enterprise AI competition may increasingly happen around permissions, integrations, memory, evaluations, audit trails and control systems, rather than model intelligence alone.
Deloitte’s broader enterprise research points toward the same problem. Only 21% of organizations surveyed in a 2026 study reported mature governance for agentic AI, even as agent adoption continued to expand.
Clay Shows What Happens When the CRM Stops Being Passive
Clay is a useful example because the product increasingly treats an account as something an agent continuously understands rather than simply a row in a database.
Clay’s Account Agents are designed to reason across information about an account, decide the next action, remember previous conclusions and then execute an allowed action. The system writes conclusions back to the account so that a later run starts with what the agent previously learned.
That is structurally different from asking an assistant:
“Research this company.”
The larger workflow becomes:
A signal arrives.
The agent understands the account.
The agent evaluates the change.
The agent decides what should happen.
The system takes an approved action.
The account state is updated.
The next decision starts from the new state.
That is much closer to an operating process than a chatbot.
Clay’s own description of its business now says companies combine data, agents, signals and workflows to build go-to-market systems, while the company says its goal is to help businesses grow autonomously with AI.
Rogo Shows the Same Shift on Wall Street
Finance provides another clear example.
Traditional financial AI was largely search driven. Ask a question about a company, retrieve documents, summarize the answer and send it back to the banker or investor.
Rogo’s Felix moves further toward delegation.
According to Rogo’s May 2026 product update, Felix can turn one instruction into deliverables including PowerPoint decks, Excel models, Word documents, dashboards and sourced research. Rogo also introduced custom agents, memory, expanded connectors and autonomous agents that can be asked through email to perform scheduled tasks and monitoring in the background.
Rogo now describes hundreds of finance-specific agents capable of executing end-to-end workflows and returning finished Excel, PowerPoint or Word outputs. The company says its shared agent library had been executed more than 430,000 times by July 2026.
This matters because Wall Street knowledge work contains enormous amounts of structured repetition.
An investment banker does not simply “write presentations.”
The banker collects data, checks sources, updates models, compares companies, revises assumptions, changes charts, prepares slides and responds to comments.
If AI only writes text, it touches one small piece of the workflow.
If an agent can operate across research, spreadsheets and presentations, much more of the workflow becomes addressable.
Regal Turns a Conversation Into an Executable Workflow
Customer service illustrates the same difference.
An AI assistant can tell a support representative what to say.
An AI agent can have the conversation.
A more advanced agent can then complete the action that the conversation requires.

Regal’s developer documentation describes agent actions that let AI systems handle calls end-to-end, including checking scheduling information, booking appointments and transferring customers to internal or external destinations. The actions can also connect to external systems through custom functions.
That final step is critical.
If a customer says, “I need to change my appointment,” an assistant might generate instructions explaining how the appointment could be changed.
An agent changes it.
The business value comes from removing the gap between understanding the customer’s intention and carrying out the required operation.
EliseAI Is Moving From Conversation Automation to Operational Work
EliseAI shows how the same idea is appearing in housing.
The New York company already uses AI across leasing, resident communication and maintenance-related workflows. In September 2026 it introduced Apollo, which it describes as an agentic AI teammate working across its property operations platform.
One example provided by the company is telling Apollo to reschedule tours when a colleague is unavailable. Another involves helping maintenance staff identify recurring unit problems.
Those examples may appear mundane compared with futuristic visions of general-purpose autonomous AI.
That is exactly why they matter.
A large part of enterprise value comes from boring operational work.
Rescheduling.
Routing.
Checking.
Following up.
Updating.
Reconciling.
Collecting.
Escalating.
The agent revolution may be built on millions of tasks that individually look unremarkable but collectively consume enormous amounts of human time.
Tabs Shows Why the CFO Office Is Becoming Agentic
Tabs offers another useful New York example.
Its products automate billing, collections and revenue operations. The company says its Billing Agent can read contracts, extract billing terms, generate invoices and synchronize information with ERP systems, while its Collections Agent can track due dates, follow up and reconcile payments.
Its current documentation makes the human-control model particularly interesting.
Tabs says teams can maintain approvals and handle exceptions where judgment is needed. Its product materials describe agents handling routine work while escalating situations that need human involvement.
That is likely to become one of the most common enterprise patterns:
Do not put a human inside every step. Put the human around the exceptions.
This is a major change from the copilot model.
In a copilot workflow, the human works and AI assists.
In a mature agent workflow, AI handles the normal path while humans manage unusual, risky or high-value decisions.
Tennr Shows Why “Autopilot” Needs Boundaries
Healthcare creates a harder version of the same problem because mistakes can have serious consequences.
Tennr describes itself as an agentic patient orchestration platform built around patient flow, payer requirements and operational decision-making. Its platform gathers and evaluates information, coordinates communications and manages workflows across different steps in the patient journey.
The company’s Autopilot feature is particularly relevant to the agent-versus-assistant distinction.
Tennr says Autopilot automatically completes high-confidence work while keeping quality controls in place.
This model is likely more useful for most businesses than the idea of giving AI unrestricted freedom.
The practical future of autonomous enterprise software is probably not:
“Let the AI do whatever it wants.”
It is:
“Let the AI independently perform actions that fall within a carefully designed operating envelope.”
Ramp Shows What Happens When Agents Get Economic Authority
Payments may represent one of the strongest lines between assistance and autonomy.
An assistant can recommend buying something.
An agent with payment authority can actually buy it.
Ramp’s Agent Cards documentation describes a system where an AI agent can receive a merchant-specific, capped single-use card, complete a purchase and then handle items such as the receipt, memo and coding. The spending authority remains limited by existing business controls.
The architecture is important because it shows how autonomy can expand without becoming unlimited.
The agent gets authority.
But the authority has boundaries.
The merchant can be restricted.
The amount can be capped.
Company permissions still apply.
That pattern could become fundamental to enterprise agent design.
Autonomy does not need to mean unlimited permission.
Original Finding #2: New York’s Agent Economy Is Becoming Vertical
One of the clearest patterns in our sample is specialization.
The leading NYC agent companies are not simply building generic assistants and hoping customers figure out what to do with them.
They are embedding agents into particular industries.
Chart 3: Industry Focus of the NYC Tech Journal Sample
| Workflow Cluster | Companies | Number |
| Finance and accounting | Rogo, Tabs, Basis, Ramp | 4 |
| Legal, compliance and enterprise governance | Norm Ai, Credal | 2 |
| Healthcare and housing operations | Tennr, EliseAI | 2 |
| GTM and customer communication | Clay, Regal | 2 |
| Institutional research and knowledge workflows | Hebbia | 1 |
This helps explain why New York is becoming important in enterprise agents.
Silicon Valley remains enormously important to foundational AI research and infrastructure.
New York has a different advantage.
It has customers.
More specifically, it has complicated customers.
Banks.
Private equity firms.
Law firms.
Hospitals.
Real estate operators.
Media companies.
Advertising agencies.
Accounting firms.
Retailers.
Large enterprise headquarters.
These organizations operate complicated workflows with high labor costs and enormous amounts of proprietary data.
That creates fertile ground for vertical agents.
Manhattan’s Employment Structure Explains the Opportunity
The economic structure of Manhattan provides another clue.
Bureau of Labor Statistics data for New York County showed approximately 2.265 million private-sector jobs in March 2026.
Professional and business services accounted for about 583,300.
Financial activities accounted for roughly 419,800.
Education and health services accounted for approximately 394,600.
Information accounted for another 189,500.
Chart 4: Selected Share of New York County Private Employment
| Sector | Employment, March 2026 | Share of Private Employment |
| Professional & business services | 583,300 | 25.8% |
| Financial activities | 419,800 | 18.5% |
| Education & health services | 394,600 | 17.4% |
| Information | 189,500 | 8.4% |
| Combined | 1,587,200 | 70.1% |
NYC Tech Journal calculation using BLS QCEW data. Figures refer to New York County, not all five boroughs.
Finance, professional/business services and education/health alone represented about 61.7% of Manhattan’s private employment in this dataset.
Add information businesses and the share reaches roughly 70.1%.
That does not mean 70% of Manhattan jobs can or should be automated.
It means Manhattan is extraordinarily concentrated in sectors where a large share of the product being sold is knowledge, coordination, information processing, communication and decision support.
Those are exactly the economic activities where AI agents can potentially create leverage.
Original Finding #3: New York Is Building Agents Around Expensive Work
There is another pattern hiding inside the sector data.
Many prominent New York agents target work performed by relatively expensive professionals.
Rogo attacks investment banking and investing workflows.
Basis targets accounting.
Norm Ai targets legal and compliance work.
Hebbia focuses heavily on institutional finance and legal research.
Tabs targets finance organizations.
Clay targets modern sales and revenue teams.
This is rational economics.
Automation becomes much easier to justify when one hour saved is expensive.
Consider a workflow that saves an employee 20 minutes.
If that task happens once a month, nobody cares.
If it occurs 200 times every day across highly paid employees, the economics look completely different.
Agent adoption is therefore likely to move fastest where four conditions appear together:
high labor cost × high frequency × structured process × measurable output.
That formula is more useful than simply asking whether a task “could use AI.”
Why the Copilot Model Eventually Hits a Ceiling
Copilots can create impressive productivity improvements, but they have a structural limitation.
The human remains the integration layer.
Imagine an employee using five AI-powered applications during a normal process.
The employee asks one system for research.
Copies the answer somewhere else.
Opens another application.
Generates a document.
Checks the document.
Copies information into the CRM.
Writes an email.
Schedules a follow-up.
Updates a spreadsheet.
The AI may have helped at several points, but the employee is still moving the workflow forward.
That means companies can end up with highly intelligent software sitting inside surprisingly manual business processes.

Agents attack that coordination cost.
The value is not simply better output.
The value is fewer handoffs.
The Best Metric Is Not Prompts Saved. It Is Handoffs Removed.
Companies frequently evaluate generative AI using weak metrics.
Number of users.
Number of prompts.
Time spent in the application.
Messages generated.
Employees trained.
These numbers can help measure adoption, but they say relatively little about business value.
For agents, a better question is:
How many human handoffs were removed from the workflow?
Suppose processing an invoice currently requires seven steps involving three employees.
If AI makes each employee 15% faster, the process improves.
If an agent safely completes five of those seven steps and sends only exceptions to a person, the operating model changes.
That is the larger opportunity.
AI Assistants Are Not Going Away
None of this means assistants become useless.
In fact, companies should be careful not to turn every AI use case into an autonomous agent.
Assistants remain the better design when the work is ambiguous, exploratory or judgment-heavy.
A CEO developing acquisition strategy probably wants an intelligent thought partner, not software independently acquiring companies.
A lawyer evaluating a novel argument may want research and drafting support while keeping control over legal judgment.
A creative director exploring a new campaign may prefer rapid idea generation rather than an autonomous system publishing whatever it creates.
The agent model works best when the desired outcome and boundaries can be clearly described.
Original Finding #4: The Real Enterprise Architecture Is Assistant + Agent + Human
The assistant-versus-agent debate is therefore slightly misleading.
Most sophisticated businesses will use both.
The more realistic operating model contains three layers.
Layer 1: Assistants for thinking
Humans use AI to explore ideas, understand information, analyze uncertainty and make decisions.
Layer 2: Agents for execution
Once the desired outcome is clear, agents handle repeatable research, coordination and operational steps.
Layer 3: Humans for judgment and exceptions
People remain responsible for unusual cases, strategic tradeoffs, relationship-sensitive decisions and actions carrying major financial, legal or reputational consequences.
The goal is not maximum autonomy.
The goal is the right autonomy.
Microsoft’s 2025 Work Trend Index captured this idea through what it called the “human-agent ratio.” Its research found 46% of surveyed leaders said their organizations were already using agents to fully automate workstreams or business processes, while 82% expected to use digital labor to increase workforce capacity within the following 12 to 18 months.
The Agent Maturity Ladder for New York Businesses
Rather than trying to jump immediately into fully autonomous systems, most companies should move through several levels.
Level 1: AI answers questions
The organization begins with secure enterprise AI for research, writing, summarization and analysis.
Risk remains relatively low because humans perform the actions.
Level 2: AI creates finished work
The system starts producing usable deliverables.
A banker gets a draft presentation.
An accountant receives a reconciliation.
A marketing team receives a complete campaign brief.
The employee still reviews and distributes the work.
Level 3: AI uses tools
The model gets controlled access to company systems.
It reads the CRM.
Queries databases.
Opens files.
Checks the ERP.
Pulls information from approved knowledge sources.
This is the point where identity and permissions become extremely important.
Level 4: AI performs actions with approval
The agent proposes an action.
A person approves.
The system executes.
This is a strong starting point for financial, legal and customer-facing operations.
Level 5: AI performs approved classes of actions automatically
Routine work proceeds without individual approval.
Humans review exceptions.
This is where meaningful operating leverage begins.
Level 6: Event-driven workflows
The agent no longer needs the employee to initiate every task.
A business event starts the process automatically.
This is much closer to autonomous operations.
Which Workflows Should NYC Companies Automate First?
The right first project is usually not the most exciting workflow.
It is the one with the cleanest economics.
A Practical Agent Opportunity Score
Score each potential workflow from one to five across the following factors.
| Factor | Question |
| Frequency | Does this happen many times each day or week? |
| Labor cost | Does expensive employee time go into it? |
| Repeatability | Does the process follow recognizable patterns? |
| Data access | Can required information be accessed reliably? |
| Error visibility | Can mistakes be detected quickly? |
| Reversibility | Can a bad action easily be corrected? |
| Measurement | Can success be measured objectively? |
| Business importance | Would improving this materially affect revenue, cost or service? |
A workflow scoring highly across most of these categories deserves investigation.
A workflow that is rare, highly subjective and impossible to measure probably does not.
Start With the “Boring Middle”
Companies often make one of two mistakes.
They either start with trivial AI tasks that create almost no economic value, or they immediately attempt to automate an extremely complicated mission-critical workflow.
The better target usually sits in the middle.
Important enough to matter.
Structured enough to automate.
Low enough risk to experiment safely.
Examples might include preparing recurring client reports, researching sales accounts, processing standard invoices, gathering documentation, preparing meeting materials, routing customer inquiries, monitoring portfolio information or handling routine follow-ups.
A 90-Day Agent Deployment Plan
Companies do not need a three-year AI transformation program before testing an agent.
A disciplined 90-day program is enough to determine whether a workflow deserves serious investment.
Days 1–15: Map the Workflow Before Buying Software
Watch employees perform the task.
Do not begin with the software.
Document where information enters the process, which systems are opened, which judgments are made, what gets copied manually, where delays occur and what defines a successful result.
Most importantly, count the handoffs.
You cannot measure improvement if you do not understand the existing workflow.
Days 16–30: Establish the Baseline
Measure current performance.
For example:
| Metric | Current Process |
| Average completion time | 42 minutes |
| Human touchpoints | 8 |
| Systems opened | 5 |
| Rework rate | 11% |
| Cost per completed workflow | $34 |
| Average waiting time | 17 hours |
Without this baseline, almost every AI pilot eventually turns into a debate about whether employees “feel” more productive.
Days 31–45: Give the Agent Read Access First
Start conservatively.
Let the agent gather information and prepare work while humans continue making changes to production systems.
This reveals whether the model understands the process before giving it meaningful authority.
Days 46–60: Introduce Approved Actions
Choose reversible operations.
Allow the agent to prepare the CRM update.
Draft the customer response.
Create the invoice.
Prepare the spreadsheet.
Generate the transaction request.
The employee approves before execution.
Measure how often the human changes what the agent proposed.
That correction rate is extremely valuable.
Days 61–75: Identify the Safe Automation Zone
Look for categories where approval is almost always routine.
Suppose employees approve 99.5% of one type of action with no changes.
That workflow may be a candidate for exception-based automation.
Another action may only receive approval 70% of the time.
Keep the human checkpoint.
Autonomy should be earned by evidence.
Days 76–90: Move From Pilot to Operating Model
At this stage, evaluate more than model accuracy.
Measure cost.
Cycle time.
Customer experience.
Failure rate.
Employee workload.
Exception frequency.
Business output.
If the economics work, document the workflow as an operating process rather than leaving it as an experiment owned by one enthusiastic employee.
The Agent KPI Dashboard Every Business Should Build
Executives need better agent metrics.
A useful dashboard might contain the following.
| KPI | Why It Matters |
| Workflow completion rate | Can the agent actually finish assigned work? |
| Human intervention rate | How often must employees rescue the process? |
| Human correction rate | How often is the output changed? |
| Exception rate | How much work reaches human review? |
| Cycle-time reduction | Is the workflow actually faster? |
| Cost per completed workflow | Are the economics improving? |
| Tool failure rate | Are integrations reliable? |
| Unauthorized-action rate | Are control boundaries working? |
| Reversal rate | How often must actions be undone? |
| Business outcome | Did revenue, cash flow, service or productivity improve? |
The final metric is the most important.

An agent that completes 10,000 tasks but creates no measurable business improvement is simply expensive automation.
The Economics of Agents Are Different From SaaS
Traditional software usually sells seats.
A company buys 500 licenses for 500 employees.
Agentic software complicates that model.
One agent may perform work previously spread across many employees.
Another may run thousands of operations every day.
The important pricing unit may therefore shift from users toward work.
Tasks completed.
Calls resolved.
Invoices processed.
Accounts researched.
Documents analyzed.
Cases routed.
Revenue collected.
This could create major changes in enterprise software procurement.
Instead of asking:
“How many people need licenses?”
CFOs may increasingly ask:
“What does this workflow cost us today, and what will it cost after automation?”
That is a much more demanding question for software vendors.
Why New York’s Regulated Industries Will Shape Agent Governance
New York may also become important for another reason.
Many of the industries adopting agents here cannot simply maximize automation and hope for the best.
Banks must think about financial controls.
Law firms need confidentiality and professional responsibility.
Healthcare companies handle sensitive information.
Insurance companies make decisions inside strict regulatory systems.
Large public companies need auditability and internal controls.
That environment forces agent builders to solve difficult enterprise problems early.
Credal provides a good example. Its documentation allows organizations to require human approval before an agent executes particular actions. Administrators can apply approval controls at the specific action or organization level.
Norm Ai presents another model. Its legal technology embeds legal reasoning into agents while its affiliated Norm Law combines those systems with attorney supervision.
The pattern is important.
In regulated environments, governance cannot be a dashboard added after the agent is built.
Governance becomes part of the product architecture itself.
Audit Trails Will Become as Important as Prompts
When AI only drafts an email, tracing every internal decision may not matter much.
When AI changes a customer account, sends an invoice, routes a patient or makes a payment-related decision, businesses need to know what happened.
Who initiated the workflow?
Which data did the agent use?
Which tools did it call?
What action did it take?
Which policy allowed that action?
Did a person approve it?
What changed inside the system afterward?
These questions turn observability from a technical feature into a management requirement.
That is especially true when an agent operates for hours rather than seconds.
Long-Running Work Changes the Risk Model
OpenAI’s 2026 research on agentic work offers one indication of how quickly task duration is expanding. The company reported that by May 2026, more than 70% of individual Codex users had asked the agent to perform at least one task estimated to require more than an hour of human work.
Basis describes a similar pattern in accounting, saying its agents can operate for hours while performing end-to-end work for accounting firms.
That introduces a new enterprise problem.
A chatbot can make one mistake.
A long-running agent can make a sequence of connected mistakes.
The longer the workflow, the more important checkpoints, evaluations and validation become.
Original Finding #5: The Competitive Moat Is Moving From Intelligence to Workflow Ownership
For the first part of the generative AI boom, model quality created enormous product differences.
That remains important.
But enterprise agents introduce another layer of competition.
Imagine two companies using similarly capable foundation models.
Company A has the better prompt.
Company B has deep integrations into the customer’s systems, understands the industry’s workflow, remembers previous work, includes audit trails, knows when to request approval and has been tested against thousands of realistic tasks.
Company B may have the stronger product even if both begin with similar underlying intelligence.
This helps explain why vertical AI companies are becoming important in New York.
The difficult problem is increasingly not:
“Can the model write?”
It is:
“Can the system reliably do this specific job inside this specific company?”
Workflow Data Can Become the New Moat
Every completed workflow creates information.
Which cases required human intervention?
Which recommendations were changed?
Which tool calls failed?
Which exceptions happened repeatedly?
Where did the employee override the agent?
What sequences produced the best outcome?
That feedback can improve future agents.
Companies that capture workflow data properly may therefore build a compounding advantage.
The agent learns not simply from documents but from the organization’s behavior.
This is one reason persistent state matters so much.
The Employee’s Job Moves Up One Level
The shift toward agents does not automatically mean every automated task becomes a lost job.
But it does change what valuable employees spend time doing.
Consider an analyst.
In the assistant era, the analyst asks AI for help researching.
In the copilot era, AI works beside the analyst.
In the agent era, the analyst delegates research and preparation.
The analyst’s job increasingly becomes deciding:
What should be investigated?
Which assumptions matter?
Is the output correct?
What does the result mean?
What should we do?
That pushes humans toward judgment, prioritization and relationships.
The uncomfortable part is that not every role contains enough higher-level work to absorb all the time automation could release. Businesses should therefore avoid pretending that agent adoption has no workforce consequences.
Leaders need to think about job redesign at the same time they think about software deployment.
Managers May Become Managers of Digital Work Too
Management could also change.
Today a manager distributes work across employees.
Tomorrow the same manager may allocate work across people and agents.
A marketing leader might have humans responsible for brand strategy and relationships while agents perform account research, campaign analysis and routine reporting.
A finance leader might have agents prepare reconciliations while accountants investigate exceptions.
A customer operations manager may oversee software handling thousands of routine conversations while the human team focuses on difficult cases.
The managerial skill becomes determining the correct division of responsibility.
What NYC CEOs Should Ask Before Approving an Agent
A useful executive review does not require deep technical knowledge.
It requires clear operating questions.
What exact work is being delegated?
If the team cannot describe the workflow clearly, it is probably too early to automate it.
What can the agent change?
Reading data is different from sending communications.
Sending communications is different from moving money.
Permission should match risk.
What happens when the agent is uncertain?
There should be a defined escalation path.
Can every important action be reconstructed?
If nobody can explain what the system did afterward, the controls are too weak.
What is the failure radius?
One incorrect draft email is small.
Automatically sending 50,000 incorrect emails is not.
Can the action be reversed?
Reversible workflows are excellent early candidates.
What business metric should improve?
Every serious deployment should answer this before implementation begins.
The ROI Formula Should Be Simple
Companies can make agent economics unnecessarily complicated.
Start with:
Annual workflow value = annual task volume × current cost per task
Then estimate:
Agent value = labor removed + cycle-time benefit + error reduction + revenue improvement − agent cost − oversight cost − implementation cost
Suppose a process occurs 100,000 times each year and costs $20 per completion.
That represents $2 million in annual process cost.
If an agent can reliably handle 70% of normal cases at $4 per completed workflow while humans handle the remaining exceptions, the opportunity becomes measurable.
Now the company has something worth testing.
What Not to Automate First
Some workflows should remain human-led for much longer.
Do not start with tasks where a single bad decision can cause enormous irreversible damage.
Do not start with processes nobody inside the company understands.
Do not automate broken operations simply because employees dislike performing them.
And do not automate work where success cannot be measured.
An agent will not repair an unclear operating model.
It may simply execute the confusion faster.
Why 2026 Looks Different From 2024
In 2024, much of enterprise AI was still about access to powerful models.
Companies were deciding which chatbot employees could use.
The next phase looks different.
Models are increasingly surrounded by tools.
Tools are connected to permissions.
Permissions enable actions.
Actions become workflows.
Workflows gain memory.
Memory allows continuity.
Continuity enables delegation.
The same underlying intelligence that once answered a question can now sit inside a system capable of completing meaningful parts of a job.
OpenAI’s enterprise data offers one indication of this shift. As of June 2026, the company reported that agentic AI use, measured through Codex tokens, represented 64% of combined ChatGPT and Codex output tokens among its enterprise customers. Because this is OpenAI customer data rather than a representative survey of all businesses, it should be viewed as a directional signal rather than a general market estimate.
Why New York Could Become the Capital of Applied Agentic AI
New York does not need to beat every other technology hub at building foundation models to win a major part of the AI economy.
It can specialize in turning models into businesses.
That is already the direction NYCEDC is pushing. The city’s AI Nexus initiative is explicitly focused on applied AI and connecting AI companies with industries that can deploy the technology.
The city’s advantage is density.
Founders building finance agents can meet bankers.
Healthcare startups can work with providers.
Legal AI companies can hire lawyers.
Real estate AI companies can test with major property operators.
Advertising AI companies sit near agencies and global brands.
Enterprise startups can sell to large corporate headquarters without leaving Manhattan.
The customer and the builder increasingly live in the same ecosystem.
That matters when building agents because workflows are full of details outsiders do not see.
The Bigger Story: Enterprise Software Is Moving From Tools to Workers
For decades, enterprise software mostly waited.
A CRM waited for someone to update it.
An ERP waited for someone to enter information.
A spreadsheet waited for an analyst.
A support platform waited for an agent.
A project management system waited for employees to move tasks.
Traditional software stored, organized and displayed information.
Agentic software introduces another possibility.
The system itself can participate in the work.
That does not mean every SaaS product becomes an autonomous worker.
But it changes what customers may eventually expect.
A CRM that only stores sales data could look increasingly passive beside a system that watches account activity and executes the next approved step.
An accounting platform that simply records transactions could look limited beside software that performs parts of the close.
A support application that simply routes tickets may struggle against software that actually resolves them.
The competitive question for enterprise software companies may become:
After the customer logs in, how much work still needs to be done manually?
What New York Business Leaders Should Do Now
Companies do not need to predict exactly how capable agents will be in 2030.
They need to understand their workflows in 2026.
Start by identifying five processes where expensive employees perform repetitive digital work.
Measure them.
Choose one.
Give an agent read access.
Test output quality.
Introduce approval-based actions.
Measure corrections.
Automate only the safest patterns.
Keep humans focused on exceptions.
Then repeat.

The businesses that succeed with agentic AI will probably not be the companies that bought the most AI products.
They will be the companies that became best at redesigning work.
Final Takeaway
The transition from AI assistants to AI agents is not primarily a change in interface.
It is a change in responsibility.
Assistants gave employees answers.
Copilots helped employees perform tasks.
Agents are beginning to accept delegated work.
Autonomous workflows can eventually respond to business events and keep processes moving with limited human involvement.
Our original analysis of 11 New York AI companies found an average Agentic Work Index score of 8.0 out of 10. All 11 publicly describe multi-step AI workflows, 73% show strong cross-system or tool capability, and 64% show meaningful operational action authority.
But only 45% showed the strongest level of persistent or background execution in our rubric, while just 18% reached our highest category for exception-based supervised autonomy.
That gap tells us where the industry is going next.
The models are learning how to do the work.
Now businesses have to learn how to let them.
For New York companies, that means the most important AI question is no longer simply:
“How can our employees use AI?”
The much bigger question is becoming:
“Which parts of our company should AI be allowed to run?”



