AI agents are forcing New York companies to answer a question that did not matter nearly as much during the first wave of generative AI: how much authority should software actually have?
The issue is no longer limited to whether an AI system can write a useful email, summarize a document, or answer an employee’s question. Modern AI agents are increasingly designed to take the next step. They can update business systems, contact customers, prepare transactions, issue refunds, change account information, investigate security alerts, create purchase orders, schedule work, and coordinate with other software.
That changes the risk completely.
A chatbot that gives a poor answer may create confusion. An AI agent that acts on a poor answer can create a financial loss, a legal commitment, a bad customer experience, a security incident, or a decision that affects a person’s job, health, or money.
This is why the most important question in enterprise AI is shifting away from, “How smart is the model?” toward a much more practical question: What is this system allowed to do without asking a human?
For New York businesses, that question is especially important. The city has one of the world’s densest concentrations of finance, insurance, health care, law, professional services, media, advertising, real estate, retail, technology, and government. These industries contain thousands of workflows where AI can create real value, but they also contain decisions where mistakes are expensive and sometimes difficult to reverse.
The strongest approach is unlikely to be full human control or full AI autonomy. Companies need something between those extremes. They need a system where AI receives more freedom when an action is routine, low-risk, easy to reverse, and well tested, while human control increases as the consequences become more serious.
That is the real meaning of human-in-the-loop AI in the age of autonomous software.
The Short Version: Human-in-the-Loop AI Is Really About Controlling Authority
Human-in-the-loop AI is often described as an arrangement where a person reviews what an AI system produces. That definition made sense when most business AI created text, recommendations, scores, or predictions.
AI agents make the concept more complicated because they are designed to act.
A traditional AI assistant might tell a customer-service employee that a refund appears appropriate. An AI agent may be able to issue the refund itself. A coding assistant might recommend a change to an application, while an agent could potentially deploy that change into production. A financial AI system might identify an investment opportunity, while an agent connected to trading infrastructure could theoretically act on it.
Those are very different levels of authority.
JPMorganChase has described this distinction in its public work on agentic AI security. The bank has emphasized that agents create additional security concerns because they can interact with systems and take actions under delegated authority. That creates a need for stronger authorization boundaries, traceable actions, controlled access, and safeguards that increase as the agent’s capabilities become more powerful.
The lesson for businesses is simple. AI risk cannot be judged only by how intelligent a model appears to be.
A highly capable model that can only read approved documents may have limited operational power. A much less capable model connected to payment systems could create significant exposure.

Companies should therefore govern agents according to the authority they receive.
Why New York Is an Important Test Market for Human-Controlled AI
New York provides an unusually strong environment for testing these ideas because so much of the city’s economy revolves around information, judgment, transactions, regulated activity, and human services.
According to New York City Comptroller employment data for June 2026, the city had roughly 4.86 million seasonally adjusted nonfarm jobs. Health and social assistance represented more than one million jobs. Professional and business services accounted for roughly 805,000, government approximately 626,000, financial activities around 522,000, educational services approximately 264,000, and the information sector about 220,000.
These numbers matter because many of the workflows inside these sectors are natural targets for AI agents.
Financial institutions process large volumes of research, compliance work, customer requests, transactions, and reporting. Hospitals and health systems manage scheduling, records, documentation, insurance processes, patient communication, and clinical information. Law firms work through contracts, discovery, research, filings, and due diligence. Professional-service firms rely heavily on documents, analysis, communication, project management, and approval chains.
The AI opportunity is therefore enormous, but so is the need for control.
New York is not only becoming a market for AI adoption. It is becoming a market for deciding how much machine authority businesses are comfortable granting.
Original NYC Tech Journal Analysis: More Than 70% of NYC Jobs Sit in Knowledge or Human-Service Sectors
To understand why human oversight matters so much in New York, NYC Tech Journal analyzed the latest detailed employment figures available from the New York City Comptroller’s August 2026 economic report.
We regrouped the major employment sectors into three broad operational categories. These categories are our own analytical framework and are not official government classifications.
Our Methodology
The first group, office and knowledge work, includes financial activities, information, and professional and business services. These are sectors where AI agents are likely to interact heavily with documents, software, analysis, communication, transactions, and structured business processes.
The second group, human-service and public-decision work, includes health and social assistance, educational services, and government. These industries contain many workflows where automation can help, but where decisions can directly affect people.
The third group, customer and physical operations, includes retail, hospitality, wholesale trade, transportation and warehousing, construction, and manufacturing.
Using the June 2026 employment figures, the result looks like this.
| NYC Tech Journal category | June 2026 employment | Share of NYC nonfarm jobs |
| Office and knowledge work | 1.546 million | 31.8% |
| Human-service and public-decision work | 1.917 million | 39.5% |
| Customer and physical operations | 1.199 million | 24.7% |
| Other sectors | 196,000 | 4.0% |
| Total | 4.858 million | 100% |
Source employment data: New York City Comptroller. Groupings and calculations are NYC Tech Journal analysis.
The first two categories together account for approximately 71.3% of New York City’s payroll employment.
This does not mean that 71% of jobs are about to be automated. It means something more useful for business leaders: a large share of New York employment sits inside sectors where AI agents will often touch information, judgment, professional responsibility, financial decisions, records, or human services.
That makes the question of oversight unusually important here.
The Wrong Question Is Whether a Human Should Be “In the Loop”
Many companies still frame the problem too simply. They ask whether a workflow should include a human or whether it can be automated.
A much better question is where the human should enter the workflow.
Consider a customer-service agent that handles refunds. One company may require an employee to review every refund before it is issued. Another company may allow the agent to automatically approve refunds under $50, require human approval between $50 and $500, and prevent the agent from processing anything larger.
Both companies technically have humans involved, but the second company has designed a much more sophisticated operating model.
It has identified a boundary.
That boundary is what makes controlled autonomy possible.
Instead of asking whether AI should be autonomous, companies should decide which actions can be autonomous, under what conditions, and up to what limit.
The Five Questions That Should Decide How Much Freedom an Agent Gets
For this analysis, NYC Tech Journal developed a simple human-oversight scoring model for common business actions.
The purpose is not to create another theoretical AI ethics framework. The goal is to give business leaders a practical way to discuss authority before an agent receives access to important systems.
Each action can be scored from zero to two across five dimensions.
Does the Action Materially Affect a Person?
An internal administrative task generally carries less risk than a decision that directly affects someone’s employment, health, financial position, legal rights, or access to a service.
A system that reorganizes internal files is different from a system that rejects an applicant or changes a customer’s financial account.
The closer the agent gets to decisions that meaningfully affect people, the stronger human control should usually become.
Can the Agent Commit the Business?
Another major question is whether the agent merely recommends an action or can actually carry it out.
Drafting a purchase order is different from sending it.
Preparing contract language is different from signing a contract.
Recommending a refund is different from processing one.
The moment software gains the ability to commit money, alter systems, sign agreements, or communicate externally on behalf of the business, its risk changes substantially.
Does the Action Involve Sensitive Data?
The information available to the agent matters as much as the action itself.
An AI system working with public marketing information is very different from one that can access medical records, financial data, employment information, legal files, customer identity information, or security credentials.
The combination of sensitive data and execution authority deserves particularly close attention.
How Easy Is the Action to Reverse?
Reversibility is one of the most useful ways to think about agent risk.
A mistaken calendar invitation is inconvenient, but easy to correct. Sending a large payment to the wrong destination is very different. Terminating an employee, exposing confidential information, signing a contract, or modifying clinical treatment can also create consequences that cannot be easily undone.
Oversight should therefore increase as reversibility decreases.
Is There a Strong Regulatory or Professional Duty?
Some workflows exist inside industries with significant regulatory, supervisory, fiduciary, professional, or safety obligations.
A routine internal administrative process may need limited oversight. A decision involving credit, insurance, securities, employment, legal advice, medical care, or public benefits may sit within a much stricter environment.
That context should affect how much authority the agent receives.
The NYC Tech Journal Human Oversight Scale
Using those five dimensions, an action can receive a score between zero and ten.
| Score | Suggested operating model | Practical meaning |
| 0–2 | Autonomous with logging | Agent can normally act inside a narrow scope |
| 3–4 | Autonomous with monitoring | Agent acts, but activity is tracked and reversible |
| 5–6 | Threshold or exception review | Agent acts only inside clearly defined limits |
| 7–8 | Human approval before execution | Agent prepares the action but cannot complete it alone |
| 9–10 | Human decision required | AI assists, while the consequential decision remains human |
This model is not intended to replace legal analysis. It is an operating framework that helps teams ask better questions.
The important change is that the conversation moves away from saying, “This AI seems safe,” and toward saying, “This specific action has this level of authority and therefore needs this level of oversight.”
Original Analysis: How Common AI Agent Actions Compare
We applied the framework to 18 actions that many New York businesses could realistically consider delegating to agents.
| Agent action | Oversight score | Suggested control |
| Search and summarize internal documents | 1 | Autonomous |
| Draft an external email for review | 1 | Autonomous |
| Update CRM notes | 3 | Monitor and audit |
| Send routine customer status messages | 4 | Monitor and keep reversible |
| Publish low-risk marketing content | 4 | Monitor with brand controls |
| Negotiate contract language without signing | 5 | Exception review |
| Issue a small refund | 6 | Dollar limits and exception review |
| Create a procurement order | 6 | Spend limits and approval rules |
| Deploy code into production | 7 | Human approval |
| Run security remediation on live systems | 8 | Human approval or strict playbooks |
| Sign a contract | 9 | Human authorization |
| Change customer bank details | 9 | Human authorization and identity checks |
| Rank or reject a job applicant | 10 | Human-controlled high-impact process |
| Approve or deny credit | 10 | Human-controlled regulated process |
| Set insurance eligibility or price | 10 | Human-controlled regulated process |
| Execute a securities trade | 10 | Strong supervisory controls |
| Issue a clinical treatment order | 10 | Clinician-controlled process |
| Deny or terminate a public benefit | 10 | Human-controlled public decision |
A clear pattern appears across the table.
Risk does not rise mainly because the AI becomes better at writing or reasoning. Risk rises when the AI gains the ability to turn its reasoning into a consequential action.
That distinction should shape enterprise AI governance.
Original Finding #1: External Write Access Is a Bigger Governance Event Than a Better Model
Companies often spend months comparing AI models.
One model may score slightly higher on an evaluation. Another may be faster. A third may handle longer documents or follow instructions more reliably.
Those differences matter, but one change can matter far more from a governance perspective: giving the agent permission to write into an external system.
Consider the difference between these two outcomes.
An AI agent says that it recommends sending an email.
Then imagine that the same agent can send the email automatically.
The model has not necessarily become smarter, yet the operational risk has changed dramatically because the agent now has the ability to represent the company externally.
The same principle applies when a refund recommendation becomes an executed refund, when a proposed database change becomes a live database update, or when a draft purchase order becomes an actual financial commitment.
This is why every move from read or recommend into write or execute should trigger a new governance review.
JPMorganChase has made a similar point in its discussion of delegated authority for AI agents. The more systems an agent can reach and the more actions it can perform, the more important identity controls, authorization boundaries, and detailed audit records become.
Morgan Stanley Offers a Useful Model for Human-Controlled AI
Morgan Stanley provides one of the clearest public examples of what thoughtful human-in-the-loop design can look like.
Its AI systems support financial advisers by helping them retrieve internal knowledge, summarize information, process meeting material, and prepare follow-up work. OpenAI’s public case study describes how Morgan Stanley built evaluation systems around real business tasks and used expert feedback to check the quality of AI outputs.
Its Debrief system is particularly relevant to the human-oversight discussion. The system can produce meeting notes and draft follow-up communication, while advisers remain responsible for reviewing and adjusting the material before it is finalized.
This design matters because the human is not forced to perform every step manually.
The AI can listen, extract information, organize details, summarize the conversation, and prepare the draft. Human attention is concentrated near the point where the work becomes an external or professional action.
That is much more efficient than placing a human between every AI step.
Human Oversight Should Not Mean Reviewing Everything
One of the easiest mistakes companies can make is assuming that safer AI requires more human approvals everywhere.
That approach often destroys the economic value of automation.
Imagine an AI agent handling tens of thousands of routine customer requests every day. If an employee must individually approve every tracking update, appointment reminder, password-reset email, or low-value account change, the business has created a new manual bottleneck.
The goal should not be maximum human involvement.
The goal should be meaningful control where meaningful consequences exist.
NIST’s AI Risk Management Framework supports this risk-based approach. It calls on organizations to define human roles, document oversight responsibilities, identify which system capabilities need stronger controls, and evaluate whether those controls actually work in practice.
This last point is easy to miss. Adding a human approval button does not automatically create effective governance.
The Rubber-Stamp Problem Could Become a Major Enterprise AI Failure
Human approval sounds reassuring until the approval becomes routine.
Imagine a compliance analyst reviewing hundreds of AI recommendations every day. If the system is normally correct, the analyst may gradually become accustomed to accepting its recommendations.
At first, each case receives careful attention. Over time, the process can turn into a repeated sequence of approvals.
The person remains technically in the loop, but the practical value of the oversight weakens.
This is one reason companies should design review systems that make disagreement possible.
A screen that simply says, “The AI recommends approval. Approve?” gives the reviewer very little useful information.
A better interface might explain that the AI is requesting a $742 refund, that the normal automated limit is $250, that the customer has submitted several recent refund requests, that shipping records indicate delivery, and that a specific policy section applies.
The difference is substantial.
The first interface asks the human to trust the machine. The second gives the human evidence that can be challenged.
Original Finding #2: Human Review Should Sit at Decision Boundaries
Our action analysis suggests that companies can automate large amounts of work without placing humans inside every intermediate step.
The better approach is to identify decision boundaries.
A contract agent, for example, could search thousands of agreements, extract clauses, compare language, identify unusual terms, summarize risks, and draft alternative wording automatically.
The important boundary comes when the business is asked to accept the agreement.
That is where stronger human control makes sense.
A similar pattern can work in finance. An AI agent may gather data, monitor markets, investigate anomalies, calculate ratios, prepare reports, and draft recommendations without requiring constant human intervention. Stronger controls can appear when the system is about to move money, execute a trade, change a client’s account, or make another consequential commitment.
Health care provides another example. AI can support documentation, retrieve records, organize information, prepare administrative work, and assist clinicians without receiving authority to make consequential treatment decisions independently.

Mount Sinai has publicly described its rollout of Microsoft Dragon Copilot in similar terms, presenting AI as a system that supports clinicians rather than replacing professional judgment. The health system also emphasized phased deployment, training, feedback, and evaluation.
Health Care Shows Why High Average Accuracy Is Not Enough
Health care also demonstrates why companies should not rely only on overall AI accuracy.
Mount Sinai researchers reported in 2025 that generative AI models sometimes changed medical recommendations when demographic or socioeconomic details changed, even though the underlying clinical situation remained the same. The research used 1,000 emergency-department cases and generated more than 1.7 million model recommendations across varied patient profiles.
Another Mount Sinai study found that AI systems could sometimes continue giving familiar answers after important facts in medical ethics scenarios had been changed. Researchers said the findings showed why human oversight remained important in nuanced and high-stakes situations.
The operational lesson reaches far beyond medicine.
An AI system can perform extremely well on average and still fail badly in unusual situations. Strong governance therefore depends not only on whether the system is usually right, but also on whether it can recognize when a case has moved outside the conditions it understands well.
The Best Agent May Be the One That Knows When to Stop
Enterprise AI teams often judge agents by completion rates.
A higher completion rate sounds better because it means the agent can finish more work without help.
That metric becomes dangerous when it encourages systems to continue acting even when evidence is poor.
A mature enterprise agent should be capable of recognizing several different kinds of uncertainty. It should notice when information conflicts, when required data is missing, when an amount exceeds its authority, when identity cannot be verified, when a policy is unclear, or when a requested action falls outside its approved permissions.
In those situations, escalation is not a failure.
It is evidence that the control system is working.
Companies may eventually find that one of the most valuable qualities in an enterprise agent is not how often it completes the task, but how accurately it recognizes the situations where it should stop.
New York’s Regulatory Environment Already Points Toward Risk-Based AI Governance
There is no single New York rule that tells every company exactly when an AI agent requires human approval.
Instead, businesses operate within a combination of employment laws, financial regulation, insurance rules, cybersecurity expectations, professional duties, government AI standards, and existing supervisory obligations.
These rules vary by industry, but many of them point toward the same basic idea: companies remain responsible for what automated systems do.
NYC Local Law 144 Makes Employment AI a High-Sensitivity Area
New York City’s Local Law 144 applies to certain automated employment decision tools.
The city’s Department of Consumer and Worker Protection explains that employers and employment agencies cannot use a covered automated employment decision tool unless required bias-audit and notice conditions have been satisfied. Enforcement began in July 2023.
The broader lesson for agent design is important.
An AI recruiting tool that schedules interviews is very different from a system that decides which candidate receives an interview.
An agent that drafts recruiter outreach is different from one that automatically rejects candidates.
As soon as software begins evaluating or materially affecting employment decisions, the governance requirements become much more serious.
This is why the label “HR agent” is almost meaningless from a risk perspective.
The real question is what the HR agent is permitted to do.
NYDFS Provides a Strong Model for AI Accountability
The New York State Department of Financial Services has created one of the clearest New York examples of risk-based AI governance.
Its 2024 circular letter covering insurers’ use of artificial intelligence and external consumer data in underwriting and pricing emphasizes governance, defined responsibilities, senior-management oversight, documentation, monitoring, testing, internal controls, and appropriate restrictions on how AI is used.
The guidance also makes an important accountability point.
Using a third-party AI provider does not remove the insurer’s responsibility for the outcome.
This principle is likely to matter far beyond insurance as agent adoption increases.
Businesses cannot outsource accountability simply by outsourcing the software.
If an agent acts inside your workflow, affects your customers, uses your data, or exercises authority on behalf of your organization, its actions remain part of your operating environment.
Wall Street Cannot Use AI to Escape Existing Supervision
The same logic applies in securities markets.
FINRA’s Regulatory Notice 24-09 makes clear that existing FINRA rules and securities laws continue to apply when firms use generative AI. The technology may be new, but the firm’s underlying obligations do not disappear.
For New York financial institutions, this matters enormously.
An AI system does not create a regulatory vacuum simply because software performed the analysis or prepared the recommendation.
Firms need to examine how existing supervisory, recordkeeping, communication, suitability, security, and compliance requirements apply to AI-enabled workflows.
It is also important to separate active rules from proposals that never became final requirements. For example, the SEC’s proposed predictive-data-analytics rules for broker-dealers and investment advisers were withdrawn in June 2025 rather than finalized.
Good AI governance should be built around the rules that actually apply rather than assumptions based on headlines.
New York City Government Is Building Similar Governance Structures
New York City’s own AI work offers another sign of where enterprise governance is heading.
The city’s AI Action Plan established initiatives involving governance, risk assessment, procurement, monitoring, employee skills, transparency, and responsible adoption. By the city’s October 2025 progress update, several foundational pieces had been completed, including an AI steering committee, guiding principles, preliminary use guidance, an AI project typology, and expanded reporting.
The city’s published AI principles emphasize accountability, validity, risk management, and responsible deployment.
Additional city legislation enacted in December 2025 created baseline standards for certain public-facing AI systems used by agencies, including risk assessment, monitoring, privacy protection, fairness, and accountability requirements.
These requirements are aimed at government systems rather than every private business in the city.
Still, the direction is significant.
Governance becomes stronger when the system has more power and when its decisions affect people more directly.
A Better Model: Five Levels of Agent Autonomy
Companies need something more useful than a simple “human versus machine” distinction.
One practical solution is to create explicit levels of agent authority.
Level 1: Read
At the first level, the agent can search, classify, summarize, compare, and analyze approved information.
It cannot modify records, communicate externally, commit money, or trigger business actions.
This is a useful starting point for companies that are still learning how their agents behave.
Level 2: Prepare
At the second level, the agent can prepare work that a person will use.
It may draft an email, create a report, prepare a financial analysis, summarize a contract, or assemble a customer-service response.
The AI performs the preparation, while a person controls whether the output moves forward.
Level 3: Execute Bounded, Low-Risk Actions
At the third level, the agent can perform carefully defined actions without receiving approval every time.
Examples might include scheduling meetings, sending standard reminders, updating selected CRM fields, routing support requests, or issuing very small service credits.
Actions should be logged and generally easy to reverse.
Level 4: Execute After Approval
At the fourth level, the AI performs almost all of the work but stops before the consequential action.
It might prepare a large refund, a production deployment, a contract revision, a purchase order, a security change, or an important external communication.
A qualified human must approve execution.
Level 5: Human Decision Required
At the highest level, AI can still perform significant supporting work.
It can gather data, calculate results, identify patterns, surface inconsistencies, prepare options, and create documentation.
However, the final decision remains with a qualified human.
Employment decisions, major credit decisions, clinical treatment choices, significant legal commitments, major financial actions, and similar high-impact cases often belong here.
The Agent Autonomy Ladder
| Level | Agent capability | Human responsibility |
| 1 | Read and analyze | Monitor data access |
| 2 | Draft and recommend | Review output |
| 3 | Execute bounded low-risk actions | Monitor exceptions |
| 4 | Prepare high-impact action | Approve execution |
| 5 | Support high-stakes decision | Make the final decision |
This framework makes one point especially clear.
An agent should not simply receive “access to Salesforce,” “access to banking,” or “access to the production environment.”
Permissions should describe actions in precise terms.
The agent may read the customer record, update shipping status, or issue refunds below a defined amount. It may be prevented from changing bank information, closing accounts, modifying identity data, or removing fraud flags.
That level of precision turns policy into real operating control.
Write the Agent’s Job Description Before Writing Its Prompt
Many companies begin agent projects by experimenting with prompts.
A better starting point is to define the agent’s job.
Before building the system, the company should understand its purpose, the systems it may access, the information it can read, the fields it may modify, the actions it can take, the value limits that apply, the conditions that require escalation, and the activities it may never perform.

The company should also identify a human owner and define how the agent can be suspended if something goes wrong.
This is similar to creating a job description for an employee.
A business would never hire someone and give that person unlimited access to every system while telling them to discover their own responsibilities.
AI agents deserve the same clarity.
The Agent Should Usually Have Less Access Than the Person Who Built It
A developer or business leader involved in creating an AI agent may have broad system access.
The deployed agent should usually receive much narrower permissions.
This follows the long-established security principle of least privilege.
A customer-service agent does not need access to employee payroll data. An HR agent does not need production database credentials. A marketing agent does not need access to company bank accounts. A contract-analysis agent does not need the ability to sign agreements.
JPMorganChase’s public work on agent security emphasizes exactly these ideas through delegated identity, authorization boundaries, system constraints, and traceable activity.
Narrow permissions reduce the number of ways that a mistake can become a serious incident.
The Two-Switch Test for High-Risk Agents
A simple screening test can help companies identify particularly risky agent designs.
Ask whether the agent has both of these capabilities.
First, can it access sensitive information?
Second, can it perform a consequential external action?
When both capabilities exist at the same time, controls should become significantly stronger.
An agent that reads financial information and can send money deserves strong restrictions. So does an agent that reads applicant data and can reject candidates, an agent that sees medical records and can trigger treatment actions, or a system that holds production credentials and can change infrastructure.
JPMorganChase has highlighted a closely related concern involving the combination of untrusted inputs, sensitive information, and external action. When these abilities come together, the potential impact increases sharply.
The lesson is that agent risk often comes from combinations of permissions rather than any single feature.
Human Review Should Focus on Exceptions
The most scalable oversight systems will probably not require humans to approve every action.
They will require human attention when something unusual happens.
Consider accounts payable.
An AI agent could potentially process a routine invoice automatically when the vendor is approved, the purchase order matches, the price sits within normal tolerance, banking details have not changed, the invoice is not duplicated, and the expense category is expected.
Human review can begin when those conditions fail.
A new vendor may require approval. A changed bank account should raise concern. A duplicated invoice, unusual location, missing purchase order, large amount, or suspicious email domain could trigger escalation.
This approach preserves automation while concentrating human attention on the cases where judgment is valuable.
Dollar Limits Should Be Built Into Agent Architecture
Financial limits provide one of the clearest ways to control agent authority.
A company might allow a customer-service agent to issue small refunds automatically while requiring approval above a specific threshold.
The same structure can work for purchasing, service credits, discounts, and other financial decisions.
| Agent action | Automatic | Human approval | Not permitted autonomously |
| Customer refund | $0–$75 | $76–$1,000 | Above $1,000 |
| Purchase | $0–$250 | $251–$5,000 | Above $5,000 |
| Service credit | $0–$100 | $101–$500 | Above $500 |
| Contract commitment | None | All | Never fully autonomous |
| Bank-detail change | None | All | Never fully autonomous |
The exact amounts will vary by organization.
What matters is that authority has a ceiling.
An agent should not receive unlimited spending power simply because the company trusts the model.
Confidence and Evidence Quality Should Also Affect Autonomy
Financial value is only one factor.
The quality of available evidence should influence whether an agent can continue.
Consider an insurance workflow where AI extracts information from documents.
If all required information is present, records agree with one another, and the case clearly fits known rules, the workflow may continue automatically.
If information conflicts, documents are missing, or the case falls near a policy boundary, the agent should escalate.
This distinction between high-confidence routine work and low-confidence consequential work is critical.
Companies should avoid using the same autonomy rule for both.
Health Care Needs an Especially Strong Boundary Between Assistance and Authority
Health care may become one of the largest opportunities for agentic AI in New York.
The city’s health and social assistance sector alone accounted for more than one million jobs in the June 2026 data.
Many health-care workflows are excellent candidates for automation.
Agents can help retrieve records, prepare documentation, coordinate referrals, complete forms, handle scheduling, prepare prior-authorization materials, support coding, organize patient communication, and complete administrative follow-up.
These activities consume enormous amounts of staff time.
The risk changes when AI moves from administrative support toward clinical authority.
FDA guidance around AI and machine-learning medical devices has emphasized the performance of the overall human-AI team, transparency, lifecycle monitoring, risk management, and safety.
For New York health systems, the most valuable model may therefore be one where AI performs much of the administrative work while clinical authority remains carefully bounded.
That approach can create significant productivity gains without treating medical judgment as just another automated workflow.
Finance May Build Some of the Most Sophisticated Approval Systems
Financial institutions face a different challenge.
The economic value of AI comes partly from speed and scale, which means companies cannot place a person between every analytical step.
Agents may eventually monitor markets continuously, reconcile accounts, inspect transactions, prepare research, analyze portfolios, investigate anomalies, generate reports, and coordinate other systems with very little manual intervention.
The control layer can become stronger as the workflow approaches consequential action.
Transferring money, changing client information, approving credit, executing securities transactions, altering risk limits, or issuing regulated communications may sit behind stricter authorization requirements.
This allows financial institutions to automate deeply without treating all actions as equally safe.
Law Firms Should Place Human Review Around Professional Judgment and Commitment
Legal work provides another strong example of selective oversight.
AI agents can search documents, summarize discovery, organize timelines, identify clauses, compare contract terms, prepare due-diligence material, monitor deadlines, and draft language.
Those activities can remove large amounts of repetitive work.
The important boundary appears when professional judgment or legal commitment begins.
An agent may identify a clause, but an attorney remains responsible for deciding what that clause means for the client.
An agent may prepare a filing, but the responsible lawyer must determine whether it is accurate and appropriate to submit.
An agent may suggest alternative contract language, while the human controls whether the business accepts that commitment.
Good human-in-the-loop design places oversight where professional responsibility lives.
HR Agents Need Particularly Clear Limits in New York
Human resources shows how quickly an ordinary administrative tool can become a high-impact decision system.
An HR agent that answers questions about vacation policy may create relatively little risk.
Scheduling interviews is also largely administrative.
Summarizing resumes introduces more judgment.
Scoring candidates changes the nature of the workflow further.
Automatically rejecting applicants creates a much more consequential system.
This is why companies should avoid broad labels such as “HR AI” or “recruiting agent” when discussing risk.
Those labels hide the important details.
The real question is whether the system merely supports an employment process or materially determines who moves forward.
Customer-Service Agents Can Receive More Freedom When the Boundaries Are Clear
Customer service will probably become one of the earliest areas where companies grant AI agents meaningful operating authority.
Many customer-service actions are frequent, structured, and reversible.
An agent might change an appointment, resend an invoice, explain order status, generate a return label, update communication preferences, or issue a small service credit.
These are strong candidates for controlled autonomy.
The same agent should not automatically receive permission to change financial information, modify identity records, waive large debts, close high-value accounts, or make significant contractual commitments.
A good architecture separates low-risk convenience from high-risk authority.
Procurement Agents Need Spending Ladders
Procurement AI becomes far more valuable once it can move beyond supplier research.
An advanced procurement agent might gather quotes, compare suppliers, negotiate standard terms, check available budget, prepare a purchase order, and coordinate delivery.
This could remove significant delay from routine buying.
However, a $50 office-supply order and a multimillion-dollar software contract should never move through the same approval path.
Procurement autonomy should increase or decrease according to purchase value, contract length, supplier risk, access to sensitive information, security impact, and business importance.
The result should be a spending ladder rather than one universal procurement rule.
IT Agents Need a Clear Boundary Between Diagnosis and Production Change
Technology and cybersecurity teams will face this problem quickly because agentic tools are increasingly capable of interacting with live infrastructure.
An agent that reads logs and identifies a likely error presents one level of risk.
An agent that restarts a production database presents another.
A coding system that proposes a software change is different from one that deploys that change to customers.
A security agent that recommends revoking a compromised account is different from one that disables access automatically.
The closer an agent gets to production execution, the more carefully companies should think about authorization, rollback, monitoring, escalation, and incident response.
Every High-Authority Agent Needs a Kill Switch
Human control is not only about approving actions.
Companies must also be able to stop agents quickly.
Every agent with meaningful execution authority should have a clearly defined shutdown process.
The company should know who can suspend the agent, whether suspension immediately removes credentials, what happens to jobs already in progress, how partially completed actions are handled, and whether the organization can identify everything the agent changed before shutdown.
Rollback matters as well.

If the agent updates thousands of records incorrectly, the company needs to know whether those changes can be reversed.
If shutting down one agent affects several connected systems or other agents, that dependency should be understood before deployment.
These are basic operational questions, but many AI pilots ignore them until something breaks.
Logging Should Record Actions, Not Just Conversations
Chat transcripts are not enough for enterprise agents.
A business needs an action trail.
The organization should be able to determine which agent performed an action, which identity it used, what system it accessed, what information it read, what external tool it called, what response the tool returned, which policy allowed the action, whether human approval was required, who approved it, and what changed afterward.
JPMorganChase has emphasized the importance of detailed execution records for autonomous and semi-autonomous AI because those records support auditing, investigation, monitoring, and incident response.
This may become one of the biggest differences between ordinary consumer AI and serious enterprise AI.
A company needs to know not merely what the agent said, but what it did.
Companies Need to Measure the Human Side of the Workflow Too
Businesses often measure model accuracy while paying very little attention to whether human oversight actually works.
That is a mistake.
NIST’s AI Risk Management Framework playbook recommends examining human oversight through measures such as overrides, complaints, downstream actions, exceptions, adjudication activity, errors, and important go-or-no-go decisions.
A practical human-in-the-loop dashboard could include the following metrics.
| Metric | What it reveals |
| Autonomous completion rate | How much work the agent finishes without intervention |
| Escalation rate | How frequently the agent reaches an authority boundary |
| Human approval rate | How often reviewers accept the AI’s recommendation |
| Human override rate | How often people identify a meaningful problem |
| False escalation rate | Whether too much easy work is being sent to people |
| Missed escalation rate | Whether risky cases escape review |
| Average review time | Cost and speed of human oversight |
| Rollback rate | How often completed actions must be undone |
| Customer complaint rate | Whether automated work causes downstream problems |
| Policy exception rate | How often teams bypass normal controls |
| Incident rate | Operational failures per unit of activity |
| Cost per completed workflow | Whether the system creates measurable value |
The human override rate deserves particular attention.
A zero override rate may look impressive, but it can indicate either an extraordinarily reliable system or a weak review process where people rarely question the AI.
Management needs to know which one it is.
Approval Latency Will Become a New Enterprise AI Metric
Another important metric is the amount of time an agent waits for a person.
Imagine that an AI completes its part of a workflow in 30 seconds but then waits eight hours for an approval.
The model is fast, yet the business process remains slow.
This is why companies that want real productivity improvements need to redesign human approval as carefully as they design the agent.
Urgent high-value decisions may require real-time review queues.
Medium-priority cases might be reviewed in batches.
Very low-risk actions may move away from pre-approval entirely and instead use retrospective monitoring and random sampling.
Once companies begin thinking this way, human-in-the-loop stops being an AI concept and becomes an operations discipline.
Original Finding #3: Human-on-the-Loop May Matter More Than Human-in-Every-Loop
As agents become more reliable, many businesses will probably shift some workflows from human-in-the-loop toward human-on-the-loop.
The difference is important.
Human-in-the-loop means that software pauses and waits for human approval before performing a particular action.
Human-on-the-loop means that software operates independently inside predefined limits while humans supervise overall performance and intervene when unusual conditions appear.
Low-risk, high-volume workflows are natural candidates for the second model.
A company does not need 10,000 employees approving 10,000 standard shipment-status messages. The agent can send those messages automatically while managers monitor complaints, failed deliveries, policy violations, unusual cases, and random quality samples.
The role of the human changes from individual approver to system supervisor.
That shift could eventually transform management itself.
Oversight Should Rise as Actions Become Harder to Reverse
Reversibility deserves far more attention than it normally receives in AI discussions.
Two AI actions can have the same probability of error but completely different consequences.
Sending an incorrect internal meeting invite is irritating but easy to fix.
Sending a large wire transfer to the wrong destination is not.
Publishing confidential information, terminating an employee, signing a binding agreement, changing medical treatment, or deleting production data may also be difficult or impossible to reverse quickly.
This gives businesses a useful operating principle.
Oversight should not be based only on the probability that the AI is wrong.
It should also be based on the cost of being wrong.
The Economics of Human Review Can Be Modeled
Human oversight also has an economic cost, which means businesses should evaluate it rather than treating more review as automatically better.
A simple starting formula is:
Expected AI risk = probability of serious error × impact of error × number of actions
Consider an AI system that performs 100,000 low-value actions each year.
If the chance of serious error is 0.01% and the average financial harm from an error is $50, expected annual direct loss would be:
100,000 × 0.0001 × $50 = $500
If manually reviewing all 100,000 actions costs $300,000, mandatory pre-approval for every action may make little economic sense unless there are other important legal, reputational, or safety concerns.
Now consider 1,000 high-value actions with a 0.5% serious-error rate and potential harm of $500,000 per failure.
The expected exposure becomes:
1,000 × 0.005 × $500,000 = $2.5 million
In that environment, meaningful human review can look economically reasonable.
This model will never capture every type of harm, but it forces businesses to consider the relationship between probability, consequence, scale, and review cost.
Build Human Review Before Building Full Autonomy
Companies should resist the temptation to launch a fully autonomous system and add safeguards afterward.
The safer approach moves in the opposite direction.
Begin with an agent that prepares work and requires human approval.
Collect real performance data.
Measure where humans disagree.
Study failures, unusual cases, and escalation patterns.
Then remove human approval from the safest categories once the evidence supports that change.
Morgan Stanley’s approach to enterprise AI offers a useful example of this philosophy. Its deployment process included evaluations tied to real-world tasks, expert review, quality testing, and gradual expansion as the firm developed greater confidence in the systems.
Autonomy should be earned through evidence rather than assumed at launch.
A 90-Day Human-in-the-Loop Plan for New York Companies
New York businesses do not need a year-long governance project before they can improve agent controls.
A focused 90-day program can establish the most important foundations.
Days 1–20: Inventory Agent Actions
Do not begin with a list of AI models.
Begin with a list of actions.
For every agent in production or development, document what it can read, write, send, modify, recommend, approve, purchase, publish, or execute.
Also record the systems and data involved.
A company with 15 AI agents may discover that those agents perform more than 100 distinct actions.
Those actions, rather than the number of models, are the real governance surface.
Days 21–40: Score Each Action
Apply a consistent framework across the company.
Evaluate human impact, authority, sensitive-data access, reversibility, and regulatory or professional significance.
Assign an autonomy level to every material action and document the reason.
The scoring does not need to be mathematically perfect.
Its main purpose is to create consistent decision-making across teams.
Days 41–60: Build the Control Layer
Translate governance policy into technical restrictions.
This may include read-only permissions, dollar limits, API restrictions, approved recipient lists, human approval gates, production boundaries, rate limits, role-based access, logging, escalation triggers, and rollback mechanisms.
A rule that exists only inside a policy document is weaker than a rule enforced directly by the system.
Days 61–75: Test Failure Conditions
Do not test agents only under ideal circumstances.
Give them incomplete information, conflicting instructions, unusual customer behavior, changed policies, unavailable systems, ambiguous requests, high-dollar cases, unexpected API responses, sensitive data, and prompt-injection attempts.
The goal is to discover when the agent continues acting even though it should stop.
That information is often more valuable than another general benchmark score.
Days 76–90: Launch With Narrow Authority
Begin with the safest approved actions.
Measure escalation rates, human overrides, complaints, rollback events, review times, policy exceptions, and cost per completed workflow.
Expand authority only when the evidence supports expansion.
This makes autonomy a controlled operating decision rather than a one-time software feature.
Every Agent Needs a Named Human Owner
There is one question senior leadership should be able to answer about every important agent.
Who owns the outcome?
The answer should not simply be the software vendor or the team that purchased the tool.
NIST’s AI Risk Management Framework emphasizes documented roles, responsibilities, accountability, and executive ownership of AI-related risk decisions.
NYDFS takes a similar approach in insurance by placing responsibility on governance bodies and senior management even when technical implementation is delegated or third-party technology is used.
An AI system without a clearly identified owner can easily become a system where everyone assumes someone else is responsible.
That is exactly the situation companies should avoid.
Vendor Agents Need the Same Standards as Internally Built Agents
Most New York companies will probably buy more AI agents than they build themselves.
That does not reduce governance responsibilities.
Before connecting a vendor agent to important business systems, companies should understand what the software can access, which actions it can initiate, how permissions are assigned, where activity logs are stored, how customer data is handled, how the underlying model changes, how incidents are investigated, and whether the customer can independently disable the system.
Broad phrases such as “enterprise grade” or “secure by design” do not answer these questions.
Businesses need to understand what the agent can actually do.
A Practical Vendor Control Checklist
| Question | Why it matters |
| Can the agent be made read-only? | Limits unintended execution |
| Can permissions be assigned by action? | Prevents unnecessary access |
| Can dollar limits be configured? | Restricts financial authority |
| Can approval be required before tool calls? | Creates meaningful control points |
| Can actions be logged independently? | Supports investigation and audit |
| Can logs be exported? | Preserves evidence outside the vendor |
| Can the agent be suspended immediately? | Supports incident response |
| Can individual tools be disabled? | Allows precise containment |
| Can data access be restricted by role? | Protects sensitive information |
| Are model or prompt changes versioned? | Makes behavior easier to trace |
| Are failed actions recorded? | Reveals hidden operational problems |
| Can sensitive actions require two approvers? | Supports separation of duties |
If a vendor cannot support most of these controls, the product may still be useful for low-risk work.
It may not be ready for high-authority enterprise use.
Agents Should Not Operate Through Human Accounts
One especially dangerous shortcut is allowing an agent to use an employee’s normal login credentials.
That weakens traceability.
A mature system should be able to distinguish between an employee performing an action directly, an agent performing an action on behalf of that employee, and an employee approving an agent’s action.
Those are different events.
JPMorganChase has highlighted delegated identity and authorization as important parts of secure agent architecture.
As agent adoption grows, managing machine identities may become as important as managing employee identities.
Multi-Agent Systems Make Human Oversight More Complicated
The challenge becomes even harder when several agents work together.
Imagine one agent performing research, another evaluating the information, another negotiating, another creating a transaction, and a fifth monitoring the result.
A company cannot place a human between every interaction without destroying the value of the system.
At the same time, removing all human control could allow errors to move from one agent to another and become more serious.
The most practical solution is again to focus on boundaries.
Agents can exchange information and complete approved internal steps automatically. Human oversight becomes stronger when the system approaches money movement, external communication, production infrastructure, sensitive records, legal commitments, or high-impact decisions.
Mount Sinai researchers studying orchestrated multi-agent systems in health care have emphasized traceability, including the ability to record which tools were called, what information those tools returned, and how the final result was assembled.
That type of traceability will become important in nearly every multi-agent environment.
Boards Should Govern Risk Appetite, Not Individual Prompts
Strong governance does not require senior executives to inspect every AI instruction or prompt.
That would be inefficient and unrealistic.
Boards and senior management should instead establish risk appetite and accountability.
Management can define enterprise standards. Business units can own specific workflows. Technical teams can enforce permissions and system boundaries. Security, legal, compliance, risk, and domain experts can challenge higher-risk deployments. Operating teams can supervise everyday performance.
NYDFS’s insurance guidance reflects a similar structure by assigning overall governance responsibilities to boards and senior management while allowing implementation to be delegated through appropriate control functions.
This keeps leadership focused on the questions that matter.
Seven Questions CEOs Should Ask About Every High-Authority Agent
Executives do not need to understand every technical detail of a large language model.
They do need to understand the authority their systems possess.
For every important agent, leadership should know what the agent can do without asking anyone, what the most damaging action it could technically perform is, which employee owns the outcome, which actions require approval, how activity can be reviewed later, how quickly the agent can be stopped, and what evidence would justify giving it more authority.
If leadership cannot answer those questions, the organization is probably not ready to scale the system.
What New York Companies Should Allow AI Agents to Do Today
For many businesses, the safest near-term strategy is not particularly complicated.
AI agents can receive broad freedom to search, organize, calculate, compare, monitor, summarize, prepare, and recommend.
They can also receive selective authority to perform low-value, reversible, well-tested actions inside clearly defined limits.
Stronger approval requirements should appear before agents can commit significant money, modify important customer records, access production systems, make consequential external statements, or affect people’s employment, health, financial position, legal rights, or access to services.
The important point is that those boundaries should not remain static.
Companies should measure performance and gradually expand authority where evidence shows that additional autonomy is justified.
The Bigger Story: Human-in-the-Loop AI Is Becoming the Operating System for the Autonomous Enterprise
The first generation of enterprise generative AI was mostly about producing answers.
An employee asked a question, received a response, reviewed it, and decided what to do next.
The next generation is about completing work.
AI agents will search systems, update records, coordinate software, contact customers, monitor exceptions, initiate transactions, and hand tasks to other agents.
That changes the governance challenge.

Businesses no longer need to decide only what AI is allowed to say.
They need to decide what AI is allowed to do.
New York may confront this ques



