AI Agent Governance: Who Is Responsible When Autonomous AI Makes a Business Decision?

Understand AI agent governance, accountability and oversight as New York companies decide who is responsible when autonomous AI makes business decisions.

AI assistants changed the way employees research, write, analyze information, and solve everyday problems. AI agents are starting to change something much bigger: who actually performs business work and who makes business decisions.

That distinction matters because an AI assistant usually gives a person an answer and waits. An AI agent can go further by opening software, checking company data, choosing an action, updating a system, sending a message, approving a request, or triggering another workflow. Once AI moves from recommending an action to carrying it out, governance becomes much more important.

For New York companies, the central question is becoming surprisingly simple: when autonomous software makes a business decision, who is responsible for the outcome?

The practical answer is not that responsibility moves from the company to the AI model. It also does not automatically move to the software vendor, the developer who configured the system, or the employee who first started using it. In most serious business settings, the organization still needs a clear chain of responsibility around what the agent is allowed to do, who approved that authority, how the decision is monitored, and who can intervene when something goes wrong.

That principle is already visible in several New York regulatory frameworks. New York City employment rules place obligations on employers using covered automated employment tools, while New York State Department of Financial Services guidance makes clear that insurers remain responsible for outcomes even when outside vendors provide the technology. Together, these rules point toward a broader idea that every company experimenting with autonomous software should understand: autonomy can be delegated, but accountability still has to be assigned to people and organizations.

This article explains how that can work in practice. It also includes original NYC Tech Journal analysis using publicly available New York employment, wage, and AI-use data to estimate where agent governance could become especially important across the city’s economy.

This article is strategic business analysis rather than legal advice. Companies using autonomous systems in regulated, high-impact, or legally sensitive workflows should have qualified counsel review their specific obligations.

The Short Version: The AI Can Make the Decision, but the Business Still Has to Own It

One of the most dangerous sentences companies may begin hearing more often is, “The AI decided.”

At first, that sounds like an explanation. In reality, it usually hides several earlier decisions that were made by people inside the company.

Someone selected the AI system. Someone decided what information it could access. Someone determined which software tools it could control. Someone decided when human approval would be required and when it would not. Someone also decided whether the agent would remain in production after the business saw how it behaved in real situations.

Good governance makes those choices visible.

A strong operating rule for New York businesses is that every autonomous workflow should have a clearly named business owner, defined limits on authority, useful decision records, escalation rules, and an identified person who has enough authority to stop the system if necessary.

New York State’s insurance guidance already follows a similar approach. The Department of Financial Services has said that boards and senior management have responsibilities for AI governance and that insurers remain responsible for outcomes when they use outside AI systems. New York City’s automated employment decision rules also place obligations on employers that choose to use covered systems.

New York State's insurance guidance already follows a similar approach. The Department of Financial Services has said that boards and senior management have responsibilities for AI governance and that insurers remain responsible for outcomes when they use outside AI systems. New York City's automated employment decision rules also place obligations on employers that choose to use covered systems.

The lesson for the broader market is straightforward. Buying an AI agent does not turn a company into a passive observer of its decisions. Once the company gives software permission to act on its behalf, governance becomes part of ordinary business management.

Why AI Agents Create a Different Governance Problem

Generative AI first entered many companies as an assistant. Employees used it to summarize documents, draft emails, write reports, generate ideas, or answer questions. A person still stood between the model and the final business action.

Agents shorten that distance.

Consider a normal customer-service chatbot. It may suggest that a customer deserves a refund, but an employee still decides whether to issue it. An agent can potentially identify the customer, check order history, decide whether the refund meets policy, process the payment, update the CRM, notify the customer, and schedule follow-up without anyone approving each step.

The software is no longer only producing information. It is executing a workflow.

The Risk Is Bigger Than Getting an Answer Wrong

Traditional discussions about AI often focus on accuracy. Accuracy remains important, but agents introduce additional forms of risk because a system can use the wrong information correctly, apply a valid policy in the wrong situation, perform a reasonable action at the wrong time, or repeat a small error hundreds of times before a person notices.

The problem therefore cannot be reduced to whether the underlying model is smart enough.

Companies have to govern the complete decision system, including the data the agent receives, the tools it can use, the rules that limit it, the approvals around it, and the monitoring that happens after deployment.

One Incorrect Answer and One Incorrect Action Are Very Different

Imagine that an AI assistant incorrectly tells an employee that a customer qualifies for a discount. If the employee checks the account before making a change, the mistake may never affect the customer.

Now imagine an autonomous pricing agent makes the same mistake and immediately changes the account, applies the discount to related services, updates the billing system, and sends revised invoices. A single model error has now become a real operational event.

The closer AI gets to real-world execution, the more governance matters.

Original NYC Tech Journal Research: Why Agent Governance Matters So Much in New York

New York is an unusually important place to study this problem because a large share of the city’s economy is built around high-value information work, professional judgment, financial activity, healthcare, business services, and other workflows that could increasingly be handled by autonomous software.

To understand the size of that governance opportunity, NYC Tech Journal analyzed publicly available employment and wage data published by the New York City Comptroller using New York State Department of Labor figures. We also reviewed newer analysis of workplace AI usage in New York.

The purpose of this analysis was not to predict how many jobs AI will automate. Current public data cannot support a credible citywide prediction of that kind.

Instead, we asked a narrower question: how much of New York City’s employment and estimated payroll sits inside sectors where high-value decisions, information processing, and professional workflows could make AI agent governance especially important?

Our Research Methodology

We started with 2024 average employment and wage data published by the New York City Comptroller using Quarterly Census of Employment and Wages information from the New York State Department of Labor.

For each major industry, we estimated the annual wage bill by multiplying average employment by the reported average annual wage. This does not give an audited payroll figure, but it provides a useful way to compare the economic weight of different sectors.

We then compared the results with December 2025 seasonally adjusted employment figures to see whether the broader employment pattern remained stable. Finally, we examined the Comptroller’s 2026 analysis of work-related Claude usage in New York to understand where current AI use appears unusually concentrated.

Table 1: NYC Tech Journal Research Method

InputPublic sourceWhat we calculatedWhat it does not measure
2024 NYC employment by sectorNYC Comptroller / NYS Department of LaborShare of city employmentAI-agent adoption
2024 average wagesNYC Comptroller / NYS Department of LaborApproximate sector wage billCompany profit or revenue
December 2025 employmentNYC ComptrollerWhether sector concentration remained stableFuture employment
2026 workplace AI analysisNYC Comptroller / Anthropic dataWhere current AI usage is concentratedAll AI platforms or all workers
New York AI-related rulesNYC and NY State sourcesExisting governance signalsUniversal private-sector AI law

There are important limits to this approach. Employment is not the same thing as AI exposure, while payroll is not the same thing as AI risk. A high-paying industry does not automatically face greater danger from autonomous software.

The numbers are useful because they show how much economic value sits inside sectors where AI agents could eventually influence important workflows.

Finding #1: Three Office-Heavy Sectors Represent About 31% of NYC Jobs but Roughly 57% of Estimated Payroll

The Comptroller reported average 2024 employment of 483,264 people in financial activities, 216,299 in information, and 751,800 in professional and business services.

Together, those sectors employed about 1.45 million people.

That equals roughly 31.3% of the total city employment represented in the dataset. Their estimated share of payroll was much larger. Using reported employment and average wage figures, NYC Tech Journal calculates that the three sectors generated approximately 56.5% of the city’s estimated wage bill.

Chart 1: Office-Sector Employment Versus Estimated Payroll

SectorShare of NYC jobsShare of estimated wage billPayroll-to-job concentration
Financial activities10.4%27.3%2.62x
Information4.7%8.5%1.82x
Professional and business services16.2%20.7%1.27x
Combined31.3%56.5%1.80x

The payroll estimate was calculated by multiplying sector employment by average annual wages and comparing the result with the citywide estimate created using total employment and average wages.

The takeaway is not that these industries are automatically more exposed to AI. The more useful conclusion is that a relatively small group of New York industries contains an unusually large share of the city’s high-value professional work.

These sectors include finance, technology, accounting, consulting, advertising, legal services, corporate operations, and other businesses that rely heavily on information processing and judgment. As agents become capable of performing more of these workflows, relatively small amounts of software could begin influencing decisions connected to a very large amount of economic value.

Finding #2: Finance Has an Especially Large Governance Footprint

Financial activities represented roughly 10.4% of employment in the 2024 dataset but approximately 27.3% of our estimated wage bill.

That produces a payroll-to-employment concentration ratio of about 2.62x.

The ratio does not mean that financial AI is 2.62 times more dangerous than AI in other sectors. It means each percentage point of employment in finance represents considerably more wage value than the city average.

That matters because financial work also contains many decision-heavy processes.

An autonomous system can move quickly from summarizing information to monitoring risk, changing an exposure limit, preparing an investment recommendation, identifying suspicious activity, approving an exception, or initiating a customer communication. As that transition happens, governance becomes part of ordinary financial risk management rather than a separate discussion about AI ethics.

Finding #3: The Employment Pattern Remained Stable in Newer Data

The 2024 numbers could have represented a temporary pattern, so we compared them with later employment figures.

December 2025 seasonally adjusted employment data showed about 505,930 people working in financial activities, 227,570 in information, and 805,350 in professional and business services. Together, those sectors employed roughly 1.54 million people.

That represented about 31.6% of total NYC nonfarm employment.

The figure is remarkably close to the 31.3% share calculated from the 2024 annual dataset.

Chart 2: NYC Office-Sector Employment Concentration

MeasurementOffice-sector employmentShare of relevant NYC employment
2024 annual average1.451 million31.3% of total employment
December 20251.539 million31.6% of total employment
December 2025 compared with private employment1.539 million36.0% of private jobs

The broader conclusion is that New York continues to have an unusually large concentration of high-value office and professional work. That concentration creates a large potential market for autonomous enterprise systems and, at the same time, a strong reason to develop governance before those systems become deeply embedded.

Finding #4: Healthcare Makes the Governance Problem Much Larger Than Wall Street

The discussion around enterprise AI in New York often begins with Wall Street, but healthcare dramatically expands the scale of the problem.

Health and social assistance employed roughly 1.10 million people in December 2025. When that sector is combined with finance, information, and professional and business services, the group represents about 61.7% of private NYC employment.

Healthcare also changes the type of risk involved.

A marketing agent producing a weak advertisement can create wasted money or reputational problems. An AI system that affects patient scheduling, clinical routing, communication, eligibility, or safety can create far more serious consequences.

That is why governance should not depend mainly on the sophistication of the model. Companies need to focus on the consequence of the action the system is allowed to take.

Finding #5: Current New York AI Use Is Already Concentrated in Technical Work

The New York City Comptroller’s 2026 analysis of Anthropic Economic Index data found that New York trails only Washington, D.C., in Claude usage relative to population.

The occupational concentration is even more interesting.

About 43% of work-related Claude conversations associated with New Yorkers were connected with computer and mathematical occupations, compared with approximately 27% nationally. Education represented about 8.5%, while arts, design, entertainment, sports, and media represented roughly 8%.

The Comptroller also calculated an unusually high relative representation factor for computer and mathematical occupations.

The important point is that today’s workplace AI use is not evenly distributed across the New York economy. Technical teams appear to be among the earliest and most intensive users.

That pattern could create an important governance challenge as companies move toward agents. The people who are best able to build an autonomous system are not automatically the people who should decide what business authority that system receives.

The Central Governance Rule: Separate Technical Ownership From Decision Ownership

One of the most useful ways to govern AI agents is to separate technical responsibility from business authority.

Every important agent should have a technical owner and a business owner.

The technical owner is responsible for whether the system is functioning as designed, while the business owner is responsible for determining what the system should be allowed to do.

Those roles should not automatically belong to the same person.

A Developer Should Not Become the Accidental Credit Committee

Imagine that an engineering team builds an AI agent that reviews invoices and decides whether customers should receive extended payment terms.

The engineers may understand the software better than anyone else in the company, but that does not mean they should decide the organization’s credit policy.

Finance should determine the eligibility rules, spending limits, exception conditions, approval thresholds, and escalation paths. Engineering should then build those decisions into the system and ensure the controls work correctly.

Governance becomes much clearer once those responsibilities are separated.

A Software Vendor Should Not Become the Accidental HR Department

The same principle applies when a business buys an AI product instead of building one internally.

A hiring platform may rank applicants using an automated system, but the employer still decides whether to use that system, how much weight the ranking receives, how humans review the output, and what happens to applicants.

New York City’s Local Law 144 shows why this distinction matters. Covered employers and employment agencies cannot simply use certain automated employment decision tools without satisfying requirements related to bias audits, publication, and notices.

Purchasing software therefore does not remove responsibility from the buyer.

What Existing New York Rules Already Tell Us About Agent Governance

New York does not currently have one universal law that answers every governance question involving autonomous enterprise agents.

Businesses do not need to wait for one.

Several existing frameworks already reveal the direction in which governance is moving.

Employment Rules Show Why Companies Need AI Inventories

New York City’s automated employment decision rules require covered tools to undergo a bias audit before certain uses, while employers must also publish specified information and provide required notices.

The larger operational lesson is that companies need to know which automated systems are being used in consequential workflows.

That sounds obvious, but it becomes difficult once employees can build agents quickly using APIs, low-code tools, and AI platforms.

A business cannot assess whether a system needs an audit, notice, legal review, or additional control if management does not know the system exists.

Governance therefore begins with inventory.

Insurance Rules Provide a Useful Organizational Model

New York State Department of Financial Services guidance for insurers provides an especially useful example of how responsibility can be structured.

The framework places oversight responsibilities with boards and senior management, expects organizations to create written policies, assign responsibilities, provide training, monitor systems, and maintain inventories of AI systems.

It also makes an important point about outside technology providers. Insurers remain responsible for the outcomes of tools they use even when those tools come from third parties.

That principle can be adapted beyond insurance.

Companies should not spend most of their time asking which person should be blamed after an AI failure. They should divide responsibility clearly before deployment so that each part of the system has an identifiable owner.

Chart 3: The Responsibility Stack for an Autonomous Decision

LayerMain questionTypical owner
Board or governing bodyWhat level of AI risk will the company accept?Board or board committee
Executive managementAre governance controls operating effectively?CEO, COO, CRO, CIO, or designated executive
Business ownerShould the agent be allowed to make this type of decision?Head of the relevant function
Legal, risk, and complianceWhich rules and limits apply?Legal, compliance, and risk teams
Technical ownerIs the system working as intended?Engineering, data, or IT
Security ownerCan the agent be manipulated, compromised, or abused?Security team or CISO
VendorDoes the product meet agreed requirements?Technology supplier
Human reviewerShould this particular high-risk action proceed?Authorized employee

The titles will vary from company to company, but the structure should remain clear.

Every material agent should sit inside an identifiable chain of authority.

Why Saying “A Human Is in the Loop” Is Not Enough

Companies often describe their AI safety plan by saying that a human remains involved.

That statement is too vague to be useful unless the organization explains what the person actually does.

Does a human review every decision or only unusual cases? Does the reviewer have enough time to understand the decision? Can the reviewer actually reject the AI recommendation? Does the employee receive enough information to challenge the system? Is the override recorded and analyzed later?

Does a human review every decision or only unusual cases? Does the reviewer have enough time to understand the decision? Can the reviewer actually reject the AI recommendation? Does the employee receive enough information to challenge the system? Is the override recorded and analyzed later?

A person who automatically clicks an approval button is technically involved, but that does not mean meaningful oversight exists.

Human Review Should Be Designed as an Operational Control

New York State public-sector rules provide an interesting signal here because they refer to meaningful human review in certain automated decision settings. Those requirements apply to government use rather than serving as a general private-sector rule, but the concept is useful for businesses.

Meaningful review should change the outcome when necessary.

That means reviewers need authority, context, time, and a clear process for escalating problems.

The company should also measure how often human reviewers disagree with the agent. If employees regularly reverse the same type of decision, the problem may lie in the workflow rather than with the people performing oversight.

Build an Autonomy Ladder Before Giving the Agent Full Control

One of the simplest governance improvements is to stop thinking about autonomy as an all-or-nothing choice.

An AI system can operate at several levels of authority.

Chart 4: A Practical AI Agent Autonomy Ladder

LevelWhat the AI can doExampleSuggested governance
0Provide informationSummarize a contractStandard access controls
1Make recommendationsSuggest whether a refund is appropriateHuman makes final decision
2Prepare an actionPrepare refund and customer messageHuman approves execution
3Execute within narrow limitsIssue refunds below $50Automatic execution with monitoring
4Operate conditionallyComplete larger workflows when defined rules are satisfiedException monitoring and limits
5Make high-impact decisionsAffect employment, eligibility, credit, or insuranceStrong specialized governance and human oversight

Companies often jump from recommendation systems to highly autonomous workflows because the technology makes that possible.

A better question is whether the additional autonomy actually creates enough business value to justify the additional risk.

If an agent can eliminate 80% of manual work while remaining at Level 2, there may be little reason to immediately move to Level 5.

The Agent’s Authority Limit May Matter More Than Its Intelligence

Companies sometimes focus heavily on which model powers an agent.

In many workflows, the more important question is how much authority the agent receives.

Consider procurement.

An agent may be able to compare suppliers, negotiate terms, create a purchase order, and initiate payment. The most effective governance control may be much simpler than the AI system itself.

The company might decide that the agent can spend up to $5,000 without human approval but cannot exceed that amount.

The rule is easy to understand, test, audit, and enforce.

Every Agent Needs an Authority Envelope

An authority envelope defines the boundaries within which an autonomous system may operate.

A customer-service agent might be allowed to issue refunds of up to $50, provide replacement products worth less than $100, and offer one month of service credit.

The same agent might be prohibited from changing contract terms, removing fraud restrictions, closing certain types of accounts, or issuing large cash reimbursements.

The distinction is essential because an AI system may technically be capable of performing an action without being authorized to perform it.

Capability should never automatically become authority.

Who Is Responsible When an AI Agent Makes the Wrong Decision?

There is no universal legal answer because responsibility depends on the specific activity, contract, industry, jurisdiction, and facts surrounding the event.

From an operational perspective, however, companies can answer a more useful question in advance: who owns each category of failure?

Table 2: Responsibility by Failure Type

FailurePrimary operational ownerSecondary owner
Wrong business ruleBusiness functionCompliance
Unreliable model outputTechnical ownerVendor
Excessive agent permissionsIT or securityBusiness owner
Required approval bypassedWorkflow ownerEngineering
Vendor update changes behaviorVendor managementTechnical owner
Problematic employment outcomeEmployer governance chainVendor where relevant
Incorrect regulated decisionRegulated businessVendor may have contractual responsibility
Missing decision recordsTechnical operationsCompliance
Prompt injection causes actionSecurityTechnical owner
Employee misuseBusiness managementHR or security
Known failures continueExecutive or business ownerRisk management

The purpose of this table is not to create a blame chart.

It is to make sure that every significant failure mode has a person or function with enough authority to prevent, detect, and correct the problem.

Every Material Agent Should Have a Decision Charter

A useful governance program should require every important autonomous workflow to have a short decision charter.

The charter should explain what business outcome the agent is trying to achieve, what actions it can take, what data it may access, which decisions require human approval, what it is prohibited from doing, and who can shut it down.

The document does not need to become a giant compliance manual.

A short, well-maintained decision charter that people actually use is more valuable than a hundred-page document nobody reads.

The Charter Should Describe the Workflow, Not Just the Model

Companies often document the technical system while barely documenting the business decision.

They may record the model name, provider, version, benchmark score, and API configuration without clearly stating what the system is authorized to do.

The priority should be reversed.

If the agent moves money, the charter should explain financial limits. If it communicates with customers, the charter should define communication authority. If it changes system access, the charter should define access rules. If it affects a person in a high-impact setting, the document should explain human review and escalation.

Governance should follow the real-world action.

The Agent Decision Record Could Become a Critical Audit Tool

Traditional software logs are designed to show what a system did.

Agent governance increasingly requires organizations to reconstruct why an important action was allowed to happen.

That means decision records may need to contain more context than ordinary technical logs.

Table 3: A Practical Decision Record for Material AI Actions

FieldWhy it matters
Agent IDIdentifies which system acted
Agent versionShows the configuration operating at the time
Date and timeEstablishes sequence
Initiating person or systemShows how the workflow started
Decision typeClassifies the action
Information sourcesShows what data was considered
Material factsHelps reconstruct the decision
Policy or rule usedConnects action to company authority
Model outputPreserves relevant recommendation information
Tool usedShows what real system was changed
Financial or customer impactMeasures consequence
Human approval statusShows whether oversight occurred
Override statusRecords disagreement
Exception codeIdentifies unusual cases
Final outcomeConnects decision to result

Not every internal AI interaction requires this level of logging.

High-impact decisions are different. If the company may someday need to explain the action to a customer, regulator, auditor, executive, or court, a generic API request log will rarely be enough.

Explainability Should Be Treated as an Operational Requirement

Companies sometimes debate whether modern AI models can truly explain their reasoning.

For day-to-day governance, that philosophical debate matters less than the company’s ability to reconstruct an important decision.

Managers should be able to answer ordinary questions.

Why did this customer receive a refund? Why was this vendor selected? Why was this transaction stopped? Why did this application receive additional review? Which facts influenced the action? Which policy allowed the agent to proceed?

Why did this customer receive a refund? Why was this vendor selected? Why was this transaction stopped? Why did this application receive additional review? Which facts influenced the action? Which policy allowed the agent to proceed?

That level of reconstruction is often more useful than asking the model to generate a long explanation after the fact.

Credit Shows Why Decision Reconstruction Matters

Credit regulation already demonstrates why businesses need clear decision records. When adverse action occurs, creditors may have obligations related to specific reasons and notice.

The broader lesson extends beyond credit.

Any autonomous system making consequential decisions should produce enough reliable information for the organization to understand how the outcome occurred.

Companies should also make sure their governance programs rely on current rules rather than outdated AI guidance, headlines, or assumptions. AI policy is changing quickly, so legal and compliance teams need to verify which requirements are actually in force.

Governance Should Follow the Use Case, Not Just the Model

Suppose a company uses one large language model for four different agents.

The first summarizes meetings. The second writes marketing copy. The third processes small customer refunds. The fourth supports a regulated eligibility decision.

The same foundation model may power all four systems, but the governance requirements should be completely different.

Model Approval Alone Is Not Enough

A company may create a list of approved AI models.

That can be useful for procurement, security, privacy, and technical standardization, but it is not sufficient for business governance.

A model that is perfectly acceptable for summarizing internal notes may not be suitable for making an unreviewed decision affecting insurance pricing, hiring, or credit.

Organizations therefore need two separate approval questions.

First, is the model approved for use inside the organization?

Second, is this specific use of the model approved for this specific decision?

The second question is usually more important.

New York’s AI Governance Thinking Also Points Toward Impact-Based Control

New York City’s AI governance principles emphasize organizational processes that establish accountability, oversee AI systems, manage risk, and support responsible deployment.

The framework also pays attention to the effect automated systems can have on access to rights, services, safety, and benefits.

Although city agency frameworks do not automatically become private-sector requirements, private companies can borrow the underlying logic.

The level of governance should rise with the potential impact of the decision.

Build a Decision-Risk Classification System

A company does not need a senior committee to approve every AI summary, meeting note, or internal draft.

That would slow adoption without meaningfully reducing risk.

A better approach is to classify workflows based on consequence.

Table 4: A Four-Tier Agent Governance Model

TierDecision typeExamplesSuggested control
Tier 1Low impactSummaries, internal research, formattingNormal automated use
Tier 2Reversible operationsScheduling, routing, small refundsBounded autonomy and monitoring
Tier 3Material business actionsPurchasing, pricing exceptions, account restrictionsApproval thresholds and stronger logging
Tier 4High-impact decisionsHiring, termination, credit, insurance, health and safetySpecialized review and strong human control

This model allows low-risk adoption to move quickly while concentrating governance resources on decisions that can create meaningful harm.

That balance is important because governance that treats every AI use as equally risky usually becomes too slow to work.

Human Approval Should Follow Consequence

Some organizations randomly review a fixed percentage of AI decisions.

That can help measure quality, but it should not become the main control mechanism.

A $10 refund and a $500,000 supplier commitment should not receive the same probability of human review.

A stronger system uses event-based approval.

The agent may operate independently while consequences remain small, but specific conditions trigger human involvement.

Chart 5: Human Oversight Should Increase With Consequence

Potential consequenceTypical autonomyHuman role
Minimal and reversibleHighPeriodic audit
Low financial impactModerate to highReview exceptions
Material customer impactModerateApprove defined thresholds
Employment or eligibility impactLowMeaningful review
Safety or regulated rightsVery lowAuthorized human decision maker
Irreversible major actionMinimalDirect approval

A procurement agent could therefore operate independently until a transaction reaches a certain value.

A collections agent might work automatically unless litigation is involved.

A customer-service agent might resolve routine refunds but escalate suspicious activity or vulnerable-customer cases.

The question should always be whether the consequence is large enough to justify stopping automation.

Every Company Needs an AI Agent Inventory

As companies deploy more agents, one of the most important governance tools may also be one of the simplest: an inventory.

New York DFS guidance already expects insurers to maintain an up-to-date inventory of AI systems that are used, developed, or recently retired.

Other companies should consider adopting the same discipline even when they are not legally required to do so.

The reason is simple.

You cannot govern systems you do not know exist.

Shadow Agents Could Become the Next Shadow IT Problem

Employees can now build small autonomous workflows using no-code tools, AI platforms, scripts, and APIs.

A salesperson might create an agent that updates CRM records. A finance analyst might create one that follows up on unpaid invoices. An HR employee might connect an AI system to recruiting tools. An operations manager might create an automation that changes inventory levels.

Each experiment can look harmless in isolation.

Together, they can create an invisible network of systems that access business data and take actions without centralized oversight.

Companies should encourage experimentation, but they also need a practical way to discover when an experiment becomes an operational system.

Inventory Capabilities, Not Just Product Names

Listing the names of AI providers is not enough.

A useful agent inventory should record the business owner, technical owner, purpose, connected systems, data accessed, tools available, autonomy level, financial limits, human approval rules, regulatory relevance, deployment date, assessment history, and shutdown contact.

The inventory becomes a map of the company’s growing autonomous workforce.

Agent Permissions Should Resemble Employee Permissions

Most companies would not give a new employee unlimited access to bank accounts, HR systems, customer records, and production software.

Agents should follow the same principle.

Each system should receive the minimum amount of authority required to complete its workflow.

If an agent only needs to read invoices, it should not automatically receive permission to modify them. If it needs to prepare a payment but not release the funds, those permissions should be separated. If the workflow only needs a small number of customer fields, there is no reason to expose the entire customer record.

The more narrowly permissions are designed, the smaller the damage that can result from mistakes, manipulation, or compromised credentials.

Governance and Security Are Becoming the Same Conversation

Agent security cannot be separated cleanly from agent governance because a security problem can quickly become a decision problem.

Prompt injection provides a simple example.

A malicious instruction hidden inside a document might cause an AI system to ignore its normal instructions. If the system can only summarize documents, the result may be a bad summary.

If the same system can send payments, change permissions, modify customer records, or contact external parties, the same manipulation becomes much more serious.

High-Risk Actions Should Be Separated From Reasoning

One of the strongest architectural choices is to separate the agent’s ability to recommend an action from its ability to execute that action.

An AI system may prepare a bank transfer without having credentials that can release the money.

It may draft an account deletion request without possessing permission to delete the account.

It may recommend changing employee access without controlling the identity system.

This approach turns security into part of the authority structure.

Vendors Should Be Governed as Part of the Decision Chain

Traditional SaaS contracts were often written for relatively predictable software.

Agentic systems can change more quickly.

A vendor may switch the underlying model, change routing logic, alter system prompts, introduce new tools, or update how the agent behaves. Those changes can affect business risk even when the product name stays the same.

New York DFS guidance provides a useful pattern by emphasizing oversight of third-party AI systems and making clear that regulated businesses remain responsible for outcomes.

Other companies can adapt the same mindset.

Table 5: Contract Questions for AI Agent Vendors

Contract areaQuestion companies should answer
Model changesCan the vendor materially change the underlying model without notice?
Company dataWhat data can the vendor access or retain?
TrainingCan customer data be used to improve future models?
LogsCan the company retrieve decision records?
TestingWhat evaluation occurs before major updates?
IncidentsHow quickly must serious failures be reported?
SubprocessorsWhich third parties participate in the system?
SecurityHow are tool permissions and credentials protected?
AuditCan relevant controls be independently reviewed?
Regulatory supportMust the vendor cooperate with inquiries?
TerminationCan logs and data be exported safely?
RollbackCan a harmful model update be reversed?

Vendor review should reflect the importance of the workflow.

A content-writing tool and an automated underwriting system should not receive the same level of due diligence.

New York’s Frontier-Model Regulation Is Important, but It Does Not Replace Enterprise Governance

New York has also moved toward regulating powerful frontier AI systems through the RAISE framework.

Those rules focus on large model developers, transparency, safety frameworks, and serious incident reporting.

They are important for the broader AI market, but most companies should understand what frontier-model regulation does not solve.

A law governing advanced model developers does not decide whether your accounts-receivable agent should be allowed to suspend a customer’s service, whether your sales agent can approve a large discount, or whether your procurement agent can commit company funds.

Those decisions still belong to the organization deploying the agent.

NIST Provides a Useful Operating Framework

Companies looking for a practical starting point can also use the NIST AI Risk Management Framework.

The framework is built around four connected functions: Govern, Map, Measure, and Manage.

For AI agents, those ideas translate naturally into business operations.

Governance means deciding who owns the system, what authority it receives, and what policies apply.

Mapping means understanding the workflow, the data used, the people affected, the tools connected to the agent, and the possible consequences.

Measurement means testing reliability, monitoring overrides, tracking errors, reviewing business outcomes, and evaluating security.

Management means changing permissions, adding approval points, fixing problems, suspending systems, and retiring agents that no longer operate within acceptable limits.

This approach is more useful than starting with broad statements about responsible AI and hoping employees can translate them into controls.

The Most Important KPI Is Not How Often Employees Use AI

Many companies measure AI adoption by counting users, prompts, generated documents, or tokens.

Those metrics show whether technology is being used.

They do not show whether autonomous decisions are being governed effectively.

Agent governance needs operational metrics.

Many companies measure AI adoption by counting users, prompts, generated documents, or tokens.

Companies should understand how many decisions an agent processes, how many actions are executed without human approval, how often employees intervene, how many mistakes occur, how often decisions are reversed, and what financial or customer impact follows.

Table 6: A Practical Agent Governance Dashboard

MetricWhy management should care
Decisions processedShows operating scale
Autonomous execution rateMeasures actual autonomy
Human approval rateShows oversight burden
Human override rateReveals disagreement with the system
Error rateMeasures reliability
Reversal rateShows how many actions must be corrected
Exception rateReveals difficult cases
Customer complaintsMeasures external impact
Average decision valueAdds economic context
Maximum decision valueShows extreme exposure
Policy violationsReveals control failures
Security incidentsMeasures adversarial risk
Time to detectionShows monitoring effectiveness
Time to rollbackMeasures operational resilience

Boards and senior leaders do not need to read individual AI logs.

They do need enough information to know whether autonomous workflows are operating inside the company’s approved limits.

Override Rates Can Reveal Problems Earlier Than Traditional AI Metrics

Human overrides deserve special attention.

Suppose employees override an agent only 1% of the time for several months. The rate then rises to 12% immediately after a vendor update.

That change may reveal a problem before customer complaints or financial losses become obvious.

The company should therefore treat overrides as useful data rather than isolated employee actions.

Teams should record what was changed, why the person disagreed with the agent, and whether similar cases keep appearing.

If the same pattern repeats, the workflow should be redesigned.

Every Material Agent Needs a Real Kill Switch

Every important autonomous system should have a reliable way to stop acting.

This sounds simple, but modern agents may operate through scheduled jobs, multiple APIs, background queues, browser automations, and several connected applications.

Turning off the visible interface may not actually stop the system.

Companies should test the shutdown process before an emergency occurs.

A proper shutdown should stop new actions, revoke sensitive permissions where necessary, cancel scheduled activity, and preserve logs for investigation.

High-risk systems may also benefit from automatic circuit breakers.

For example, an agent could stop automatically if refund volume suddenly spikes, error rates exceed a threshold, suspicious payment activity appears, or an unusually large number of sensitive actions occurs in a short period.

Test Agents Like Business Processes, Not Chatbots

Chatbot testing often focuses on whether an answer is correct.

Agent testing should simulate complete business workflows.

A procurement agent should be tested against duplicate invoices, fake suppliers, missing approvals, altered bank details, conflicting instructions, and unusual contract conditions.

A customer-service agent should encounter fraud attempts, angry customers, vulnerable users, requests outside policy, and transactions that exceed its authority.

A hiring workflow should be tested against relevant fairness and accommodation issues.

The purpose is to discover where the complete system breaks, not merely whether the model can answer a question correctly.

Test Whether the Agent Knows When to Stop

One of the most valuable behaviors in a high-risk agent is refusing to continue.

Companies spend enormous effort training autonomous systems to complete tasks. They should spend equal effort defining situations where the agent must stop and escalate.

A good test set therefore includes cases in which the correct outcome is not a completed workflow.

The correct outcome is: do not execute this action; send it to a human.

Build an Agent Incident Process Before the First Serious Failure

Traditional incident response usually focuses on cyberattacks, system outages, and data breaches.

Autonomous systems introduce another category: the decision incident.

An agent might apply an incorrect refund rule for several hours even though the software itself never crashes. Technically, everything works as designed.

The business outcome is still wrong.

Companies need an incident process that can handle this type of event.

The investigation should determine which decisions were affected, how many customers or transactions were involved, what information the agent used, which system version was operating, whether the actions can be reversed, whether the agent stayed inside its formal authority, and whether the authority itself was designed too broadly.

The organization should also ask whether monitoring could have detected the problem earlier.

A disciplined post-incident review can improve agent governance more than another generic AI-awareness presentation.

Boards Should Govern Risk Appetite Rather Than Individual Prompts

Board involvement does not mean directors should review prompt templates or debate model parameters.

Their role should remain at the level of material business risk.

The board should understand where autonomous systems can create major financial, regulatory, customer, employment, safety, or reputational consequences.

Useful questions include where AI agents are already allowed to operate without direct human approval, how much money they can commit, which workflows affect employees or customers, which systems could create regulatory exposure, how quickly management can stop an agent, and whether serious exceptions are reaching senior leadership.

These questions keep governance focused on business responsibility rather than technical fashion.

The Biggest Governance Failure May Be Unclear Ownership

Imagine that an agent causes a serious business mistake.

Operations says engineering built the system. Engineering says the business team created the rules. The business team says the vendor recommended the configuration. The vendor says the customer selected the settings. Legal says nobody asked for review.

At that point, the company does not only have an AI problem.

It has an ownership problem.

Every material autonomous workflow should therefore have one clearly named executive or business leader who owns the business outcome.

That person does not personally become responsible for every technical failure or legal issue. The purpose is to ensure that someone has the authority to demand changes, reduce autonomy, escalate problems, or stop the system.

A Practical Responsibility Model for AI Agents

Traditional RACI charts can become overly complicated.

For agents, responsibilities should remain closely connected to the real business decision.

Table 7: Simple Agent Responsibility Model

ActivityBusiness ownerTechnical teamLegal / complianceSecurityVendor
Define business objectiveAccountableConsultedConsulted——
Set decision authorityAccountableImplementsConsultedConsulted—
Choose technical modelConsultedAccountableConsultedConsultedSupports
Approve sensitive data useConsultedResponsibleAccountableConsultedConsulted
Set financial limitsAccountableImplementsConsulted——
Test the workflowAccountableResponsibleConsultedSecurity testingSupports
Monitor performanceAccountableResponsibleReviews major exceptionsReviews security eventsSupports
Approve major changesAccountableResponsibleConsultedConsultedProvides notice
Stop deploymentAccountable executiveExecutes shutdownCan escalateCan isolate systemSupports

The objective is not bureaucracy.

The objective is making sure that every major question beginning with “who decides?” has a clear answer.

Governance Should Eventually Become Part of the Technical Platform

Written policies will not scale if companies eventually operate hundreds of agents.

The controls themselves need to move into infrastructure.

Financial limits can be enforced technically. Sensitive actions can require approval tokens. High-risk transactions can generate special logs. Certain data can be blocked from selected systems. Agents can be restricted to approved APIs, while dangerous workflows can require two-person authorization.

Governance becomes much stronger when unauthorized behavior is technically difficult or impossible rather than merely prohibited in a policy document.

A Better Architecture Is Often “Agent Proposes, System Enforces”

Companies should avoid relying entirely on a language model to remember every business rule.

A stronger architecture separates flexible AI reasoning from predictable enforcement.

The agent might decide that a customer deserves a $2,500 refund. A separate policy system can check whether the agent’s authority is limited to $500. If the limit is exceeded, the request automatically moves to a human reviewer.

The AI remains flexible enough to understand the situation.

The business rule remains predictable.

That combination will often be safer than placing every control inside a prompt and hoping the model follows it perfectly.

What New York Companies Should Do During the Next 90 Days

Companies do not need to build an enormous AI bureaucracy before experimenting with autonomous workflows.

They do need a basic control system.

A practical program can be built in three stages.

Table 8: A 90-Day AI Agent Governance Plan

PeriodMain objectiveCore deliverable
Days 1-30DiscoverAgent inventory and workflow map
Days 31-60ClassifyRisk tiers, owners, and authority limits
Days 61-90ControlLogging, approvals, monitoring, and incident process

Days 1-30: Find Every Agent That Can Act

Begin with systems connected to important company software.

Review CRM systems, finance tools, HR platforms, customer-service applications, procurement software, identity systems, internal automation tools, and AI platforms.

Ask employees what they have built.

Discovery should not feel punitive. If workers believe admitting an experiment will lead to punishment, shadow agents will remain hidden.

The company should focus first on systems that can actually perform actions. A chatbot that only answers questions should not receive the same attention as an agent that can change customer records or approve spending.

Days 31-60: Define Ownership and Limits

Assign a business owner and technical owner to every material agent.

Classify the workflow by risk level, define prohibited actions, establish financial and operational thresholds, and decide when legal, compliance, privacy, or security review is required.

This is also the right time to map applicable regulation.

Companies should not assume that one general AI policy covers hiring, credit, insurance, healthcare, or other specialized activities.

Days 61-90: Turn the Rules Into Controls

The final stage is operational.

Add useful logging, approval gates, permission restrictions, monitoring dashboards, incident procedures, and shutdown capabilities.

Then perform a serious simulation.

Assume that one high-impact agent made 10,000 incorrect decisions overnight.

Can the company identify the affected transactions? Can it stop the system immediately? Can management reconstruct what happened? Can it identify which version was responsible? Can the business reverse the actions?

If the answer is no, the governance system is not ready for high levels of autonomy.

Governance Must Be Fast Enough That Employees Will Actually Use It

Poorly designed governance can create its own risk.

If every small AI experiment requires months of reviews, multiple committees, and extensive paperwork, employees will either stop innovating or find ways around the process.

Companies should separate experimentation from production authority.

Employees should be able to test low-risk agents inside controlled environments relatively quickly. Stronger review should begin when the system receives access to production data, customer records, payments, HR decisions, or important business systems.

A useful principle is simple:

Experiment quickly, but grant real authority carefully.

Governance Should Become Stronger as Agents Gain More Context

An agent connected to one system has limited knowledge.

As companies integrate agents with CRM data, financial information, email, HR platforms, supplier systems, internal documents, contracts, and operational databases, those systems become more useful.

They also become more powerful.

An agent that understands customer history, pricing, inventory, contracts, payment behavior, and internal policy may be able to make excellent decisions.

It may also be capable of making much larger mistakes.

Data access and decision authority should therefore expand together rather than separately.

Multi-Agent Systems Make Accountability Even More Important

Governance becomes more complicated when one AI agent delegates work to another.

Imagine that a sales agent identifies an opportunity. A pricing agent determines a discount. A legal agent checks contract terms. A finance agent estimates profitability. Another system updates the CRM and prepares documents.

If something goes wrong, asking which agent made the decision may not produce a useful answer.

The stronger approach is to govern the complete workflow above the individual agents.

The orchestration layer should clearly define which agent can recommend, which can approve, and which can execute.

One agent should not quietly expand another agent’s authority.

A finance agent should not approve a legal commitment simply because another agent requested it.

This is the software equivalent of separation of duties.

The Bigger Story: AI Governance Is Becoming Organizational Design

The first phase of enterprise AI governance focused heavily on acceptable-use rules, model selection, privacy, and security.

That made sense when humans remained clearly responsible for turning AI output into action.

Agentic AI changes the problem.

Once software begins performing work rather than simply helping employees perform work, businesses must decide how authority will be divided between people and machines.

That is not only a technology question.

It is an organizational-design question.

Companies Are Creating a New Kind of Digital Worker

An AI agent may have a goal, permissions, responsibilities, memory, software tools, performance metrics, and the ability to complete work independently.

Those characteristics begin to resemble a job.

Traditional companies already understand how to govern human employees. Employees receive defined roles. Their access usually reflects their responsibilities. Spending limits exist. Certain decisions require approval. Managers review performance. Serious issues can be escalated.

Autonomous software needs a digital version of the same structure.

The organization should know what the agent’s job is, who manages it, what systems it can access, how much authority it has, which decisions it cannot make, how performance is measured, and how the system is removed when it is no longer trusted.

Original Finding: New York’s Governance Opportunity Is Bigger Than Current AI Adoption Numbers Suggest

The employment analysis reveals an important point.

New York does not need every worker to start using autonomous AI before agent governance becomes economically significant.

Financial activities, information, and professional and business services represent only about 31% of jobs in our 2024 analysis but approximately 57% of estimated wage value.

When education and health are added, the combined footprint rises to roughly 56.9% of employment and approximately 70.9% of estimated wage value.

That does not mean 70.9% of New York wages are directly exposed to AI agents.

It means a very large share of the city’s labor economy sits inside sectors where information processing, financial decisions, professional judgment, education, healthcare, and other high-consequence services matter enormously.

That makes New York a particularly important testing ground for enterprise agent governance.

The right time to build these controls is before autonomous workflows become invisible pieces of everyday operations.

The Final Principle: Never Let Autonomy Become Anonymous

The biggest governance failure may not come from a spectacularly intelligent AI making an extraordinary mistake.

It could come from something much more ordinary.

A small decision becomes automated because it saves employees time. The agent performs well enough that nobody worries about it. Permissions gradually expand. The vendor updates the model. Transaction volume grows. Human review becomes less frequent.

Months later, the company discovers that software has quietly been making thousands of material business decisions every week without anyone clearly owning the outcome.

Strong governance is designed to prevent that situation.

Every important autonomous decision should be connected to a visible chain of authority that includes the business rule, the technical system, the responsible business owner, the monitoring process, and ultimately the organization that chose to deploy the agent.

The software can perform the work.

It can gather information, reason through options, communicate with other systems, and execute approved actions.

Every important autonomous decision should be connected to a visible chain of authority that includes the business rule, the technical system, the responsible business owner, the monitoring process, and ultimately the organization that chose to deploy the agent.

The company still has to decide what the software is allowed to do.

For New York businesses moving from AI assistants toward autonomous operations, that distinction may become one of the most important management principles of the next decade.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top