For years, customer service automation had a bad reputation for a simple reason: customers could tell when a company was trying to keep them away from a human.
The old chatbot asked people to pick from a menu. The old phone system asked them to press 1, press 2, enter an account number, repeat the account number, explain the problem, and then explain the same problem again after finally reaching an employee.
AI agents are starting to change that model.
The important change is not that software can write more natural replies. A chatbot that sounds human but still cannot fix a problem is simply a more pleasant dead end. The bigger shift is toward AI systems that can understand what a customer wants, look at relevant company information, use tools, move through a workflow, and bring in a human when the issue becomes too risky or complex.
New York provides an unusually useful place to watch this change.
The city is home to major businesses in finance, insurance, healthcare, retail, travel, telecom, software, wellness, and commerce. These companies deal with huge customer volumes, expensive service operations, strict rules, demanding customers, and complicated internal systems. At the same time, New York has become home to a growing group of companies building the technology behind AI-powered customer service.
Our analysis suggests that the transition is already well underway, but it is happening in a more careful way than the phrase “autonomous customer service” might suggest.
In an original NYC Tech Journal review of public disclosures from ten New York-based or strongly New York-anchored customer-facing companies, all ten showed public evidence of production AI being used in customer experience or service. Nine showed evidence that AI was being connected to customer context or business systems. Eight disclosed some measurable result or level of usage.
But only two provided clear public evidence, under our strict definition, of a live AI system taking a meaningful customer action rather than mainly answering, recommending, routing, summarizing, or assisting a human.
That gap tells us where customer service is really going.
The next phase is not about putting a large language model in a chat window. It is about building a reliable service system that can move from answering to understanding, then from understanding to acting, while knowing when to stop.
The Short Version: Customer Service AI Is Moving From Conversation to Resolution
The first generation of customer service bots was built around information retrieval. Customers asked a narrow question, and the system tried to return a matching answer.
AI agents are built around a different goal: completing the job.
Consider a customer who says:
“My card was charged twice for the same order. I need the extra charge reversed.”
A basic chatbot might explain the company’s refund policy.
A better generative chatbot might explain the policy using friendly, natural language and correctly identify which part applies.
A true service agent would go further. With proper permissions, it could authenticate the customer, inspect the transactions, confirm whether there was a duplicate charge, determine whether the request meets the refund rules, start the permitted workflow, give the customer a reference number, update the CRM, and send the case to a human if something does not look right.
That is a major difference.
| Capability | Traditional chatbot | Generative chatbot | AI customer service agent |
| Understand natural language | Limited | Strong | Strong |
| Answer FAQ questions | Yes | Yes | Yes |
| Use conversation context | Limited | Usually | Usually |
| Read customer/account context | Rare | Sometimes | Often |
| Call business systems | Rare | Sometimes | Core capability |
| Complete multi-step workflows | No | Limited | Yes |
| Take approved actions | No | Limited | Yes |
| Decide when to escalate | Rule-based | Better | Context-aware |
| Work with human agents | Basic transfer | Possible | Designed into workflow |
| Main goal | Answer | Better answer | Resolution |
This distinction matters because customers rarely contact a business because they want a conversation. They contact it because something needs to happen.
They want the flight changed. They want to know where an order is. They want the duplicate payment fixed. They want an appointment moved. They want their benefits explained. They want access restored. They want the problem to disappear.
The winning customer service agent will therefore not be the one that talks most like a person.

It will be the one that safely gets the most useful work done.
Why the Human Layer Is Not Going Away
There is a temptation to frame AI customer service as a competition between software and people. Consumer data suggests that would be a mistake.
Verizon’s 2025 global customer-experience study surveyed 5,000 consumers and 500 senior executives. It found that 88% of consumers were satisfied with interactions handled mostly or fully by humans, compared with 60% for AI-driven interactions. Almost half of consumers, 47%, said the inability to reach a live person when needed was their biggest frustration with automated service.
That does not mean customers reject AI.
It means customers reject automation that traps them.
The important design question is therefore not, “Can AI handle this interaction?” A better question is, “At which parts of this interaction does AI create value, and at which point does a human create more value?”
Customer Service Has Three Very Different Kinds of Work
A large service operation usually contains work that falls into three broad groups.
The first group is repetitive and low-risk. A customer wants to know whether a payment arrived, how a membership works, when a shipment is expected, what a benefit covers, or how to change a simple setting.
These interactions are strong candidates for AI-led service.
The second group has a known workflow but carries more risk. A customer wants to modify an account, change a reservation, request a refund, dispute a fee, replace something, or make another change that affects money or access.
AI can often help here, but permissions, verification, limits, and approval rules become much more important.
The third group is messy. The customer may be angry. The policy might not clearly cover what happened. Fraud may be possible. A medical, financial, legal, or safety issue might be involved. The customer may simply need a person who can understand an unusual situation and make a judgment.
Trying to automate all three groups in the same way is one of the fastest ways to damage customer trust.
Original Research: The NYC Customer Service AI Evidence Map
To understand how far New York businesses have actually moved beyond chatbots, NYC Tech Journal built an original evidence map of ten customer-facing companies with strong New York ties.
This is not a survey of executives, and it is not intended to represent every company in New York. Instead, it is a structured review of what businesses themselves, regulatory filings, company reports, and named technology partners have publicly disclosed.
Our Methodology
We reviewed public information available through September 12, 2026, with a focus on company releases, investor materials, regulatory filings, official reports, product documentation, and named customer case studies.
The ten-company sample covers telecom, financial services, travel, wellness, healthcare, fintech, retail, commerce, and digital health. The companies reviewed were Verizon, American Express, JPMorgan Chase, JetBlue, ClassPass, Oscar Health, Rho, Macy’s, Whop, and Noom.
Each company was coded across six questions.
| Research variable | What counted as evidence |
| Production service AI | Public evidence AI is being used in a real customer-service or customer-experience workflow |
| Customer or system context | Public evidence AI can use customer, account, transaction, reservation, product, health, or other operating context |
| Human control | Explicit human handoff, human assistance, agent copilot, review, or similar control |
| Public result | A disclosed metric such as usage, containment, response time, cost, capacity, CSAT, or deflection |
| Voice/IVR | Live or publicly announced customer-service voice/IVR AI |
| Live action execution | Clear evidence the live system can do something meaningful for the customer beyond returning information or recommendations |
We applied the final category conservatively. For example, making a recommendation did not count as executing an action. Neither did an announcement saying that action-taking capabilities were planned.
“Not publicly disclosed” also does not mean a company lacks a capability. It means we did not count a capability unless we found clear public evidence for it.
Another limitation matters. Several performance numbers come from technology-vendor case studies. Those can provide useful operational detail, but they are marketing sources rather than independent audits. We therefore label them as reported results and avoid treating them as universally comparable.
Chart 1: What Our 10-Company New York Sample Publicly Shows
| Capability | Companies with clear public evidence | Share of sample | Visual |
| Production customer-service/CX AI | 10 of 10 | 100% | ██████████ |
| Customer or system context | 9 of 10 | 90% | █████████░ |
| Quantified result or usage | 8 of 10 | 80% | ████████░░ |
| Explicit human control/assist layer | 7 of 10 | 70% | ███████░░░ |
| Voice/IVR live or announced | 3 of 10 | 30% | ███░░░░░░░ |
| Explicit live action execution | 2 of 10 | 20% | ██░░░░░░░░ |
Source: NYC Tech Journal original analysis of company disclosures and named vendor case studies through September 12, 2026.
This is the most important result in our research.
AI in customer service is no longer an experiment among the companies we examined. Context-aware systems are also becoming normal.
But fully executable customer service remains much less common in public disclosures.
The Evidence Behind the Dataset
| Company | Publicly disclosed direction | Context/system use | Human layer | Public metric | Action evidence |
| Verizon | AI assistant, routing and agent support | Yes | Yes | Qualitative improvements publicly described | Yes, including account-related tasks |
| American Express | AI service tools and planned conversational agents | Yes | Yes | About 1M AI app-search inquiries monthly disclosed | Limited/transitioning |
| JPMorgan Chase | Production AI including customer-service and call-center efficiency | Yes | Yes | No directly comparable customer outcome used in our analysis | Not clearly disclosed |
| JetBlue | AI-powered digital service and virtual agents | Yes | Yes | 45% containment; 73,000 workforce hours reported in Q1 2023 | Not clearly disclosed |
| ClassPass | Generative chat/email service and agent assistance | Yes | Yes | 95% cost reduction reported by vendor | Not clearly disclosed |
| Oscar Health | Oswell and Superagent | Yes | Yes | 85% of member questions completed; 67% response-time reduction | Yes |
| Rho | AI customer-service copilot | Yes | Yes | 95% CSAT while handling 12% more contacts without added headcount, vendor-reported | Not clearly disclosed |
| Macy’s | Conversational shopping assistant | Yes | Human connection kept central | Higher engagement/conversion disclosed | Recommendation-focused |
| Whop | AI support | Yes | Human support exists | 65–70% deflection reported by vendor | Not clearly disclosed |
| Noom | AI customer support | Not sufficiently disclosed for our context test | Separate human support/coaching | Deflection rose from 59% to 68%, vendor-reported | Not clearly disclosed |
The company evidence comes from a mix of first-party disclosures and named customer case studies. Verizon has described AI-powered customer assistance and routing; American Express says its service teams are already using AI while it prepares to pilot conversational AI agents for its legacy IVR; JPMorgan Chase has identified customer service and call-center efficiency among its important production AI areas.
JetBlue’s technology partner ASAPP reports a 45% containment rate and 73,000 workforce hours saved in Q1 2023. ClassPass’s partner Decagon reports a 95% support-cost reduction while describing customer-context-driven chat and email automation. These are useful case data, but they remain vendor-reported figures and should be read that way.
Oscar provides unusually clear first-party evidence of action-oriented AI. Its 2025 impact report says Oswell uses medical records, plan information, and doctor visits to help with tasks such as prescription refills and test-result explanations. Oscar says Oswell completes 85% of member questions with high accuracy and quality, while its AI tools reduced Care Guide response times by 67% during open enrollment.
Macy’s provides another useful contrast. Its 2026 filings describe Ask Macy’s as an AI-powered conversational shopping assistant, say the system is expanding to help store colleagues, and note that AI and automation are being pushed across the customer journey while human connection remains central.
Finding #1: Context Is Becoming Normal Before Autonomy Is
Nine of the ten companies in our sample provided enough public evidence for us to code some form of customer or business context into their AI service experience.
That is important because context is the point at which AI starts becoming useful.
A generic model can explain a refund policy.
A context-aware model can understand which product a customer bought, when the purchase happened, whether it was delivered, which policy applies, what previous support interactions occurred, and what next step is allowed.
This is the foundation of agentic customer service.
Yet context alone does not create autonomy.
The hard part begins when a system receives permission to change something.
Finding #2: The Real Bottleneck Is Permission to Act
Only two of our ten companies met the strict test for clear public evidence of live action execution.
This is not surprising.
Answering a customer incorrectly is bad. Moving money incorrectly, canceling the wrong reservation, changing a health-related setting, modifying an account without valid permission, or granting an invalid refund can be much worse.
This is why the future customer-service stack will likely be built around graduated authority.
An AI agent might have full authority to answer a policy question, permission to execute a low-dollar refund below a set threshold, permission to prepare but not submit a larger refund, and no authority at all to resolve suspected fraud.
That is more practical than asking whether a company should “turn autonomy on.”
Finding #3: Human Escalation Is a Product Feature, Not a Failure
Seven of ten companies in our strict coding showed a clear human support, handoff, review, or copilot layer connected to the disclosed AI workflow.
That pattern aligns with consumer research.
Verizon’s study found that 47% of consumers named difficulty reaching a person as their top frustration with automated customer service. The lesson for New York companies is simple: a good escalation path should be designed before launch, not added later after customers complain.
A useful AI agent should know when confidence is falling. It should recognize when a customer is repeating themselves. It should identify high-risk topics. It should package the relevant history for the human employee rather than simply dumping the customer into a new queue.
The best handoff is not “I can’t help you. Contact support.”
It is: “I need a specialist to complete this safely. I have already passed them the order, payment, and steps we tried, so you will not need to start over.”
Do Not Use “Deflection” as the Only Measure of Success
One of the biggest problems we found while reviewing public AI customer-service results was that companies and vendors use different measures.
Some publish containment. Others publish deflection. Others publish questions completed, response-time reductions, cost changes, customer satisfaction, or workforce capacity.
Those numbers should not be blended into one artificial industry average.
Chart 2: Publicly Reported Automation Measures Are Not Apples to Apples
| Company | Reported measure | Reported result | What it means |
| JetBlue | Containment | 45% | Share of relevant interactions contained in the automated experience in the cited period |
| Noom | Deflection | 68% | Vendor-reported support deflection after deployment |
| Whop | Deflection | 65–70% | Vendor-reported range |
| Oscar Health | Questions completed | 85% | Oscar’s stated share of member questions completed by Oswell |
| ClassPass | Cost reduction | 95% | Vendor-reported support cost reduction |
JetBlue’s reported results include a 45% containment rate reached in May 2023, along with 280 seconds saved per conversation and 73,000 workforce hours saved during Q1 2023.
Noom’s named vendor case study reports that deflection moved from 59% to 68% within about 30 days. That is a nine-percentage-point gain, or approximately a 15.3% relative increase from the starting level.
Chart 3: Noom’s Reported Deflection Change
| Stage | Deflection | Visual |
| Before cited improvement | 59% | █████████████████████████████░░░░░░░░░░░ |
| After | 68% | ██████████████████████████████████░░░░░░░ |
| Change | +9 percentage points | +15.3% relative |
The problem with making deflection the main goal is that a company can improve the number while making the customer experience worse.
Imagine a bot that makes it deliberately difficult to reach a person. Its apparent containment rate might rise. Customer frustration may rise with it.
The better question is not, “How many customers did AI keep away from agents?”
It is, “How many customer problems did the company solve correctly, quickly, and at a sensible cost?”
The New Customer Service KPI Should Be Resolved Work
This changes the dashboard a company should build.
Average handle time still matters. Cost still matters. Deflection still matters. But each needs to sit beside measures that tell leadership whether the customer actually received a good outcome.
| KPI | What to measure | Why it matters |
| Resolution rate | Issues fully resolved | Measures actual work completed |
| First-contact resolution | Issues that do not require another contact | Exposes fake “containment” |
| Repeat contact rate | Same issue returns within 3–7 days | Catches hidden failures |
| CSAT | Customer satisfaction after interaction | Keeps efficiency tied to experience |
| Human escalation rate | Share transferred to people | Helps tune automation boundaries |
| Escalation quality | Whether context transfers correctly | Measures handoff quality |
| Incorrect-answer rate | Materially wrong responses | Core AI reliability metric |
| Action failure rate | Failed or incorrectly executed actions | Critical once agents can use tools |
| Policy violation rate | Responses/actions outside company rules | Measures control quality |
| Cost per resolved issue | Full service cost divided by resolutions | Better than cost per conversation |
| Time to resolution | Start-to-finish time | Reflects what customers feel |
| Human rework rate | AI work employees must redo | Finds hidden operational cost |
A company that saves 20% per conversation but creates 30% more repeat contacts has not improved customer service.

It has moved costs around.
Verizon Shows Why Agent Assist and Customer-Facing AI Are Converging
Verizon offers a useful picture of how the customer-service stack is changing because the company is applying AI on both sides of the conversation.
It has publicly described an AI-powered customer experience that can direct people toward the right expert and help employees work with relevant company knowledge. Verizon has also discussed AI for routing, sentiment, supervisor escalation, and personalized assistance.
That combination matters.
The old approach treated self-service and human service as separate channels.
The emerging model treats them as one system.
The Customer Should Not Care Who Did Each Step
Imagine a Verizon customer starts in an app because a bill is unexpectedly high.
An AI system can identify the account, understand the billing question, gather relevant plan and usage information, and explain the likely cause. If the customer needs a change the agent cannot safely complete, the case can move to an employee who already has the context.
From the customer’s point of view, the important measure is not whether AI “contained” the interaction.
It is whether the bill problem was solved without unnecessary effort.
American Express Is Showing Why Voice May Be the Next Major Interface
American Express said in its 2026 Chairman’s Letter that its customer-service teams are already using AI tools. It also said its AI-powered app search handles roughly one million inquiries per month and that it plans to pilot conversational AI agents to modernize its legacy interactive voice response experience.
Voice matters because many of the hardest customer interactions still happen by phone.
Customers call when something is urgent, confusing, emotional, expensive, or difficult to solve in an app.
Traditional IVR was designed to classify those callers.
Agentic voice AI is being designed to understand them.
Voice Agents Need More Than Speech
A realistic voice agent needs at least four layers.
First, it needs to understand what the person is saying, including interruptions, incomplete sentences, accents, background noise, and corrections.
Second, it needs context. Knowing that someone said “the charge yesterday” is not useful unless the system can safely identify the relevant account and transaction.
Third, it needs tools. If all it can do is speak, the customer will eventually hit the same wall as an old IVR.
Fourth, it needs a safe human transition.
That final layer becomes especially important in financial services, where authentication, fraud, disputes, payments, credit, and account controls create consequences far beyond a normal FAQ.
JetBlue Shows Why Customer Service Agents Need Operational Context
Airlines are a brutal test for customer-service automation.
A customer may start with a simple question about a departure time. Minutes later, a weather event can change the flight, gate, connection, bag plan, seat availability, rebooking options, and travel plans for thousands of people.
Static FAQs are poorly suited to this environment.
JetBlue’s partnership with ASAPP provides a useful long-running example. ASAPP reports that JetBlue reached a 45% containment rate, increased digital adoption fivefold, saved about 280 seconds per conversation, and saved 73,000 workforce hours in Q1 2023.
More recent discussions around JetBlue’s customer experience have increasingly focused on AI systems working alongside employees rather than simply answering basic digital questions.
Disruption Is Where Agents Become Valuable
Suppose a New York traveler is flying from JFK to Los Angeles and a cancellation occurs.
A simple bot can tell the traveler the flight is canceled.
An agentic system can potentially inspect the reservation, loyalty status, available inventory, connecting travel, company rules, and alternative flights. It can then present valid options and, if authorized, carry out the selected change.
That jump from explaining what happened to helping complete the recovery is the core of agentic service.
ClassPass Shows Why Personal Context Changes the Economics
ClassPass is another important New York example because customer-service questions frequently depend on account history.
A member may ask about a booking, missed class, cancellation rule, credit balance, studio policy, membership status, or historical reservation. A generic FAQ is often not enough.
Decagon’s ClassPass case study says the AI uses member account information and historical reservations to provide personalized support across chat and email. The vendor reports a 95% cost reduction and says launch-time deflection was far above ClassPass’s original expectations.
ClassPass has also used AI assistance for human employees.
This is important because the economics do not need to come entirely from replacing customer interactions. If AI shortens the time employees spend searching for history, policies, or the right response, the same human team can work on harder cases.
The Hidden Benefit Is Service Capacity
Many executives begin an AI project by calculating how many tickets can be automated.
That is understandable but incomplete.
There is another source of value: the difficult cases humans can now spend more time resolving.
If AI removes basic questions but leaves employees with the exact same targets and workflows, the business may waste that benefit.
A smarter company redesigns human work at the same time.
Oscar Health Shows What Happens When Customer Service Becomes Actionable
Oscar Health may be the clearest case in our sample of the boundary between customer service and an actual personal agent.
Oscar’s impact report says Oswell uses information including medical records, plan details, and doctor visits. The company says it can help manage prescription refills, review symptoms, explain test results, and connect members to virtual care. Oscar also says Oswell completes 85% of member questions with high accuracy and quality, while its AI tools reduced Care Guide response times by 67% during open enrollment.
This is a much deeper model than a health-insurance FAQ chatbot.
The system has context, specialized information, and a path toward action.
Regulated Industries May Need More AI Controls, Not Less AI
Healthcare also demonstrates why autonomy cannot simply mean “let the model decide.”
The higher the stakes, the more important it becomes to define exactly what the agent is allowed to read, say, recommend, update, and initiate.
This suggests an important principle for New York’s financial, insurance, healthcare, and legal businesses:
Agentic AI should increase operational control, not reduce it.
A well-built workflow can make permissions explicit. It can record which information the agent used, which tool it called, what happened, and when the interaction moved to a person.
That audit trail may eventually be one of the strongest arguments for agentic systems in regulated environments.
Rho Shows Why the First Agent May Sit Beside the Employee
Not every business needs to start with autonomous self-service.
New York fintech Rho took a more controlled route using an AI copilot integrated with its support workflow. Maven AGI’s case study says Rho maintained a 95% customer satisfaction score while supporting 12% more monthly contacts without increasing headcount.
The important part of this model is that the employee remains the operating layer.
AI can gather context, suggest useful information, draft replies, summarize interactions, and reduce searching.
The employee remains responsible for the customer relationship and final judgment.
Copilots Can Be the Training Ground for Agents
This path has another benefit.
Before a business gives an AI permission to execute a workflow automatically, it can watch what the AI recommends while humans remain in control.
Teams can study where the system succeeds. They can find bad edge cases. They can improve knowledge sources. They can measure which proposed actions humans accept or reject.
In other words, agent assist can become a data-gathering stage for future autonomy.
Macy’s Shows That Customer Service and Shopping Are Starting to Merge
Retail creates a different type of service problem.
Customers do not always arrive with a complaint. They may need help choosing a product, comparing options, finding the correct size, building an outfit, or deciding what goes well with something they already own.
Macy’s says its Ask Macy’s AI-powered conversational shopping assistant is being expanded to support in-store colleagues as well as digital customers. The company has also said it is advancing AI and automation across the customer journey while keeping human connection central.
Earlier 2026 disclosures said people engaging with Ask Macy’s showed higher conversion rates. That relationship should not be read as proof that the assistant caused every increase, because people choosing to use an assistant may already be more engaged shoppers.

Still, the direction is strategically important.
Customer support, sales assistance, search, and product discovery are beginning to converge.
The Best Agent May Generate Revenue as Well as Reduce Cost
This changes the AI business case.
If a retail assistant helps shoppers find the right item faster, reduces returns caused by poor choices, answers product questions, and creates higher conversion, the system is not simply a customer-service expense.
It becomes part of revenue generation.
That is likely to matter greatly in New York’s retail, travel, hospitality, media, beauty, fashion, and financial-services sectors.
Whop Shows How Support Data Can Become Product Data
Support conversations contain information that many companies fail to use.
Customers tell a business where a product is confusing, where onboarding breaks, which policies are unclear, which payment flow creates problems, and which feature repeatedly fails.
Brooklyn-based commerce platform Whop’s Decagon case study is interesting for this reason. The vendor reports 65–70% support deflection and says insights from customer conversations helped identify product issues, including payment-related friction.
This points toward a broader role for service agents.
The customer-service agent should not only resolve issues. It should help the company understand why those issues keep happening.
Great AI Service Should Reduce Future Service Demand
Suppose 8,000 customers contact support in a month about the same checkout confusion.
Automating those 8,000 conversations is useful.
Finding the reason they contacted support and fixing the checkout problem is much more valuable.
The most mature customer-service AI systems will therefore connect service operations with product, engineering, marketing, finance, and operations.
The aim should not be to automate more tickets forever.
It should be to eliminate preventable tickets.
New York Is Also Building the Agent Infrastructure
The New York story is not limited to companies buying customer-service AI.
The city also has a significant group of technology companies building it.
Regal has built enterprise voice agents designed to work with CRM data and business workflows. The company says its system handles millions of conversations and can connect agents with tools and APIs so they can take actions rather than simply answer.
ASAPP, headquartered in New York, launched a multi-agent customer-service platform in 2026 designed around coordinated AI systems and human employees.
Parloa, which has a New York presence, is building voice-focused AI agents that businesses can connect with internal systems, test through simulations, and evaluate before deployment.
PolyAI also maintains a New York office and competes in enterprise voice AI.
This combination matters.
New York is not simply a market where companies are experimenting with customer service agents. It is becoming a place where large customers, regulated industries, experienced operators, AI vendors, investors, and enterprise buyers can all interact.
The Economics Are Particularly Powerful in New York
Customer service is labor-intensive.
The latest BLS New York metro occupation-specific figure we use for our simple capacity model is from May 2023, when the New York-Newark-Jersey City area had an estimated 139,430 customer service representative jobs, with a mean annual wage of $51,200. We use that older local occupation figure rather than pretending it is a 2026 wage estimate.
More recent BLS data show the overall New York metro mean wage across occupations reached $41.50 per hour in May 2025, highlighting the generally high cost base in the region.
Original Analysis: The 100-Agent Capacity Model
Using the 2023 customer-service representative mean annual wage simply as a transparent base, a 100-person service organization represents about $5.12 million in annual base wages before benefits, managers, software, facilities, recruiting, training, turnover, and other costs.
That does not mean AI creates $5.12 million in savings. It also does not mean a 20% capacity gain should produce 20 layoffs.
It gives leaders a way to understand how valuable time can be.
| Capacity equivalent | Wage-equivalent value using 100 × $51,200 | What it could mean operationally |
| 10% | $512,000 | More volume without equivalent hiring |
| 20% | $1,024,000 | More capacity for difficult cases |
| 30% | $1,536,000 | Meaningful redesign of service operations |
| 40% | $2,048,000 | Large change requiring workforce planning |
Source: NYC Tech Journal calculation using May 2023 BLS New York metro mean annual customer-service-representative wage. This is a capacity illustration, not a forecast of cash savings.
This distinction is critical.
If AI frees 20% of a team’s time, the business might reduce contractor spending. It might avoid future hiring. It might extend support hours. It might improve service levels. It might assign people to retention, VIP support, fraud cases, escalations, or revenue-generating conversations.
The correct business case depends on what the company does with the freed capacity.
Which Customer Service Workflows Should New York Businesses Automate First?
The best first use case is usually not the workflow with the greatest possible savings.
It is the workflow with the best combination of high volume, clear rules, reliable data, measurable outcomes, and limited downside when something goes wrong.
A company should rank workflows accordingly.
| Workflow | Automation fit | Why |
| Order/status questions | Very high | Clear data, high volume, low judgment |
| Appointment information | Very high | Structured workflow |
| Basic account questions | High | Strong if identity is verified |
| Product information | High | Good knowledge-retrieval use case |
| Billing explanation | High | Valuable if transaction context is reliable |
| Routine reservation changes | Medium-high | Requires tools and rule checks |
| Small refunds within set policy | Medium-high | Good with strict thresholds |
| Subscription changes | Medium-high | Requires authenticated actions |
| Complex disputes | Medium-low | Higher judgment and financial risk |
| Fraud complaints | Low for full autonomy | High consequence |
| Medical safety questions | Low for unrestricted autonomy | High consequence |
| Threats, crises, vulnerable customers | Human-led | Needs judgment and care |
This is where many agent projects go wrong.
Leadership chooses the largest contact category and says, “Automate this.”
The better approach is to break that category into individual intents.
“Billing” might include invoice copies, payment status, duplicate charge questions, payment failures, refunds, price disputes, fraud, and hardship.
Those are not one workflow.
They should not receive the same level of AI authority.
Build Around an Authority Ladder
One of the simplest ways to deploy AI agents safely is to create levels of authority.
Level 1: Read
The agent can read approved knowledge and provide information.
This is close to advanced self-service and carries relatively low operational risk.
Level 2: Read Customer Context
The agent can use authenticated account information to make its response specific.
Now it can say, for example, “Your order shipped yesterday” instead of explaining how shipping normally works.
Level 3: Recommend
The agent can calculate or identify the next valid action but cannot execute it.
A customer or employee must approve the next step.
Level 4: Prepare
The agent can fill forms, prepare updates, draft a refund, build a change request, or assemble the needed workflow.
A human approves execution.
Level 5: Execute Within Rules
The system can perform low-risk actions when every required condition is satisfied.
Examples might include rescheduling within a permitted window, issuing a small refund, updating an address after verification, or changing a low-risk preference.
Level 6: Escalate
The agent recognizes that it has reached its authority limit and sends the case to the correct human with full context.
This should not be treated as the lowest level.
Good escalation is one of the most important capabilities in the entire system.
A Better Agent Architecture: Answer → Understand → Act → Verify → Handoff
Businesses often spend too much time choosing the model and too little time designing the workflow.
A reliable service agent needs more than an LLM.
Step 1: Answer
The agent needs a controlled knowledge layer.
Policies, FAQs, product documentation, prices, service rules, account instructions, escalation policies, operating hours, and other materials need owners, dates, and version control.
If employees cannot agree which document contains the correct policy, AI will not magically fix the problem.
Step 2: Understand
The agent needs enough context to understand what is actually happening.
This may come from a CRM, order platform, reservation system, payment system, help desk, health platform, subscription database, product telemetry, or other operating system.
Access should be limited to what each workflow requires.
More data is not automatically better.
Step 3: Act
Tools turn a chatbot into an agent.
Each tool should perform a narrow operation with defined inputs, permissions, output, and limits.
Instead of giving an AI broad database access, create controlled functions such as “check_order_status,” “prepare_refund,” “change_appointment,” or “update_shipping_address.”
The smaller the permission surface, the easier the system is to test.
Step 4: Verify
Never assume a tool call succeeded because the model requested it.
The workflow should confirm the outcome.
If the agent issues a permitted refund, it should verify that the transaction system created the refund and then provide the real reference information to the customer.
This closes a dangerous gap between saying something happened and confirming that it happened.
Step 5: Handoff
Human transfer should preserve the conversation summary, authentication state where permitted, relevant records, tool results, detected intent, and reason for escalation.
The human should enter with a head start.
If a customer has to repeat the entire story, the company has failed to create a unified service experience.
A Practical 90-Day Customer Service Agent Plan
A New York company does not need to begin with a year-long transformation.
A disciplined team can learn a large amount in 90 days if the first scope is narrow.
Weeks 1–2: Build the Baseline
Do not start by buying a model.
Start by understanding the work.
Take at least several weeks of customer interactions and group them by intent. Measure volume, average handling time, repeat contacts, transfers, wait time, resolution rate, customer satisfaction, and estimated cost.
Find the ten largest reasons customers contact the company.
Then identify which problems are clear enough to automate.
Weeks 3–4: Fix the Knowledge Layer
Gather the content employees actually use.
Remove old documents. Resolve contradictory policies. Assign owners. Add effective dates. Define which source wins when documents disagree.
This work is boring.
It is also one of the most important parts of the project.
Weeks 5–6: Launch Low-Risk Answers
Start with informational tasks.
Do not give the agent broad action permissions yet.
Run the system against historical conversations before putting it in front of customers. Compare its replies with the answers that experienced employees believe are correct.
Weeks 7–8: Add Context and One Controlled Action
Connect the agent to the minimum customer data needed for the selected workflow.
Then add one narrow action.
A subscription company might allow a plan change under clearly defined conditions. A retailer might allow an order cancellation before fulfillment. A service business might allow appointment rescheduling.
Do not build ten tools because the demo looks impressive.
Prove one tool is safe.
Weeks 9–10: Attack the System Before Customers Do
Test strange requests.
Try incomplete sentences. Angry customers. Conflicting instructions. Requests to override policy. Fraud-like scenarios. Attempts to reveal another customer’s information. Requests that require a human.
Measure whether the agent refuses the correct things as carefully as whether it completes the correct things.
Weeks 11–12: Release Gradually
Start with a limited share of eligible traffic.
Compare AI-led interactions with a control group whenever possible.
Watch repeat contact, CSAT, resolution, transfer quality, action errors, customer complaints, and human rework.
Expand because the data supports expansion, not because management announced an “AI transformation.”
The 90-Day Operating Table
| Period | Main job | Exit test |
| Weeks 1–2 | Map contacts and economics | Top intents and baseline KPIs known |
| Weeks 3–4 | Clean knowledge | Trusted source set established |
| Weeks 5–6 | Test low-risk answers | Accuracy reaches agreed threshold |
| Weeks 7–8 | Add context + one action | Tool works reliably under rules |
| Weeks 9–10 | Red-team and simulate | Failure modes documented and controlled |
| Weeks 11–12 | Limited production rollout | Resolution improves without unacceptable risk |
Human Approval Should Be Designed Into the Workflow
The worst form of human oversight is a label that says, “A human is in the loop,” without defining what that human actually does.
Approval should be linked to the risk of the action.
| Action type | Suggested operating model |
| Explain public policy | AI can usually answer |
| Explain authenticated account information | AI with access controls |
| Small reversible change | AI can execute inside rules |
| Larger financial adjustment | AI prepares; human approves |
| Policy exception | Human decides |
| Fraud suspicion | Specialist takes over |
| Legal threat | Route to approved team |
| Health or safety emergency | Immediate appropriate escalation |
| High-value customer exception | Human judgment |
| Unknown situation | Fail safely to human |
The important word is reversible.

An action that can easily be undone can often tolerate greater automation than one that creates permanent or high-cost consequences.
Give the Agent a Budget, Not Unlimited Freedom
One practical governance tool is an action budget.
A retail agent might be permitted to issue up to a small fixed amount in store credit under specific circumstances. Anything above the limit requires an employee.
A travel agent might rebook within the same ticket class and price band but ask for approval if the change creates an additional cost.
A financial-service agent may be allowed to explain a transaction but never move money without stronger controls.
The exact numbers differ by business.
The principle is the same: autonomy should have boundaries that software can enforce.
Measure the Cost of Bad Automation
AI business cases often measure the savings created by successful automation while ignoring the expense of failures.
That creates a distorted picture.
Suppose an AI agent handles 100,000 interactions for less money than a human team. That looks excellent until 8,000 customers have to contact the business again, 2,000 require employee rework, and hundreds receive incorrect actions.
A more useful equation is:
Net AI service value = labor capacity created + revenue impact + faster resolution + avoided contacts − AI costs − rework − failure costs − escalation costs − risk costs
This forces leadership to evaluate the entire customer journey.
Why “Cost Per Resolved Contact” Is Better Than “Cost Per Contact”
Traditional service teams often look at cost per ticket or cost per conversation.
AI can make that metric misleading.
Suppose the old model required one $8 interaction to solve a problem.
An AI system reduces the first interaction to $2, but 30% of customers then need a second interaction costing another $8.
The apparent cost reduction is much larger than the real one.
Tracking cost per resolved issue makes gaming harder.
It also aligns the finance team, service team, and customer around the same outcome.
AI Agents Will Change Customer Service Jobs Before They Eliminate Customer Service
It is easy to predict that AI will remove service roles.
The more immediate shift is likely to be a change in what remains for people.
When repetitive information requests disappear, human queues become harder.
Employees may see a larger share of angry customers, unusual exceptions, fraud, complicated disputes, high-value accounts, policy gaps, and sensitive situations.
That creates a new workforce problem.
Human Agents Will Need More Judgment, Not Less
A company cannot remove easy work and keep training employees as if their jobs remain the same.
They will need stronger product knowledge, greater decision authority, more emotional skill, better investigation tools, and clearer escalation rules.
Pay structures may eventually need to change too.
If people are handling only the hardest 20–40% of cases, their role is no longer equivalent to traditional first-line support.
New Customer Service Roles Will Appear
The service organization is also likely to develop new operational specialties.
Someone needs to review failed AI interactions.
Someone needs to maintain policies.
Someone needs to study agent behavior.
Someone needs to test new actions.
Someone needs to connect support insights back to product teams.
Someone needs to investigate why the agent escalated 14% of one workflow but 3% of another.
In many organizations, customer operations may begin to look more like a software operation.
Small New York Businesses Should Not Copy Enterprise Architectures
A Manhattan law firm, Brooklyn ecommerce company, Queens medical practice, restaurant group, or 50-person SaaS company does not need to build the same infrastructure as a national bank.
For smaller companies, simplicity matters more.
The first goal should be to remove repetitive work that has a clear answer and a clear financial value.
Start With the Channel That Already Has Good Data
If 70% of support arrives through email, starting with an expensive voice-agent project makes little sense.
If customers already use web chat heavily, start there.
If phone calls dominate because the issue is urgent or complicated, voice may deserve attention earlier.
Follow customer behavior rather than AI hype.
Buy Before You Build—Unless Customer Service Is Your Product
Most businesses do not need to build their own foundation model.
They need good workflow design, clean data, system integrations, monitoring, evaluation, and controls.
A custom build becomes more attractive when the workflows are highly unusual, service itself is a major competitive advantage, proprietary context is extremely valuable, or the company already has a strong internal AI engineering team.
For everyone else, building every layer from scratch can turn a customer-service project into a research program.
How to Evaluate an AI Customer Service Vendor
Do not run a vendor demo using the vendor’s perfect example questions.
Use your own failures.
Give the system anonymized versions of real customer conversations, including cases that confused your best employees.
Then test the entire workflow.
| Question | Why it matters |
| What data can the agent access? | Defines privacy and context |
| How are permissions controlled? | Defines action risk |
| Can we limit individual tools? | Prevents excessive authority |
| How is knowledge updated? | Prevents stale answers |
| What happens when sources disagree? | Tests uncertainty |
| Can it cite internal sources to employees? | Improves trust and review |
| How does human transfer work? | Critical customer experience |
| What history reaches the human? | Prevents repetition |
| How are actions logged? | Needed for investigation |
| Can we replay failed conversations? | Needed for improvement |
| How are evaluations run? | Shows whether quality is measurable |
| Can we test changes before release? | Reduces production risk |
| How is customer data retained? | Important for privacy and compliance |
| What happens if a connected system fails? | Tests operational resilience |
| How quickly can the agent be disabled? | Essential emergency control |
A polished voice should rank low on the list.
A reliable action log should rank high.
Privacy Becomes Harder When Agents Become More Useful
The better an agent understands the customer, the more useful it can become.
That also increases the amount of sensitive information it may touch.
A generic FAQ bot might need no personal information.
A real account agent may need identity data, transaction history, order history, subscriptions, preferences, health information, payment information, or internal notes.
Data access therefore needs to follow the workflow.
An order-status agent needs order information.
It does not need the entire customer database.
Voice Agents Need Extra Legal and Operational Care
Businesses deploying voice agents should have legal counsel review the relevant federal, state, industry, recording, privacy, and consumer-contact rules for their particular use case.
For outbound calls, the FCC has confirmed that AI-generated voices fall within the TCPA’s rules governing artificial or prerecorded voices, making consent requirements particularly important.
New York’s current wiretapping law generally defines unlawful wiretapping around the absence of consent from either a sender or receiver, but businesses should not treat a short summary as legal advice. Rules can depend on the parties, locations, communication method, industry, and purpose.
The technical lesson is simpler.
If voice AI becomes a serious service channel, compliance cannot be added after launch.
It has to be part of the architecture.
The Biggest Mistake Is Automating a Broken Process
AI does not remove operational confusion.
It can multiply it.
If a company has three conflicting refund policies, employees may currently make different choices.
An AI agent can make the wrong choice much faster and at much larger scale.
Before automating a process, write down exactly how a skilled employee is expected to handle it.
Which system should they check?
Which policy controls?
What exceptions exist?
Which amounts can they approve?
When should they stop?
Who owns the escalation?
If the business cannot answer those questions, the workflow is not ready for autonomy.
The Second Biggest Mistake Is Making the Bot Too Hard to Escape
A company may see human escalation as a cost.
Customers see it as insurance.
They are more willing to use automation when they know there is a way out.
Verizon’s consumer research makes the risk clear: inability to reach a live human is one of the biggest frustrations consumers report with automated experiences.
A visible human path does not necessarily increase human volume.
If the AI works well, most customers will not need it.
But knowing it exists changes trust.
The Third Mistake Is Optimizing the Agent in Isolation
Customer service is connected to the rest of the business.
If customers repeatedly ask why a delivery date is wrong, improving the chatbot answer may not be the best solution.
Fixing the delivery estimate may be.
If thousands of customers ask how to cancel a subscription because the cancellation setting is difficult to find, the product team should fix the interface.
This is why AI service analytics may become almost as valuable as AI service automation.
The system can turn millions of customer conversations into a continuously updated map of where the business is creating friction.
What New York Business Leaders Should Do Now
The first job is not to “deploy agents.”
It is to decide what work an agent should own.
Choose one customer journey that has enough volume to matter but is clear enough to control. Map the information, systems, decisions, permissions, and exceptions involved from beginning to end.
Then define what success means before selecting technology.
If the current first-contact resolution rate is 62%, document it.
If customers contact the company twice on average for a particular issue, document it.
If a workflow costs $14 per resolution, document it.
If CSAT is 78%, document it.
Without a baseline, almost any AI result can be presented as a success.
Build the Human Path at the Same Time
Do not build the AI experience first and the escalation experience later.
Define them together.
Which cases automatically go to humans?
Which signals suggest frustration?
What information moves with the case?
Does the customer keep their place in line?
Can the employee see what the AI already tried?
Can the human reverse an AI action?
Can the employee report an incorrect AI recommendation in one click?
These details determine whether the system feels like one service operation or several disconnected ones.
Give One Executive Ownership of Resolution
Agentic customer service crosses technology, customer experience, operations, security, privacy, finance, and often legal.
That makes ownership easy to dilute.
Someone needs to own the final customer outcome.
Not model accuracy.
Not automation rate.
Not AI adoption.
Resolution.
Five Predictions for New York Customer Service Agents
1. Voice Will Become Much More Important
Chat was the easiest place to deploy generative AI.
Voice is where a great deal of high-value service still happens.
American Express’s stated move toward conversational AI in IVR and the rise of New York voice-agent companies such as Regal point toward a broader shift.
Voice will not win because speaking is futuristic.
It will win when customers can call, explain a complicated problem normally, and have the system actually resolve it.
2. Companies Will Stop Buying “One Bot”
The service problem is too broad for one unrestricted agent.
A better architecture may include specialized agents for billing, reservations, order status, fraud triage, technical support, retention, and other areas, coordinated through a routing layer.
This makes testing and permissions easier.
The billing agent does not need every capability of the fraud system.
The returns agent does not need access to payroll data.
Specialization creates useful boundaries.
3. QA Will Move From Sampling to Near-Continuous Review
Traditional contact centers often review a small fraction of conversations because human quality checks are expensive.
AI can evaluate far larger shares of interaction data.
The important step will be reviewing not just tone or script compliance but whether the AI used the right source, selected the correct workflow, called the right tool, respected its authority level, and achieved the intended outcome.
4. Customer Service Data Will Flow Back Into Product Teams
Service conversations are one of the richest sources of customer research inside a company.
AI makes that data easier to structure.
Companies will increasingly be able to ask:
What caused the most preventable support demand this week?
Which new feature created confusion?
Which policy generates the most anger?
Which failed payment reason is rising?
Which product description causes the most pre-purchase questions?
Which agent workflow has suddenly started escalating more cases?
The support organization will become a real-time sensing system for the business.
5. “AI or Human?” Will Become the Wrong Question
Customers do not wake up wanting an AI agent.
They usually do not wake up wanting a human agent either.
They want the problem fixed.
The mature operating model will therefore route each part of the job to whoever—or whatever—can complete it best.
AI might authenticate the customer, gather records, identify the issue, calculate options, and prepare the workflow.
A person might decide the exception.
AI might then execute the approved action, update every system, create the summary, and send confirmation.
That is neither “human customer service” nor “AI customer service.”
It is customer service redesigned around the work.
The Bigger Story: Customer Service Is Becoming Executable
Chatbots changed the front door of customer service.
AI agents could change what happens after the customer walks through it.
The evidence from New York already shows the early shape of that change. In our ten-company analysis, production AI was widespread, contextual integration was common, measurable operational results were frequently disclosed, and human involvement remained important. Yet clear public evidence of live, meaningful action execution was still much less common.
That gap is where the next stage of competition will happen.
The companies that win will not be those with the chatbot that writes the most human-sounding paragraph.
They will be the companies that connect language understanding to reliable knowledge, clean customer data, carefully controlled tools, measurable workflows, sensible human judgment, and strong operational controls.
For New York businesses, this creates a large opportunity.
The city combines high service costs, complicated industries, demanding customers, major enterprises, strong AI talent, and a growing customer-agent vendor ecosystem. Those conditions make it a natural place for agentic customer service to move from demonstration to daily operations.
But the goal should remain simple.
A customer arrives with a problem.
The system understands it.
The right work happens.
The customer does not need to explain everything twice.
A human appears when human judgment matters.

And the business learns enough from that interaction to make the next one better—or prevent it from being needed at all.
That is the real move beyond chatbots.
It is not about making customer service more automated.
It is about making customer service more capable.



