Artificial intelligence in New York hospitals is no longer a future idea. It is already reading medical images, listening to doctor-patient conversations, predicting which patients may become seriously ill, helping staff respond to insurance denials, finding people who may qualify for clinical trials, and quietly handling work that once took doctors and hospital employees hours.
But the real story is very different from the idea that AI is about to replace doctors.
The biggest change happening inside New York hospitals is much more practical. Hospitals are finding small parts of healthcare where machines can sort information, recognize patterns, create first drafts, or bring urgent cases to a human’s attention faster. Doctors, nurses, radiologists, and other specialists are still making the important decisions.
That distinction matters.
A radiology algorithm may flag a possible blood clot, but a radiologist still interprets the scan. An ambient AI tool may create a draft medical note, but a clinician reviews it. A prediction system may identify a patient with a high risk of delirium, but a clinical team decides what that patient’s symptoms mean and what should happen next.
New York is becoming one of the most useful places in the country to watch this change because the city has almost every type of healthcare organization in one market. It has huge academic systems such as NYU Langone Health, Mount Sinai, and NewYork-Presbyterian. It has specialized institutions such as Memorial Sloan Kettering Cancer Center and Hospital for Special Surgery. It has Montefiore Einstein serving the Bronx, Northwell Health operating across a huge regional network, and NYC Health + Hospitals running the country’s largest municipal healthcare system.
These organizations are not following one AI playbook. Some are building their own models. Others are buying AI platforms from outside companies. Some are doing both.
That gives us an unusually useful opportunity to see where hospital AI is actually working, where it is still being tested, and where expectations may be running ahead of the evidence.
Original Research: Building the NYC Hospital AI Deployment Map
To understand where AI adoption has actually reached meaningful clinical use, NYC Tech Journal reviewed publicly available information covering eight major hospital systems with substantial New York City care operations.
The systems included were NYU Langone Health, Mount Sinai Health System, Northwell Health, NewYork-Presbyterian, Montefiore Einstein, Memorial Sloan Kettering Cancer Center, Hospital for Special Surgery, and NYC Health + Hospitals.

We then classified public evidence across seven areas: imaging and radiology, clinical documentation, clinical prediction, patient access and communication, hospital operations, pathology and advanced diagnostics, and clinical trial or research workflow.
How We Classified AI Use
A hospital received an Operational/Active Rollout classification only when public evidence showed that a technology was already being used, implemented, deployed, or actively rolled out.
A Pilot classification was used when the system clearly described the technology as a pilot, Phase 1 project, limited test, or similar early deployment.
Research papers were not automatically treated as hospital deployments. This is important because New York academic medical centers publish a huge amount of AI research, but an algorithm working in a study does not mean it has become part of routine patient care.
We also counted each system only once within each category. A hospital running 15 imaging algorithms therefore did not receive 15 times the weight of a hospital running one. This reduces the advantage that organizations with larger public-relations operations would otherwise receive.
Finally, a blank category means we did not find strong enough public evidence for that category in the material reviewed. It does not mean the hospital has no AI activity there.
The dataset is a disclosure-based map of visible AI adoption as of August 31, 2026, rather than a ranking of which hospital has the “best AI.”
NYC Tech Journal Hospital AI Deployment Matrix
| Health system | Imaging | Clinical documentation | Prediction / decision support | Patient access / communication | Operations | Pathology / advanced diagnostics | Trial workflow |
| NYU Langone | Operational | Operational | Operational | — | — | — | — |
| Mount Sinai | Operational | Operational | Operational | Operational | Operational | — | — |
| Northwell Health | Operational | Active rollout | — | — | Operational | — | — |
| NewYork-Presbyterian | Operational | Operational | Operational | Operational | — | — | — |
| Montefiore Einstein | Operational | Operational | Operational | — | — | — | — |
| Memorial Sloan Kettering | Operational | Pilot | — | — | Operational | Operational | Operational |
| Hospital for Special Surgery | — | Operational | — | — | — | — | — |
| NYC Health + Hospitals | Operational | Pilot | — | Pilot | Operational | — | — |
The underlying evidence is substantial. NYU Langone says roughly 1,500 doctors now use ambient AI following a year-long pilot and large-scale rollout. Mount Sinai describes AI systems operating across clinical care, imaging, patient access, and claims work. Northwell has been deploying ambient documentation while already operating imaging AI. NewYork-Presbyterian says it has 120 AI initiatives underway across clinical and nonclinical areas. Montefiore has deployed ambient documentation and publicly describes AI-supported clinical and imaging workflows. MSK has documented clinical AI deployments as well as formal AI governance. HSS is implementing ambient AI at enterprise scale. NYC Health + Hospitals has documented AI use in radiology, revenue cycle management, and several pilots.
Chart 1: Where AI Is Most Visible Across the Eight-System Sample
| AI workflow | Systems with operational/active rollout | Pilot only | Total systems with visible activity | Share of sample with visible activity |
| Clinical documentation | 6 | 2 | 8 | 100% |
| Imaging and radiology | 7 | 0 | 7 | 87.5% |
| Clinical prediction / decision support | 4 | 0 | 4 | 50% |
| Operations / revenue / workforce | 4 | 0 | 4 | 50% |
| Patient access / communication | 2 | 1 | 3 | 37.5% |
| Pathology / advanced diagnostics | 1 | 0 | 1 | 12.5% |
| Clinical trial workflow | 1 | 0 | 1 | 12.5% |
This creates one of the most important findings in our research: clinical documentation has become the most widely visible AI use case in our New York hospital sample. Every system reviewed has either publicly implemented ambient clinical documentation or disclosed a pilot.
That does not mean documentation AI is more mature than radiology AI. Radiology has a much longer history of machine learning and a mature regulatory path. The FDA’s current AI-enabled device list includes a large number of radiology products that have gone through applicable premarket requirements for their intended uses.
Instead, the data suggests that generative AI has found its first truly horizontal hospital workflow.
Every hospital creates clinical notes. Every specialty has documentation. Almost every doctor feels some amount of administrative pressure. That makes documentation one of the few AI opportunities that can potentially spread from primary care to orthopedics, oncology, emergency medicine, surgery, and dozens of other areas without requiring an entirely different business problem in each department.
Radiology went deep first.
Ambient documentation is going wide.
Radiology Is Where Hospital AI Learned to Become Useful
If someone had walked into a New York hospital ten years ago and asked where artificial intelligence would first become part of normal medical work, radiology would have been one of the safest guesses.
Medical images create exactly the kind of problem machine learning can handle well. Hospitals produce huge numbers of X-rays, CT scans, MRIs, mammograms, ultrasounds, and other images. Some contain subtle patterns that are hard to spot. Some require urgent attention. Others simply need to move through a crowded work queue faster.
AI can help with all three.
AI Usually Does Not Replace the Radiologist
The most useful way to understand radiology AI is as an extra set of eyes combined with a traffic controller.
Imagine 100 scans waiting to be read. An AI system may analyze them and identify one with signs that could suggest a stroke, pulmonary embolism, brain bleed, or another urgent problem. That scan can then receive additional attention quickly rather than simply waiting in chronological order.
Other tools help measure structures, identify suspicious regions, improve image quality, or compare patterns with previous studies.
But there is an important reason hospitals should be careful about describing these tools as “AI radiologists.”
Northwell provides one of the best real-world examples.
Northwell’s Pulmonary Embolism Data Shows Both the Power and the Limit
A 2026 study looked at the clinical implementation of an FDA-cleared AI tool used to detect pulmonary embolism, or blood clots in the lungs, across a Northwell network.
Researchers examined 32,501 CT pulmonary angiography exams from 29,492 patients. The overall agreement between the AI system and radiologists was 97.79%.
At first glance, that sounds like an argument for turning more interpretation over to machines.
The more interesting finding says the opposite.
When the AI and radiologist disagreed, expert review favored the radiologist in almost 89% of those disagreements. Among confirmed pulmonary embolism cases, interpreting radiologists uniquely identified 483 cases, or about 15%, while the AI alone uniquely identified 26, or about 0.8%.
That is an extremely useful result for healthcare executives.
The correct question is not, “Is AI better than a radiologist?”
The better question is, “What combination of AI and radiologist catches problems faster and more reliably than either workflow would alone?”
The Best Radiology AI Strategy Is Usually Augmentation
Hospitals should therefore measure AI imaging systems in context.
Does the technology reduce the time before an urgent scan reaches the right specialist? Does it find clinically meaningful cases that might otherwise be delayed? Does it lower repetitive work? Does it reduce errors without creating too many false alarms? Does performance stay stable across hospitals, scanners, patient groups, and time?
A tool can have impressive overall accuracy and still struggle with specific kinds of cases.
Northwell’s pulmonary embolism study found that agreement changed depending on features such as whether the embolism was acute or chronic and where it was located. That is exactly why overall accuracy alone can be misleading.
Imaging AI Is Spreading Beyond Simple Detection
New York hospitals are also using AI to improve what happens before and after image interpretation.
NYU Langone has described AI being used to shorten scan times, identify subtle abnormalities, and improve imaging workflows, while emphasizing that tools are tested and validated before use. Montefiore’s cancer programs say AI is used to help identify and track patients at increased risk for lung cancer, while NewYork-Presbyterian describes imaging and cardiovascular AI as important parts of its broader digital strategy.
At Memorial Sloan Kettering, the problem becomes even more specialized. MSK has said AI is routinely used to improve radiation treatment planning, where software helps clinicians work with medical images as they decide precisely where radiation should be delivered.
The next generation of radiology AI will therefore be less about a single “detect disease” button. It will increasingly sit across the imaging workflow: improving acquisition, organizing worklists, identifying urgent cases, measuring anatomy, comparing scans, preparing reports, and supporting treatment planning.
Clinical Documentation Has Become the Breakout Generative AI Use Case
Radiology may be the most mature clinical AI category, but documentation is where generative AI is spreading with remarkable speed.
The reason is simple.
Doctors did not go into medicine because they wanted to spend large parts of their day typing.
Yet modern healthcare requires detailed documentation. Clinicians record symptoms, medical history, examination findings, diagnoses, decisions, medications, orders, follow-up plans, and many other details. Those notes are important for patient care, communication, billing, regulation, quality measurement, and legal records.
The problem is that documentation takes time.
Ambient AI tries to give some of that time back.
What Ambient Clinical AI Actually Does
An ambient system listens to the conversation between a clinician and patient, with appropriate disclosure or consent processes depending on the system and workflow.
It then uses speech recognition and generative AI to turn that conversation into a draft clinical note.
The doctor does not need to type every sentence manually. Instead, the job shifts toward reviewing, correcting, and approving the draft.
That difference sounds small until it is multiplied across hundreds or thousands of clinicians and millions of visits.
NYU Langone Has Moved Beyond the Small Pilot Stage
NYU Langone offers one of the clearest signs that ambient AI has crossed into large-scale use.
In summer 2026, NYU Langone Chief Clinical Officer Oren Cahlon said approximately 1,500 doctors were already using ambient AI. He said NYU had first piloted the technology for a year and then moved to a large-scale rollout.
That sequence matters.
Hospitals should not interpret the rapid growth of ambient AI as a reason to skip testing. NYU’s example suggests the opposite: pilot first, learn how the system behaves, understand the workflow, and then expand.
The technology becomes valuable when the hospital changes more than the note-writing tool. It has to fit inside the electronic health record, work with existing visits, create usable output, allow easy review, and avoid giving clinicians another separate application to manage.
Montefiore Started With About 150 Primary Care Doctors
Montefiore Einstein announced an initial DAX Copilot rollout beginning with roughly 150 primary care physicians, with plans to expand across additional clinicians and specialties.
The system also highlighted something that can easily get ignored when hospital leaders become excited about productivity: patients need to understand what is happening.
Montefiore said transparency and consent were important parts of implementation and planned to give patients information and obtain consent during check-in.
That is a strong operational lesson.
Patient trust should be designed into the workflow rather than added after complaints begin.
Northwell Is Treating Ambient AI as Enterprise Infrastructure
Northwell’s deployment shows how quickly documentation AI is moving from departmental experiment to infrastructure project.
The health system announced plans to deploy Abridge across its 28 hospitals, while Northwell serves more than three million patients annually and is undertaking a major Epic electronic health record transformation. By May 2026, Northwell’s chief digital officer said the rollout was expanding from an initial group of clinicians, with the broader effort expected to continue alongside the Epic transition.
That timing is not accidental.
Ambient AI becomes more valuable when it is deeply integrated with the system where clinicians already work. A brilliant transcription model that forces doctors to copy and paste information between applications can create new problems while supposedly fixing old ones.
HSS Shows Why Specialty-Specific AI Will Matter
Hospital for Special Surgery is another important case because orthopedic notes are not the same as primary care notes.
HSS announced an enterprise Abridge implementation for an organization caring for around 200,000 patients a year and performing more than 40,000 surgeries annually. The collaboration specifically discussed making documentation better suited to orthopedic care, including information about imaging, laterality, procedures, risks, and treatment options.
This points toward the next stage of ambient AI.
Generic note generation will increasingly become table stakes.
The competitive question will be whether AI understands how a cardiologist documents a visit differently from an orthopedic surgeon, oncologist, psychiatrist, emergency physician, or pediatrician.
Mount Sinai Is Taking the Assistant Beyond Transcription
Mount Sinai announced a system rollout of Microsoft Dragon Copilot, combining ambient listening, natural-language technology, generative AI, clinical documentation, information retrieval, and administrative assistance.
This points toward a much larger change.
The first generation of ambient AI answers, “Can the computer write my note?”
The next generation asks, “Once the computer understands the visit, what other work can it prepare?”
That might include identifying follow-up tasks, surfacing relevant information, preparing orders for review, creating patient instructions, drafting letters, or helping organize the next steps in care.
The note could eventually become only one output from a much wider clinical assistant.
Original Analysis: Ambient AI Has Reached Every System in Our Sample
Our eight-system sample produces an unusually strong result.
| Ambient documentation status | Number of systems | Share |
| Operational or active rollout | 6 | 75% |
| Publicly documented pilot | 2 | 25% |
| No public activity identified | 0 | 0% |
The two systems we classify as pilot-stage are Memorial Sloan Kettering and NYC Health + Hospitals rather than organizations with no ambient activity.
MSK’s published governance work described two ambient AI pilots. NYC Health + Hospitals has publicly discussed ambient AI pilots, and internal board material describes a Phase 1 ambient-listening initiative focused on reducing physician documentation burden.
This is not proof that every New York hospital is about to use the same product.

It is evidence of something more useful: ambient documentation has achieved remarkable problem-solution fit across very different hospital models.
It is appearing in public hospitals, academic medical centers, specialty institutions, regional networks, cancer centers, and Bronx primary care.
Few generative AI applications in healthcare can currently make the same claim.
Clinical Prediction May Create More Value Than It Gets Attention For
Ambient documentation gets headlines because doctors immediately understand the benefit.
Predictive AI is quieter.
These systems examine information already being generated by the hospital and try to identify patterns that suggest something important may happen next.
A patient may be at increased risk of deterioration. Another may be likely to return to the hospital. Someone may be developing delirium. Another patient may have a combination of symptoms, laboratory values, medications, and notes that deserves faster clinical attention.
This is where AI starts moving from administrative assistance toward direct clinical decision support.
Mount Sinai’s Delirium Program Is One of New York’s Strongest Real-World Examples
Mount Sinai’s delirium system stands out because the hospital has published actual post-deployment results instead of simply announcing that it has “an AI model.”
The model looks at structured electronic health record data along with information contained in clinicians’ notes. It identifies hospitalized patients with a high risk of delirium so trained staff can evaluate them.
The study involved more than 32,000 patients admitted to The Mount Sinai Hospital.
After deployment, the monthly delirium detection rate rose from 4.4% to 17.2%, roughly a fourfold increase. Mount Sinai also reported reductions in potentially inappropriate medication use among older adults.
This is what strong hospital AI evidence looks like.
The model has a specific job. It fits into an existing workflow. A human team receives the information. The hospital measures what changed after deployment.
The goal is not “use AI.”
The goal is “find patients who need attention earlier.”
NYU Is Using AI to Look Forward, Not Just Backward
NYU Langone says it is using AI for predictive work and electronic health record reminders, allowing systems to examine patient information and surface risks or steps clinicians might otherwise miss. NYU’s work on models such as NYUTron has also explored predicting hospital readmissions using the language contained in clinical notes.
That matters because medical records contain two kinds of valuable information.
There is structured information such as age, laboratory values, diagnoses, medications, and vital signs.
Then there is language.
A clinician’s note might contain subtle descriptions of weakness, confusion, living conditions, symptoms, family concerns, or changes over time. Modern language models can potentially use this unstructured information in ways older rule-based systems could not.
Prediction Creates an Alert-Fatigue Problem
There is also a major danger.
Hospitals already have alerts.
A lot of them.
If every AI model sends another notification, staff may start ignoring the entire system.
That means the most accurate prediction model is not always the most useful hospital product.
Imagine two systems. Model A detects slightly more high-risk patients but sends twice as many false alerts. Model B misses a few more cases but sends highly focused alerts at the exact point where a clinical team can take action.
Model B may create more real-world value.
Hospitals therefore need to measure what happens after an AI alert. Was it seen? Was it useful? Did it change care? How long did the response take? Was the alert correct? Did staff begin dismissing alerts after repeated false positives?
Without those answers, a hospital may have a technically impressive algorithm and an operationally useless product.
Pathology Shows What the Next Wave of Medical AI Could Look Like
Radiology receives more AI attention because medical imaging became digital earlier.
Pathology is now following a similar path.
As glass slides become high-resolution digital images, computer vision can examine enormous numbers of cells, identify patterns, quantify biomarkers, and help pathologists focus on important areas.
Memorial Sloan Kettering is particularly interesting here because cancer care creates a strong reason to connect pathology, imaging, genomics, clinical data, and treatment decisions.
MSK Has Taken Some Pathology AI Into Clinical Use
MSK has described DeepLIIF, an AI-supported pathology system, as deployed clinically. The institution has said the system was expected to support analysis across more than 100,000 whole-slide images per year. MSK has also developed other AI systems spanning surgery and cancer care.
At the same time, MSK continues to publish much more advanced pathology models that are still being researched and validated.
That distinction is crucial.
A research paper saying an AI model matched or exceeded expert performance does not mean that every MSK pathologist is now using it for patient decisions.
For healthcare buyers, investors, and journalists, this is one of the easiest mistakes to make when studying medical AI.
Research Performance Is Not Deployment Performance
A hospital model has to survive much more than a research test.
Real patients may be different from the training dataset. Equipment varies. Documentation changes. Disease prevalence changes. Clinicians use technology in unexpected ways. Software gets updated. Hospital systems merge. Data feeds break.
A model can be excellent in a controlled study and disappointing in production.
That is why evidence from thousands of real clinical cases, such as Northwell’s pulmonary embolism analysis or Mount Sinai’s delirium deployment, is particularly useful. Those studies tell us what happens after AI meets the messy reality of hospital care.
AI Is Moving Into the Hospital’s Back Office Too
Some of the highest-value hospital AI may never directly diagnose anyone.
Hospitals are enormously complicated businesses. They schedule appointments, manage beds, submit insurance claims, appeal denials, code visits, answer phone calls, coordinate employees, process patient messages, purchase supplies, identify open clinical trials, and manage huge amounts of paperwork.

AI can work on all of these problems.
Mount Sinai Is Using AI Against Insurance Denials
Mount Sinai has described an AI process that helps prepare appeals when insurers deny claims.
The system can search the patient’s electronic record for information related to the denial and help draft an appeal. Mount Sinai reported a 3 percentage-point improvement in the overturn rate, which it associated with about $5 million in revenue impact.
This is strategically important because it shows why hospitals may fund administrative AI faster than some clinical AI.
The return can be measured directly.
If a tool costs $1 million and helps recover several million dollars that would otherwise have been lost, the business case is easier to defend than an AI platform promising a vague improvement in “innovation.”
NYC Health + Hospitals Is Finding Similar Revenue Opportunities
NYC Health + Hospitals has reported AI embedded into revenue-cycle workflows to improve revenue capture and reduce claims denials. In a public board presentation, the system reported $1.7 million in realized payments from corrected claims and another $20 million identified opportunity.
This is another reminder that the future of hospital AI will not be decided only in radiology reading rooms or operating theaters.
Finance departments may become major AI buyers.
So will scheduling teams, call centers, coding groups, patient-access departments, human resources teams, and supply-chain organizations.
AI Is Starting to Change How Patients Enter the System
Mount Sinai has also described AI-supported patient-facing tools, including an AI symptom checker and a voice system called Ava that operates within scheduling phone lines. NewYork-Presbyterian has publicly described AI-assisted patient in-basket messaging, where the system helps clinicians draft responses for human review.
NYC Health + Hospitals has discussed pilots for AI-supported patient messaging as well.
This may become a very large category because hospitals receive enormous amounts of communication that does not necessarily require a doctor to start from a blank page.
The important phrase is start from a blank page.
Generative AI often creates the most immediate value when it produces a first draft that a trained person can review.
AI Can Also Help Hospitals Find the Right Patient for the Right Trial
Cancer centers face a different information problem.
A single patient might potentially qualify for one of many clinical trials, each with detailed eligibility rules. Staff may need to read medical histories, laboratory results, diagnoses, treatment histories, genetic findings, and trial criteria before determining whether the patient is worth screening further.
That takes time.
MSK has deployed an AI-supported clinical trial matching system through a collaboration with Triomics. MSK noted that manual prescreening can take as much as 45 minutes per patient in some workflows.
For a major cancer center with many trials and patients, even modest improvements can compound quickly.
The broader lesson goes beyond oncology.
Some hospital tasks are difficult not because a human lacks knowledge but because the relevant information is spread across thousands of documents and databases.
That is exactly the type of information-retrieval problem where well-designed AI can help.
Chart 2: The Most Important Numbers Behind New York Hospital AI
One challenge in healthcare AI is that announcements frequently contain adjectives instead of results.
Words such as “transformative,” “revolutionary,” and “groundbreaking” tell hospital buyers almost nothing.
So NYC Tech Journal separated some of the harder numbers we found from the broader marketing language.
| Organization | Publicly reported AI evidence | What the number actually tells us |
| NYU Langone | About 1,500 physicians using ambient AI after a year-long pilot | Documentation AI has reached large-scale clinician adoption inside at least one major NYC system. |
| Mount Sinai | Delirium detection increased from 4.4% to 17.2% in a study of more than 32,000 patients | A predictive model changed a measurable clinical process after deployment. |
| Mount Sinai | AI-assisted denial appeals associated with roughly $5 million of revenue impact | Back-office AI can have a directly measurable financial case. |
| Northwell | 32,501 CT pulmonary angiography exams analyzed; AI-radiologist agreement 97.79% | Large-scale clinical evidence shows both strong AI performance and the continued value of expert review. |
| NewYork-Presbyterian | 120 AI initiatives underway across clinical and nonclinical uses | AI management is becoming a portfolio problem rather than a single-project problem. |
| Montefiore | Initial ambient AI rollout covered about 150 primary care physicians | Hospitals can begin with a focused group before expanding into more specialties. |
| MSK | Governance program covered 26 AI models, 2 ambient pilots, and 33 nomograms in its reported first year | Large institutions increasingly need formal systems to register and monitor AI. |
| HSS | Enterprise ambient initiative supports an institution caring for roughly 200,000 patients annually | Specialty hospitals are adopting the same core technology but adapting it to specialty workflows. |
| NYC Health + Hospitals | Revenue-cycle AI reported $1.7 million realized and $20 million additional identified opportunity | Public-sector AI can be justified through operational and financial outcomes, not only clinical innovation. |
These figures should not be added together to create a hospital “AI score.”
The measurements are too different. One describes users, another clinical accuracy, another revenue, and another portfolio size.
That itself is a valuable finding.
Hospital AI still lacks a common measurement language.
The next stage of maturity will require health systems to move from counting AI projects toward consistently reporting outcomes.
NewYork-Presbyterian Shows Why AI Governance Is Becoming Infrastructure
One of the biggest changes in 2026 is not a new medical model.
It is the rise of software for managing other AI.
NewYork-Presbyterian says it has 120 AI initiatives underway, covering clinical and nonclinical use cases. In August 2026, it announced adoption of Signal 1’s AI Management System to help manage and monitor AI technologies across the institution.
The platform is intended to support standardized risk assessment, monitoring, bias detection, and auditing across predictive AI, generative systems, and emerging AI agents.
This sounds less exciting than a model detecting cancer.
It may ultimately be more important.
Hospitals Are Moving From AI Projects to AI Portfolios
Managing one AI model can be relatively simple.
Managing 100 is not.
Hospital leaders eventually need answers to basic questions.
Which systems are active? Who owns them? What patient data do they use? Which vendor built them? Which model version is running? What decisions do they influence? How was each one tested? Has performance changed? What happens when the vendor updates the model? Which patient groups were included in validation? How quickly can the system be turned off?
A spreadsheet may work when the hospital has five AI tools.
It becomes dangerous when there are dozens or hundreds.
MSK Is Building the Same Kind of Governance Muscle
MSK’s published experience offers another useful benchmark.
During the first year described in its responsible AI governance work, the institution registered and monitored 26 AI models, handled two ambient AI pilots, and reviewed 33 nomograms. MSK’s researchers argued that governance and quality assurance are possible at scale but require structured risk assessment and lifecycle management.
That last phrase matters: lifecycle management.
Hospitals cannot test AI once and assume it remains good forever.
Patients change. Clinical practices change. software changes. Models get updated. Data pipelines drift. New locations are added.
Good AI governance therefore looks less like approving a medical device once and more like continuous quality management.
Regulators Are Moving Toward the Same Lifecycle Idea
Federal regulation is also moving in this direction.
The FDA’s AI-enabled medical device list is designed to provide transparency around authorized AI-enabled devices and states that products listed have met applicable premarket requirements, including review related to their intended use, safety, and effectiveness.
But AI can change after launch in ways older medical equipment usually could not.
That creates a new problem.
In August 2025, the FDA issued final guidance covering predetermined change control plans for AI-enabled device software. The idea is that developers can describe planned modifications, how those changes will be developed and validated, and how their impact will be assessed while maintaining safety and effectiveness.
For hospital buyers, the strategic lesson is bigger than the regulation itself.
Do not evaluate only the model you are buying today.
Evaluate the process by which tomorrow’s version will be changed, tested, monitored, and communicated.
What New York Hospitals Should Measure Before Expanding an AI Tool
The New York systems moving furthest with AI appear to have one thing in common: the strongest examples focus on a defined workflow rather than “AI transformation” as an abstract goal.
Hospitals considering a new deployment should use the same discipline.
Start With the Problem, Not the Model
A hospital should be able to describe the problem in one clear sentence before buying technology.
“Doctors spend too much time writing notes” is a problem.
“Urgent CT scans are not always prioritized quickly enough” is a problem.
“Too many valid insurance claims are denied because supporting information is difficult to assemble” is a problem.
“We need generative AI” is not a problem.
This distinction protects hospitals from spending money simply because a technology is fashionable.
Establish a Baseline Before the Pilot
Suppose a hospital wants to test ambient documentation.
It should measure the current documentation workload before turning the AI on. That may include time spent finishing notes, same-day note completion, after-hours electronic health record use, clinician satisfaction, coding quality, and correction rates.
Without a baseline, an executive may hear that clinicians “love the AI” without knowing whether it saved five minutes a day or an hour.
The same rule applies to diagnostic AI.
Measure turnaround time, missed cases, false alerts, escalation speed, clinician workload, and patient outcomes before implementation.
Then compare like with like.
Measure Errors, Not Just Success
Generative AI creates an unusual measurement challenge because its output can look professional even when something is wrong.
A draft note may sound perfectly medical while including a symptom the patient never reported.
That means hospitals need error categories rather than one vague accuracy score.
They should know how often the model invents information, leaves out important information, puts facts in the wrong part of a note, confuses who said something, creates incorrect medication details, or produces language that requires substantial editing.
NYU Langone’s research into AI-generated patient messages offers a useful warning. In one study, AI-assisted drafts could achieve similar accuracy in some comparisons but were also longer and more likely to contain complex language. AI can therefore solve one communication problem while creating another. The output still needs thoughtful human review.
Test Across the Population New York Actually Serves
A model working well in a narrow dataset is not enough for New York City.
NYC hospitals serve patients across languages, ages, races, income levels, immigration backgrounds, insurance types, medical complexity, disabilities, and neighborhoods.
An ambient system must work with different accents and communication styles.
A clinical prediction model should be checked across relevant demographic and clinical groups.
A patient-message assistant should not write at a level that is too difficult for the people receiving it.
A scheduling voice system should not become less effective for patients who speak English differently from the people used to train it.
Diversity can make New York AI deployment harder.
It can also make New York one of the most valuable places in the world to learn how healthcare AI performs outside a narrow laboratory environment.
A Practical 90-Day Hospital AI Scorecard
Hospital executives do not need 100 metrics for every early deployment.
They need a small set connected to the problem being solved.
| Area | Example measure before launch | Measure during pilot | Expansion question |
| Clinical safety | Current error or missed-event rate | AI-assisted error, miss, and false-positive rates | Does the tool improve safety without creating a new serious failure mode? |
| Clinician workload | Minutes spent on task | Minutes saved after review and correction | Is the time saving large enough to matter? |
| Workflow adoption | Current process completion | Percentage of eligible staff actually using AI | Are clinicians voluntarily continuing to use it? |
| Output quality | Existing note, report, or task quality | Corrections required per AI output | Does quality hold as use increases? |
| Patient experience | Satisfaction and complaints | AI-related satisfaction, consent refusals, complaints | Do patients understand and accept the workflow? |
| Equity | Results by relevant patient group | Performance differences by language, age, sex, race or clinical subgroup where appropriate | Are gaps appearing that require redesign? |
| Financial result | Current cost / revenue loss | Cost saved, revenue recovered, or capacity created | Does the business case survive real operating costs? |
| Reliability | Current system downtime | Failures, latency, integration problems | Can the hospital depend on the technology during normal operations? |
The most important column is the final one.
A pilot should be designed around an expansion decision.

Hospitals should know in advance what evidence would make them scale the product, change it, continue testing, or stop.
What Health-Tech Companies Selling to New York Hospitals Should Learn
The New York market is sending a clear message to AI startups.
Hospitals do not need another impressive demo.
They need a solution that survives contact with real healthcare.
Integration Is Becoming More Important Than Raw Model Performance
Many hospital AI vendors can now access strong underlying language or computer-vision models.
That reduces the value of simply having “AI.”
The difficult work shifts into integration.
Can the product work inside Epic or another electronic health record? Can clinicians launch it without another login? Can outputs flow into the correct place? Can the hospital audit what happened? Can the system support existing identity, security, privacy, and access rules?
Northwell’s ambient rollout occurring alongside its Epic transformation is a good example of how tightly these decisions are becoming connected. Montefiore also emphasized integration with Epic in its ambient deployment.
The winning healthcare AI vendor may therefore not have the most dramatic demo.
It may have the least disruptive implementation.
Vendors Need Evidence From Real Workflows
“We achieved 95% accuracy” is increasingly insufficient.
Hospital buyers should ask 95% of what, measured where, against which reference standard, for which population, and with what consequences when the model is wrong.
The Northwell pulmonary embolism data is a perfect illustration. A 97.79% overall agreement number sounds excellent, yet looking closely at disagreement cases revealed important situations where human interpretation remained essential.
Startups should expect more buyers to demand this deeper evidence.
The Best Sales Pitch May Be a Smaller Promise
Healthcare AI companies often describe platforms capable of transforming an entire hospital.
That can make the procurement process harder.
A narrower promise can be stronger.
“We reduce the work required to appeal this class of denied claims.”
“We draft outpatient clinical notes for these specialties.”
“We move suspected pulmonary embolism scans higher in the radiologist worklist.”
“We identify possible trial candidates from these data sources.”
Specificity creates measurability.
Measurability creates trust.
And trust creates expansion.
The Human-in-the-Loop Model Is Not a Temporary Compromise
People sometimes describe human review as something healthcare AI needs only because today’s models are not good enough.
That misunderstands medicine.
Healthcare contains judgments that involve uncertainty, patient preferences, incomplete information, unusual cases, ethics, and responsibility.
A model may estimate risk. It does not know the entire life of the patient sitting in front of the doctor.
A documentation model can summarize a conversation. It cannot take professional responsibility for whether the medical record is correct.
An imaging algorithm may highlight an abnormality. It does not automatically understand every competing diagnosis and clinical context.
The strongest New York examples therefore do not remove experts from the workflow.
They change what the experts spend their time doing.
Northwell’s pulmonary embolism results show this clearly. Mount Sinai’s delirium model alerts a trained team rather than automatically diagnosing and treating the patient. Ambient AI creates draft documentation rather than silently signing the medical record.
That is likely to remain the dominant model for high-stakes healthcare AI.
The Bigger Opportunity Is to Redesign the Workflow Around AI
Simply placing AI on top of an old process can limit its value.
Imagine that an ambient tool saves a doctor five minutes creating a note, but the physician still spends ten minutes moving information manually between systems.
The hospital has automated one piece of a broken workflow.
The more strategic question is what the entire process should look like now that AI can understand some of the information flowing through it.
For a patient visit, the future workflow could begin before the doctor enters the room. AI might organize relevant history and recent results. During the conversation, ambient technology could capture information. Afterward, the system could draft documentation, identify possible follow-up tasks, prepare plain-language instructions, and surface actions for clinician approval.
The doctor would still control the clinical decisions.
But the computer would handle more of the information movement surrounding those decisions.
That is much more important than automated note-taking alone.
Why New York Could Become a Major Healthcare AI Test Market
New York has several structural advantages that make it unusually important for healthcare AI.
The first is scale.
NewYork-Presbyterian alone reports more than two million annual visits and more than 4,000 inpatient beds across its broader system. Northwell reports more than three million patients annually. NYU Langone says it records more than 12 million outpatient visits annually across its wider network.
The second advantage is specialization.
A startup can encounter world-class cancer care at MSK, orthopedic medicine at HSS, complex academic medicine at Mount Sinai, NYU, and NewYork-Presbyterian, major public-sector delivery at NYC Health + Hospitals, and community-focused care across the Bronx through Montefiore.
The third advantage is diversity.
If healthcare AI only works for a narrow slice of the population, New York is likely to expose the weakness quickly.
That can make deployment more difficult, but it can also create better products.
New York’s Greatest Advantage May Be Its Ability to Validate AI
The city should not try to win healthcare AI by simply having the largest number of startups or pilots.
A more defensible position would be becoming the place where medical AI proves that it actually works.
That means developing strong methods for real-world validation.
Can a model work across Manhattan and the Bronx? Across private and public hospitals? Across community clinics and academic centers? Across multiple languages? Across different imaging equipment? Across different patient populations?
A product that survives that test becomes much more valuable nationally.
The Next Battle Will Be Over AI Monitoring
One of the clearest signals from our research is that AI governance is becoming its own technology category.
NYP’s decision to adopt a dedicated AI management system is especially important because the hospital is no longer treating every AI application as an isolated technology purchase.
This is likely to spread.
Hospitals will need a live inventory of AI systems, their owners, risk categories, versions, data sources, validation evidence, performance results, and incidents.
They will also need to know when something changes.
A model that was safe at launch can become less useful if clinical practice shifts or the patient population changes. A vendor update may improve one capability and weaken another. An integration error can suddenly feed incomplete information into an otherwise strong model.
The FDA’s current approach to planned modifications in AI-enabled device software reflects the same broad truth: AI requires management across time, not just evaluation on purchase day.
What Hospitals Should Not Do With AI
New York’s current wave of deployments also makes several mistakes easier to see.
Hospitals should not buy an AI product simply because another famous medical center announced it.
Two organizations can use the same software and produce completely different results because their workflows, staff, patients, technology, and leadership are different.
Hospitals should also avoid treating clinician enthusiasm as the only success measure. A tool can feel useful but introduce subtle documentation errors. Another can technically save time while increasing downstream work for nurses, coders, or other departments.
Financial savings alone are not enough for clinical tools either.
A system that saves money while increasing patient risk has failed.
At the same time, hospitals should not demand impossible evidence before testing low-risk administrative tools. The level of governance should match the level of possible harm.
Drafting a scheduling message and predicting life-threatening clinical deterioration should not pass through identical approval processes.
The goal is not maximum bureaucracy.
It is proportional control.
What NYC Tech Journal Will Be Watching Next
The most important hospital AI story over the next several years will probably not be whether a computer can beat a doctor on one carefully designed benchmark.
Those demonstrations will continue.

The more important question is whether hospitals can build an operating model that allows hundreds of AI capabilities to work safely together.
Ambient AI Will Expand Beyond Notes
Documentation is the starting point because it solves an obvious problem.
But once an AI system can understand a clinical conversation, there is a natural path toward helping with the work created by that conversation.
Expect vendors to move toward follow-up preparation, order support, coding assistance, patient instructions, care-plan updates, prior authorization, referral workflows, and other tasks.
The hospital will need to decide where drafting ends and autonomous action begins.
That boundary will become one of healthcare AI’s biggest governance questions.
Radiology AI Will Become Less Visible
This may sound strange, but successful radiology AI could eventually attract less attention.
Technology becomes invisible when it becomes infrastructure.
Hospitals do not issue press releases every time a normal imaging reconstruction system runs.
As AI becomes embedded into scanners, worklists, image processing, triage, reporting, and treatment planning, clinicians may simply treat many algorithms as part of the imaging environment.
That would be a sign of maturity rather than declining importance.
More AI Projects Will Be Judged on Financial Results
Mount Sinai’s insurance-appeal results and NYC Health + Hospitals’ revenue-cycle numbers show why administrative AI will attract serious investment.
Hospitals operate under intense financial pressure.
Products that reduce denials, automate repetitive work, improve scheduling, free clinical capacity, or decrease staff burden can create a business case without needing to claim they have reinvented medicine.
This is one reason some of the biggest healthcare AI companies may eventually be built around workflows that patients barely notice.
Hospitals Will Begin Retiring AI Too
A mature AI strategy is not just about adding tools.
It is also about removing them.
As portfolios grow, hospitals will discover overlapping products, weak models, unused software, outdated experiments, and tools whose benefits do not justify their costs.
The ability to shut down an AI project will therefore become just as important as the ability to launch one.
That is healthy.
Healthcare does not need the maximum possible amount of artificial intelligence.
It needs the right amount in the right places.
The Bottom Line
New York hospitals are already showing what practical healthcare AI looks like.
It does not look like a robot replacing a physician.
It looks like AI moving an urgent scan higher in a radiology worklist. It looks like a doctor speaking naturally with a patient while software prepares the first draft of the note. It looks like a model warning a clinical team that a hospitalized patient may be developing delirium. It looks like software finding documentation needed for an insurance appeal. It looks like a cancer center using algorithms to analyze pathology slides or identify possible clinical-trial candidates.
Our original review of eight major New York hospital systems found visible imaging AI activity at seven, while all eight have either operational ambient documentation or publicly documented pilots. Half showed operational evidence in clinical prediction or hospital operations, while specialized uses such as pathology and trial matching remain more concentrated.
That distribution tells us where the market is today.
Radiology is the mature clinical foundation. Documentation is the breakout generative AI application. Prediction is beginning to prove measurable clinical value. Administrative AI is developing a clearer financial case. Specialized areas such as pathology and trial matching show where deeper medical AI may go next.
But the most important shift may be happening above all of these applications.
New York hospitals are starting to build systems for managing AI itself.
As hospitals move from five algorithms to 50 and eventually hundreds, model selection will no longer be the hardest problem. The harder job will be determining which systems are safe, which are useful, whether they continue to perform, how humans should oversee them, and when they should be changed or removed.
That is the real transformation underway.
The hospitals most likely to succeed with AI will not be the ones that deploy it everywhere first.
They will be the ones that become exceptionally good at knowing where AI genuinely makes healthcare better—and where it does not.



