AI agents look simple when you watch a demo.
You type a request. The agent thinks for a few seconds. It searches for information, writes something, updates a system, runs code, or completes another task. It can feel as if one smart AI model did everything.
That is not what is happening underneath.
A serious AI agent is becoming a small software system of its own. It needs a model to reason. It needs data to understand the business. It needs memory so it does not forget everything after each task. It needs tools to interact with other software. It may need a secure computer where it can run code. It needs permissions that determine what it can touch. It needs monitoring so somebody can understand what it actually did.
If the agent works on an important business process, it may also need durable databases, audit logs, identity controls, testing systems, and ways to stop dangerous actions before they happen.
That collection of infrastructure is what we call the agentic AI stack.
New York is becoming an interesting place to study this stack because several companies based in or operating deeply from the city are building different pieces of it. They include large infrastructure businesses such as MongoDB and Datadog, specialist AI infrastructure companies such as Modal and Pinecone, newer companies such as Daytona and Runlayer, security platforms such as Noma Security and AppViewX, and agent-focused companies such as Tavily and Emergence AI.
The important story is not simply that New York has AI startups.
The bigger story is that New York is developing infrastructure for what happens around the model.
That may become extremely important as businesses move from AI that answers questions to AI that actually performs work.
The Short Version: The AI Model Is Becoming Only One Part of the System
For most of the generative AI era, companies focused heavily on models.
Which model is smartest?
Which model has the largest context window?
Which model is cheapest?
Which model is fastest?

Those questions still matter. But once an AI system starts taking actions, a different set of questions becomes just as important.
The Questions Change Once AI Can Act
Where does the agent get trustworthy information? What does it remember? What tools can it call? Where does its code run? Which credentials does it use? What happens when hundreds of agents act at the same time?
Companies also need to know how they can trace a bad decision, understand a failed workflow, or immediately stop an agent that begins doing the wrong thing.
The Model Is No Longer the Whole Architecture
This change is already showing up in the technical standards around agents. Anthropic introduced Model Context Protocol, or MCP, as an open way for AI applications to connect with tools and data. MCP later became part of the Linux Foundation’s Agentic AI Foundation.
The important idea is bigger than one protocol.
The AI industry is gradually standardizing the plumbing around agents.
The next competition may therefore be less about who can put another chatbot on a website and more about who owns the infrastructure that allows thousands or millions of AI agents to operate reliably.
New York companies are already competing across that infrastructure.
What Is the Agentic AI Stack?
A normal software application usually follows a path that engineers understand well.
A person clicks something. The application runs predictable code. The code reads or writes data. A result appears.
Agents change that pattern.
Agents Can Choose Their Own Execution Path
An agent can decide which path to take during a task.
It might call one tool for one request and seven tools for another. It may search the web, retrieve information from a database, generate code, run that code, examine the result, revise the code, call another agent, update Salesforce, and then send a message.
The same starting request can therefore produce different execution paths.
That makes the infrastructure underneath the agent much more important.
A Simple View of the Agentic AI Stack
| Layer | What it does | Why an agent needs it |
| Model | Reasons and generates | Gives the agent intelligence |
| Orchestration | Controls steps and workflow | Determines what happens next |
| Context and tools | Connects the agent to outside systems | Lets the agent gather information and take action |
| Data and memory | Stores facts, state and history | Prevents the agent from starting from zero |
| Runtime and sandbox | Provides somewhere to execute | Lets agents safely run code and processes |
| Observability and evaluation | Records what happened and measures quality | Makes failures visible |
| Identity and security | Determines who or what can do what | Limits dangerous actions |
| Durable state | Keeps critical information correct and available | Allows long-running agents to recover and scale |
The Layers Are Already Starting to Overlap
Companies do not always fit neatly into only one layer.
A database company can become an agent-memory company. An observability platform can become an agent-governance platform. A cloud runtime can become an agent sandbox. A machine-identity company can become an AI-agent identity company.
Infrastructure Companies Are Expanding Sideways
This convergence matters because enterprise buyers usually do not want twenty separate platforms.
They want fewer systems that can cover multiple parts of the stack while still giving them strong control.
That creates an opportunity for infrastructure companies that already own one critical layer to expand into adjacent areas.
Original NYC Tech Journal Research: Mapping New York’s Agent Infrastructure Cluster
To understand where New York is actually building strength, we created a small original dataset rather than relying on a generic list of AI startups.
Our Research Methodology
The dataset was built from publicly available company websites, technical documentation, product announcements, corporate location pages, legal pages, and official press releases reviewed through September 17, 2026.
A company qualified for the sample if two conditions were met.
Requirement One: A Meaningful New York Connection
We looked for clear evidence such as headquarters in New York, a corporate New York address, a major New York office, or a clear operating base in the city.
Requirement Two: Direct Relevance to Agent Infrastructure
The company’s product needed to support at least one major infrastructure function required by AI agents in production.
We then coded each company across seven functional categories. A category received a mark only when public product materials directly supported that classification.
This is therefore a functional coverage study. It is not a valuation ranking, market-share estimate, or claim that these are the only relevant agent infrastructure companies in New York.
The NYC Agentic Infrastructure Dataset
| Company | NYC evidence | Main role in the agent stack |
| Modal | New York office and team presence | AI compute, sandboxes, execution and infrastructure observability |
| Daytona | New York corporate address | Stateful, isolated computers for agents |
| Pinecone | New York City headquarters | Retrieval, knowledge and vector infrastructure |
| MongoDB | U.S. headquarters in New York | Operational data, retrieval and persistent agent memory |
| Cockroach Labs | Headquartered in New York City | Durable state, vector memory and resilient operational data |
| Tavily | New York address and office presence | Real-time web context and search for agents |
| Runlayer | Strong NYC operating and hiring presence | MCP gateway, tool access, policy, identity and agent control |
| Datadog | Global headquarters in New York City | Agent observability, tracing, experiments and governance |
| AppViewX | New York headquarters | Agent identity and machine identity security |
| Noma Security | New York address | Agent discovery, security, access control and governance |
| Emergence AI | New York address | Agent coordination, memory, verification and governed autonomy |
Where the NYC Sample Is Most Concentrated
| Infrastructure layer | Companies in our 11-company sample | Share of sample |
| Identity, policy and security | 7 | 64% |
| Data, retrieval and memory | 4 | 36% |
| Tools and live context | 4 | 36% |
| Observability, testing and evaluation | 4 | 36% |
| Durable state and resilience | 3 | 27% |
| Orchestration and control plane | 2 | 18% |
| Compute and sandbox execution | 2 | 18% |
Visual View
Identity / policy / security ███████ 7
Data / retrieval / memory ████ 4
Tools / live context ████ 4
Observability / testing / evals ████ 4
Durable state / resilience ███ 3
Orchestration / control plane ██ 2
Compute / sandbox execution ██ 2
This chart does not mean security makes up 64% of New York’s entire agent infrastructure market.
It means security or policy functions appeared in seven of the eleven companies in our selected sample.
That distinction matters.
But the pattern is still useful.
Original Finding #1: New York’s Advantage May Be the Infrastructure Around the Model
The most interesting result from our sample is what it does not show.
New York’s clearest concentration is not foundation-model training.
Instead, the companies in this sample are heavily focused on making AI usable inside real organizations.
New York Is Strong Where Enterprise AI Gets Difficult
That means context, databases, execution environments, security, identity, monitoring and control.
This fits New York unusually well.
Financial Firms Need Controlled Intelligence
A Wall Street bank does not only need an intelligent agent.
It needs an agent that can access the right market information, respect permissions, keep a record of its actions and avoid sending confidential information somewhere it should not go.
Healthcare Needs Strong Boundaries
A hospital does not merely need a model that understands medical language.
It needs controlled data access, durable records, identity, auditability and clear boundaries around what an automated system can do.
Law Firms and Agencies Have Similar Problems
A law firm needs controls around client documents and approvals.
An advertising company may need an agent that connects to campaign tools, analytics systems and creative workflows without receiving unlimited authority.
As agents move into these environments, the value shifts from pure intelligence toward controlled intelligence.
That appears to be an area where New York can build deeply.
Original Finding #2: Security Is Moving Into the Runtime
Security used to happen mostly around applications.
Companies protected the network. They protected databases. They protected APIs. They authenticated users.
Agents create a different problem because the dangerous decision can happen inside the workflow itself.
A Workflow Can Be Technically Valid and Still Be Wrong
Imagine an employee asks an agent to research a customer and prepare an account plan.
During that process, the agent might read Salesforce, call a web-search tool, access internal documents, run code and draft an email.
Every one of those steps can be technically valid while the full chain is still wrong.
Security Must Understand Context
That means security systems increasingly need to understand the identity of the agent, its purpose, the requested tool, the user’s authority and the sequence of actions leading to the request.
Runlayer is building infrastructure around this problem through its MCP gateway and agent controls.
Noma Security is approaching the problem through agent discovery and access control.
AppViewX is extending machine identity into agent identity.
These are different products, but they point in the same direction.
The agent itself is becoming a security principal.
The Infrastructure Layer Becomes More Important as Agents Gain Authority
A chatbot that can only answer questions has a limited blast radius.
An agent that can read confidential files has more.
An agent that can update records has even more.
An agent that can move money, deploy software, approve claims or contact customers has much more.
Authority Determines Infrastructure Requirements
This creates a simple rule for enterprise AI.
The more authority you give the agent, the more infrastructure you need around it.
| Agent authority | Example | Infrastructure requirement |
| Level 1: Answer | Summarize a document | Model, retrieval |
| Level 2: Research | Search multiple sources and compare them | Retrieval, live web access, citations, monitoring |
| Level 3: Recommend | Suggest a business decision | Memory, data quality, evaluation, audit trail |
| Level 4: Prepare action | Draft CRM updates or transactions | Tool access, permissions, identity, review |
| Level 5: Execute | Change records, send messages, run code | Runtime security, identity, sandbox, audit, rollback |
| Level 6: Operate continuously | Manage a workflow over hours or days | Durable memory, monitoring, budgets, policy engine, kill switch |
A Prompt Is Not an Authorization System
Many companies make the mistake of moving from Level 1 to Level 5 while barely changing their architecture.
That is dangerous.
A prompt can tell an agent what it should do.
It cannot replace technical controls that determine what the agent is actually allowed to do.
The Compute Layer: Where Agents Actually Run
Agents sometimes need more than an API.
They need a computer.
A coding agent may need to clone a repository, install packages, edit files and run tests. A research agent might execute Python. A data agent may need to run temporary workloads.
That creates demand for fast, temporary and isolated computing environments.
Modal Is Building Infrastructure for Agent Execution at Scale
Modal is one of the clearest New York examples.
The company describes its platform as cloud infrastructure built specifically for AI workloads.
Its infrastructure includes container execution, scheduling, storage and other primitives that are increasingly useful for agents.
Sandboxes Are Becoming a Production Primitive
Modal’s Sandbox product is especially relevant.
It allows developers to create isolated environments where agents can run generated code and other untrusted workloads.
The deeper idea is important.
An agent may increasingly get its own temporary computer in the same way a web request receives temporary compute today.
Why Sandboxes Matter
Giving an AI model direct access to the machine running your main application is risky.
The model can make mistakes.
A malicious prompt can try to convince it to run harmful commands.
Generated code may contain bugs.
Third-party packages may behave unexpectedly.
The Brain and the Hands Can Be Separated
A sandbox limits the blast radius.
The agent’s reasoning system can live in one environment while the code it generates executes somewhere isolated.
That architectural separation can make security, failure handling and observability easier to manage.
Daytona Is Building a Computer for Every Agent
Daytona is attacking a closely related problem.
Its core concept is a programmatic computer that an AI agent can create, pause, fork, snapshot and destroy.
Why Agent Computers Are Different From Normal Cloud Servers
Traditional cloud environments were largely designed for applications deployed by people.
Agent environments may need to start and disappear much more frequently.
They may also need to fork into multiple possible paths.
Agents Can Explore Several Futures at Once
An agent may start an environment, attempt a task, take a snapshot, fork the state into four alternative paths, test each path and retain only the successful one.
Humans rarely work like that.
Agents can.
This is why infrastructure built specifically for agents may eventually look different from traditional cloud infrastructure.
The Data and Memory Layer: Agents Need More Than a Giant Prompt
The next major layer is memory.

People often talk about agent memory as if it means keeping a chat transcript.
Production memory is much more complicated.
Agents Need Several Kinds of Memory
An agent may need to remember what a customer requested yesterday, what actions were taken, which results were successful, what permissions applied and what the current state of the business system is.
Simply inserting everything into a prompt becomes expensive and noisy.
That gives databases and retrieval systems a central role in the agent stack.
MongoDB Is Moving From Application Database to Agent Data Platform
MongoDB is one of the largest New York infrastructure companies and is increasingly positioning its products around AI agents.
Persistent Memory Needs Real Databases
Production agents need information that survives beyond one model request.
MongoDB’s direction reflects that need through persistent memory, search and retrieval capabilities.
Agents Need Both Knowledge and Current Truth
An autonomous system often needs two types of information.
It needs semantic knowledge:
“Which earlier case looks similar to this one?”
It also needs current operational truth:
“What is this customer’s account status right now?”
Keeping those things closer together can reduce the number of systems developers have to synchronize.
That makes databases increasingly important to agent architecture.
Pinecone Is Becoming an Agent Knowledge Layer
Pinecone occupies another critical part of the stack.
The company built its reputation around vector databases.
Retrieval Quality Becomes More Important When Agents Act
Vector search lets applications retrieve information based on meaning instead of relying only on exact words.
That became important for retrieval-augmented generation.
Agents increase its importance further.
Bad Retrieval Can Cause Bad Actions
An autonomous system needs good context before it makes a decision.
If retrieval is wrong, every later step can also become wrong.
That means retrieval quality becomes part of operational safety.
The long-term opportunity is therefore larger than storing vectors.
It is becoming the knowledge layer that many agents use before taking action.
Cockroach Labs Is Treating Agent Memory as Mission-Critical State
Cockroach Labs adds another angle.
The New York-headquartered company has spent years building distributed SQL infrastructure around consistency, resilience and scale.
Those properties matter even more when agents start changing real systems.
Agents Do Not Only Read Data
They also modify it.
Imagine ten agents working on related transactions at the same time.
They need to know which action happened first. They need consistent state. They cannot silently overwrite important changes.
Traditional Database Features Become Agent Features
Once agents begin performing real operational work, ordinary database properties take on new meaning.
Durability becomes memory.
Transactional consistency becomes coordination.
Access controls become agent permissions.
Audit logs become explainability infrastructure.
Original Finding #3: “Memory” Is Splitting Into Several Markets
Our research suggests the phrase agent memory is too broad.
At least four different problems are hiding inside it.
The Four Memory Layers
| Memory type | Question it answers | Likely infrastructure |
| Working memory | What is happening in this task right now? | Agent runtime or orchestration |
| Semantic memory | What past information is relevant? | Vector retrieval and knowledge systems |
| Operational state | What is true in the business right now? | Transactional database |
| Historical memory | What did the agent previously do and why? | Database, logs and observability |
Buyers Should Not Treat These as the Same Thing
A vector database and a transactional database can both be called “agent memory.”
They solve different problems.
Many production systems will need both.
The Live Context Layer: Agents Need Fresh Information
Models know a great deal, but their internal knowledge alone is not enough for many business workflows.
A sales agent needs current information about an account.
A financial research agent needs recent events.
A procurement agent may need current vendor data.
A competitive-intelligence agent needs information that changed this week.
Tavily Is Building Search for Agents Rather Than Humans
Tavily is focused on connecting AI systems to fresh web information.
That sounds similar to ordinary search, but the user is different.
Machine Search Has Different Requirements
A person wants a search-results page.
An agent needs structured information it can immediately process inside a workflow.
It may need source details, extracted page content, relevance signals and clean output that can move directly into another step.
That makes agent search its own infrastructure category.
Machine-to-Machine Payments Could Change Tool Access
One particularly interesting development is the idea that agents may eventually discover and pay for infrastructure dynamically.
Instead of a person manually creating every API account and configuring every service, an agent could potentially discover a service during a workflow, pay for a request and use it immediately.
That points toward a future where infrastructure itself becomes more autonomous.
The Tool Layer Is Becoming a Platform of Its Own
Agents become truly useful when they can act in software.
That requires tools.
An agent may need access to Gmail, Slack, GitHub, Salesforce, Workday, a database, an internal API or thousands of other systems.
MCP Is Helping Standardize Tool Connections
Model Context Protocol is one attempt to create a standard way for agents and AI applications to connect with tools and data.
That standardization can make integrations easier.
It does not remove the security problem.
Easier Connections Can Create More Governance Problems
The easier it becomes to connect agents to tools, the faster organizations can end up with hundreds of unmanaged connections.
That creates a market for gateways, policy engines and control planes.
Runlayer Is Building a Control Plane Between Agents and Tools
Runlayer sits directly in this gap.
Its platform provides infrastructure around MCP, including gateway controls, access rules, runtime security and observability.
The Integration Problem Gets Large Quickly
Without a gateway, many agents can end up connecting independently to many tool servers.
If a company has X agents and Y tool systems, the number of direct relationships can grow toward X × Y.
A centralized gateway can conceptually push the management problem closer to X + Y.
Original NYC Tech Journal Analysis: Direct Integration vs Gateway Model
| Agents | Tool systems | Direct relationships | Gateway relationships | Theoretical reduction |
| 5 | 10 | 50 | 15 | 70.0% |
| 10 | 20 | 200 | 30 | 85.0% |
| 20 | 30 | 600 | 50 | 91.7% |
| 50 | 50 | 2,500 | 100 | 96.0% |
| 100 | 100 | 10,000 | 200 | 98.0% |
The practical architecture will be more complicated than this simplified model.
But the shape of the problem is what matters.
As both agent count and tool count rise, direct integration complexity grows much faster.
That makes control planes increasingly valuable.
Observability: Companies Need to See What the Agent Actually Did
Traditional application monitoring asks familiar questions.
Did the server crash?
How long did the request take?
What was the error rate?
Agents introduce new questions.
Agent Failures Are Often Workflow Failures
Why did the agent choose this tool?
Which agent handed the task to another agent?
How many steps did the workflow take?
Which model call caused the cost spike?
Why did the agent repeat the same action six times?
Did the final task succeed even though every individual API call technically succeeded?
Those are different observability problems.
Datadog Is Moving Directly Into Agent Observability
Datadog is one of New York’s largest infrastructure software companies.
Its move into AI agent monitoring reflects how observability itself is changing.
Traces Now Need to Explain Decisions
Traditional traces show how software requests move through systems.
Agent traces increasingly need to show reasoning paths, tool calls, model interactions and handoffs.
The Final Outcome Matters More Than Individual API Calls
A workflow can contain twenty technically successful API calls and still produce a failed business result.
That means agent monitoring needs to connect technical events to actual task completion.
This is why agent observability will become more than a debugging tool.
Agent Observability Will Also Become an Economic System
Agents create variable costs.
A normal software function may execute once.
An agent may call a model, search four systems, call another model, run code, retry a failed step and then invoke another agent.
The user still sees one task.

The infrastructure sees many paid actions.
Cost Per Task Matters More Than Token Cost
This distinction will become extremely important.
A cheap model inside a wasteful workflow can cost more than an expensive model inside a well-designed one.
Companies Need Workflow Economics
Teams should measure how much a successful business outcome costs.
That includes model usage, search, compute, tool calls, retries and human intervention.
Without that view, companies cannot tell whether agent automation is truly becoming more efficient.
The KPI Dashboard Every Enterprise Agent Program Should Build
Executives should avoid measuring agent adoption with only numbers such as monthly active users or prompt volume.
Those metrics show activity.
They do not show whether autonomous work is actually working.
A Better Agent KPI Dashboard
| KPI | What it tells you |
| Task completion rate | Percentage of delegated jobs actually finished |
| Human rescue rate | How often a person must take over |
| Cost per completed task | Full model, search, compute and tool cost |
| Median steps per successful task | Workflow efficiency |
| Tool failure rate | Reliability of external systems |
| Retry rate | How often the agent gets stuck |
| Unauthorized action blocks | How frequently policy prevented dangerous behavior |
| Retrieval failure rate | How often the agent lacked useful context |
| Human approval rate | How often proposed actions are accepted |
| Rollback rate | How often completed actions must be undone |
| Time to completion | Whether automation really makes work faster |
Adoption Is Not the Same as Value
A company can have thousands of prompts per month and still have very little business value.
A mature agent program should instead be able to show that work is being completed faster, more reliably or at lower cost.
If it cannot, the company probably has an AI demo rather than an operating system for autonomous work.
Identity Is Becoming One of the Most Important Layers
There is another difficult question underneath every autonomous action.
Who did it?
Suppose an employee tells an agent to prepare a customer refund.
The employee has one identity.
The agent has another.
The software service used to process the refund has another.
Delegation Creates an Identity Chain
If another agent becomes involved, there may be another identity.
A useful audit record cannot simply say that “John’s account” performed the action.
It needs to show the chain of delegation.
Agent Identity May Become as Important as Employee Identity
Traditional companies have human employees and machines.
The autonomous enterprise adds another class of actor.
The agent.
That actor needs identity, permissions, expiration rules and logs.
AppViewX Is Extending Machine Identity to AI Agents
AppViewX has long worked on certificates, PKI and machine identity.
Its move into agent identity is therefore a logical extension.
Agents Need Their Own Credentials
If an agent uses credentials borrowed permanently from a human account, companies lose visibility and control.
A stronger system gives the agent its own identity.
Agent Credentials Should Be Narrow and Temporary
The ideal permission model is usually not “this agent can access Salesforce.”
It is closer to:
“This agent can read specific Salesforce records for the next ten minutes because this authorized employee delegated this specific task.”
That is a much stronger security model.
Noma Security Is Treating Agent Access as Its Own Security Problem
Noma Security is taking a broader agent-security approach.
The platform focuses on discovering, governing, testing and protecting AI systems and agents.
Discovery Comes Before Control
Companies cannot secure agents they do not know exist.
As employees begin creating their own AI workflows, unofficial agents and MCP servers can spread quickly.
Enterprises Need an Agent Inventory
Before companies can apply policies, they need to know which agents exist, what systems they reach and what identities they use.
This may become the AI equivalent of asset management in traditional cybersecurity.
Emergence AI Is Working on Verified Autonomy
Emergence AI is working farther up the autonomy stack.
Its work focuses on areas such as memory, coordination, verification and safe execution.
Long-Running Agents Create New Problems
An agent operating for three minutes is very different from an agent running for three days.
State accumulates.
Memory changes behavior.
Agents coordinate.
Errors can compound.
Long-Horizon Autonomy Is an Infrastructure Problem
Once agents run for long periods, companies need stronger checkpointing, monitoring, verification and recovery.
This is why multi-agent and long-horizon systems cannot rely only on better prompting.
They need production infrastructure.
Original Finding #4: The Stack Is Converging
Our company map shows that the clean layers described earlier are already beginning to merge.
MongoDB is not only a database. It is moving deeper into vector search and persistent agent memory.
Cockroach Labs is not only a SQL database. It is adding agent-oriented access and retrieval features.
Datadog is not only monitoring servers. It is increasingly monitoring agent behavior.
AppViewX is extending machine identity into agent identity.
Runlayer combines connectivity, policy, security and observability.
Pinecone is moving from raw vector infrastructure toward a wider knowledge layer.
Modal is moving beyond compute toward sandboxes, storage, execution controls and observability.
Consolidation Is Likely
The final enterprise agent stack may be smaller than today’s architecture diagrams suggest.
Companies do not want twenty-seven different infrastructure vendors if five platforms can cover the same requirements.
The Winners May Expand From Strong Core Layers
A company that already owns critical database infrastructure can move into memory.
A security company can move into agent identity.
An observability company can move into governance.
A compute company can move into execution control.
This is how software platforms are formed.
Public Scale Signals Around the Agent Stack
It would be misleading to put very different infrastructure companies into one revenue or customer ranking because many do not disclose comparable numbers.
A better approach is to look at the different scale signals they publicly report.
These Signals Are Not Directly Comparable
SDK downloads, API requests, customers and sandbox executions measure different things.
They should not be placed into one competitive ranking.
What They Do Show
Together, they suggest that the infrastructure beneath agents is no longer a small experimental category.
Tool connectivity is growing.
Machine-oriented search is growing.
Agent execution is scaling.
Databases are repositioning around agents.
Observability is expanding into AI workflows.
Security vendors are treating agents as a new class of system.
The surrounding infrastructure is becoming a market of its own.
What the Production Agent Stack Could Actually Look Like
Consider a New York investment company building a research agent.
The employee gives the agent a company name and asks for a detailed investment brief.
The Workflow Touches Many Layers
The model provides reasoning.
Tavily could supply current public web information.
Pinecone or MongoDB could provide relevant internal knowledge.
CockroachDB or MongoDB could hold durable application state.
Runlayer could control tool access.
Modal or Daytona could provide isolated execution when code needs to run.
Datadog could trace the workflow.
AppViewX or Noma could add identity or security controls.
One Agent Can Depend on Many Infrastructure Systems
The company would not necessarily use all of these vendors.
The point is that a production system needs the functions represented by these layers.
The model alone is not the architecture.
A Simplified Enterprise Agent Flow
USER
│
▼
IDENTITY + PERMISSIONS
│
▼
AGENT / ORCHESTRATION
│
├────► MODEL
│
├────► INTERNAL DATA + MEMORY
│
├────► LIVE WEB CONTEXT
│
├────► MCP / TOOL GATEWAY
│
└────► SECURE SANDBOX / COMPUTE
│
▼
BUSINESS SYSTEMS
Every step
│
▼
OBSERVABILITY + SECURITY + AUDIT
This is why companies should stop thinking about the AI model as the whole system.
The model is one service inside a larger operating layer.
How New York Companies Should Build Their Agent Stack
The wrong approach is to buy every interesting AI infrastructure product and then search for something to automate.

Start with the workflow.
Step 1: Write the Agent’s Job Description
Describe exactly what the agent is responsible for.
“Help the sales team” is too broad.
“Research a target account, gather five verified business signals, compare them against existing CRM data, create a draft account brief and send it to a sales manager for approval” is much better.
Architecture Should Follow the Job
The second version tells you what infrastructure is required.
The agent needs web access.
It needs CRM access.
It needs retrieval.
It needs a place to store results.
It needs an approval boundary.
It needs logs.
That is how architecture should be designed.
Step 2: Separate Reading From Acting
Create two categories of tools.
The first category reads information.
The second category changes something.
Read Permissions and Write Permissions Should Not Be Equal
An agent reading a Salesforce opportunity is very different from an agent modifying the opportunity value.
Do not treat those permissions equally.
This is one of the simplest ways to reduce agent risk.
Step 3: Design Memory Intentionally
Decide what the agent should remember.
Do not automatically save every interaction forever.
Give Every Memory Type a Purpose
Some information belongs in task memory.
Some belongs in a retrieval system.
Some belongs in the company’s normal database.
Some should disappear after the workflow ends.
Memory should have an owner, retention period and reason for existing.
Step 4: Put Untrusted Execution in a Sandbox
If an agent can generate or run code, isolate it.
Do not give experimental code unnecessary access to production infrastructure.
Keep Network and Credential Access Narrow
Use limited network rules, temporary credentials and strong environment boundaries.
Generated code should not automatically inherit the same privileges as trusted production software.
Step 5: Centralize Tool Access Before Agent Count Explodes
Connecting one agent directly to five APIs can feel manageable.
Connecting one hundred agents across one hundred systems can become an infrastructure project of its own.
Standardize Authentication Early
Create a common approach for authentication, tool discovery, authorization and logs.
This becomes harder to retrofit after dozens of teams have already built their own integrations.
Step 6: Instrument the Workflow Before Giving It More Authority
You should be able to reconstruct what an agent did before allowing it to perform high-impact actions.
Log the Full Execution Path
Capture which model was used, which tools were called, which documents were retrieved, how long each step took, how much the workflow cost and what final action occurred.
Without this information, improving an agent becomes guesswork.
Step 7: Add Human Approval at Decision Boundaries
Not every step needs human review.
Review should sit where authority changes.
Humans Should Control the Important Boundaries
Reading a document may not need approval.
Drafting an invoice may not need approval.
Actually issuing a payment may.
The goal is not human review everywhere.
It is human control where the downside becomes meaningful.
A 90-Day Agent Infrastructure Plan
A company does not need to build the complete autonomous enterprise in one quarter.
It needs one production workflow that teaches the organization how the stack actually works.
Days 1–30: Build a Narrow Workflow
Choose one task with clear inputs and outputs.
Prefer a repetitive process that currently requires people to search across several systems but where mistakes can still be reviewed.
Instrument Everything From the Beginning
Connect only the data the agent actually needs.
Log every important step.
Do not begin with full autonomy.
Days 31–60: Add Controlled Actions
Allow the agent to prepare actions before executing them.
Measure how often humans approve its proposed actions without changes.
Approval Rate Is a Useful Readiness Signal
If employees repeatedly rewrite the agent’s work, the system may not be ready for autonomy.
If approval is consistently high, the organization can begin testing limited automated actions.
Days 61–90: Automate the Safest Decisions
Look at real production traces.
Find actions with high success rates, low disagreement and low downside if reversed.
Increase Autonomy With Evidence
Automate those actions first.
Keep ambiguous or high-risk decisions behind approval.
This gradually increases autonomy using evidence rather than enthusiasm.
Buy vs Build: What Should Enterprises Own?
Companies should not build every layer themselves.
The better question is where custom infrastructure creates strategic value.
A Practical Buy-vs-Build Framework
| Capability | Usually buy | Sometimes build |
| GPU and elastic compute | Yes | Only at unusual scale |
| Generic sandbox infrastructure | Yes | If execution environment is core IP |
| Vector retrieval | Usually | If retrieval itself is your product |
| Transactional database | Usually | Rarely |
| Web search infrastructure | Usually | Only for highly specialized domains |
| MCP/tool gateway | Increasingly | Large platform teams may build internally |
| Agent orchestration | Mixed | Often customized around workflow |
| Business-specific evaluation | No | Usually worth owning |
| Permissions policy | Platform plus custom rules | Business rules must remain yours |
| Audit and observability | Platform | Custom dashboards on top |
| Domain knowledge | No | This is often core company value |
Buy the Plumbing, Own the Business Logic
Your competitive advantage probably does not come from writing another container scheduler.
It may come from knowing exactly how your company’s underwriting, legal review, portfolio research, procurement or customer-support workflow should operate.
That is the part worth owning.
The New York Advantage: Enterprise Density
New York’s agent infrastructure opportunity becomes more interesting when you consider the customers surrounding these companies.
The city is dense with industries where software actions have consequences.
Finance.
Insurance.
Healthcare.
Advertising.
Media.
Retail.
Real estate.
Professional services.
These Industries Create Hard Infrastructure Problems
They have expensive knowledge work, complicated software stacks and strong requirements around control.
That combination creates a valuable local feedback loop.
Difficult Customers Can Produce Better Infrastructure
Agent companies can learn from demanding enterprise customers.
Infrastructure companies can learn from agent builders.
Security companies can see new attack surfaces.
Large enterprises can hire from the same engineering ecosystem.
That creates the possibility of a real agent infrastructure cluster.
Why Finance Could Be an Especially Important Test Market
Financial services may become one of the most useful environments for testing serious agent infrastructure.
Financial Agents Need More Than Intelligence
An agent can research companies, prepare models, review filings, monitor portfolios, investigate transactions, reconcile data or help with compliance.
But nearly every useful action touches sensitive information.
Finance Forces the Full Stack to Work Together
The agent needs high-quality retrieval.
It needs fresh information.
It needs permissions.
It needs auditability.
It may need strong transactional data.
It may need a human approval boundary.
It may need a record of exactly which information supported a decision.
That requires much of the agentic stack.
Legal AI Creates Similar Demands
Law firms present another example.
An agent researching a legal issue needs access to authoritative sources.
A document agent needs permission to access the correct client files.
A drafting agent needs version history.
Legal Work Makes Context Boundaries Important
A workflow agent may interact with document-management systems, email, billing software and internal knowledge.
Confidential information cannot freely move between all these systems.
The Hard Part Is Moving From Knowledge to Action
Generating text is only one piece.
The difficult technical problem is controlling what happens when the agent takes that information and acts on another system.
That is an infrastructure problem.
Healthcare Raises the Standard Even Further
Healthcare adds another layer.
An administrative agent may be able to schedule appointments, answer patient questions, route documents or prepare records.
Sensitive Data Raises the Infrastructure Requirement
Clinical decisions and patient data require tighter controls.
The more sensitive the environment, the more valuable identity, policy, memory, monitoring and human approval become.
Safety Infrastructure Is Business Infrastructure
These capabilities should not be viewed as optional compliance overhead.
They are the systems businesses need before autonomy can safely spread.
What CTOs Should Ask Agent Infrastructure Vendors
The agent infrastructure market is becoming crowded quickly.
Buyers should not evaluate products based only on impressive demos.
Ask What Happens When Something Goes Wrong
Can every action be traced to both the agent and the user who delegated the task?
Can permissions be scoped to a specific tool or action?
Can credentials expire?
Can the agent’s network access be restricted?
Can historical runs be reconstructed?
Can data retention rules be configured?
Can spending be limited?
Can the system recover if a workflow fails halfway through?
Can an administrator stop an agent immediately?
The Most Important Question Is About Failure
What happens when the model behaves unexpectedly?
If the answer is simply that the model usually follows the prompt, the architecture is not ready for serious autonomous work.
The Most Important Design Principle: Assume the Model Will Eventually Make a Bad Decision
Businesses should not build agent infrastructure around the assumption that AI will always behave correctly.
Traditional software engineering does not work that way.
Mature Infrastructure Assumes Failure
Databases assume machines can fail.
Networks assume packets can disappear.
Security systems assume credentials can be stolen.
Distributed systems assume components will eventually break.
Agent infrastructure should make a similar assumption.
One Bad Decision Should Not Become a Major Incident
At some point, the model will misunderstand an instruction.
It will retrieve the wrong information.
It will choose the wrong tool.
It will call the correct tool with the wrong argument.
It may repeat an action.
It may encounter malicious content designed to influence its behavior.
Good agent infrastructure is what prevents one incorrect decision from becoming an organizational incident.
The Agentic AI Stack May Become the Next Major Enterprise Software Layer
There is a familiar pattern in technology.
A new application wave appears.
At first, developers build nearly everything themselves.
Then common problems become visible.
Infrastructure companies form around those problems.
We Have Seen This Pattern Before
Cloud computing created large markets for databases, monitoring, APIs, identity, security, observability and developer platforms.
Mobile created its own infrastructure ecosystem.
Agentic AI appears to be moving through the same cycle.
Prompts Were Only the Beginning
Developers first focused on prompts.
Then they needed frameworks.
Now they need memory.
They need sandboxes.
They need tool connectivity.
They need durable state.
They need identity.
They need monitoring.
They need evaluation.
They need governance.
They need ways to control autonomous actions.
That is how a technology stack forms.
The Bigger Story: New York Could Become a Control Layer for the Autonomous Enterprise
Our original sample points toward a useful conclusion.
New York may not need to dominate the foundation-model race to become one of the most important cities in agentic AI.

It can build what enterprises need after the model becomes intelligent enough.
New York Companies Are Filling Different Parts of the Same Stack
Modal and Daytona are building places where agents can execute.
Pinecone, MongoDB and Cockroach Labs are building ways for agents to retrieve, remember and maintain state.
Tavily connects agents to fresh information.
Runlayer connects them to tools while adding controls.
Datadog helps companies see what agents are doing.
AppViewX and Noma are building identity, policy and security around autonomous systems.
Emergence AI is exploring the orchestration and verification required for more complex autonomy.
The Real Opportunity Is Controlled Autonomous Work
These products solve different technical problems.
Together, they reveal the same business shift.
Software is changing from something people operate into something that can increasingly operate on behalf of people.
Intelligence Is Only the Beginning
Once that happens, the enterprise needs to know what the agent can see.
What it can remember.
What it can use.
What it can change.
Where it can execute.
How much it can spend.
Which identity it operates under.
Why it made a decision.
How the decision can be reviewed.
And how the system can be stopped when something goes wrong.
That infrastructure is becoming the real operating layer of agentic AI.
For New York, that may be a much larger opportunity than building another chatbot.
It is the opportunity to build the systems that make autonomous software safe enough, reliable enough and useful enough to run real businesses.



