A Winning Enterprise-Grade AI Operating Model Builds Trust and Oversight
By Bryan Dougherty, President, Product and Technology of Arcesium

Everybody is trying to crack the code for successful AI. The frontier of AI operationalization is moving from copilots and digital assistants to autonomous agents executing multi-step operations, and humans are evolving from direct operators to supervisors of outcomes. This next phase of AI transformation is opening a gap between firms capturing measurable value and firms still running pilots.
As agents take on more of the execution, the constraint does not disappear. It moves to the edges of the workflow, to the moment an agent’s output must be checked and judged against what the business needs.
Firms need to guard against taking the adoption curve at too high speed. Opinions vary on how to operationalize AI, particularly on build versus buy, choice of models, and internal ownership. These are negotiables, based on need and resources. There are, however, nonnegotiable fundamentals that can help firms graduate from experimentation, catch up with the leaders, or course-correct issues holding back their progress.
- Deterministic Architecture for a Probabilistic Technology
The AI governance framework is how humans dictate how AI is to reason. LLMs are probabilistic by design; that is what makes the technology transformative. But asset managers like their mathematical equations to have only one answer. In some use cases, firms do want probabilistic AI agents reasoning live: Probabilistic AI is best suited for both the lowest-risk tasks, like summarization, and the most sophisticated tasks, like generating an investment hypothesis with a human in the loop. Stuck in the middle are numerous operational functions where AI agents can substantially impact resource-saving efficiency, and where the middle and back office have zero tolerance for errors in performance and attributions, data quality, NAV calculations, P&L calculations, reconciliations, and cash balances.
What “deterministic” means in practice, and how it differs from simple automation, is worth stating precisely. Most agents already run deterministically: The steps are fixed, the APIs are defined, and the sequence does not change between runs. The model’s job when it runs is narrow: interpret the request, resolve parameters, and decide which tool to call. Everything else — invoking that tool, holding the session together, executing the work — happens in the runtime around it. The work itself is deterministic; the model is just the layer that decides what to point it at.
Agents produce the same output from the same input on every run, and the logic becomes code that can be read and reviewed, rather than a prompt that must be trusted. Almost every agent running in the background, scheduled or event-triggered, should run this way. The only work that genuinely requires a model at runtime is unstructured data processing: reading a document, an email, or free text. Everything else is deterministic work that does not need to be rediscovered on every run.
This deterministic architecture choice is critical to mitigating live model risk, or the risk that agents produce bad information that cascades down the firm’s risk, compliance, accounting, and treasury functions, causing costly havoc.
- Agentic AI Needs Its Own Software Development Lifecycle
Every mature engineering organization runs on a software development lifecycle: a defined path from idea to production, in which each stage carries different owners, different reviewers, and different controls. Agentic AI in financial services does not yet have an equivalent, and much of the industry’s governance conversation remains focused on why oversight matters rather than on what governs an agent, stage by stage, from design to retirement. This scaffolding, called the agent harness, determines how an agent gets built, tested, deployed, monitored, and eventually retired.
Split the harness into two control planes, since collapsing them is precisely what keeps deployments stalled short of production. The first is tech-controlled: the plumbing that has to be right regardless of which expert is using it. The most effective agents reason over a single system of record, with the same controls, audit trails (including full version history), and accountability framework as every other institutional system. Every model interaction, accessed data point, query, action, and outcome is recorded in frozen logs.
The same containment philosophy should extend to what an agent is permitted to touch: Agents should hold no unique privileges, so an agent can only do what the person who triggered it is already entitled to do, and there is no path to data or actions a human could not reach directly. That inherited access can be narrowed further, constraining a broadly entitled user’s agent to only the functions its specific task requires. Any action that would change a system of record should pass through human approval, and agents should execute inside a secured sandbox that limits the systems they can reach at runtime. The reach of an agent is deliberately bound, which is why its actions can be trusted.
Only 22% of asset managers say they are very confident they could pass an independent AI governance audit today, which reflects the tech-controlled layer failing to exist rather than an argument against building it. Strong governance does not slow deployment; it earns the green light for production in the first place, by providing structured, reliable authoring environments, preventing catastrophic operational rollbacks, and keeping deployments active.
- Subject Matter Expert Enablement Is the Real Differentiator
The agent harness has two control planes for a reason. The first, tech-controlled, is the plumbing: the system of record, the audit trails, the access boundaries. The second is expert-controlled, and it is the one most firms have not built at all. This is where a subject-matter expert, not an engineer and not the model, decides what “correct” looks like: which exceptions get escalated, which outputs need a second look, and which edge cases the workflow should refuse to touch. Getting the split right produces a harness in which technology guarantees the same thing happens every time, and expertise guarantees the right thing was defined in the first place. Getting it wrong, by letting engineering own both planes or by leaving the expert layer undefined, produces a system that is auditable but not trustworthy. No one with the relevant judgment signed off on what it is allowed to decide.
That second plane matters because bottlenecks do not disappear as AI takes on more of the work; they move to the edges. As agents take on more execution, the challenge shifts to reviewing their work and ensuring it serves the business. This is what makes expert enablement, rather than model selection, the real point of leverage.
A recent National Bureau of Economic Research working paper, “Writing Code vs. Shipping Code,” by Demirer, Musolff, and Yang, tracked more than 500,000 GitHub developers across three generations of AI tools. It found that AI coding tools made developers up to 240% more productive. Software releases increased by about 30%. Actual usage of that new software, however, stayed flat. Every generation made writing code faster; none of them moved the human work of deciding whether the code is right and whether anyone wants it.
As production becomes cheaper, judgment becomes more valuable. That is why builder capabilities matter. When a subject-matter expert can define an agent workflow directly, in their own language, instead of routing their judgment through an engineer who must translate it secondhand, that judgment stops being a bottleneck and moves the business forward.
A Winning Enterprise-Grade AI Operating Model Builds Trust and Oversight
As Andrew W. Lo, director of the MIT Laboratory for Financial Engineering, has observed, “Designing systems that are inherently accountable is the most important challenge to overcome to unlock widespread AI adoption in the financial industry.” The emerging AI transformation winners will not simply be those that deploy agents first but those who turn agents into measurable value swiftly. Building the harness, both its tech-controlled and expert-controlled halves, enables them to scale AI with complete auditability, explainability, and risk management.
Sophisticated model intelligence will not be the differentiator for asset managers. A unified data architecture, a governance harness that puts the right decisions in the right hands, and domain-specific context married to deterministic controls will be.