Morgan Stanley has demonstrated that the path to faster, more accurate financial operations may run through deliberately limiting artificial intelligence autonomy, not expanding it. By deploying an agentic system called FIXR to automate profit and loss reconciliation — one of the most accuracy-critical and deadline-driven workflows in banking — the firm has halved the time required for the task while keeping human decision-makers firmly embedded in every step of the process. The result challenges a prevailing assumption that more agent autonomy inevitably yields greater efficiency.
Every trading day, Morgan Stanley’s desks execute transactions across cash equities, debt investments, and other instruments. At the close of each day, controllers must reconcile P&L data across the firm’s Finance, Risk, Operations, and Trade Capture systems. Hundreds of thousands of attributes routinely fail to match, requiring manual investigation of each discrepancy, or “break.” Controllers must make decisions on adjustments and secure sign-off before the numbers reach the trading desk — all under a hard morning deadline. Previously, this process could consume up to six hours for a single book. FIXR now completes the same work in two to three hours. Across roughly 100 controllers, that translates to approximately 1,500 hours saved per week.
“It’s much more like a co-worker than a copilot,” said Todd Johnson, Managing Director at Morgan Stanley, speaking at a recent VB AI Impact event. The internal system, he explained, goes beyond the straightforward tasks typically associated with generative AI deployments. “We think that’s where the opportunity is to really unlock more complex work in the organization.”
FIXR: How Morgan Stanley’s P&L Reconciliation Agent Works
After nightly P&L calculations complete, FIXR automatically analyzes breaks and proposes resolutions based on rules it has learned from past controller behavior. The system coordinates multiple specialized agents working in concert. One agent interprets historical guidance to develop start-of-day resolutions. A second learns from controller actions and documents the rules they apply. A third converts repeated patterns into durable, automated logic that the system can apply without case-by-case review.
Over time, FIXR can auto-clear certain breaks it has encountered before, suggest solutions for less familiar discrepancies, request human assistance when confidence is low, and flag items for investigation. When the same type of break is resolved through an identical method repeatedly, the system codifies that pattern into a firm rule. Critically, humans never leave the loop. Controllers review, approve, or correct every recommendation, and those decisions are fed back to improve the next day’s run. The agent learns daily — from the same people whose work it augments — what it got right and what it got wrong, then codifies that knowledge as it iterates.
“You still preserve that element of human accountability even as you start to automate,” Johnson said. “Over time you’ll see more and more of those items resolved in an automatic way.”
He emphasized that autonomy requires a high degree of trust. Enterprises will not realize efficiency gains if every agent output is checked by a human. The human-agent feedback loop was essential to solving the challenge of controlled, measured, repeatable automation. “We recognized that all that intelligence that’s sitting in the mind of a controller is gonna be difficult to get all into an agent on day one,” Johnson noted.
Process-First Design and Extensibility
Before any AI was introduced, Johnson’s team conducted a thorough process intelligence assessment, mapping and mining workflows to identify where automation would deliver the greatest advantage. The question was not simply “where can we add AI?” but rather: is the right solution an agent, traditional automation, or a straightforward re-engineering of an inefficient step? “If we can fix that first before we add agents to the problem, then we really will be transforming the opportunity,” Johnson said.
The P&L sign-off process contained numerous manual steps suitable for automation. By having agents take over these time-consuming tasks, controllers are freed for more value-added analysis and deeper risk consideration. But time savings alone were not the deciding factor. Johnson’s team chose P&L reconciliation because hundreds of controllers perform this work globally — across the Americas, Europe, and Asia — making it a highly scalable use case. The strategy: start with a single use case, prove its value, extend it, and then roll out more broadly across the organization.
Deterministic by Design
A deliberate architectural choice underpins the system’s reliability. Johnson explained that the team limited how much of the workflow depended on the model’s judgment. “If you have an opportunity to make things very prescribed and repeatable, that’s cheaper in terms of token consumption, it’s more repeatable in terms of controls — and have the LLM do the stuff where you don’t need that kind of deterministic workflow,” he said. As the system accumulates controller feedback on a given break type, Morgan Stanley converts that pattern into a fixed rule rather than leaving it to the model’s discretion. This hybrid approach keeps the large language model focused on cases that genuinely require its interpretive capacity, while routine matches are handled through deterministic logic.
Governance: Are Agents Code or Digital Employees?
The question of how to govern agents is emerging as a fundamental issue in the agentic era. Johnson argues that agents are “probably a little bit of both,” and that nuance is essential for oversight. Technical teams remain responsible for infrastructure protections such as firewalls and encryption. But a new dynamic arises around what Johnson calls the “performance element”: the humans using agents are accountable for their outputs because the agent is aiding their business work. A senior controller working with a junior controller does not relinquish responsibility simply because assistance is involved. “One of our strong principles in our AI governance generally is that there always has to be human accountability, even if there’s a degree of automation,” Johnson said.
That accountability, however, is not static. Johnson noted with candor that agentic AI requires ongoing training because models are constantly evolving. “You’re never gonna be able to say: ‘We’ve done all the evaluation and testing that we need to do. Let’s just let it go.’ You’re going to have to have a constant view as it evolves over time.”
What Enterprise AI Can Learn from Morgan Stanley’s Approach
Morgan Stanley’s experience resonates with patterns emerging across enterprise AI deployments. In a recent VB Pulse survey, nearly three-quarters of 87 respondents reported seeing little to no return on investment from custom model fine-tuning, describing a “sandbox graveyard” of AI projects that proved too costly to sustain. The survey findings, while directional, suggest that a process-first, buy-and-blend approach may be more sustainable than chasing bespoke models. Governance challenges also loom large: 38% of respondents cited the absence of a single accountable owner as their biggest barrier to production AI, and only two of the 87 enterprises surveyed had active monitoring and alerting in place to detect model failures.
Morgan Stanley’s method — start with a rigorous process assessment, keep humans in the loop, convert repeated judgments into deterministic rules, and hold a single thread of accountability — offers a replicable template for organizations wrestling with similar challenges. The firm did not attempt to build an all-knowing agent from the outset. Instead, it built a system that learns incrementally from the expertise already present in its workforce, codifying that knowledge into rules that reduce both time and cognitive load over time.
What This Means for Enterprises Deploying Agents Today
The most important takeaway from Morgan Stanley’s experience is that agent autonomy is not a binary setting or a goal to maximize. The right level of autonomy depends on the risk profile of the task, the maturity of the organization’s understanding of its own workflows, and the degree of trust that has been built between human operators and the system. For enterprises evaluating agentic AI, the lesson is to invest first in process intelligence, design for incremental learning, and architect systems that push deterministic work away from the model rather than toward it. That approach — less autonomy, more structure, continuous human feedback — is what cut a six-hour reconciliation process to three and saved 1,500 hours per week. It is a model worth watching as the agentic era moves from pilot projects into the core operations of the enterprise.