Latest Posts

Stay in Touch With Us

Got a story worth telling? Send it our way. We read every tip that lands in our inbox.

Livebriefs

  /  All News   /  The Investment Research Stack Should Outlast Model Churn 

The Investment Research Stack Should Outlast Model Churn 

  

By Angana Jacob, Head of Research Data at Bloomberg

Angana Jacob

AI is increasing the throughput of research faster than most firms are redesigning the infrastructure underneath it. 

Researchers can move more quickly from idea generation and coding through testing, validation and production. But AI is not simply accelerating an otherwise unchanged research process. Model choice, compute economics, context, routing and tooling increasingly sit inside the research loop itself, shaping what analysis can be attempted, how much reasoning is affordable and which workflows are economically viable. 

One way to frame the shift is through latency. HFT compressed the path from market event to observation, decision and order. AI is compressing a much broader research loop: from question to evidence, hypothesis, test, interpretation and decision. That compression matters across holding periods, from intraday strategies to long-horizon fundamental research. The complication is that the models, compute economics and deployment choices used to shorten the loop are themselves changing rapidly. 

At the same time, firms are expanding across strategies and asset classes while model capability, pricing and deployment options continue to shift. They are therefore having to make long-lived research infrastructure decisions around technology assumptions they should expect to revise repeatedly. 

That uncertainty puts a premium on what has to remain coherent across those changes: point-in-time data, financial semantics, historical relationships and machine-readable context. For quants, these disciplines have always determined whether research is credible. Under agentic AI, they remain central to validity and reproducibility while increasingly determining whether research can scale economically.¹ 

Investment AI Has a Context Integrity Problem 

Quant research has always depended on reconstructing the information set actually available at the time. Revised economic or earnings data, post-event estimates or a present-day survivor universe can leak future information into a backtest through look-ahead or survivorship bias. 

AI amplifies the consequences of those familiar biases. As hypothesis generation, feature construction and backtesting accelerate, firms can run more experiments across more datasets and longer automated research chains. Increasing research throughput also increases false-discovery capacity if temporal leakage or inconsistent research controls can enter and persist across those experiments. 

Investment meaning itself is deeply contextual. An observation rarely has significance in isolation; much of the signal lies in its relationship to expectations, exposures and the information set around it. 

For historical research, that context is multidimensional: what existed, how it was connected and what was known at the time. Companies acquire and divest businesses, securities and classifications change, index membership evolves, and estimates, ratings, supply chains and ownership relationships move through time. 

The point is to reconstruct not just the observations, but the investment world that gave them meaning. That historical world should not change depending on which model is looking at it.  

The missing layer in many financial AI architectures is therefore a point-in-time investment context graph: a computable model of market context as it existed at any point in time.  

This distinction becomes more important as protocols such as MCP simplify how agents connect to data and tools. MCP can remove substantial integration friction, but it does not resolve the underlying semantics. Two sources can expose tools through the same protocol while still encoding different meanings, mappings or historical conventions. Integration interoperability is not semantic interoperability. 

Consider an agent investigating why European automakers sold off after a Chinese policy announcement. Retrieving the policy statement, company filings, Chinese vehicle sales, equity prices and analyst commentary is relatively straightforward. Producing defensible research requires reconstructing which companies had material China exposure at the time, which subsidiaries and joint ventures belonged to which listed parents, which supplier and commodity relationships were relevant, how currencies moved, which securities were in the relevant indices, and what analyst expectations actually were before the event. 

An AI system can retrieve valid observations from authoritative sources and still reconstruct an invalid market state if the surrounding context is semantically or temporally inconsistent. 

That makes provenance, temporal metadata and explicit market semantics part of the research infrastructure. They make explicit for agents what a human researcher might otherwise infer: what something is, how it relates and when that relationship was valid. 

For historical research, markets are not represented by one static context graph, but by an evolving sequence of financial states. A temporal knowledge graph can encode that history, allowing the system to reconstruct the market structure and information set that existed at the point in time being analysed. 

Reliable reconstruction requires three forms of integrity: factual integrity, whether the observation is correct; relational integrity, whether it is connected to the correct entity or exposure; and temporal integrity, whether both the fact and relationship were valid and available at the time. 

Most discussions of grounded AI focus on the first. Investment research requires all three. Otherwise, a system can be grounded in individually authoritative sources while still reconstructing a market state that never existed, effectively introducing look-ahead bias at the context layer. 

This creates an important distinction between the reasoning process and the context it operates on. The reasoning may remain probabilistic, but different models should still be reasoning over the same reproducible historical world. What existed, how it was connected and what was known at the time should not depend on which model happens to be analysing it. 

The same requirement extends from research into production. When a historical signal moves into a live workflow, entities, features, revisions and relationships need to retain the same meaning. Research-to-production continuity is therefore partly a state-consistency problem, not simply a question of deploying the model that performed well in backtesting. 

As Strategies Multiply, Investment Context Has to Converge 

This consistency becomes more consequential as investment firms expand across asset classes, instruments, time horizons and research styles in search of diversification, capacity and uncorrelated return streams. Fragmented research infrastructure becomes more expensive as the mandate broadens. 

The same Chinese policy event illustrates why context cannot remain desk-specific. It may affect European automakers through demand, suppliers through production volumes, industrial metals through expected consumption, currencies through trade exposure, and credit spreads through changes in earnings and balance-sheet risk. Teams may express that view through equities, credit, commodities or FX, but they are analysing the same underlying economic event. 

As the strategy mix broadens, that duplication becomes less of a local inconvenience and more of a firm-wide cost. Different desks can reasonably use different models, features and investment horizons. What becomes harder to justify is each rebuilding its own representation of the same companies, securities and economic relationships. 

An equity analyst, macro researcher and credit quant may ask different questions and reach different conclusions from the same event. They should still be able to start from a common understanding of what the entities are, how they are connected and what was known at the time. Shared investment context does not mean centralising the research process; it gives differentiated teams a common grounding layer on which to build their own forecasts, features, positions and decisions. 

AI makes the economics of fragmentation harder to ignore. A firm can generate strong productivity gains inside individual teams while still duplicating data engineering, governance, model access and inference costs across the organisation. If each new strategy requires another version of the same underlying context, local productivity does not translate cleanly into firm-wide scale. 

The opportunity is for each additional strategy to reuse more of what the firm has already built. That is where shared context begins to create operating leverage: teams remain differentiated, but the cost of expanding into new research areas does not rise in lockstep with the number of desks or workflows. 

The context should converge. The models and tools above it should remain fluid. 

Model Fluidity Is Becoming an Architectural Requirement 

Capability, cost, deployment constraints and control requirements are changing too quickly for firms to assume that today’s preferred model will remain the right one for every task. 

A frontier model may be appropriate for reasoning across conflicting evidence around an unfamiliar company, credit or cross-asset event. A smaller model may be sufficient for classifying earnings commentary, extracting terms from filings or tagging entities and exposures. Open-weight models do not need to outperform frontier models across the board to change the build-versus-buy decision. If they are good enough for a specific workflow, the additional control can make them attractive, particularly where firms are working with sensitive or accumulated proprietary research. 

The firms spending the most on frontier inference will not necessarily extract the most research value from AI. Much of the advantage may come from deciding where expensive reasoning matters, where smaller models are sufficient, and where deterministic software should do the work instead. 

That same fluidity applies to the economics around the model. Falling inference costs can make previously uneconomic workflows viable, while better models and agent tooling can reduce the need for bespoke infrastructure. As those trade-offs move, architectural decisions that looked rational six months earlier can quickly become outdated. 

This makes stable data and context more important, not less. Model fluidity only works if switching models changes the reasoning engine, not the financial world it sees. Entity identity, historical relationships, definitions and point-in-time semantics need to remain dependable enough that firms can distinguish a change in model performance from a change in the underlying representation. 

Whatever orchestration framework or access standard, including MCP, sits between them, model substitution and routing will only work cleanly if the data semantics, context, research logic and deterministic tools can be reused across models. 

That also requires a clear boundary between what needs probabilistic reasoning and what should remain deterministic. Interpretation and judgement belong in models; calculations, data transformations, factor construction and historical joins are usually better handled through code and analytical libraries. Models can decide what needs to be done without recreating work that software can perform exactly. 

That boundary makes it easier to change the model mix as capability and economics shift. 

Token Economics Will Expose Weak Data Architecture 

Model fluidity can improve economics by routing work to the right model. But there is another source of cost that receives less attention: the data layer itself. 

An agent confronted with an ambiguous identifier may make additional retrieval calls. Inconsistent schemas can force it to inspect more documentation; missing entity relationships can trigger repeated searches; conflicting definitions can send the workflow back through earlier reasoning. Poorly structured context can also push large amounts of irrelevant information into the model simply to establish what the system is looking at. 

Bad context therefore creates an inference tax. 

As research loops become longer and agents perform more retrieval, reasoning and tool use, data architecture begins to influence token economics as well as research reliability.² Illustratively, if a workflow requires three model calls carrying roughly 10,000 input tokens each, it processes around 30,000 input tokens. If ambiguity turns that into eight calls carrying closer to 20,000 tokens each, input consumption alone can exceed 150,000 tokens before accounting for output or other model costs. The precise multiplier will vary, but the mechanism is straightforward: ambiguity compounds through agentic loops. The same compounding applies to reliability: as probabilistic steps accumulate, so do the opportunities for error to enter and propagate through the workflow.  

Token usage therefore links directly back to context integrity. As agentic research moves into production, firms will have to control both model choice and the inference imposed by weak data architecture. Otherwise, AI productivity gains can disappear in the cost of scaling them. 

The Research Stack Is Being Built Under Technological Uncertainty 

Taken together, these trade-offs leave investment firms making architectural commitments while model capability, deployment choices and inference economics continue to move. What makes this transition unusually difficult is that parts of the infrastructure now participate continuously in the research output and its economics. The challenge is to invest deeply in the data, context and institutional intelligence that can compound without locking the research stack too tightly to model and deployment choices that may age quickly. 

That makes extensibility a more demanding test of architecture. New models, datasets and strategies should be able to build on what is already there even as the technology above it changes. The research stack should accumulate value in its data, context and core infrastructure, while keeping models and tooling replaceable as capability, economics and deployment choices continue to shift.  

Source notes 

¹ In February 2026, participants in the Bank of England and FCA Artificial Intelligence Consortium discussed AI ROI in financial services. Some argued that access to and governance of underlying data may be a more significant differentiator for ROI than speed of AI deployment. The minutes explicitly note that these are participants’ views rather than formal Bank or FCA policy conclusions. 

² The Bank of England’s July 2026 Financial Stability Report notes that more complex agentic workflows and longer reasoning chains can require larger volumes of input and output tokens, raising usage-based costs and potentially making some applications prohibitively expensive at scale. 

   

You don't have permission to register