When a takeover is announced, agents at Balyasny extract the economic and legal terms, review closing conditions, and estimate whether and when it will close. Chief AI Officer Charlie Flanagan says this research package now takes less than a day, including human review, compared with three to five days previously. The agent’s run takes about 30 minutes.
The assignment continues after the first assessment. Agents update deal probabilities as filings and press releases arrive. Balyasny’s central Applied AI team supplies infrastructure that investment teams customize around their strategies. Investors retain judgment over the conclusions they use.
My comparison of these accounts examines how funds specify the work they delegate and decide whether to accept the result. Investment expertise enters through data and labels, definitions and instructions, executable tools, and acceptance criteria. Bridgewater provides the broadest technical record in this selection, along with a live investment program.
Public disclosures through October 1, 2026. Sources and dates at the end.
I’m launching a new workshop, Saturday, October 10, 10:00 AM-2:30 PM ET, ML for Trading in the Age of AI Agents, on how coding and research agents change strategy research.
Tomorrow’s free lesson, How AI Agents Change the ML for Trading Workflow, introduces the workflow. Register for the lesson to receive a workshop discount.
Investment models predate research agents
Hedge funds were using machine learning before conversational LLMs. In a 2019 interview, Man AHL’s Anthony Ledford described ML in daily and intraday strategies, execution algorithms, and smart order routing, including deployed text-based strategies. AQR states that it actively uses ML across its investment process. Its November 2024 portfolio study illustrates learning nonlinear relationships between stock predictors and returns, separately from a managed portfolio’s performance.
Research agents add activities around those models. They can prepare data, investigate documents, write experiments, and revisit an analysis. The distinction helps locate responsibility. An agent that implements a candidate signal still depends on the research criteria used to accept it and the portfolio process used to trade it.
Millennium retains assignments and employee permissions
Millennium reports more than 1,600 personalized assistants, called Digital Twins, in use across the firm. Employees delegate recurring research, meeting preparation, and scheduled work through email. AI Product Manager Diana Meditz gives a specific example. Her assistant reviews overnight AI news and relevant inbox updates, then delivers a morning report without another prompt. Each assistant has its own identity and retained context. It starts with no access, receives permissions system by system, and cannot exceed its employee’s entitlements. The count measures adoption across employees; the firm supplies no measured gain in research quality.
Balyasny’s deal agent revisits an assessment when evidence changes. Millennium’s morning assistant carries a standing assignment into the next scheduled run. A team borrowing either approach has to specify what persists. For deal research, that includes the current assessment and evidence that could change it. For a morning report, it includes the employee’s instructions, relevant sources, and delivery schedule. Retained conversation alone would leave those requirements ambiguous.
Millennium separately announced a project with Anthropic to build, pilot, and optimize a digital risk analyst. It would explain daily changes in risk positions and retain context across questions to support human risk managers. The risk analyst remains a development announcement, separate from the deployed Digital Twins.
D. E. Shaw’s Neil Katz described a similar infrastructure approach in 2024: Assistants, an LLM Gateway, and DocLab for querying documents. Teams could connect models to their own data and software. D. E. Shaw and Balyasny both provide common infrastructure with local customization, enabling investment teams to specify work for their own strategies.
Citadel replicates papers; Man generates signal candidates
Citadel’s Ken Griffin described an agent that reads a finance paper, reproduces its results, and produces out-of-sample results. He reported an average of two to three hours per paper, against a human process taking six to eight weeks. The transcript supplies no paper list, failure rate, or accounting of human review time. Griffin also said the automation had not reduced headcount. Citadel had more problems for its existing people to pursue.
Man’s AlphaGPT starts earlier in research. Separate components propose hypotheses, translate them into code using internal tools and databases, and evaluate candidate signals against statistical, risk, and economic criteria. Man reports its greatest success so far in systematic equity research. Its investment committee reviews hypotheses and economic reasoning, while technology teams review implementation and tests before live use. The account establishes candidates passing internal research thresholds, without attributing live returns to AlphaGPT.
The comparison exposes a shared implementation problem. Citadel starts with a published method that the experiment must preserve. Man starts with a proposed relationship that the generated code must measure. Man explicitly warns that an agent can propose one idea and implement another. Its example asks whether stocks with more buy orders than sell orders tend to outperform. To implement that idea, a researcher would need to specify whether “more” means order counts or submitted share volume, over which interval, and when the information becomes available. A successful backtest using a different definition would still answer a different research question. Man also identifies the multiple-testing risk from producing many candidate signals and describes expanding monitoring infrastructure to maintain oversight quality as signal volume grows.
Man also describes using Claude Skills to express workflows in plain English. Its partnership announcement with Anthropic names AlphaGPT as an existing tool.
Two Sigma lowers the cost of preparing research ideas
Two Sigma’s Ben Wellington explains a different output. An LLM can generate descriptions of companies or events, which researchers can analyze as datasets. He uses CEO microexpressions during earnings calls as an example of an idea whose technical preparation had previously made it expensive to explore. The example illustrates a research possibility, without identifying a deployed predictor. Researchers supply the question and test whether the generated data contains predictive information. Wellington warns that poorly designed automation can make their results increasingly homogeneous.
Wellington reports feature-research insights arriving in days where the work previously took months, without supplying a controlled timing sample. CTO Jeff Wecker describes adapting AI tools to internal platforms. Cheaper preparation can reopen ideas researchers had previously set aside. Citadel’s account similarly describes using existing staff to investigate more problems.
Generated research data introduces a historical-information problem. Two Sigma’s Jin Choi explicitly warns about trusting backtests that predate an LLM’s knowledge cutoff. A model labeling an old earnings call may know events that occurred afterward. For a team adopting the method, the source date, model version, and prompt version belong in the research record. Those records help investigate leakage, although recording them alone cannot remove knowledge already in the model.
Bridgewater uses investor judgment in data, plans and tests
Bridgewater describes AIA Labs’ ambition as an artificial investor capable of the range of activities human investors perform. The firm’s account also describes AI research and tools feeding back into Pure Alpha, with experienced investors supplying feedback. Its resources include macroeconomic data, recorded investment reasoning, and investor expertise. Its public studies show different ways of using those resources.
Investor labels define what is worth reading
Bridgewater and Thinking Machines examined six document-filtering and segmentation tasks on data cleared for public release. Relevance depends on the investor’s question and responsibilities. The researchers first improved prompts by asking experts to articulate distinctions such as financially relevant versus useful to a macro investor. Expert prompting raised frontier-model accuracy from roughly 50% into the mid-to-high 70s.
Training examples supplied further judgment. The team initially trained on non-expert vendor labels, then sent examples where the model disagreed with its training label to investors for correction. Final evaluation used a held-out set. A trained Qwen3-235B model reached 84.7% average accuracy, versus 78.2% for the best prompted frontier comparator. The 13.8-fold saving concerns inference cost per task, excluding annotation, training, and integration. Ablations show that training choices also affect accuracy.
The method uses expert attention to settle contested examples and expert instructions to define the task. The result measures agreement with investor labels on the disclosed tasks. It does not isolate the effect of labels from the training procedure or compare with a closed model fine-tuned on the same data.
PAT turns a research question into inspectable computation
The Pocket Analyst Tool, or PAT, performs exploratory research. Bridgewater’s applied AI team describes an internally deployed analyst used by hundreds of investors. PAT retrieves proprietary documents and series, clarifies the question, and plans dataframe outputs with declared schemas and dependencies. Its execution architecture enforces validation and caches intermediate calculations. User corrections can become human-audited benchmarks and reviewed changes to context or implementation.
PAT makes the intended analysis explicit before generating code. That addresses the same relationship between research intent and implementation raised by Citadel and Man. The analyst can inspect the plan and revise the analysis. The presenters scope PAT as exploratory research, separate from trading. I covered the architecture in How Bridgewater Engineers a Research Agent.
SQL execution can succeed while the calculation is wrong
A text-to-SQL study by researchers at UIUC and Bridgewater AIA Labs, in collaboration with Thinking Machines, shows how successful execution can mislead. The researchers combine expert-checked training data with rewards that examine query behavior.
In the study’s example, a generated query averages line-item prices when the question requires average order totals. The result happens to match because every order in that database has one line item. With multi-item orders, the queries can return different averages. In a pilot training run, 32.8% of positive execution-match rewards went to queries that failed the additional equivalence check. The authors downweighted those rewards using bounded verification of query equivalence. Task definitions therefore enter the training objective as well as the prompt. This is database-query research, with no disclosed internal rollout.
A forecaster can add information while trailing the market
The AIA Forecaster report studies probability forecasts. Multiple agents search news, a supervisor investigates disagreements with further searches, and a statistical correction calibrates the combined forecast. The authors report no statistically significant difference from expert superforecasters on the two ForecastBench sets with those comparisons. That result concerns those tests, not general forecasting equivalence.
On MarketLiquid, the report gives rounded Brier scores of 0.126 for AIA and 0.111 for market consensus, where lower is better. A fitted combination scored 0.106 under leave-one-out evaluation. The benchmark covers 322 events at five forecast dates each, yielding 1,610 questions. These events resolved in April- May 2025. A forecast can contribute information even when it performs worse alone than the forecast already available. This results from combining probabilities before portfolio sizing, costs, and risk constraints.
My Agent Engineering: Build, Evaluate, and Deploy AI Agents workshop (4th session on November 21) covers the architecture described by Bridgewater's AIA Forecaster and teaches multi-agent system design, evaluation, and deployment. Use the link for an early-bird discount through October 31.
AIA Macro: returns and investment history
Reuters dates Bridgewater’s AI-strategy work to 2018 and AIA Macro’s inception to late 2023. Bridgewater’s July 2024 statement then described a vehicle with external-client capital that combines tabular learning, reasoning tools, and LLMs. It specified $1.6 billion in AUM raised in nominal terms.
Reuters reports AIA Macro’s 16.4% January-September gain beside Pure Alpha’s 18.4% for the same interval. One unnamed source supplied the AIA figures, including 13.2% annualized since inception and over $4.5 billion in current AUM. Current AUM and the earlier nominal raised-AUM disclosure have different measurement bases. The report does not specify the fee basis, risk exposures, or tool-level return attribution.
The same-period comparison describes two strategy outcomes. According to Bridgewater, Pure Alpha also uses research and tools developed with AIA Labs. It cannot serve as an AI-disabled control. Nor does the public record connect AIA Macro’s return to PAT, the document filter, the SQL model, or this particular forecaster.
What a research team must specify before delegating work
For a team adapting these methods, the work product determines the first check. A document filter needs examples of what this investor would read and what it must not discard. A deal-monitoring agent needs the evidence behind its current probability and rules for revisiting it. A paper-replication or signal agent needs an inspectable link from the stated hypothesis to the variables and code that test it.
Balyasny explains how it tests those dependencies. Flanagan describes evaluating models both alone and within the full agent environment, with users’ tools, files, and requirements. The tests examine numerical mistakes, missed coverage, unsupported conclusions, and retrieval failures. When an economics result improved unexpectedly, the team reran the evaluation and checked the tasks and scoring before accepting it. A team adapting this method can test the research deliverable with its actual dependencies, alongside the model’s isolated performance.
Cheaper preparation and implementation can expand the research agenda. The constraint may then become the team’s ability to distinguish useful results from errors and statistical accidents. A useful adoption test measures reliable research accepted for use, at a total cost that includes expert review. Where the output informs forecasts or portfolio decisions, evaluate its incremental contribution to the existing process.
Sources
Dates below identify publication or announcement, except where a recording, submission or investment event is specified. Dates without a year are in 2026.
Balyasny: Deal monitoring and shared infrastructure (March 6); deal research and evaluation, Charlie Flanagan interview (September 17).
Citadel: Ken Griffin interview and transcript (recorded June 2; published July 9).
Man Group: AlphaGPT’s research workflow (November 13, 2025); Anthropic partnership, naming the existing AlphaGPT (February 11); Man AHL’s earlier ML applications (2019).
Millennium: Diana Meditz on Digital Twins (September); digital risk analyst development announcement (August). The pages disclose months, without exact days.
Two Sigma: Feature-research productivity and historical-data risks (January 21); Ben Wellington on generating research datasets (July 9).
Bridgewater, investor labels and PAT: AIA Labs’ research program; expert-judgment study and original PDF and Thinking Machines’ account (June 30); PAT presentation page (Bridgewater reports recording on May 19) and original video (published July 24).
Bridgewater, SQL and forecasts: UIUC/AIA Labs text-to-SQL study (August 27); AIA Forecaster technical report, ForecastBench comparisons and MarketLiquid results, Table 7 (first submitted November 10, 2025).
Bridgewater, investment history and returns: External-capital disclosure (July 1, 2024); Reuters’ first-half account and strategy history; Reuters’ January-September returns (October 1).
Historical comparisons: AQR’s active use of ML and stock-portfolio study (November 2024); D. E. Shaw’s Neil Katz on shared AI infrastructure (September 10, 2024).


