Skip to main content
QUICK REVIEW

[Paper Review] Designing Agentic AI-Based Screening for Portfolio Investment

Mehmet Caner, Agostino Capponi|arXiv (Cornell University)|Mar 24, 2026
Stock Market Forecasting Methods0 citations
TL;DR

The paper builds an agentic AI framework with two LLM-based stock screeners (fundamental and sentiment) plus high-dimensional precision-matrix weighting, introduces sensible screening, and shows Sharpe-ratio gains on S&P 500 data (2020–2024 and 2015–2024).

ABSTRACT

We introduce a new agentic artificial intelligence (AI) platform for portfolio management. Our architecture consists of three layers. First, two large language model (LLM) agents are assigned specialized tasks: one agent screens for firms with desirable fundamentals, while a sentiment analysis agent screens for firms with desirable news. Second, these agents deliberate to generate and agree upon buy and sell signals from a large portfolio, substantially narrowing the pool of candidate assets. Finally, we apply a high-dimensional precision matrix estimation procedure to determine optimal portfolio weights. A defining theoretical feature of our framework is that the number of assets in the portfolio is itself a random variable, realized through the screening process. We introduce the concept of sensible screening and establish that, under mild screening errors, the squared Sharpe ratio of the screened portfolio consistently estimates its target. Empirically, our method achieves superior Sharpe ratios relative to an unscreened baseline portfolio and to conventional screening approaches, evaluated on S&P 500 data over the period 2020--2024.

Motivation & Objective

  • Motivate portfolio construction as a two-stage process: stock screening and weight optimization.
  • Develop a multi-agent AI architecture combining fundamental and sentiment analysis for stock screening.
  • Introduce the concept of sensible screening and theoretical Sharpe-ratio consistency.
  • Empirically validate the framework against benchmarks and alternative methods across market regimes.

Proposed method

  • Three-layer architecture with two specialized AI agents (LLM-S for fundamentals, FinBERT for sentiment) and a consensus rule to select assets.
  • A high-dimensional precision matrix estimation procedure (e.g., nodewise regression, residual nodewise regression, POET, deep learning-based methods, nonlinear shrinkage) to determine optimal weights.
  • Treat the number of assets in the portfolio as a random variable realized by screening and prove that the screened portfolio’s squared Sharpe ratio consistently estimates the target under sensible screening (Theorem A.1).
  • Rolling retraining schedules (annual for LLM-S, monthly for FinBERT) to capture slow narratives and fast news, with a two-out-of-three consensus rule.
  • Comparison against multiple baselines (purely quantitative, single-agent LLMs, conventional screening, and human-in-the-loop systems) to isolate the value of the multi-agent design.

Experimental results

Research questions

  • RQ1How does integrating LLM-based fundamental screening and sentiment analysis impact portfolio performance relative to benchmarks?
  • RQ2Can a high-dimensional precision-matrix approach yield superior weight formation for screens with a random cardinality?
  • RQ3Does the concept of sensible screening ensure consistent Sharpe-ratio estimation when screening introduces errors?
  • RQ4What is the performance of the agentic AI framework across different market regimes and over extended horizons?

Key findings

  • The multi-agent AI framework achieves Sharpe ratios higher than the market benchmark in most method–objective combinations.
  • Over 2020–2024, the best agentic configuration attains an annualized Sharpe ratio of 1.1867 (up from 0.6324 for the S&P 500), an 88% improvement over the index.
  • Over 2015–2024, the Agentic AI architecture dominates, with a peak Sharpe ratio of 0.9429 versus 0.7298 for the market benchmark.
  • The best agentic configuration yields an annualized return of 36.34% over the five-year window, compared to 19.99% for the best purely quantitative baseline.
  • In many configurations, human analyst inputs degrade performance due to biases, highlighting AI-based screening’s value.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.