AI Portfolio Research

Historical Research Evidence

A full development record for model performance, signal diagnostics, execution, horizon selection, tail risk, and model specifications. SPY remains visible wherever portfolio outcomes are compared.

Evidence statusDevelopment only

Portfolio performance

Exact daily overlapping-sleeve accounting. Use the cost selector to compare zero-cost sensitivity, declared implementation costs, and stress cases.

SPY always visible

Cumulative return

All four models and full-investment SPY

Drawdown

Focused model versus SPY

Return distribution and execution

Daily return shape, turnover, trade counts, and cost sensitivity are kept separate from signal statistics.

Daily return distribution

Focused model versus SPY

Daily turnover

Twenty-session overlapping sleeves

Signal quality

Cross-sectional rank IC, top-40 precision, decile ordering, yearly behavior, and excess returns are evaluated before portfolio compounding.

Decision-level information coefficient

Same-date Spearman correlation

Mean return by score decile

Decile 10 contains the highest model scores

Horizon search and matched controls

The original 5-60 session screen is shown beside the later exact non-overlapping 40-126 session attribution. Point estimates are not presented as an optimal horizon.

Equal-budget horizon screen

Gross CAGR, implementation CAGR, and rank IC

Selection contribution

Frozen Ridge top 40 versus matched random and schedule-matched SPY

40 and 50 sessions show the clearest selection contribution versus sector-matched random portfolios.
Every tested attribution horizon trails schedule-matched SPY before transaction costs.
No horizon has passed the preregistered standard for a unique, benchmark-beating optimum.

Tail-risk A/B experiment

The fully invested Ridge-rank control is compared with the frozen downside treatment and SPY. Reduced tail exposure is not enough if winner rejection destroys returns.

Treatment inactive

Cumulative return

Control, tail treatment, and SPY

Risk and classification outcomes

Precision, recall, CVaR, drawdown, and rejected winners

Selected model specifications

Fold-level feature counts and selection stability. This section does not substitute selection frequency for SHAP.