Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

2356 · Jul 201919922001200920172026
48 results for Russell 1000

Membership in the Russell 1000 and 2000 Indices is based on a ranking of market capitalization in May. Each index is separately value weighted such that firms just inside the Russell 2000 are comparable in size to firms just outside (i.e. at the bottom of the Russell 1000) but have much higher index weights. These feat…

2015-09-01abs ↗pdf ↗

AI predicts stock winners with 2.43 Sharpe ratio, but returns are highly concentrated.

problem Predicting stock returns with AI, focusing on identifying top winners.
method Deployed a state-of-the-art LLM to autonomously search the web for stock attractiveness, avoiding look-ahead bias.
result AI can generate alpha by identifying top winners, but returns are highly concentrated.

Deep fundamental factor models are developed to automatically capture non-linearity and interaction effects in factor modeling. Uncertainty quantification provides interpretability with interval estimation, ranking of factor importances and estimation of interaction effects. With no hidden layers we recover a linear fa…

2019-03-18abs ↗pdf ↗

New method accurately reconstructs Russell 3000 index, revealing crowded portfolios.

problem Crowding in index portfolios during reconstitution events.
method Developed a Python package for accurate index reconstruction using CRSP US Stock data.
result Annual Russell 3000 portfolios are more crowded than quarterly ones, suggesting lower transaction costs.

DPLS improves asset pricing by capturing non-linear risk factor structures.

problem Estimating asset pricing models with non-linear risk factor structures.
method Deep Partial Least Squares (DPLS) for dynamic and flexible factor modeling.
result DPLS models outperform linear models in asset pricing, capturing non-linear risk factor interactions.

Paper uses machine learning to analyze stock market anomalies, predicting drift direction and portfolio performance.

problem Capturing dynamics of Post-Earnings-Announcement Drift (PEAD) using machine learning.
method Uses Extreme Gradient Boosting (XGBoost) with genetic algorithm optimization to analyze PEAD dynamics.
result Demonstrates how PEAD dynamics are influenced by different factors across sectors and quarters.

Pseudo-Anosov subgroups in surface bundles over tori are convex cocompact.

problem Understanding the structure of pseudo-Anosov subgroups in surface bundles over tori.
method Using the Birman exact sequence to show convex cocompactness.
result Finitely generated, purely pseudo-Anosov subgroups are convex cocompact in surface bundles over tori.

In this article we study Weinstein structures endowed with a Lefschetz fibration in terms of the Legendrian front projection. First we provide a systematic recipe for translating from a Weinstein Lefschetz bifibration to a Legendrian handlebody. Then we present several applications of this technique to symplectic topol…

2016-10-21abs ↗pdf ↗

Investigates the relationship between US money supply and asset indices over 2001-2019.

problem Determining the relationship between US money supply and asset indices growth.
method Information entropy methodology applied to US asset indices (Property, Russell 2000, S&P 500, NASDAQ) over 2001-2019.
result Growth in US broad money supply is the main determinant of US asset indices growth, especially the NASDAQ and Russell 2000.

The paper provides conditions for amalgamation of certain subgroups and preserves convexity properties.

problem Conditions for amalgamation of subgroups in hierarchically hyperbolic groups.
method Study of amalgamation conditions and preservation of convexity properties.
result Conditions under which amalgamation preserves hierarchical quasiconvexity and strong quasiconvexity.

Study reveals a hidden cost in derivatives markets through option-implied discount factors.

problem The hidden cost in derivatives markets, not visible in price space.
method Minute-level NBBO data on options, reduced-form specification linking carry gap to implementation risk, trading frictions, and financial conditions.
result An annualized carry gap exists, linked to implementation risk and financial conditions.

Study of graphs interpolating curve and pants graphs, providing formulae and geometry classifications.

problem Understanding the large-scale geometry of graphs connecting curve and pants graphs.
method Developed explicit formulae for quasi-flat ranks and classified geometries using twist-free graphs of multicurves.
result Explicit formulae for quasi-flat ranks and classification of geometries into hyperbolic, relatively hyperbolic, and thick cases.

Stable subgroups identified in genus two handlebody group.

problem Characterizing stable subgroups in genus two handlebody group.
method Proving genus two handlebody group is hierarchically hyperbolic, using quasi-isometric embedding properties and Hamenstädt-Hensel construction.
result Stable subgroups identified and characterized.

A very simple heuristic approach to the unfolding problem will be described. An iterative algorithm starts with an empty histogram and every iteration aims to add one entry to this histogram. The entry to be added is selected according to a criteria which includes a χ2χ^2 test and a regularization. After a relatively s…

2014-10-17abs ↗pdf ↗

The paper optimizes portfolios with transaction costs in a large asset universe.

problem Optimizing portfolios with transaction costs in a large asset universe.
method Mean-variance optimization with nonconvex penalty for proportional and quadratic transaction costs.
result The proposed models show satisfactory performance and highlight the importance of transaction costs.

In the last twenty-five years (1990-2014), algorithmic advances in integer optimization combined with hardware improvements have resulted in an astonishing 200 billion factor speedup in solving Mixed Integer Optimization (MIO) problems. We present a MIO approach for solving the classical best subset selection problem o…

2015-07-11abs ↗pdf ↗

Feature extraction from financial data is one of the most important problems in market prediction domain for which many approaches have been suggested. Among other modern tools, convolutional neural networks (CNN) have recently been applied for automatic feature selection and market prediction. However, in experiments …

2018-10-21abs ↗pdf ↗

WeSpeR speeds up non-linear shrinkage for high-dimensional weighted covariance.

problem Computing non-linear shrinkage formulas for high-dimensional weighted sample covariance.
method Derive extit{WeSpeR} algorithm using asymptotic sample spectrum properties.
result Significantly speeds up non-linear shrinkage in dimensions higher than 1000.

LLMs struggle to generate random numbers from statistical distributions, leading to biased results in applications.

problem LLMs' inability to generate random numbers accurately from specified distributions.
method Dual-protocol design: Batch Generation and Independent Requests, benchmarking 11 models across 15 distributions.
result Sampling fidelity degrades with distributional complexity and horizon, leading to systematic biases in downstream applications.

Paper extends Cohen's method to compute Jones polynomial for certain braid subfamilies.

problem Computing Jones polynomial for specific knot families.
method Using weighted adjacency matrices and determinants for certain subfamilies of braid groups.
result Jones polynomial can be computed in polynomial time for certain subfamilies of braid groups.

Deep Neural Networks (DNNs) are vulnerable to adversarial attacks, especially white-box targeted attacks. One scheme of learning attacks is to design a proper adversarial objective function that leads to the imperceptible perturbation for any test image (e.g., the Carlini-Wagner (C&W) method). Most methods address targ…

2019-05-25abs ↗pdf ↗

New econometric results for financial duration models under varying tail behaviors.

problem Estimation and inference challenges in financial durations models with random event counts.
method Analysis of likelihood estimators for ACD models, focusing on tail behavior and stationarity.
result Asymptotic normality breaks down for tail indices smaller than one, leading to mixed Gaussian estimators with non-standard rates of convergence.

Stock prediction has always been attractive area for researchers and investors since the financial gains can be substantial. However, stock prediction can be a challenging task since stocks are influenced by a multitude of factors whose influence vary rapidly through time. This paper proposes a novel approach (Word2Vec…

2019-02-13abs ↗pdf ↗

We introduce a new regression framework, Gaussian process regression networks (GPRN), which combines the structural properties of Bayesian neural networks with the non-parametric flexibility of Gaussian processes. This model accommodates input dependent signal and noise correlations between multiple response variables,…

2011-10-19abs ↗pdf ↗

Paper designs a lossy compression method for lossless prediction.

problem Ensuring high performance on predictive tasks with minimal data.
method Characterizes bit-rate requirements for invariant transformations, designs unsupervised objectives for neural compressors.
result Achieves substantial rate savings on ImageNet compared to JPEG without compromising classification performance.

This paper presents a data-driven approach to model planar pushing interaction to predict both the most likely outcome of a push and its expected variability. The learned models rely on a variation of Gaussian processes with input-dependent noise called Variational Heteroscedastic Gaussian processes (VHGP) that capture…

2017-04-10abs ↗pdf ↗

NYSE stock prices show persistent correlations over years, exploitable through arbitrage strategies.

problem Predicting and exploiting long-term price correlations in NYSE stocks.
method Analyzed 1000 NYSE stocks over 5 years, measured discrepancies from Brownian motion, and tested arbitrage strategies.
result 45% of a stock's 1-hour returns variance is explained by cross-correlations with other stocks, especially during high volatility periods.

We present a system and a set of techniques for learning linear predictors with convex losses on terascale datasets, with trillions of features, {The number of features here refers to the number of non-zero entries in the data matrix.} billions of training examples and millions of parameters in an hour using a cluster …

2011-10-19abs ↗pdf ↗

DMIDAS improves long-term forecasting accuracy in healthcare and electricity data.

problem Challenging long-term forecasting accuracy and computational complexity.
method Smoothness regularization and mixed data sampling techniques integrated into NBEATS architecture.
result Improves prediction accuracy by 5% on long forecasting horizons (1000 timestamps) compared to state-of-the-art models.