New statistical factors improve portfolio risk estimation.
problem Improving estimation of portfolio risk using new statistical factors.
method Matrix factor models and statistical methods (partial F test, double selection LASSO).
result New statistical factors add explanatory power in asset pricing.
Factor models are a class of powerful statistical models that have been widely used to deal with dependent measurements that arise frequently from various applications from genomics and neuroscience to economics and finance. As data are collected at an ever-growing scale, statistical machine learning faces some new cha…
Develops a framework for identifying mispriced assets through attention factors for statistical arbitrage.
problem Identifying mispriced assets in statistical arbitrage trading.
method Uses conditional latent factors learned from firm characteristic embeddings to identify time-series signals and form a trading strategy.
result Achieves an out-of-sample Sharpe ratio above 4 on the largest U.S. equities over a 24-year period.
New framework for interpretable firm characteristics factors.
problem Creating statistically efficient and economically interpretable factors from firm characteristics.
method Grouping related characteristics and deriving one factor per group, combining economic intuition with data-driven clustering.
result Parsimonious, transparent factors outperform benchmarks in out-of-sample tests.
New tests for identifying the number of latent factors in short panels with small time dimensions.
problem Determining the number of latent factors in short panels with small time dimensions.
method Eigenvalue tests based on variance-covariance matrices of asset returns, with assumptions on spherical errors or instrumental variables for factor betas.
result Established asymptotic distributional results and proposed a novel statistical test for weak factors.
Study tests if equity factors explain Bitcoin's risk and returns.
problem Explaining Bitcoin's risk and return with equity factors.
method Applied statistical methods to test Fama-French factors on Bitcoin's excess returns.
result Fama-French factors have explanatory power on Bitcoin's risk and returns.
Survey on factor models and their applications in econometrics.
problem Estimating low-rank structures in high-dimensional models.
method Low-rank recovery techniques for factor model estimation.
result New insights into factor model applications in econometrics.
Paper compares neural networks and classical statistics for dementia prediction, highlighting interpretability of classical methods.
problem Tackles the challenge of interpreting risk factors for dementia prediction.
method Compares neural networks and classical statistics for dementia prediction.
result Classical statistics provide clearer interpretation of risk factors compared to neural networks.
Optimizes Bayesian priors for matrix factorization without posterior inference.
problem Selecting optimal priors for Bayesian models in machine learning.
method Prior predictive distribution and virtual statistics matching user-provided or observed data statistics.
result Analytically determines hyperparameters for Poisson factorization models.
Study improves interpretability in generative models by disentangling latent variables in scientific datasets.
problem Extracting generative factors from complex, high-dimensional datasets in unsupervised or semi-supervised settings.
method Introducing Aux-VAE, a novel architecture within the VAE framework, which disentangles latent variables by guiding them with auxiliary variables.
result Aux-VAE achieves disentanglement with minimal modifications to the standard VAE loss function, validated on multiple datasets.
Divide-and-conquer method speeds sparse factorization for large matrices.
problem Sparse factorization of large matrices for statistical learning.
method Statistical problem formulation, divide-and-conquer approach, stagewise learning.
result Efficient algorithm with lower complexity than existing methods.
Proposes a new tensor factorization model for better link prediction in knowledge graphs.
problem Lack of information in treating missing and non-existing relations equally in tensor factorization models.
method Introduces a binary tensor factorization model with probit link to address the issue.
result Shows improved prediction accuracy and interpretability compared to existing models.
We introduce a new factor model for log volatilities that performs dimensionality reduction and considers contributions globally through the market, and locally through cluster structure and their interactions. We do not assume a-priori the number of clusters in the data, instead using the Directed Bubble Hierarchical …
Machine learning helps estimate risk premiums of stocks without knowing their factors.
problem Estimate risk premiums of stocks without knowing their underlying factors.
method Used elastic-net machine learning to project stock returns onto peers and construct replicate portfolios.
result Unique stocks have higher SARP and excess returns than ubiquitous stocks.
The paper analyzes statistical arbitrage using a factor model of equity returns.
problem Analyzing and trading statistical arbitrage strategies in equity markets.
method Conditional factor model, state space framework, online risk premia estimation, mean reversion trades.
result The model outperforms other methods in statistical arbitrage trading strategies over a 29-year period.
We develop a monitoring procedure to detect changes in a large approximate factor model. Letting r be the number of common factors, we base our statistics on the fact that the (r+1)-th eigenvalue of the sample covariance matrix is bounded under the null of no change, whereas it becomes spiked under cha…
Bayes factors and relative belief ratios are compared as measures of statistical evidence.
problem Which measure of evidence is more appropriate: Bayes factors or relative belief ratios?
method Comparison of Bayes factors and relative belief ratios, considering properties and restrictions.
result Relative belief ratio has better properties as a measure of evidence.
T-Rex uses EM to fit robust factor models in noisy data.
problem Robustly fitting factor models in high-dimensional data with heavy tails and outliers.
method Expectation-Maximization (EM) algorithm based on Tyler's M-estimator for elliptical distributions.
result Demonstrates robustness in direction-of-arrival estimation and subspace recovery.
New theory for PCA under weak latent factors, improving inference and testing.
problem Statistical inference for PCA with weak latent factors and cross-sectional dependence.
method Comprehensive estimation and inference theory for PCA under nearly minimal factor strength, non-asymptotic.
result Asymptotic normality of PCA-based estimator for N≍T with SNR growth rate. The paper derives statistics of multi-factor functions from their Fourier transforms.
problem Deriving statistics of multi-factor functions from Fourier transforms.
method Developed an m-Coefficient/Index Annihilation Theorem to analyze the moments of a function from its Fourier transform.
result The mth moment of a function becomes a series of terms, each with precisely m Fourier coefficients, and the indices sum to zero.
Method estimates exogenous and endogenous factors from event times.
problem Estimating factors influencing event occurrence.
method Combines inhomogeneous Poisson and Hawkes processes, fits using free energy minimization.
result Four regimes identified based on factor detection.
Data-driven factor graphs improve BCJR detection robustness.
problem Implementing BCJR detection with accurate channel model knowledge.
method Learn factor graph using machine learning from labeled data.
result BCJRNet learns to implement BCJR detection from small training sets.
Develops a statistical model for SOFR term structure in incomplete markets.
problem Incomplete liquidity and completeness in SOFR derivatives market.
method Statistical model incorporating macroeconomic factors and jumps in SOFR rates.
result Model is well-suited for risk management and derivatives pricing.
Enhances risk model with new statistical factors.
problem Missing information in existing risk models.
method Maximum likelihood estimation to refine and add new factors.
result Captures structure missed by original model.
Study finds whitepaper narratives do not predict market factor structure.
problem Predicting market behavior from cryptocurrency whitepaper claims.
method Zero-shot NLP classification combined with CP tensor decomposition of market data.
result Weak alignment between whitepaper claims and market statistics and latent factors.
Method controls extrapolation in prediction profiles for statistical and machine learning models.
problem Avoiding invalid predictions due to extrapolation in prediction profiles.
method Genetic algorithm optimization over constrained factor regions.
result Optimal factor settings without constraint are often invalid and extrapolated.
Improved fantasy football performance predictor using human feedback.
problem Lack of external factors in traditional statistical models.
method Combining statistical data with human feedback from various sources.
result Model outperformed regular statistical predictors by over 300 points.
Unified framework for disentangled representations using mechanistic independence.
problem Identifiability of disentangled latent factors under statistical dependencies.
method Introduces mechanistic independence to characterize latent factors by their actions on observed variables, proposing various independence criteria.
result Establishes conditions for identifiability of latent subspaces without statistical assumptions.
The present paper provides a study of high-dimensional statistical arbitrage that combines factor models with the tools from stochastic control, obtaining closed-form optimal strategies which are both interpretable and computationally implementable in a high-dimensional setting. Our setup is based on a general statisti…
Introduces m-connecting imset and factorization for ADMG models.
problem Handling latent confounding in DAG models.
method Introduces m-connecting imset and m-connecting factorization criterion for ADMG models.
result Equivalence of m-connecting factorization criterion to global Markov property.
New algorithm improves clustering accuracy without sacrificing scalability.
problem Improving clustering accuracy for large datasets.
method Nonnegative low-rank semidefinite programming with Burer-Monteiro factorization.
result Significantly smaller mis-clustering errors compared to existing methods.
Recommender systems relying on latent factor models often appear as black boxes to their users. Semantic descriptions for the factors might help to mitigate this problem. Achieving this automatically is, however, a non-straightforward task due to the models' statistical nature. We present an output-agreement game that …
We present a very fast algorithm for general matrix factorization of a data matrix for use in the statistical analysis of high-dimensional data via latent factors. Such data are prevalent across many application areas and generate an ever-increasing demand for methods of dimension reduction in order to undertake the st…
Non-negative matrix factorization (NMF) is a new knowledge discovery method that is used for text mining, signal processing, bioinformatics, and consumer analysis. However, its basic property as a learning machine is not yet clarified, as it is not a regular statistical model, resulting that theoretical optimization me…
NeuralFactors uses deep learning to improve factor analysis in equity modeling.
problem Enhancing classical factor models for better risk forecasting and portfolio construction.
method Introduces a novel machine-learning approach (NeuralFactors) that outputs factor exposures and returns, trained using variational autoencoders.
result NeuralFactors outperforms prior approaches in log-likelihood performance and computational efficiency.
EFS uses LLMs to optimize sparse portfolios by evolving alpha factors.
problem Sparse portfolio optimization in dynamic market regimes.
method Evolutionary feedback loop with LLM-generated alpha factors.
result Significantly outperforms baselines in diverse datasets.
Risk statistic is a critical factor not only for risk analysis but also for financial application. However, the traditional risk statistics may fail to describe the characteristics of regulator-based risk. In this paper, we consider the regulator-based risk statistics for portfolios. By further developing the propertie…
Based on a new atomic norm, we propose a new convex formulation for sparse matrix factorization problems in which the number of nonzero elements of the factors is assumed fixed and known. The formulation counts sparse PCA with multiple factors, subspace clustering and low-rank sparse bilinear regression as potential ap…
We empirically analyze the price and liquidity responses to trade signs, traded volumes and signed traded volumes. Utilizing the singular value decomposition, we explore the interconnections of price responses and of liquidity responses across the whole market. The statistical characteristics of their singular vectors …
A network-based approach identifies financial factors from asset interactions, explaining market dynamics.
problem Characterizing joint financial asset behavior through underlying drivers.
method Modeling market as coupled iterated maps, where asset returns depend on past returns and interactions.
result Stable patterns of co-movement (financial factors) emerge from asset interactions, explaining asset variance.
We propose a 4-factor model for overnight returns and give explicit definitions of our 4 factors. Long horizon fundamental factors such as value and growth lack predictive power for overnight (or similar short horizon) returns and are not included. All 4 factors are constructed based on intraday price and volume data a…
New algorithm learns FMDP structure while minimizing regret.
problem Regret minimization in FMDPs with unknown structure.
method Optimism in face of uncertainty principle combined with statistical structure learning.
result First algorithm to learn FMDP structure while minimizing regret.
In a very high-dimensional vector space, two randomly-chosen vectors are almost orthogonal with high probability. Starting from this observation, we develop a statistical factor model, the random factor model, in which factors are chosen at random based on the random projection method. Randomness of factors has the con…
Deep learning improves Bayes factor computation for likelihood-free models.
problem Computing Bayes factors for likelihood-free models is challenging.
method Proposes a deep learning estimator of Bayes factors using simulated data.
result Establishes consistency of the Deep Bayes Factor estimator.
Unified Bayesian framework improves clinical trial hypothesis testing.
problem Lack of transparency and inability to quantify evidence in traditional P-values.
method Interval null hypothesis framework combined with Bayes factor-based tests.
result Bayesian interval hypothesis testing ensures frequentist error control and interpretability.
Latent Dirichlet allocation (LDA) is useful in document analysis, image processing, and many information systems; however, its generalization performance has been left unknown because it is a singular learning machine to which regular statistical theory can not be applied. Stochastic matrix factorization (SMF) is a res…
Corporate bond factor research is flawed due to measurement errors and ex-post filtering.
problem Replication crisis in corporate bond factor research.
method Analysis of 108 signals across nine thematic clusters, correction of transaction prices and return filtering.
result Majority of previously documented factors do not produce statistically significant alphas after correction.
Improved statistical inference for adaptive Thompson Sampling.
problem Statistical inference challenges in Thompson Sampling.
method Inflating posterior variance in Thompson Sampling.
result Asymptotically normal estimates of arm means with logarithmic regret increase.