Paper proposes a new deflation varimax method for vintage factor analysis.
problem Finding a scientifically meaningful low-dimensional representation of data.
method Deflation varimax procedure for orthogonal matrix rotation.
result The proposed method achieves minimax optimal factor loading estimation.
LSTMs outperform DFM in nowcasting COVID-19 economic variables.
problem Timely estimation of macroeconomic variables during the pandemic.
method Comparison of LSTM and DFM performance on three variables (export values, volumes, and services exports).
result LSTMs outperformed DFM in two-thirds of variable/quarter combinations.
Paper extends quantile factor analysis with probabilistic methods for better economic policy and financial condition prediction.
problem Improving accuracy in economic and financial condition prediction.
method Probabilistic quantile factor analysis with regularization and variational approximations.
result The probabilistic estimator outperforms a recent loss-based estimator in many cases.
Factor analysis provides linear factors that describe relationships between individual variables of a data set. We extend this classical formulation into linear factors that describe relationships between groups of variables, where each group represents either a set of related variables or a data set. The model also na…
Factor Engine simplifies financial factor computation and analysis in Python.
problem Efficient computation and analysis of financial factors.
method Modular, extensible Python library with decorators, integrates with data science ecosystem.
result Mispricing factors computed by Factor Engine and Stata implementation are highly similar.
Survey of factor analysis, PCA, variational inference, and VAE.
problem Dimensionality reduction and generative modeling of data.
method Variational inference, factor analysis, probabilistic PCA, and VAE.
result Derivation and explanation of ELBO, EM, and closed-form solutions.
We introduce a factor analysis model that summarizes the dependencies between observed variable groups, instead of dependencies between individual variables as standard factor analysis does. A group may correspond to one view of the same set of objects, one of many data sets tied by co-occurrence, or a set of alternati…
NCFA uses deep learning and causal discovery to analyze complex data.
problem Analyzing complex, interdependent data with causal relationships.
method NCFA combines latent causal discovery and variational autoencoders.
result NCFA outperforms standard VAEs in sparsity, complexity, and causal interpretability.
We propose a nonparametric Bayesian factor regression model that accounts for uncertainty in the number of factors, and the relationship between factors. To accomplish this, we propose a sparse variant of the Indian Buffet Process and couple this with a hierarchical model over factors, based on Kingman's coalescent. We…
We present a novel factor analysis method that can be applied to the discovery of common factors shared among trajectories in multivariate time series data. These factors satisfy a precedence-ordering property: certain factors are recruited only after some other factors are activated. Precedence-ordering arise in appli…
VarFA efficiently estimates student skill levels with uncertainty for adaptive testing.
problem Efficiently estimating student skill levels with uncertainty for adaptive testing.
method VarFA uses variational inference to extend factor analysis models for educational data.
result VarFA efficiently handles large datasets and produces uncertainty estimates.
Deep model learns complex latent codes without assuming factor structure.
problem Learning latent codes with complex, non-factorial distributions.
method Deep generative factor analysis with beta process prior and stochastic EM algorithm.
result Preliminary results show model can approximate complex distributions.
Sensitivity analysis for individualized effects in OTRs with binary risk factors.
problem Addressing omitted confounding in individualized effects of OTRs.
method Simulation-based sensitivity analysis to simulate unmeasured confounders.
result Benchmarking the strength of omitted confounding for binary risk factors.
This work connects LLE, factor analysis, and probabilistic PCA through a stochastic perspective.
problem Exploring the theoretical connection between LLE, factor analysis, and probabilistic PCA.
method Solving the stochastic linear reconstruction of LLE using expectation maximization.
result LLE, factor analysis, and probabilistic PCA are shown to be connected through a stochastic perspective.
Method for factor analysis in short panels without assuming sphericity or Gaussianity.
problem Factor analysis in short panels without assuming sphericity or Gaussianity.
method Pseudo maximum likelihood method and asymptotically uniformly most powerful invariant test.
result Systematic risk explains a large part of cross-sectional total variance in bear markets but is not spanned by observed factors.
NeuralFactors uses deep learning to improve factor analysis in equity modeling.
problem Enhancing classical factor models for better risk forecasting and portfolio construction.
method Introduces a novel machine-learning approach (NeuralFactors) that outputs factor exposures and returns, trained using variational autoencoders.
result NeuralFactors outperforms prior approaches in log-likelihood performance and computational efficiency.
Proposes CC-NMDF for analyzing manifold-valued data.
problem Nonlinear structure in manifold-valued data requires new analysis methods.
method Curvature-corrected nonnegative manifold data factorization (CC-NMDF) with an iterative algorithm.
result Demonstrates CC-NMDF on real-world diffusion tensor MRI data.
Proposes a new machine learning-based method for conjoint analysis.
problem Testing the importance of factors in conjoint analysis with interactions.
method Conditional randomization test based on machine learning algorithms.
result Validates the importance of factors in conjoint analysis without model specification.
Factor analysis or sometimes referred to as variable analysis has been extensively used in classification problems for identifying specific factors that are significant to particular classes. This type of analysis has been widely used in application such as customer segmentation, medical research, network traffic, imag…
Study develops sector rotation models using factor and fundamental analysis.
problem Understanding and predicting sector shifts in financial markets.
method Systematic sector classification, factor analysis, and fundamental metrics evaluation.
result Developed predictive models with notable predictive capabilities.
Proposes a robust factor analysis for matrix data.
problem Robust factor analysis for matrix data with heavy-tailed or contaminated data.
method Bilinear factor analysis based on the matrix-variate t distribution. result Significantly higher breakdown point than traditional methods.
Optimal tensor PCA for estimating factors and loadings in high-dimensional panel data.
problem Estimating factors and loadings in high-dimensional panel data with non-negligible correlations.
method Tensor Principal Component Analysis (TPCA) for estimating factors and loadings in a tensor factor model.
result Simple TPCA is optimal for strong factors and can be improved for weak factors with alternating least-squares iterations.
Paper relaxes factor analysis for noisy data, improving robustness.
problem Challenges in finding robust low dimensional approximations for data with heteroskedastic noise.
method Introduces a relaxed version of Minimum Trace Factor Analysis (MTFA) as a convex optimization method.
result Effective at not overfitting to heteroskedastic perturbations and addressing common issues in factor analysis.
New method aggregates GDS analyses of randomly selected interaction models to identify important factors in screening experiments.
problem Erroneous conclusions from main-effects models in screening experiments.
method Gauss-Dantzig Selector Aggregation over Random Models (GDS-ARM).
result Identifies important factors by aggregating GDS analyses of randomly selected interaction models.
DMSTF models spatio-temporal data with deep Markov priors.
problem Analyzing nonlinear multimodal spatio-temporal dynamics.
method Deep Markov spatio-temporal factorization with stochastic variational inference.
result DMSTF outperforms other methods in predictive performance and clustering.
Study uses ML and causal analysis to predict student performance factors.
problem Understanding socio-academic and economic factors affecting student performance.
method Employed machine learning techniques and causal analysis on 1,050 student profiles.
result Ridge Regression achieved robust predictions with MAE of 0.12 and MSE of 0.024.
Authors improve accuracy analysis for portfolio optimization with multiple timescale factors.
problem Asymptotic accuracy of portfolio optimization approximations for general utility functions and two timescale factors.
method Construct sub- and super-solutions to fully nonlinear problem.
result Rigorous justification of accuracy for portfolio optimization with general utility functions and two timescale factors.
This paper compares two stock factor models in China's A-share market.
problem Contradicting results in existing research on stock factor models.
method Empirical analysis using China's A-share data from 2005-2020, orthogonalizing redundant factors, and 25-group portfolio returns calculation.
result The five-factor model outperforms the three-factor model in explaining excess return rates.
New method for factor analysis using nuclear and ℓ0 norms.
problem Finding a low-rank plus sparse decomposition from noisy covariance matrix.
method Formulated an optimization problem with nuclear norm, ℓ0 norm, and KL divergence. Used alternating minimization algorithm. result Algorithm effectively decomposes covariance matrices in synthetic and real datasets.
Dynamic factor analysis reveals insights into Philippine stock market dynamics.
problem Understanding complex stock market dynamics.
method Dynamic factor model using Kalman method and maximum likelihood estimation.
result Common factors extracted from the model represent market trends and volatility.
This paper uses Factored Latent Analysis (FLA) to learn a factorized, segmental representation for observations of tracked objects over time. Factored Latent Analysis is latent class analysis in which the observation space is subdivided and each aspect of the original space is represented by a separate latent class mod…
H-GAT improves stock selection by capturing complex higher-order stock relations and integrating both technical and fundamental analysis.
problem Stock selection difficulty and lack of comprehensive analysis.
method Higher-order Graph Attention Network (H-GAT) that incorporates both technical and fundamental analysis.
result H-GAT outperforms existing methods in stock selection metrics.
A new property fixes look-ahead bias in backtesting and trading pipelines.
problem Fixing look-ahead bias in backtesting and trading pipelines.
method Developed a pipeline calculus separating availability from reference time, and a type-and-effect system for the value-independent fragment.
result The check scales linearly and catches all leaks, including those missed by differential and tiling detectors.
We analyze linear factor models for asset pricing panels.
problem Characterizing cross-sectional and inter-temporal properties of returns and factors.
method Conditional means and covariances, review of Kozak and Nagel (2024) conditions.
result Low-dimensional factor portfolios can span efficient portfolios in unbalanced panels.
Factor analysis has proven to be a relevant tool for extracting tissue time-activity curves (TACs) in dynamic PET images, since it allows for an unsupervised analysis of the data. Reliable and interpretable results are possible only if considered with respect to suitable noise statistics. However, the noise in reconstr…
Matrix factorizations and their extensions to tensor factorizations and decompositions have become prominent techniques for linear and multilinear blind source separation (BSS), especially multiway Independent Component Analysis (ICA), NonnegativeMatrix and Tensor Factorization (NMF/NTF), Smooth Component Analysis (Smo…
Study finds whitepaper narratives do not predict market factor structure.
problem Predicting market behavior from cryptocurrency whitepaper claims.
method Zero-shot NLP classification combined with CP tensor decomposition of market data.
result Weak alignment between whitepaper claims and market statistics and latent factors.
Industry-scale recommendation systems have become a cornerstone of the e-commerce shopping experience. For Etsy, an online marketplace with over 50 million handmade and vintage items, users come to rely on personalized recommendations to surface relevant items from its massive inventory. One hallmark of Etsy's shopping…
In 2002, the UCR time series classification archive was first released with sixteen datasets. It gradually expanded, until 2015 when it increased in size from 45 datasets to 85 datasets. In October 2018 more datasets were added, bringing the total to 128. The new archive contains a wide range of problems, including var…
Sparse GFA identifies disease factors in FTD subgroups.
problem Heterogeneity in neurological disorders hinders understanding and treatment.
method Sparse Group Factor Analysis (GFA) with regularised horseshoe priors.
result Identified latent disease factors differentially expressed in FTD subgroups.
On a periodic basis, publicly traded companies are required to report fundamentals: financial data such as revenue, operating income, debt, among others. These data points provide some insight into the financial health of a company. Academic research has identified some factors, i.e. computed features of the reported d…
Study finds significant premium for low-beta stocks in firm-level idiosyncratic return distributions.
problem Understanding the role of common idiosyncratic quantile factors in asset pricing.
method Quantile factor analysis to extract common idiosyncratic quantile factors with asymmetric pricing effects.
result Significant premium for innovations to the lower-tail factor: high-beta stocks outperform low-beta stocks by around 7-8% per year.
Enhanced AI analysis predicts S&P 500 stock dynamics using various financial metrics.
problem Predicting S&P 500 stock performance with complex interplay of factors.
method Advanced financial metrics, machine learning, and integration of traditional and modern analytics.
result Enhanced predictive accuracy in market behavior and investment strategies.
Style Miner generates stable and significant style factors for time series analysis.
problem Finding significant and stable explanatory factors in high-dimensional time series data.
method Proposes a reinforcement learning method to balance explanatory power and stability constraints.
result Outperforms existing methods by a large margin and achieves a 10% gain in R-squared explanatory power.
The paper develops a new model for high-dimensional spatial arbitrage pricing.
problem Estimating spatial interactions in high-dimensional asset pricing.
method Integrates spatial interactions with multi-factor analysis using generalized shrinkage Yule-Walker (SYW) estimation.
result Established asymptotic properties for high-dimensional spatial arbitrage pricing models.
Paper develops models to forecast private equity fund cash flows.
problem Limited literature on illiquid alternative asset cash flow forecasting.
method Develops benchmark model and two novel approaches (direct vs. indirect) using LSTM/GRU models and macroeconomic indicators.
result Direct model performs better and aligns with actual cash flows, but indirect model's performance is less clear.
Quantitative Investment, built on the solid foundation of robust financial theories, is at the center stage in investment industry today. The essence of quantitative investment is the multi-factor model, which explains the relationship between the risk and return of equities. However, the multi-factor model generates e…
New method explains high-dimensional sphere data with latent factors.
problem Understanding intricate dependence structure in high-dimensional sphere data.
method Exploratory factor analysis of the projected normal distribution with a fast alternating expectation profile conditional maximization algorithm.
result Uniformly excellent results on various data types, including tweets, brain imaging, and cancer gene expression.