Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

82163245326 · Jun 202019922001200920172026
48 results for underlying factors

Method ranks generative models without needing latent factor supervision.

problem Challenges in selecting generative models for qualities like disentanglement.
method Ranking generative models based on training dynamics, without requiring labels for latent factors.
result Method correlates with supervised disentanglement metrics and can predict downstream performance.

Estimates linear model from noisy covariates and instruments using spectral regularization.

problem Estimating a linear model from many noisy covariates and instruments.
method Two-stage least squares with spectral regularization of canonical correlations.
result Upper and lower bounds on estimation error, proving optimality of the method with noisy data.

Sparse NMF with archetypal regularization aims to robustly represent data points.

problem Representing data points as sparse linear combinations of archetypes.
method Sparse NMF with archetypal regularization, introducing strong and weak robustness.
result Theoretical robustness guarantees hold under minimal assumptions.

We develop a new model and algorithms for machine learning-based learning analytics, which estimate a learner's knowledge of the concepts underlying a domain, and content analytics, which estimate the relationships among a collection of questions and those concepts. Our model represents the probability that a learner p…

2013-03-22abs ↗pdf ↗

Rigidity of Wasserstein spaces over Riemannian manifolds

problem Isometric rigidity of L2 Wasserstein spaces over Riemannian manifolds
method Showing L2 Wasserstein spaces are isometrically rigid if and only if their underlying manifolds do not admit a Euclidean de Rham factor
result Isometry of L2 Wasserstein spaces over non-Euclidean manifolds

This study examines the evolving causal structure of equity risk factors.

problem Redundancy and risk contagion in multi-factor strategies during financial crises.
method Causal structure learning methods applied to US equity market data over 29 years.
result Statistically significant sparsifying trend of causal structure during normal times, but densification during financial stress.

We propose a framework for constructing factor models for alpha streams. Our motivation is threefold. 1) When the number of alphas is large, the sample covariance matrix is singular. 2) Its out-of-sample stability is challenging. 3) Optimization of investment allocation into alpha streams can be tractable for a factor …

2014-06-13abs ↗pdf ↗

FACTM combines FA with correlated topic modeling for structured data integration.

problem Integrating structured data modalities like text and single cell sequencing.
method Bayesian FACTM model combining FA and correlated topic modeling with variational inference.
result FACTM outperforms other methods in identifying clusters in structured data and integrating them with simple modalities.

Generative models improve causal effect estimation from observational data.

problem Estimating causal effects from observational data, especially when confounding factors are present.
method Proposes a progressive sequence of Variational Auto-Encoder models to learn underlying factors and causal effects.
result Empirical results show superior performance compared to state-of-the-art approaches.

Proposes a method for tensor completion with sparse factors and missing data.

problem Recovering nonnegative data from noisy observations with missing values.
method Sparse nonnegative Tucker decomposition with 0\ell_0 norm for sparsity, maximum likelihood estimation, and error bounds.
result The method outperforms existing tensor-based or matrix-based methods in nonnegative tensor data completion.

Real-world datasets are often biased with respect to key demographic factors such as race and gender. Due to the latent nature of the underlying factors, detecting and mitigating bias is especially challenging for unsupervised machine learning. We present a weakly supervised algorithm for overcoming dataset bias for de…

2019-10-26abs ↗pdf ↗

Many important schemes in signal processing and communications, ranging from the BCJR algorithm to the Kalman filter, are instances of factor graph methods. This family of algorithms is based on recursive message passing-based computations carried out over graphical models, representing a factorization of the underlyin…

2020-01-31abs ↗pdf ↗

The paper tackles tensor factorization and completion from noisy data.

problem Sparse nonnegative tensor factorization and completion from partial and noisy observations.
method Minimizes the sum of maximum likelihood estimation and tensor 0\ell_0 norm with nonnegativity constraints.
result Error bounds and minimax lower bounds are established for the proposed model.

MGLM models all possible language channel factorizations for improved multilingual generation.

problem Generating multilingual text with flexibility and quality.
method Generative joint distribution model over language channels, marginalizing all possible factorizations.
result MGLM outperforms traditional models in multilingual generation tasks.

Motivated by an application in computational biology, we consider low-rank matrix factorization with {0,1}\{0,1\}-constraints on one of the factors and optionally convex constraints on the second one. In addition to the non-convexity shared with other matrix factorization schemes, our problem is further complicated by a c…

2014-01-23abs ↗pdf ↗

We derive simple return models for several classes of bond portfolios. With only one or two risk factors our models are able to explain most of the return variations in portfolios of fixed rate government bonds, inflation linked government bonds and investment grade corporate bonds. The underlying risk factors have nat…

2010-11-14abs ↗pdf ↗

Kernel Three-Pass Regression Filter improves forecasting efficiency for nonlinear dependencies.

problem Forecasting with high-dimensional predictors and latent factors.
method Developed a new estimator, Kernel Three-Pass Regression Filter (K3PRF), to address nonlinear dependencies.
result Empirically shows significant improvement in long-term forecasting performance.

Most machine learning methods require careful selection of hyper-parameters in order to train a high performing model with good generalization abilities. Hence, several automatic selection algorithms have been introduced to overcome tedious manual (try and error) tuning of these parameters. Due to its very high sample …

2020-01-16abs ↗pdf ↗

We give an online algorithm and prove novel mistake and regret bounds for online binary matrix completion with side information. The mistake bounds we prove are of the form O~(D/γ2)\tilde{O}(D/γ^2). The term 1/γ21/γ^2 is analogous to the usual margin term in SVM (perceptron) bounds. More specifically, if we assume that there i…

2019-06-17abs ↗pdf ↗

A new method uses hyperspherical latent spaces to disentangle data with periodic structures.

problem Disentangling data with periodic or cyclic underlying factors in Euclidean space.
method Diffusion Variational Autoencoder with a modified Evidence Lower Bound.
result The method can recover periodic true factors effectively.

A study finds that only a few factors explain corporate bond risk, rendering extensive bond factor literature redundant.

problem The redundancy of extensive bond factor literature in explaining corporate bond risk premia.
method Bayesian Model Averaging Stochastic Discount Factor analysis of 18 quadrillion models.
result A Bayesian Model Averaging SDF explains risk premia better than low-dimensional models, with an out-of-sample Sharpe ratio of 1.5 to 1.8.

Much research has been devoted to the problem of estimating treatment effects from observational data; however, most methods assume that the observed variables only contain confounders, i.e., variables that affect both the treatment and the outcome. Unfortunately, this assumption is frequently violated in real-world ap…

2020-01-29abs ↗pdf ↗

A new method for analyzing multi-source, multi-way data reduces dimensionality and reveals shared and individual structures.

problem Analyzing multi-source, multi-way data from different high-throughput technologies.
method Multiple Linked Tensor Factorization (MULTIFAC) extending CP decomposition with L2 penalties and EM algorithm for incomplete data.
result MULTIFAC approximates underlying signal, identifies shared and unshared structures, and imputes missing data.

We consider the problem of identifying current coupons for Agency backed To-be-Announced (TBA) Mortgage Backed Securities. In a doubly stochastic factor based model which allows for prepayment intensities to depend upon current and origination mortgage rates, as well as underlying investment factors, we identify the cu…

2015-10-07abs ↗pdf ↗

We propose a deep factorization model for typographic analysis that disentangles content from style. Specifically, a variational inference procedure factors each training glyph into the combination of a character-specific content embedding and a latent font-specific style variable. The underlying generative model combi…

2019-10-02abs ↗pdf ↗

We address the curse of dimensionality in dynamic covariance estimation by modeling the underlying co-volatility dynamics of a time series vector through latent time-varying stochastic factors. The use of a global-local shrinkage prior for the elements of the factor loadings matrix pulls loadings on superfluous factors…

2016-08-30abs ↗pdf ↗

We show how random matrix theory can be applied to develop new algorithms to extract dynamic factors from macroeconomic time series. In particular, we consider a limit where the number of random variables N and the number of consecutive time measurements T are large but the ratio N / T is fixed. In this regime the unde…

2012-01-31abs ↗pdf ↗

New method disentangles shared and private latent factors in multimodal data.

problem Challenges in disentangling shared and private latent factors in multimodal data.
method Proposes a modification to existing multimodal Variational Autoencoders (MMVAE) to better handle modality-specific variation.
result Demonstrates improved robustness of modified MMVAE to modality-specific variation.

The best-known and most commonly used distribution-property estimation technique uses a plug-in estimator, with empirical frequency replacing the underlying distribution. We present novel linear-time-computable estimators that significantly "amplify" the effective amount of data available. For a large variety of distri…

2019-03-04abs ↗pdf ↗