New method aggregates GDS analyses of randomly selected interaction models to identify important factors in screening experiments.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new method ranks and selects features without model fitting.
TF-MoDISco (Transcription Factor Motif Discovery from Importance Scores) is an algorithm for identifying motifs from basepair-level importance scores computed on genomic sequence data. This technical note focuses on version v0.5.6.5. The implementation is available at https://github.com/kundajelab/tfmodisco/tree/v0.5.6…
New method disentangles shared and private latent factors in multimodal data.
The paper identifies the minimum mean-variance spanning set and its importance in asset evaluation.
Customer Satisfaction is the most important factors in the industry irrespective of domain. Key Driver Analysis is a common practice in data science to help the business to evaluate the same. Understanding key features, which influence the outcome or dependent feature, is highly important in statistical model building.…
Advertising and feed ranking are essential to many Internet companies such as Facebook and Sina Weibo. Among many real-world advertising and feed ranking systems, click through rate (CTR) prediction plays a central role. There are many proposed models in this field such as logistic regression, tree based models, factor…
In this paper, we employ machine learning techniques to analyze seventeen seasons (1999-2000 to 2015-2016) of NBA regular season data from every team to determine the common characteristics among NBA playoff teams. Each team was characterized by 26 predictor variables and one binary response variable taking on a value …
In recent years, real estate industry has captured government and public attention around the world. The factors influencing the prices of real estate are diversified and complex. However, due to the limitations and one-sidedness of their respective views, they did not provide enough theoretical basis for the fluctuati…
This study examines the evolving causal structure of equity risk factors.
Economic factors significantly influence stock returns, as shown by attribution analysis.
Study analyzes correlation structure in two-factor Hull-White model for XVA calculations.
BSFP method reveals latent patterns in multi-omic data for predicting lung function in HIV-associated OLD.
Deep fundamental factor models are developed to automatically capture non-linearity and interaction effects in factor modeling. Uncertainty quantification provides interpretability with interval estimation, ranking of factor importances and estimation of interaction effects. With no hidden layers we recover a linear fa…
BeMF improves recommendation reliability in recommender systems.
New statistical factors improve portfolio risk estimation.
Marginal maximum likelihood (MML) estimation is the preferred approach to fitting item response theory models in psychometrics due to the MML estimator's consistency, normality, and efficiency as the sample size tends to infinity. However, state-of-the-art MML estimation procedures such as the Metropolis-Hastings Robbi…
During the past few years Boolean matrix factorization (BMF) has become an important direction in data analysis. The minimum description length principle (MDL) was successfully adapted in BMF for the model order selection. Nevertheless, a BMF algorithm performing good results from the standpoint of standard measures in…
Factorization Machine (FM) is a widely used supervised learning approach by effectively modeling of feature interactions. Despite the successful application of FM and its many deep learning variants, treating every feature interaction fairly may degrade the performance. For example, the interactions of a useless featur…
Estimates crypto risk premia using hidden factors and finds significant integration with traditional markets.
Individual risk models need to capture possible correlations as failing to do so typically results in an underestimation of extreme quantiles of the aggregate loss. Such dependence modelling is particularly important for managing credit risk, for instance, where joint defaults are a major cause of concern. Often, the d…
Paper explores subdifferential chain rules for matrix factorization and related machine learning models.
We consider the problem of learning a linear factor model. We propose a regularized form of principal component analysis (PCA) and demonstrate through experiments with synthetic and real data the superiority of resulting estimates to those produced by pre-existing factor analysis approaches. We also establish theoretic…
Develops a framework for identifying mispriced assets through attention factors for statistical arbitrage.
This study quantifies systemic importance in global banks using a continuous framework that amplifies localized shocks.
We study the problem of online influence maximization in social networks. In this problem, a learner aims to identify the set of "best influencers" in a network by interacting with it, i.e., repeatedly selecting seed nodes and observing activation feedback in the network. We capitalize on an important property of the i…
The study examines cross-border lending behavior from G7 countries, showing changes in driving factors after the 2008 financial crisis.
Much research has been devoted to the problem of estimating treatment effects from observational data; however, most methods assume that the observed variables only contain confounders, i.e., variables that affect both the treatment and the outcome. Unfortunately, this assumption is frequently violated in real-world ap…
Deep weight factorization improves neural network training through smooth optimization of sparse penalties.
Hedonic models predict 84-92% of U.S. real estate prices, highlighting environmental factors' impact.
In this paper we propose a computationally efficient algorithm for on-line variable selection in multivariate regression problems involving high dimensional data streams. The algorithm recursively extracts all the latent factors of a partial least squares solution and selects the most important variables for each facto…
This paper investigates factors influencing SGD minima.
Dyadic Data Prediction (DDP) is an important problem in many research areas. This paper develops a novel fully Bayesian nonparametric framework which integrates two popular and complementary approaches, discrete mixed membership modeling and continuous latent factor modeling into a unified Heterogeneous Matrix Factoriz…
Unified framework for nonconvex matrix completion with linearly parameterized factors.
Proposes CC-NMDF for analyzing manifold-valued data.
Machine learning explainability limits identifying causal variables.
We address the problem of unsupervised disentanglement of discrete and continuous explanatory factors of data. We first show a simple procedure for minimizing the total correlation of the continuous latent variables without having to use a discriminator network or perform importance sampling, via cascading the informat…
We propose a family of novel hierarchical Bayesian deep auto-encoder models capable of identifying disentangled factors of variability in data. While many recent attempts at factor disentanglement have focused on sophisticated learning objectives within the VAE framework, their choice of a standard normal as the latent…
Bayesian hierarchical tensor factorization model for international trade flows
Intangible investment becomes a strong predictor of stock returns over time.
Develops a deep multi-factor model for factor investing with clear financial insights.
Multiresolution analysis and matrix factorization are foundational tools in computer vision. In this work, we study the interface between these two distinct topics and obtain techniques to uncover hierarchical block structure in symmetric matrices -- an important aspect in the success of many vision problems. Our new a…
Automates model comparison in probabilistic programming.
Proposes FARM model combining latent factor and sparse regression.
This paper pretends to analyze the importance which the natural advantages and local resources are in the manufacturing industry location, in relation with the "spillovers" effects and industrial policies. To this, we estimate the Rybczynski equation matrix for the various manufacturing industries in Portugal, at regio…
The Importance Weighted Auto Encoder (IWAE) objective has been shown to improve the training of generative models over the standard Variational Auto Encoder (VAE) objective. Here, we derive importance weighted extensions to AVB and AAE. These latent variable models use implicitly defined inference networks whose approx…
Matrix factorization is a key tool in data analysis; its applications include recommender systems, correlation analysis, signal processing, among others. Binary matrices are a particular case which has received significant attention for over thirty years, especially within the field of data mining. Dictionary learning …
The paper tackles three financial issues: time resolution, nonstationarity, and latent factors.