Nonlinear kernel regression models are often used in statistics and machine learning because they are more accurate than linear models. Variable selection for kernel regression models is a challenge partly because, unlike the linear regression setting, there is no clear concept of an effect size for regression coeffici…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Bayesian framework explains diverse explanatory values.
A measure of relative importance of variables is often desired by researchers when the explanatory aspects of econometric methods are of interest. To this end, the author briefly reviews the limitations of conventional econometrics in constructing a reliable measure of variable importance. The author highlights the rel…
New models reduce discrimination in machine learning without sacrificing explanatory bias.
PROD method improves high-dimensional regression by handling strong correlations.
FFRK automatically extracts features for spatial interpolation without external variables.
Style Miner generates stable and significant style factors for time series analysis.
XGL uses global explanations to guide human supervision in machine learning.
Paper proposes a new method for learning business process representations.
Study predicts customer data sharing in Open Banking and explains key factors.
Knockoffs method selects financial factors, controlling false discoveries.
There has recently been a surge of work in explanatory artificial intelligence (XAI). This research area tackles the important problem that complex machines and algorithms often cannot provide insights into their behavior and thought processes. XAI allows users and parts of the internal system to be more transparent, p…
RelatIF selects more intuitive training examples for explaining model predictions.
A new PCR method using SVD with sparse regularization.
Paper uses interbank contagion to predict U.S. bank defaults, finding it highly explanatory.
This text discusses several popular explanatory methods that go beyond the error measurements and plots traditionally used to assess machine learning models. Some of the explanatory methods are accepted tools of the trade while others are rigorously derived and backed by long-standing theory. The methods, decision tree…
OutlierTree detects outliers using decision trees and provides explanations.
Paper improves deep learning convergence rates for low-dimensional data.
When response variables are nominal and populations are cross-classified with respect to multiple polytomies, questions often arise about the degree of association of the responses with explanatory variables. When populations are known, we introduce a nominal association vector and matrix to evaluate the dependence of …
Although interactive learning puts the user into the loop, the learner remains mostly a black box for the user. Understanding the reasons behind queries and predictions is important when assessing how the learner works and, in turn, trust. Consequently, we propose the novel framework of explanatory interactive learning…
Principal component regression (PCR) is a two-stage procedure that selects some principal components and then constructs a regression model regarding them as new explanatory variables. Note that the principal components are obtained from only explanatory variables and not considered with the response variable. To addre…
Statistical detection of a rare class of objects in a two-class classification problem can pose several challenges. Because the class of interest is rare in the training data, there is relatively little information in the known class response labels for model building. At the same time the available explanatory variabl…
Realized moments of higher order computed from intraday returns are introduced in recent years. The literature indicates that realized skewness is an important factor in explaining future asset returns. However, the literature mainly focuses on the whole market and on the monthly or weekly scale. In this paper, we cond…
The problem of inferring the direct causal parents of a response variable among a large set of explanatory variables is of high practical importance in many disciplines. Recent work exploits stability of regression coefficients or invariance properties of models across different experimental conditions for reconstructi…
Proposes models to better represent ordinal data with non-unimodal distributions.
Scientists interact with deep learning models to avoid misleading results.
The paper develops a method to model high-dimensional data with many variables and weak signals.
Develops method to assess feature importance in black-box models for unconditional distribution.
We consider the problem of sparsity-constrained -estimation when both explanatory and response variables have heavy tails (bounded 4-th moments), or a fraction of arbitrary corruptions. We focus on the -sparse, high-dimensional regime where the number of variables and the sample size are related through $…
We study the problem of identifying a probability distribution for some given randomly sampled data in the limit, in the context of algorithmic learning theory as proposed recently by Vinanyi and Chater. We show that there exists a computable partial learner for the computable probability measures, while by Bienvenu, M…
We investigate entropy as a financial risk measure. Entropy explains the equity premium of securities and portfolios in a simpler way and, at the same time, with higher explanatory power than the beta parameter of the capital asset pricing model. For asset pricing we define the continuous entropy as an alternative meas…
According to Dennett, the same system may be described using a `physical' (mechanical) explanatory stance, or using an `intentional' (belief- and goal-based) explanatory stance. Humans tend to find the physical stance more helpful for certain systems, such as planets orbiting a star, and the intentional stance for othe…
Enhances patient failure prediction using dynamic survival models.
New method selects direct causal parents from large sets of variables.
Interpretable representations improve explainable AI by translating complex data into understandable concepts.
A new framework for time series analysis using state-space learning.
PiNets provide faithful explanations for neural networks.
We analyze correlations among stock returns via a series of widely adopted parameters which we refer to as explanatory variables. We subsequently exploit the results to propose a long only quantitative adaptive technique to construct a profitable portfolio of assets which exhibits minor drawdowns and higher recoveries …
Subset selection in multiple linear regression aims to choose a subset of candidate explanatory variables that tradeoff fitting error (explanatory power) and model complexity (number of variables selected). We build mathematical programming models for regression subset selection based on mean square and absolute errors…
Improved neural network model for predicting latent budgets in compositional data.
We consider the problem of predicting several response variables using the same set of explanatory variables. This setting naturally induces a group structure over the coefficient matrix, in which every explanatory variable corresponds to a set of related coefficients. Most of the existing methods that utilize this gro…
We propose a K-sparse exhaustive search (ES-K) method and a K-sparse approximate exhaustive search method (AES-K) for selecting variables in linear regression. With these methods, K-sparse combinations of variables are tested exhaustively assuming that the optimal combination of explanatory variables is K-sparse. By co…
This paper compares two stock factor models in China's A-share market.
Multimodal sensory data resembles the form of information perceived by humans for learning, and are easy to obtain in large quantities. Compared to unimodal data, synchronization of concepts between modalities in such data provides supervision for disentangling the underlying explanatory factors of each modality. Previ…
The title is self-explanatory. We aim to give an easy to read and self-contained introduction to the field of harmonic manifolds. Only basic knowledge of Riemannian geometry is required. After we gave the definition of harmonicity and derived some properties, we concentrate on Z. I. Szabó's proof of Lichnerowicz's conj…
Deep learning extracts terrain texture covariates for geostatistical modeling.
New research shows input-gradients can be manipulated without changing model's core function, challenging their use for model interpretation.
In this paper we present a nonparametric method for extending functional regression methodology to the situation where more than one functional covariate is used to predict a functional response. Borrowing the idea from Kadri et al. (2010a), the method, which support mixed discrete and continuous explanatory variables,…