Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

24487296 · Jun 202019922001200920182026
48 results for Explanatory Analytics

BAKR method improves kernel regression for nonlinear and binary classification.

problem Challenges in variable selection for nonlinear kernel regression models.
method Proposes a novel framework for Bayesian approximate kernel regression with effect size analogs.
result BAKR provides a computationally efficient method for nonlinear regression and binary classification.

Bayesian framework explains diverse explanatory values.

problem Understanding and predicting human preferences for explanations.
method Developed a Bayesian account to integrate various explanatory values.
result Core values from psychology, statistics, and philosophy emerge from a common framework.

Hybrid econometrics and ML for food policy priority analysis.

problem Constructing reliable measures of variable importance in econometrics.
method Conventional econometrics combined with advanced machine learning algorithms.
result Demonstrated the applicability of hybrid approach in policy priority issues.

Interactive learning explained to users improves trust and model understanding.

problem Lack of user understanding and trust in interactive learning models.
method Proposes a framework where learners explain interactive queries and predictions to users, using visual explanations.
result Boosts predictive and explanatory powers of and user trust in learned models.

PROD method improves high-dimensional regression by handling strong correlations.

problem Violation of Irrepresentable Condition in LASSO for high-dimensional data.
method PROD procedure based on orthogonal decomposition of design matrix.
result PROD enhances performance of high-dimensional penalized regression.

FFRK automatically extracts features for spatial interpolation without external variables.

problem Spatial interpolation challenges, especially nonstationarity and lack of explanatory variables.
method Feature-Free Regression Kriging (FFRK) method that extracts geospatial features.
result FFRK outperforms classical methods in predicting heavy metal concentrations.

Style Miner generates stable and significant style factors for time series analysis.

problem Finding significant and stable explanatory factors in high-dimensional time series data.
method Proposes a reinforcement learning method to balance explanatory power and stability constraints.
result Outperforms existing methods by a large margin and achieves a 10% gain in R-squared explanatory power.

XGL uses global explanations to guide human supervision in machine learning.

problem Improving model quality through human-machine interaction.
method XGL employs global explanations to guide human selection of informative examples.
result XGL avoids overselling the model's quality and performs comparably to other strategies.

Paper proposes a new method for learning business process representations.

problem Challenges in capturing all useful information in business process data.
method Combines Gramian Angular Fields and Convolutional Neural Networks for representation learning.
result Demonstrates effectiveness of the approach through visualization and multiple process prediction tasks.

Study predicts customer data sharing in Open Banking and explains key factors.

problem Predicting and explaining customer data sharing in Open Banking environments.
method Hybrid data balancing strategy with ADASYN and NEARMISS, XGBoost models, SHAP, CART.
result 91.39% accuracy for inflow and 91.53% for outflow predictions, revealing influential features.

RelatIF selects more intuitive training examples for explaining model predictions.

problem Influence functions identify outliers as explanatory examples, leading to poor explanations.
method RelatIF separates global and local influence, optimizing for local relative to global effects.
result Examples selected by RelatIF are more intuitive than those from influence functions.

A new PCR method using SVD with sparse regularization.

problem Lack of response variable information in traditional PCR.
method One-stage SVD approach with two loss functions and sparse regularization.
result Obtains principal component loadings with response variable information.

Study on identifying probability distributions from random data, showing computable partial learners exist.

problem Identifying probability distributions from random data samples.
method Algorithmic learning theory approach, focusing on computable probability measures and high oracles.
result Characterization of oracles that compute explanatory learners for computable probability measures.

Paper uses interbank contagion to predict U.S. bank defaults, finding it highly explanatory.

problem Predicting U.S. bank defaults using interbank contagion.
method Regression and neural network models were used to analyze U.S. commercial bank data.
result Interbank contagion is highly explanatory in default prediction, often outperforming established metrics.

Paper improves deep learning convergence rates for low-dimensional data.

problem Sub-optimal rates in deep learning due to unrealistic assumptions on intrinsic dimension.
method Introduced an entropic notion of intrinsic dimension for exponential families and demonstrated improved convergence rates.
result Test error scales as O~(n2β2β+dˉ2β(λ))\tilde{\mathcal{O}}\left(n^{-\frac{2β}{2β+ \bar{d}_{2β}(λ)}}\right), improving on best-known rates.

When response variables are nominal and populations are cross-classified with respect to multiple polytomies, questions often arise about the degree of association of the responses with explanatory variables. When populations are known, we introduce a nominal association vector and matrix to evaluate the dependence of …

2011-09-12abs ↗pdf ↗

Principal component regression (PCR) is a two-stage procedure that selects some principal components and then constructs a regression model regarding them as new explanatory variables. Note that the principal components are obtained from only explanatory variables and not considered with the response variable. To addre…

2014-02-26abs ↗pdf ↗

Proposes models to better represent ordinal data with non-unimodal distributions.

problem Real-world ordinal data often have non-unimodal conditional probability distributions.
method Develops approximately unimodal likelihood models to better represent non-unimodal CPDs.
result Proposed models can effectively represent both unimodal and nearly unimodal CPDs.

The paper develops a method to model high-dimensional data with many variables and weak signals.

problem Modeling high-dimensional dependent data with many explanatory variables and low signal-to-noise ratio.
method Penalized regression for high-dimensional data, factor modeling of residuals, high-dimensional white noise testing, projected Principal Component Analysis.
result Established asymptotic properties of the proposed method for high-dimensional data.

Scientists interact with deep learning models to avoid misleading results.

problem Deep neural networks can misinterpret data and achieve high performance by exploiting confounding factors.
method Introduce explanatory interactive learning (XIL) where scientists revise models based on explanations.
result XIL helps prevent misleading results and encourages model trust.

Develops method to assess feature importance in black-box models for unconditional distribution.

problem Lack of methods to analyze feature importance in black-box models for unconditional distribution.
method Approximation method to compute feature importance curves for unconditional distribution.
result Produces sparse and faithful results, computationally efficient.

We investigate entropy as a financial risk measure. Entropy explains the equity premium of securities and portfolios in a simpler way and, at the same time, with higher explanatory power than the beta parameter of the capital asset pricing model. For asset pricing we define the continuous entropy as an alternative meas…

2015-01-06abs ↗pdf ↗

New algorithm discovers causal relationships from observational data efficiently.

problem Inferring direct causal parents from a large set of variables.
method Orthogonal structure search approach, scaling to large graphs, guarantees for nonlinear relationships.
result Significant improvements over existing methods in causal discovery from observational data.

This paper finds that realized kurtosis predicts stock variance better than realized skewness for daily returns.

problem The explanatory power of realized skewness for daily stock returns is limited.
method An extensive empirical analysis of realized skewness and realized kurtosis on daily stock returns and variance.
result Realized kurtosis shows significant forecasting power for stock variance, while realized skewness is less effective for daily returns.

Interpretable representations improve explainable AI by translating complex data into understandable concepts.

problem Many explainers use interpretable representations but overlook their full potential and assumptions.
method An in-depth analysis of interpretable representations for tabular, image, and text data, identifying strengths, weaknesses, and desiderata.
result Linear model quantifies interpretable concepts' influence on black-box predictions, revealing their explanatory properties and manipulability.

Paper tackles robust MM-estimation for high-dimensional data with heavy tails or arbitrary corruption.

problem Sparsity-constrained MM-estimation with heavy-tailed or corrupted data.
method Defines Robust Descent Condition (RDC) and uses Robust Hard Thresholding (IHT) with gradient estimators.
result Robust Hard Thresholding is minimax optimal for kk-sparse high-dimensional linear and logistic regression with heavy tails or arbitrary corruption.

A new framework for time series analysis using state-space learning.

problem Ineffectiveness of traditional Kalman filtering in handling big data and multiple explanatory variables.
method State Space Learning (SSL) framework using statistical learning for high-dimensional regression.
result SSL outperforms traditional methods in subset selection and forecasting accuracy.

Defines devices and agents based on behavior, using computational theory.

problem Differentiating between systems described by mechanical and intentional stances.
method Formal definition of devices and agents, using Bayes' rule to calculate subjective probability based on behavior.
result Bayesian approach to distinguishing between mechanical and intentional systems.

Improved neural network model for predicting latent budgets in compositional data.

problem Predicting response variables in compositional data with non-negativity constraints.
method LBA-NN, a feed forward neural network model that incorporates K-means clustering for interpretation.
result LBA-NN outperforms traditional LBA in prediction accuracy, specificity, recall, and mean square error.

This paper compares two stock factor models in China's A-share market.

problem Contradicting results in existing research on stock factor models.
method Empirical analysis using China's A-share data from 2005-2020, orthogonalizing redundant factors, and 25-group portfolio returns calculation.
result The five-factor model outperforms the three-factor model in explaining excess return rates.

The title is self-explanatory. We aim to give an easy to read and self-contained introduction to the field of harmonic manifolds. Only basic knowledge of Riemannian geometry is required. After we gave the definition of harmonicity and derived some properties, we concentrate on Z. I. Szabó's proof of Lichnerowicz's conj…

2010-07-03abs ↗pdf ↗

Deep learning extracts terrain texture covariates for geostatistical modeling.

problem Improving prediction accuracy in geostatistical modeling using terrain texture data.
method Deep learning approach to automatically derive optimal terrain texture covariates from SRTM 90m DEM.
result Deep learning-derived covariates have strong explanatory power (R-squared around 0.6) for geochemical data.

The paper corrects bias in predictions used as explanatory variables in regression models.

problem Bias in predictions used as explanatory variables in regression models.
method Instrumental variables constructed from multiple splits of the original data.
result The proposed method recovers estimates close to the true values, even in small samples.

Optimizes subset selection in multiple linear regression models.

problem Choosing a subset of variables for regression models to balance fit and complexity.
method Developed mathematical programming models and algorithms for subset selection, tested with branch-and-bound and iterative heuristic approaches.
result Proposed models and algorithms efficiently find optimal or near-optimal solutions.

New research shows input-gradients can be manipulated without changing model's core function, challenging their use for model interpretation.

problem Current methods for model interpretability using input-gradients are flawed due to their arbitrary manipulability.
method Investigated by reinterpreting logits as unnormalized log-densities, proposing novel approximations for score-matching.
result Improving alignment between implicit density model and data distribution enhances gradient structure and explanatory power.

Time series of counts arise in a variety of forecasting applications, for which traditional models are generally inappropriate. This paper introduces a hierarchical Bayesian formulation applicable to count time series that can easily account for explanatory variables and share statistical strength across groups of rela…

2014-05-15abs ↗pdf ↗

Proposes a method to capture similarities and covariances between related tasks in multivariate regression.

problem Predicting multiple response variables with shared explanatory variables and capturing within-group similarities.
method Uses multivariate linear mixed models to estimate coefficients and errors, modeling within-group similarities through joint estimation of covariance matrices.
result The proposed MrRCE method outperforms natural competitors and alternative estimators in various model settings.