Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

86171257342 · Jun 202019922001200920172026
48 results for ANOVA approximation

Two ANOVA-based algorithms boost random Fourier feature models for function approximation.

problem Approximating high-dimensional functions with low-order interactions.
method Utilizes ANOVA decomposition to learn low-order functions and index sets of important variables.
result Significantly reduces approximation error compared to existing methods.

The paper uses transformed ANOVA to identify important fire detection variables.

problem Identifying key variables for forest fire detection.
method Developed a complete orthonormal system for standard normal distribution, applied Z-score transformation, and used ANOVA approximation.
result Attribute ranking reveals important variables for fire detection.

The paper explores SHAP scores and their connection to functional ANOVA, highlighting challenges in approximations.

problem Estimating SHAP scores and understanding their limitations.
method Using the connection to functional ANOVA, the paper outlines the challenges in SHAP approximations and their relation to feature distribution and ANOVA terms.
result Challenges in SHAP approximations are primarily due to feature distribution and the number of ANOVA terms estimated.

We solve the ANOVA decomposition for categorical inputs.

problem Lack of a closed-form expression for ANOVA decomposition with categorical dependent variables.
method Bridge functional analysis with discrete Fourier analysis to derive a closed-form decomposition.
result Closed-form decomposition for categorical inputs without assumptions.

Bayesian-TPNN improves ANOVA-TPNN for detecting higher-order components.

problem Difficulty in incorporating higher-order components in ANOVA-TPNN due to computational and memory constraints.
method Bayesian inference procedure for functional ANOVA model with TPNN basis functions.
result Bayesian-TPNN detects higher-order components with reduced computational cost.

Paper introduces efficient methods for estimating cross-partial derivatives and sensitivity indices.

problem Efficiently estimating cross-partial derivatives and sensitivity indices in complex models.
method Using randomized points and constraints, the paper develops estimators with optimal convergence rates and low bias.
result The estimators achieve optimal rates of convergence and do not suffer from the curse of dimensionality.

Meta-ANOVA simplifies complex models for better interpretability.

problem Complex models are hard to interpret, limiting their use in fields needing accountability.
method Transforms black-box models into interpretable ANOVA models by screening unnecessary interactions.
result Meta-ANOVA provides an interpretable model for any prediction model, proving asymptotic consistency.

PED-ANOVA efficiently calculates HP importance in arbitrary subspaces.

problem Understanding the role of different hyperparameters in arbitrary subspaces.
method Derive a novel f-ANOVA formulation for arbitrary subspaces and use Pearson divergence (PED) for a closed-form calculation of HP importance.
result Demonstrates successful identification of important HPs in different subspaces.

Unified framework for feature-based explanations using ANOVA and game theory.

problem Differences between feature-based explanations methods limit their applicability.
method Introduces a unified framework combining fANOVA and cooperative game theory.
result Uncovered similarities and differences between various explanation techniques.

We provide a unified view of additive explanations for dependent inputs.

problem Challenges in obtaining a tractable representation and estimating the decomposition for dependent inputs.
method Combining Hilbert space methods with generalized functional ANOVA, we build an explicit decomposition Riesz Basis.
result Proposed a simple yet powerful algorithm to estimate the decomposition from data.

Improved SVM classification with interpretable features from scattered data.

problem Classification of scattered data points in high-dimensional spaces.
method Truncated ANOVA decomposition for sparse feature selection; use of trigonometric or wavelet feature maps.
result Better classification accuracy and interpretability with 1\ell_1-norm regularization.

This work uses ANOVA to understand how different factors contribute to test error in machine learning models.

problem Understanding why overparametrized models generalize well despite potentially fitting noise.
method Analysis of variance (ANOVA) to decompose test error into components of variance.
result The interaction between training samples and initialization can dominate variance, and there are phase transitions in variance behavior.

Improved Gaussian process models for interpretable predictions.

problem Complex responses require high-dimensional interaction terms in additive Gaussian processes.
method Orthogonal additive kernel (OAK) with orthogonality constraint on additive functions.
result OAK models achieve similar or better predictive performance with fewer terms, retaining interpretability.

New ML algorithms improve model interpretability without sacrificing performance.

problem Lack of interpretability in complex machine learning models.
method Developed new algorithms based on fANOVA framework, including GAMI-Lin-T and GAMI-Net.
result GAMI-Lin-T and GAMI-Net perform comparably to EBM and better in interpretability.

Itô processes are the most common form of continuous semimartingales, and include diffusion processes. This paper is concerned with the nonparametric regression relationship between two such Itô processes. We are interested in the quadratic variation (integrated volatility) of the residual in this regression, over a un…

2006-11-09abs ↗pdf ↗

We introduce a semi-parametric Bayesian model for survival analysis. The model is centred on a parametric baseline hazard, and uses a Gaussian process to model variations away from it nonparametrically, as well as dependence on covariates. As opposed to many other methods in survival analysis, our framework does not im…

2016-11-02abs ↗pdf ↗

Hypothesis testing is one of the most common types of data analysis and forms the backbone of scientific research in many disciplines. Analysis of variance (ANOVA) in particular is used to detect dependence between a categorical and a numerical variable. Here we show how one can carry out this hypothesis test under the…

2019-03-01abs ↗pdf ↗

Shallow trees in ensemble models make models more interpretable and sometimes better.

problem Lack of transparency in high-performing tree ensemble models.
method Developed an interpretation algorithm to convert tree ensembles into functional ANOVA representations. Proposed strategies to enhance interpretability.
result Shallow trees in ensemble models can lead to better generalization performance and improved interpretability.

A new measure of causal influence quantifies intrinsic contributions in DAGs.

problem Quantifying intrinsic causal contributions in Directed Acyclic Graphs (DAGs).
method Recursive decomposition of node contributions, structure-preserving interventions, Shapley symmetrization.
result A measure of intrinsic causal contribution that is invariant to node relabeling.

Study federates measurement of demographic disparities from quantile sketches.

problem Misalignment of fairness goals with siloed data collection and privacy regulations.
method Federated auditing of demographic parity through score distributions, using Wasserstein--Frechet variance and quantile summaries.
result Proposes a one-shot, communication-efficient protocol to estimate global disparity and its decomposition.

A new method quickly identifies key variables and interactions.

problem Identifying key variables and interactions in high-dimensional data.
method Kernel trick for sparse orthogonal decomposition in O(# covariates) time.
result Outperforms existing methods for large, high-dimensional data sets.

A new GP framework for discovering unknown functions and hypergraph structure.

problem Discovering unknown functions and hypergraph structure in data.
method Interpretable Gaussian Process framework for Type 3 problems.
result Polynomial complexity for data-driven discovery of unknown functions and hypergraph structure.

The focus of this paper is the efficient computation of counterparty credit risk exposure on portfolio level. Here, the large number of risk factors rules out traditional PDE-based techniques and allows only a relatively small number of paths for nested Monte Carlo simulations, resulting in large variances of estimator…

2016-08-03abs ↗pdf ↗

TSL learns separable models to avoid signal cancellation and off-support extrapolation.

problem Signal cancellation and off-support extrapolation in additive models.
method Tensor Separation Learning (TSL) via stagewise greedy procedure with orthogonal refitting.
result TSL avoids information loss caused by marginalizing higher-order interactions.

New decompositions misattribute differences between populations, even when outcomes are identical.

problem Misattribution of differences between populations using common functional decompositions.
method Extending the Kitagawa-Oaxaca-Blinder decomposition to nonlinear functional decompositions.
result Functional ANOVA and Accumulated Local Effects can misattribute differences even when outcomes are identical in two populations.

A new method selects variables efficiently for fast and accurate dynamic system identification.

problem Efficiently selecting variables for scalable Gaussian processes.
method Forward variable selection using Karhunen-Loève decomposition and Gibbs sampling.
result Method yields competitive accuracies and inference times for dynamic systems.

Among interpretable machine learning methods, the class of Generalised Additive Neural Networks (GANNs) is referred to as Self-Explaining Neural Networks (SENN) because of the linear dependence on explicit functions of the inputs. In binary classification this shows the precise weight that each input contributes toward…

2019-08-16abs ↗pdf ↗

We extend multi-way, multivariate ANOVA-type analysis to cases where one covariate is the view, with features of each view coming from different, high-dimensional domains. The different views are assumed to be connected by having paired samples; this is a common setup in recent bioinformatics experiments, of which we a…

2009-12-16abs ↗pdf ↗

Neuroimaging research has predominantly drawn conclusions based on classical statistics, including null-hypothesis testing, t-tests, and ANOVA. Throughout recent years, statistical learning methods enjoy increasing popularity, including cross-validation, pattern classification, and sparsity-inducing regression. These t…

2016-03-06abs ↗pdf ↗

We propose a sequential learning policy for noisy discrete global optimization and ranking and selection (R\&S) problems with high dimensional sparse belief functions, where there are hundreds or even thousands of features, but only a small portion of these features contain explanatory power. We aim to identify the spa…

2015-03-18abs ↗pdf ↗

SAE-FiRE extracts key financial info from long documents, improving earnings surprise predictions.

problem Predicting earnings surprises from long, redundant financial documents.
method Sparse Autoencoder feature selection to filter out noise and identify key dimensions.
result SAE-FiRE significantly outperforms baseline approaches in financial datasets.