GraphLIME explains GNN models by selecting key features locally.
problem Explaining the effectiveness of GNN models is challenging due to complex nonlinear transformations.
method GraphLIME uses HSIC Lasso for nonlinear feature selection in GNN models.
result GraphLIME provides more descriptive explanations than existing methods.
SISR improves feature attribution in complex payoff schemes.
problem Distorted feature attributions due to non-additive payoff functions and high-dimensional feature spaces.
method Sparse Isotonic Shapley Regression (SISR) learns a monotonic transformation to restore additivity and enforces L0 sparsity.
result SISR achieves strong support recovery and stable attributions across various payoff schemes.
This work explains RL policies using causal models, revealing important patterns and failures.
problem Understanding why RL policies succeed or fail in complex, high-dimensional systems.
method Developed a nonlinear Causal Model Reduction framework to learn simplified causal models from RL policy actions and rewards.
result The approach can uncover important behavioral patterns and failure modes in trained RL policies.
This paper surveys various methods for dimensionality reduction and nearest neighbor search.
problem Efficiently reducing high-dimensional data to lower dimensions while preserving essential information.
method Linear and nonlinear random projections, including sparse random projections, random Fourier Features, and Random Kitchen Sinks.
result Various methods for dimensionality reduction and nearest neighbor search are explained and compared.
Introduces nonlinear splittings on fibre bundles for generalizing connections.
problem Generalizing connections on fibre bundles.
method Definition and properties of nonlinear splittings, including affine, homogeneous, and principal splittings.
result Curvature map defined for nonlinear splittings, linking to nonholonomic systems and magnetic Lagrangian systems.
Method estimates bivariate causal models using normalising flows and variational Gaussian process regression.
problem Lack of explainability in AI models, especially in causal mechanisms.
method Combination of normalising flows for density estimation and variational Gaussian process regression for post-nonlinear models.
result Method better explains cause-effect pairs than simple additive noise models.
New analysis explains pathology of deep Gaussian processes.
problem Pathology of deep Gaussian processes reduces learning capacities with increased layers.
method Study nonlinear dynamic systems corresponding to DGPs, derive recurrence relations.
result Provide tighter bounds and rate of convergence for dynamic systems.
The geometry of the target space of an N=(2,2) supersymmetry sigma-model carries a generalized Kahler structure. There always exists a real function, the generalized Kahler potential K, that encodes all the relevant local differential geometry data: the metric, the B-field, etc. Generically this data is given by nonlin…
Extends importance sampling to nonlinear models using adjoint operators.
problem Lack of tools for identifying important data points in nonlinear models.
method Introduces adjoint operator for nonlinear maps, generalizes norm and leverage scores.
result Generalized scores provide approximation guarantees for nonlinear mappings.
This is a survey article on recent progress of comparison geometry and geometric analysis on Finsler manifolds of weighted Ricci curvature bounded below. Our purpose is two-fold: Give a concise and geometric review on the birth of weighted Ricci curvature and its applications; Explain recent results from a nonlinear an…
In this article we will analyse how to compute the contribution of each input value to its aggregate output in some nonlinear models. Regression and classification applications, together with related algorithms for deep neural networks are presented. The proposed approach merges two methods currently present in the lit…
We prove a priori estimates for a class of transverse fully nonlinear equations on Sasakian manifolds and give some geometric applications such as the transversion Calabi-Yau theorem for transverse balanced and (strongly) Gauduchon metrics. We also explain that similar results hold on compact oriented, taut, transverse…
We explain a simple construction of solutions to a family of PDE's in two dimensions which includes that defining zero scalar curvature Kahler metrics, with two Killing fields, and the affine maximal equation.
Despite outstanding contribution to the significant progress of Artificial Intelligence (AI), deep learning models remain mostly black boxes, which are extremely weak in explainability of the reasoning process and prediction results. Explainability is not only a gateway between AI and society but also a powerful tool t…
Deep SSMs use neural networks to identify complex systems.
problem Identifying nonlinear systems with high uncertainty.
method Deep state space models with neural networks.
result Deep SSMs outperform traditional methods on benchmarks.
The construction of a linear connection on a pullback bundle from a connection on a vector bundle is explained in terms of fiberwise linear approximation. This procedure clarifies the geometric meaning of the linearized connection as well as the associated parallel transport and curvature.
The study explains how market-makers' hedging affects stock volatility during gamma-squeeze events.
problem Endogenous volatility amplification in option markets during gamma-squeeze events.
method Developed a theoretical framework linking hedging behavior and market turbulence, incorporating beta-normalized volatility.
result Low-beta stocks amplify volatility more during gamma-squeeze events.
This review explores methods to explain deep neural networks and their applications.
problem Understanding the decision-making process of deep neural networks.
method Overview of interpretability methods, theoretical foundations, and comparative evaluations.
result Demonstrates the effectiveness of explainable AI in various applications.
Paper explains distance-based classifiers using neural network structures.
problem Making distance-based classifiers explainable.
method Uncovering latent neural network structures in distance-based classifiers.
result Novel explanation approach outperforms baselines.
Theory explains how deep nets learn features from data.
problem Understanding how deep neural networks learn features from data.
method Developed a noise-nonlinearity phase diagram and a mechanical theory.
result Links feature learning across layers to generalization.
Bell's theorem shows quantum correlations can't be explained by classical causal models, even with some measurement dependence.
problem Quantum correlations violate classical causal models.
method Using causal networks, the study bounds the level of measurement dependence and derives nonlinear Bell inequalities.
result Quantum correlations can't be explained by classical causal models even with some measurement dependence.
Proposes a new derivative concept for nonlinear DRO problems.
problem Optimizing nonlinear functions in probability space with distributionally robust optimization.
method Introduces Gateaux derivative for smoothness and proposes a Frank-Wolfe algorithm.
result Validates theoretical results on portfolio selection problems with numerical validation.
GNNs improve semi-supervised node regression, but why? We explain.
problem Understanding when and why GNNs succeed in semi-supervised node regression.
method Aggregate-and-readout model encompassing message passing architectures, least-squares estimation over GNNs with linear graph convolutions and a deep ReLU readout.
result Sharp non-asymptotic risk bound separating approximation, stochastic, and optimization errors.
Theory explains deep nonlinear networks' plateaus and transitions.
problem Understanding long plateaus and feature acquisition transitions in deep nonlinear networks.
method Derived an exact identity for Frobenius norms, classified activation functions, and reduced matrix flow to a scalar ODE.
result Escape time law τ⋆=Θ(ε−(r−2)) for deep nonlinear networks, where r is the number of bottleneck layers. This paper uses NLDT to find interpretable control rules from complex DRL policies.
problem Complex, non-interpretable policies from black-box AI methods.
method Evolutionary optimization of NLDT for hierarchical control rules.
result Interpretable control rules with similar performance to black-box DRL.
Multi-view data are increasingly prevalent in practice. It is often relevant to analyze the relationships between pairs of views by multi-view component analysis techniques such as Canonical Correlation Analysis (CCA). However, data may easily exhibit nonlinear relations, which CCA cannot reveal. We aim to investigate …
New measure LMN explains neural network grokking.
problem Delayed generalization after memorization in neural networks.
method Defined LMN to measure network complexity, showing LMN correlates with test losses linearly.
result LMN reveals intriguing XOR network behavior and is a promising complexity measure.
Explains non-lorentzian theories and their dynamics.
problem Understanding non-lorentzian kinematics and dynamics.
method Review of kinematical spacetimes, construction of particle dynamics actions, discussion of gravity theories and field theories.
result Introduction and analysis of non-lorentzian gravity and field theories.
The paper investigates causal relationships in heart failure prediction using machine learning.
problem Understanding the causal relationships between clinical variables and heart failure.
method Proposes a new computational framework for causal structure discovery (CSD) of mixed-type clinical variables for binary disease outcomes.
result Feature importance from nonlinear classifiers strongly correlates with causal strength of variables, but not differentiating cause and effect.
Method extracts features from signals for classification with explainability.
problem Lack of interpretability in signal classification models.
method Combining scattering transform and multiclass logistic regression with zeroth-order optimization.
result Uncovered the meaning of scattering transform coefficients.
The paper explains financial volatility using simple news-driven models.
problem Financial volatility's fat-tailed and clustered nature.
method Simple models of traders' attitudes towards news.
result Simple models explain fat-tailed and clustered volatility.
The book explains deep learning theory and how networks learn nontrivial representations.
problem Understanding and optimizing deep neural networks.
method Developed RG flow to characterize signal propagation, solved layer-to-layer equations, and analyzed representation learning.
result Predictions of trained networks are nearly-Gaussian, with depth-to-width ratio controlling deviations.
DF2M uses deep neural networks within a factor model for high-dimensional functional time series forecasting.
problem Forecasting high-dimensional functional time series with explainability and accuracy.
method Bayesian nonparametric model based on Indian Buffet Process and multi-task Gaussian Process, incorporating a deep kernel function.
result DF2M provides better explainability and superior predictive accuracy compared to conventional deep learning models.
Graph neural networks (GNNs) have emerged as a powerful tool for nonlinear processing of graph signals, exhibiting success in recommender systems, power outage prediction, and motion planning, among others. GNNs consists of a cascade of layers, each of which applies a graph convolution, followed by a pointwise nonlinea…
Deep learning searches for nonlinear factors for predicting asset returns. Predictability is achieved via multiple layers of composite factors as opposed to additive ones. Viewed in this way, asset pricing studies can be revisited using multi-layer deep learners, such as rectified linear units (ReLU) or long-short-term…
We give a physical derivation of generalized Kahler geometry. Starting from a supersymmetric nonlinear sigma model, we rederive and explain the results of Gualtieri regarding the equivalence between generalized Kahler geometry and the bi-hermitean geometry of Gates-Hull-Rocek. When cast in the language of supersymmetri…
We solve the long standing problem of finding an off-shell supersymmetric formulation for a general N = (2, 2) nonlinear two dimensional sigma model. Geometrically the problem is equivalent to proving the existence of special coordinates; these correspond to particular superfields that allow for a superspace descriptio…
Nonlinear methods such as Deep Neural Networks (DNNs) are the gold standard for various challenging machine learning problems, e.g., image classification, natural language processing or human action recognition. Although these methods perform impressively well, they have a significant disadvantage, the lack of transpar…
RFMs transition from linear to nonlinear under specific input-label correlation.
problem Understanding the transition from linear to nonlinear behavior in RFMs.
method Analyzing RFMs under spiked covariance designs, characterizing the interaction between anisotropy and input-label correlation.
result The RFM generalization error is governed by the strength of input-label correlation, leading to a clear nonlinear advantage above a specific boundary.
New method improves model explainability and accuracy with low computational cost.
problem Improving model explainability and accuracy in classification models.
method Distributionally robust optimization to learn sparse ensembles of rule sets.
result Improves model performance on various metrics compared to competing methods.
Many natural systems, such as neurons firing in the brain or basketball teams traversing a court, give rise to time series data with complex, nonlinear dynamics. We can gain insight into these systems by decomposing the data into segments that are each explained by simpler dynamic units. Building on switching linear dy…
Machine-learning models have been recently used for detecting malicious Android applications, reporting impressive performances on benchmark datasets, even when trained only on features statically extracted from the application, such as system calls and permissions. However, recent findings have highlighted the fragili…
Study moduli spaces of elliptic PDEs using derived C∞-geometry.
problem Representability of moduli spaces of solutions of elliptic PDEs.
method Derived C∞-geometry, stacks of relative jets, nonlinear Fredholm analysis. result Moduli stack of solutions is relatively representable by quasi-smooth derived C∞-schemes. Complex nonlinear models such as deep neural network (DNNs) have become an important tool for image classification, speech recognition, natural language processing, and many other fields of application. These models however lack transparency due to their complex nonlinear structure and to the complex data distributions…
New decompositions misattribute differences between populations, even when outcomes are identical.
problem Misattribution of differences between populations using common functional decompositions.
method Extending the Kitagawa-Oaxaca-Blinder decomposition to nonlinear functional decompositions.
result Functional ANOVA and Accumulated Local Effects can misattribute differences even when outcomes are identical in two populations.
A new method quickly identifies key variables and interactions.
problem Identifying key variables and interactions in high-dimensional data.
method Kernel trick for sparse orthogonal decomposition in O(# covariates) time.
result Outperforms existing methods for large, high-dimensional data sets.
Explains Bernstein theorems for various geometric PDEs.
problem Bernstein problem for minimal surface, Monge-Ampère, and special Lagrangian equations.
method Expository review of existing theorems and systems.
result Discussion of Bernstein theorems for different geometric PDEs.
Several machine learning models, including neural networks, consistently misclassify adversarial examples---inputs formed by applying small but intentionally worst-case perturbations to examples from the dataset, such that the perturbed input results in the model outputting an incorrect answer with high confidence. Ear…