Develops statistical guarantees for neural networks with regularization.
problem Lack of comprehensive mathematical theories for neural networks.
method General statistical guarantee for least-squares with regularizers.
result Prediction error increases sub-linearly in layers, logarithmically in parameters.
Paper proves mathematically that poisoning datasets can be detected.
problem Detecting data poisoning attacks in datasets.
method Mathematical definition and Conformal Separability Test.
result Dataset poisoning can be effectively detected.
Exponential functionals of Brownian motion have been extensively studied in financial and insurance mathematics due to their broad applications, for example, in the pricing of Asian options. The Black-Scholes model is appealing because of mathematical tractability, yet empirical evidence shows that geometric Brownian m…
DKPS provides guarantees for synthetic data from Transformer models, improving downstream tasks.
problem Lack of labeled data for building performant AI models.
method Data Kernel Perspective Space (DKPS) for mathematical analysis of synthetic data quality.
result Concrete statistical guarantees for the quality of transformer model outputs.
ViTaX provides formal guarantees for targeted explanations in safety-critical systems.
problem Need trustworthy explanations for safety-critical deep neural networks.
method Formal reachability analysis for targeted, semifactual explanations.
result First method to provide formally guaranteed explanations of model resilience.
CLIP controls neural network stability by bounding Lipschitz constants.
problem Neural networks lack mathematical guarantees of stability, especially to adversarial examples.
method Develops a variational regularization method (CLIP) to control the Lipschitz constant of neural networks.
result CLIP provides a tighter bound on the actual Lipschitz constant compared to layer-wise methods.
Approximate dynamic programming is a popular method for solving large Markov decision processes. This paper describes a new class of approximate dynamic programming (ADP) methods- distributionally robust ADP-that address the curse of dimensionality by minimizing a pessimistic bound on the policy loss. This approach tur…
This study optimizes crypto-market trading conditions without assuming convexity.
problem Optimizing crypto-market trading conditions without convexity.
method Rigorous mathematical analysis of constant function market makers under quasilinear trade functions.
result Quasilinear trade functions can replicate convex functions' robustness against arbitrage.
Improves data-driven reachability estimation for complex systems.
problem Estimating reachable states in complex dynamical systems with unknown parameters.
method Uses Christoffel functions and conformal prediction to improve sample efficiency and robustness.
result Guaranteed convergence to the true reach set with improved sample efficiency and robustness.
The paper provides statistical guarantees for sparse deep learning.
problem Understanding the potential and limitations of sparse deep learning.
method Develops statistical guarantees for different types of sparsity in sparse deep learning.
result Statistical guarantees for sparse deep learning with mild dependence on network widths and depths.
Survey of mathematical foundations for reinforcement learning.
problem Design and analysis of modern reinforcement learning algorithms.
method Organizes mathematical structures from probability, optimization, and operator theory.
result Unified mathematical entry point for researchers in various fields.
We prove Transformers can learn diverse Gröbner bases.
problem Training Transformers for Gröbner basis computation.
method Prove generality of dataset generation algorithm; propose extended algorithm.
result Datasets are sufficiently general for diverse Gröbner bases learning.
In this study we introduce a new technique for symbolic regression that guarantees global optimality. This is achieved by formulating a mixed integer non-linear program (MINLP) whose solution is a symbolic mathematical expression of minimum complexity that explains the observations. We demonstrate our approach by redis…
The resilience of low-degree Rademacher chaos is studied, providing probabilistic lower bounds.
problem Understanding how much a Rademacher chaos can withstand adversarial sign-flips without significant probability changes.
method Probabilistic lower-bound guarantees for the resilience of Rademacher chaos of arbitrary degree.
result Probabilistic lower-bound guarantees for the resilience of Rademacher chaos of arbitrary degree, especially meaningful for constant degree.
Optimizes group testing for COVID-19 to reduce test numbers.
problem Minimizing tests for accurate infection detection.
method Bayesian approach with genetic algorithms and sub-modularity.
result Greedy-adaptive method provides theoretical guarantees.
CIVET method provides robustness guarantees for VAEs under adversarial attacks.
problem Certified probabilistic guarantees for VAEs in safety-critical applications.
method Bounding worst-case VAE error by error on support sets at the latent layer.
result CIVET outperforms state-of-the-art methods in robustness and standard performance.
Develops statistical guarantees for image-to-image regression models.
problem Current image-to-image regression models lack statistical guarantees for model mistakes and hallucinations.
method Uncertainty quantification techniques with rigorous statistical guarantees for image-to-image regression problems.
result Derives uncertainty intervals around each pixel with formal mathematical guarantees.
This paper uses SLT to ensure learning guarantees in CD detection.
problem Lack of learning guarantees in CD detection algorithms.
method Adapting SLT assumptions to CD scenarios to ensure learning guarantees.
result Ensured learning guarantees in CD detection algorithms.
The book explores alternatives to worst-case analysis for algorithm performance.
problem Providing strong worst-case guarantees for many algorithms is impossible.
method Surveying and detailing various nuanced analysis approaches.
result More nuanced analysis approaches are needed for fundamental problems.
Transformers learn topic structure through embedding and attention mechanisms.
problem Understanding how transformers capture semantic structure in text.
method Combination of mathematical analysis and experiments on Wikipedia and synthetic data.
result Embedding and attention layers encode topic structure in transformers.
Tangles improve clustering in various datasets.
problem Clustering diverse datasets efficiently and accurately.
method Tangles aggregate cuts to identify dense structures, leading to soft cluster characterization.
result Tangle framework generates hierarchical soft dendrograms for cluster exploration.
Deep learning models viewed through tame geometry for convergence guarantees.
problem Understanding convergence guarantees in deep learning models.
method Introducing tame geometry concepts and tools for nonsmooth nonconvex settings.
result Illustrates tame geometry as a natural framework for AI systems, especially deep learning.
Method guarantees coherent factuality for language model outputs in reasoning tasks.
problem Ensuring correctness of language model outputs in reasoning tasks.
method Developed a conformal-prediction-based method applied to subgraphs within a deducibility graph.
result Achieved coherent factuality across target coverage levels, 90% on stricter definition.
Currently, pension providers are running into trouble mainly due to the ultra-low interest rates and the guarantees associated to some pension benefits. With the aim of reducing the pension volatility and providing adequate pension levels with no guarantees, we carry out mathematical analysis of a new pension design in…
This paper provides a PAC-Bayesian bound for CVaR in machine learning.
problem Learning algorithms minimizing CVaR of empirical loss.
method Generalization bound of PAC-Bayesian type, reducing CVaR estimation to expectation estimation.
result The bound is small when empirical CVaR is small, providing concentration inequalities for CVaR.
Unified framework for removing unwanted information from machine learning models.
problem Removing undesirable features or data points from machine learning models while preserving utility.
method Information-theoretic regularization approach for data point and feature unlearning.
result Unified mathematical framework with provable guarantees for both data point and feature unlearning.
We augment adversarial training (AT) with worst case adversarial training (WCAT) which improves adversarial robustness by 11% over the current state-of-the-art result in the ℓ2 norm on CIFAR-10. We obtain verifiable average case and worst case robustness guarantees, based on the expected and maximum values of the…
Develops a framework for causal structure learning using both interventional and observational data.
problem Lack of identifiability of causal structures with only observational data.
method Bilevel polynomial optimization (Bloom) framework for causal structure discovery from interventional and observational data.
result Bloom framework provides convergence and optimality guarantees, surpassing other learning algorithms in experiments.
New mathematical framework proves the effectiveness of reducing neural network sizes.
problem Selecting optimal neural network sizes to avoid overfitting.
method Adaptive group Lasso applied to one-hidden-layer feedforward networks.
result Adaptive group Lasso is consistent and can accurately reconstruct network sizes.
The paper deals with the problem of reconstructing the topological structure of a network of dynamical systems. A distance function is defined in order to evaluate the "closeness" of two processes and a few useful mathematical properties are derived. Theoretical results to guarantee the correctness of the identificatio…
In this paper we explore an identity in distribution of hitting times of a finite variation process (Yor's process) and a diffusion process (geometric Brownian motion with affine drift), which arise from various applications in financial mathematics. As a result, we provide analytical solutions to the fair charge of va…
We propose a spectral clustering method based on local principal components analysis (PCA). After performing local PCA in selected neighborhoods, the algorithm builds a nearest neighbor graph weighted according to a discrepancy between the principal subspaces in the neighborhoods, and then applies spectral clustering. …
An important problem in machine learning and statistics is to identify features that causally affect the outcome. This is often impossible to do from purely observational data, and a natural relaxation is to identify features that are correlated with the outcome even conditioned on all other observed features. For exam…
In many real-world applications of machine learning, data are distributed across many clients and cannot leave the devices they are stored on. Furthermore, each client's data, computational resources and communication constraints may be very different. This setting is known as federated learning, in which privacy is a …
Subspace clustering is the problem of clustering data points into a union of low-dimensional linear/affine subspaces. It is the mathematical abstraction of many important problems in computer vision, image processing and machine learning. A line of recent work (4, 19, 24, 20) provided strong theoretical guarantee for s…
New method uses model comparison signals to improve LLM evaluation accuracy.
problem Limited benchmark sizes and model stochasticity in evaluating LLMs' mathematical reasoning.
method Combines standard labeled outcomes with model comparison signals to design a statistically efficient evaluation framework.
result Semiparametric estimator achieves the semiparametric efficiency bound and substantially improves ranking accuracy.
In this paper the multivariate fractional trading ansatz of money management from Ralph Vince (Portfolio Management Formulas: Mathematical Trading Methods for the Futures, Options, and Stock Markets, John Wiley & Sons, Inc., 1990) is discussed. In particular, we prove existence and uniqueness of an optimal f of the res…
Synthetic data can be used to ask more questions and accelerate discovery with provable validity guarantees.
problem Valid inference with synthetic data
method Task exchangeability
result Provable validity guarantees for synthetic data inference
New SAE algorithm proves feature recovery for LLMs with theoretical guarantees.
problem Achieving interpretable features in large language models (LLMs).
method Proposed a statistical framework and bias adaptation technique for sparse autoencoders (SAEs).
result Proved correct recovery of all monosemantic features under specific data sampling.
New federated learning algorithms improve model aggregation robustness.
problem Improving model aggregation in federated learning.
method Complete mathematical convergence analysis and novel aggregation algorithms.
result Derived novel algorithms that modify model architecture based on client contributions.
We show how the success of deep learning could depend not only on mathematics but also on physics: although well-known mathematical theorems guarantee that neural networks can approximate arbitrary functions well, the class of functions of practical interest can frequently be approximated through "cheap learning" with …
Bayesian quadrature uses probabilistic models for estimating intractable integrals.
problem Estimating intractable integrals in complex models.
method Probabilistic, model-based approach using Gaussian processes.
result Comprehensive review and systematic taxonomy of Bayesian quadrature methods.
Scientists and engineers rely on accurate mathematical models to quantify the objects of their studies, which are often high-dimensional. Unfortunately, high-dimensional models are inherently difficult, i.e. when observations are sparse or expensive to determine. One way to address this problem is to approximate the or…
Improved ridge estimators avoid tuning parameters for high-dimensional data.
problem Difficulty in calibrating tuning parameters for ridge estimators.
method Developed modified ridge estimators that eliminate tuning parameters.
result Modified ridge estimators outperform standard methods in prediction accuracy.
Subset selection in multiple linear regression aims to choose a subset of candidate explanatory variables that tradeoff fitting error (explanatory power) and model complexity (number of variables selected). We build mathematical programming models for regression subset selection based on mean square and absolute errors…
Sparsity helps reduce diffusion model costs.
problem High computational costs in diffusion models.
method Introduced sparsity concept to reduce input dimensionality.
result Sparsity reduces computational complexity to intrinsic data dimension.
A new machine learning model uses matrix exponentials for universal approximation.
problem Developing a robust and efficient machine learning model.
method Introduces a novel architecture using matrix exponentials as the only nonlinearity.
result The model achieves universal approximation properties and outperforms other models on benchmark tasks.
Mathematical model audits social media algorithms to prevent bias.
problem Algorithmic filtering can bias users' decisions and societal norms.
method Formalized mathematical framework for auditing social media algorithms.
result Data-driven statistical auditing procedure to regulate algorithmic bias.