Boosted additive models reveal new insights and potential pathologies.
problem Theoretical understanding of boosted additive models (BAMs) and their convergence behavior.
method Study of solution paths of BAMs and derivation of convergence results.
result Uncovering pathologies of boosting for certain additive model classes.
This review clarifies XAI for regression models and establishes new theoretical insights.
problem Lack of XAI techniques for regression models, especially in safety-critical applications.
method Clarifies conceptual differences, establishes theoretical insights, provides demonstrations, discusses challenges.
result Novel theoretical insights and demonstrations of XAI for regression models.
Theoretical framework explains why few epochs are enough for LLM fine-tuning.
problem Understanding why few epochs are sufficient for LLM fine-tuning.
method Combining early stopping theory with attention-based Neural Tangent Kernel (NTK) for LLMs.
result Formalizes convergence rate of attention-based fine-tuning with respect to sample size.
Formalizes concepts as latent variables in hierarchical models for high-dimensional data.
problem Lack of formalization and theoretical insights for learning discrete concepts from high-dimensional data.
method Formalizes concepts as latent causal variables in a hierarchical model, formulates conditions for concept identification.
result Conditions for identifying latent hierarchical models in unsupervised data, handling complex structures and high-dimensional data.
New insights into how data transformations affect self-supervised clustering.
problem Impact of data transformations on self-supervised clustering convergence.
method Theoretical and empirical analysis of various data transformations.
result Certain transformations help in faster convergence of self-supervised clustering.
Mutual information has been successfully adopted in filter feature-selection methods to assess both the relevancy of a subset of features in predicting the target variable and the redundancy with respect to other variables. However, existing algorithms are mostly heuristic and do not offer any guarantee on the proposed…
We explore non-acyclic GFlowNets in discrete settings.
problem Training and understanding non-acyclic GFlowNets in discrete environments.
method Relaxing acyclicity assumption, simpler theoretical framework, novel theoretical insights, experimental validation.
result Theoretical and experimental validation of non-acyclic GFlowNets in discrete environments.
This paper improves model training by using a reference model to guide target model training.
problem Improving generalization and data efficiency in model training.
method DRRho risk minimization framework based on Distributionally Robust Optimization (DRO).
result DRRho risk minimization improves generalization and data efficiency compared to training without a reference model.
The paper explores stability and generalization of deep GCNs.
problem Understanding the stability and generalization of deep GCNs from a theoretical perspective.
method Theoretical analysis of stability and generalization properties of deep GCNs.
result The stability and generalization of deep GCNs are influenced by the maximum absolute eigenvalue of the graph filter operators and the depth of the network.
This work analyzes Fréchet regression using comparison geometry, providing theoretical and practical insights.
problem Analyzing data on complex structures like manifolds and graphs.
method Theoretical analysis through comparison geometry, focusing on existence, uniqueness, and stability of the Fréchet mean.
result Key results on the existence, uniqueness, and stability of the Fréchet mean, along with statistical guarantees for nonparametric regression.
Investigates training challenges for morphological neural networks.
problem Training morphological neural networks with gradient descent is difficult.
method Examines differentiation and backpropagation for morphological networks, using Bouligand derivative.
result Provides theoretical insights and guidelines for initialization and learning rates.
The 1/3 Financial Rule helps prevent household bankruptcy through balanced spending, savings, and debt repayment.
problem Reducing household bankruptcy risk through effective financial planning.
method Mathematical modeling, game theory, behavioral finance, and technological analysis.
result The 1/3 Financial Rule emerges as a robust solution for supporting household financial stability.
This paper provides theoretical insights into why and how deep learning can generalize well, despite its large capacity, complexity, possible algorithmic instability, nonrobustness, and sharp minima, responding to an open question in the literature. We also discuss approaches to provide non-vacuous generalization guara…
New algorithm recovers sparse measures in polynomial time.
problem Recovering sparse measures from Fourier moments.
method Polynomial-time recovery method inspired by mean-field theory.
result Improves upon convex relaxation methods in specific parameter regime.
New insights into correntropy-based regression reveal robustness and unified approaches.
problem Learning robust regression functions under additive noise.
method Minimum distance estimation and conditional mean, mode, median functions.
result Unified approach to conditional mean, mode, and median functions.
This thesis presents some geometric insights into three different types of two player prediction games -- namely general learning task, prediction with expert advice, and online convex optimization. These games differ in the nature of the opponent (stochastic, adversarial, or intermediate), the order of the players' mo…
This work explores the generalization properties of diffusion models, providing theoretical and empirical insights.
problem Theoretical understanding of diffusion models' generalization capabilities remains underdeveloped.
method Theoretical exploration and quantitative analysis of generalization gaps in diffusion models.
result Established polynomially small generalization error (O(n−2/5+m−4/5)) for diffusion models, avoiding the curse of dimensionality. New insights into using momentum for non-convex optimization.
problem Improving training of non-convex models like deep neural networks.
method Developed a Lyapunov analysis of SGD with momentum using stochastic primal averaging.
result Precise conditions under which SGD+M outperforms SGD and optimal hyper-parameter schedules.
Limit order books (LOBs) match buyers and sellers in more than half of the world's financial markets. This survey highlights the insights that have emerged from the wealth of empirical and theoretical studies of LOBs. We examine the findings reported by statistical analyses of historical LOB data and discuss how severa…
The VAE's reconstruction ability is studied using PAC-Bayes theory.
problem Understanding the performance of VAEs for unseen data.
method PAC-Bayes theory is applied to analyze VAE's reconstruction error.
result Generalization bounds on VAE's reconstruction error are provided.
Despite widespread interest and practical use, the theoretical properties of random forests are still not well understood. In this paper we contribute to this understanding in two ways. We present a new theoretically tractable variant of random regression forests and prove that our algorithm is consistent. We also prov…
New group theory insights on knot surgery results.
problem Understanding non-simply connected 3-manifolds from Dehn surgery.
method Group theoretic analysis of Property P conjecture variations.
result New group theoretic perspectives on Dehn filling.
Theoretical analysis of MCR for improving imputation quality in partially observed data.
problem Improving model generalization in partially observed settings.
method Theoretical analysis of Measure Consistency Regularization (MCR) for neural network distance.
result MCR's generalization advantage is not always guaranteed and can be monitored through a duality gap.
Sparse NMF with archetypal regularization aims to robustly represent data points.
problem Representing data points as sparse linear combinations of archetypes.
method Sparse NMF with archetypal regularization, introducing strong and weak robustness.
result Theoretical robustness guarantees hold under minimal assumptions.
Kernel PCA helps analyze multivariate extremes and clusters them effectively.
problem Analyzing the dependence structure of multivariate extremes.
method Kernel PCA as a method for clustering and dimension reduction.
result Kernel PCA preimages effectively identify clusters in multivariate extremes.
The paper explores theoretical insights into WGANs for better understanding and stability.
problem Stabilizing the training process of GANs.
method Theoretical analysis and statistical convergence study of WGANs.
result Theoretical properties and convergence of WGANs are clarified.
The design of codes for communicating reliably over a statistically well defined channel is an important endeavor involving deep mathematical research and wide-ranging practical applications. In this work, we present the first family of codes obtained via deep learning, which significantly beats state-of-the-art codes …
The study examines denoising and noisy-input regression under distribution shift, revealing double descent behavior and insights for data augmentation.
problem Understanding denoising in machine learning, especially under noisy inputs and distribution shift.
method Theoretical analysis of supervised denoising and noisy-input regression, considering low-rank data and proportional regime.
result The test error exhibits double descent under general distribution shift, indicating that overfitting the noise can be benign, tempered, or catastrophic.
Proposes using gradients as features for efficient deep learning adaptation.
problem Efficient deep representation learning for different tasks.
method Designs a linear model incorporating gradients and activations of a pre-trained network.
result Shows strong results across various tasks and datasets.
New insights into balancing reward and fairness in stochastic MAB.
problem Balancing reward and fairness in stochastic multi-armed bandits.
method Formulated a penalization framework and proposed a hard-threshold UCB-like algorithm.
result Asymptotic fairness, nearly optimal regret, better reward-fairness tradeoff.
The study examines how the number of noise samples affects diffusion models' performance.
problem Understanding the balance between generalization and memorization in diffusion models.
method Theoretical analysis and empirical experiments with Denoising Score Matching (DSM) using random features.
result Precise expressions for test and train errors under specific conditions reveal the mechanisms of generalization and memorization.
New insights explain speedup saturation in distributed learning with large batches and delays.
problem Understanding and optimizing speedup in distributed learning with large batches and delays.
method Theoretical analysis of strongly convex, convex, and non-convex settings, considering data sparsity.
result Identification of a data-dependent parameter explaining speedup saturation in both batch size and gradient staleness.
This work analyzes SGGMs, offering convergence insights and practical design tips.
problem Theoretical convergence analysis for SGGMs with a system of coupled SDEs.
method Non-asymptotic convergence analysis for three graph generation paradigms.
result Unique factors affecting convergence in SGGMs and practical hyperparameter selection.
In this paper, we present and illustrate some new tools for rigorously analyzing training data selection methods. These tools focus on the information theoretic losses that occur when sampling data. We use this framework to prove that two methods, Facility Location Selection and Transductive Experimental Design, reduce…
The paper provides formulas for volatility in various models, including rough volatility.
problem Calibrating SPX and VIX options with rough volatility models.
method Developed explicit formulae using Malliavin calculus for Gaussian processes.
result New insights on joint calibration of SPX and VIX options.
By simulating the easy-to-hard learning manners of humans/animals, the learning regimes called curriculum learning~(CL) and self-paced learning~(SPL) have been recently investigated and invoked broad interests. However, the intrinsic mechanism for analyzing why such learning regimes can work has not been comprehensivel…
Study reveals mutual information is crucial for understanding algorithm performance in stochastic convex optimization.
problem Uncertainty in capturing the exceptional performance of learning algorithms using existing information-theoretic generalization bounds.
method Examined the relationship between mutual information and generalization in stochastic convex optimization.
result Mutual information is necessary for true risk minimization in stochastic convex optimization, indicating existing bounds fall short.
EVODiff optimizes DM inference by reducing conditional entropy, improving image generation.
problem Slow and inaccurate inference in diffusion models.
method Entropy-aware variance optimization for efficient inference.
result Significant improvement in image generation quality and efficiency.
The paper provides a theoretical framework for machine learning.
problem Lack of rigorous theory guiding machine learning experiments.
method Bayesian statistics and Shannon's information theory.
result Theoretical insights applicable across various machine learning settings.
New model learns from random graph samples to estimate graph parameters.
problem Scalability issues in graph learning methods for large graphs.
method Develops a graph classification model working on randomly sampled subgraphs.
result Validates mini-batch learning on graphs and provides generalization bounds.
Reinforcement learning is a promising approach to learning robotics controllers. It has recently been shown that algorithms based on finite-difference estimates of the policy gradient are competitive with algorithms based on the policy gradient theorem. We propose a theoretical framework for understanding this phenomen…
The pathwise coordinate optimization is one of the most important computational frameworks for high dimensional convex and nonconvex sparse learning problems. It differs from the classical coordinate optimization algorithms in three salient features: {\it warm start initialization}, {\it active set updating}, and {\it …
BELIEF framework interprets GLMs using binary linear models.
problem Understanding and interpreting generalized linear models (GLMs) with binary outcomes.
method Developed a framework called binary expansion linear effect (BELIEF) to interpret GLMs through transparent linear models.
result BELIEF framework reveals perfect predictors in complete separation scenarios.
Researchers analyze inverse optimal transport, deriving theoretical and empirical insights.
problem Understanding the inverse problem of inferring cost matrices from optimal couplings.
method Formalized and analyzed using entropy-regularized optimal transport, with theoretical and empirical contributions.
result Characterization of the manifold of cross-ratio equivalent costs and derivation of an MCMC sampler.
New insights and algorithms improve prediction models with time-series privileged information.
problem Efficient learning of nonlinear prediction models with limited data.
method Generalization of LuPI to nonlinear tasks, using random features and representation learning.
result Theoretical and empirical evidence supports the use of privileged time-series information for nonlinear prediction.
New insights into surface energy reduction.
problem Energy behavior of degenerating submanifolds.
method Analyzing regularized Riesz energy for closed submanifolds.
result Energy blows up as submanifolds degenerate.
Proposes RVP to address theoretical concerns of V-REx for OOD generalization.
problem Theoretical concerns about V-REx's motivation and utility.
method Risk Variance Penalization (RVP) modifies V-REx's regularization.
result RVP discovers a robust predictor and finds invariant predictors under certain conditions.
Paper analyzes sample complexity of polynomial neural networks.
problem Understanding the sample complexity of polynomial neural networks.
method Extends previous literature to polynomial neural networks and analyzes sample complexity.
result Obtains novel results on sample complexity of polynomial neural networks.