The paper proves convex risk minimization selects a unique conditional probability model.
problem General conditional probability estimation in various settings.
method Convex risk minimization and empirical risk minimization.
result The unique conditional probability model is selected by convex risk minimization.
A new method for ML classification without using norms.
problem Machine Learning classification problems without norm criteria.
method Quantum mechanics-like probability states and eigenvalues for classification.
result Avoids norm usage for classification, estimating y y y distribution. The study optimizes bounds for comparing training and population loss.
problem Optimizing bounds for comparing training and population loss.
method Derives generic information-theoretic and PAC-Bayesian generalization bounds using convex comparator functions.
result The tightest possible bound is obtained with the comparator being the convex conjugate of the CGF of the bounding distribution.
Jiang et al. (2020) found no uniformly tight generalization bounds for neural networks in the overparameterized setting.
problem Finding uniformly tight generalization bounds for neural networks in the overparameterized setting.
method Examined more than a dozen generalization bounds, proving that no bounds can be uniformly tight in the overparameterized setting.
result No generalization bounds can be uniformly tight in the overparameterized setting.
New bounds on machine learning model generalization error moments.
problem Understanding the performance of machine learning models.
method Information-theoretic bounds on the moments of the generalization error of learning algorithms.
result Proposed bounds on generalization error moments and their high-probability bounds.
The study identifies conditions for algorithms to have tight generalization bounds.
problem Understanding which algorithms have tight generalization bounds.
method Analyzing conditions that preclude tight generalization bounds and identifying stable algorithms.
result Stable algorithms have tight generalization bounds, while unstable ones do not.
New lower bounds nearly match existing upper bounds for boosted classifiers.
problem Understanding the generalization performance of boosted classifiers.
method Margin-based lower bounds on boosted classifiers.
result Lower bounds nearly match the k k k th margin bound, settling the generalization performance of boosted classifiers. New bounds show limitations of sample-wise information-theoretic generalization.
problem Limitations of sample-wise information-theoretic generalization bounds.
method Analysis of existing bounds and derivation of new bounds.
result No sample-wise information-theoretic bounds exist for expected squared generalization gap.
New bounds show large language models can generalize beyond training data.
problem Generalization of large language models beyond training data.
method Compression bound derivation using prediction smoothing and SubLoRA.
result Large language models can discover generalizable regularities.
New bounds for nearly-linear networks without training.
problem Generalization of neural networks close to linearity.
method Perturbation of linear networks to derive bounds.
result First non-vacuous bounds for neural nets.
New bound on machine learning model performance using Jensen-Shannon information.
problem Understanding the performance of machine learning models.
method Proposes a new information-theoretic bound on generalization error.
result Shows that the new bound can be tighter than mutual information-based bounds under certain conditions.
The study improves PAC-Bayesian bounds for adversarial generative models.
problem Improving generalization bounds for adversarial generative models.
method Extending PAC-Bayesian theory to generative models, developing bounds for Wasserstein and total variation distances.
result New training objectives for Wasserstein and Energy-Based GANs.
This paper refines learning algorithm bounds, proving sharper generalization and lower bounds.
problem Improving generalization bounds for uniformly stable learning algorithms.
method Developed a new concentration inequality for weakly correlated random variables and proved sharper generalization and lower bounds.
result The new concentration inequality implies a stronger generalization bound than previous results.
New bound for neural networks with full-rank weights, independent of network width.
problem Understanding generalization of neural networks with full-rank weight matrices.
method Using Koopman operators to derive a tighter generalization bound for full-rank weight matrices.
result The bound is tighter than existing norm-based bounds when condition numbers are small.
The paper improves SVM margin-based generalization bounds.
problem Improving generalization bounds for SVMs.
method Revisiting and improving classic generalization bounds in terms of margins, complementing with a nearly matching lower bound.
result Almost settles the generalization performance of SVMs in terms of margins.
Introduces bounded scale measure and generalizes property A.
problem Defining property A for large scale spaces with bounded geometry.
method Introduces bounded scale measure, shows its coarse invariance, and generalizes property A.
result Definition of property A for large scale spaces with bounded scale measure is a coarse invariant.
The paper bounds generalization error for iterative learning with bounded updates.
problem Generalization error of iterative learning algorithms with bounded updates for non-convex loss functions.
method Information-theoretic techniques, reformulating mutual information as update uncertainty, variance decomposition.
result Improved generalization error bounds for iterative learning algorithms with bounded updates.
New bounds for large language models using token properties.
problem Vacuous generalization bounds for large language models.
method Martingale properties and Monarch matrices.
result Non-vacuous generalization bounds for LLMs up to 70B parameters.
Paper introduces new bounds linking data compressibility to generalization error.
problem Establishing data-dependent generalization bounds.
method Variable-size compressibility framework linking generalization error to compression rate of input data.
result New bounds depend on empirical data measure, subsuming existing PAC-Bayes and intrinsic dimension bounds.
New tighter generalization bounds for deep networks like CNNs and ResNets.
problem Establishing tighter bounds for deep neural networks' generalization error.
method Introducing a new characterization of Lipschitz properties and margin-based data-dependent error bounds.
result Significantly tighter generalization bounds for deep neural networks, including CNNs and ResNets.
New distribution-dependent inequalities improve generalization bounds.
problem Improving generalization bounds for learning models.
method Proposed four types of conditions for probabilistic boundedness and bounded differences, derived several distribution-dependent extensions of Hoeffding's and McDiarmid's inequalities.
result Tighter generalization bounds for functions not satisfying existing conditions.
New bound matches exact generalization error for quadratic Gaussian problem.
problem Understanding generalization error in quadratic Gaussian problems.
method Information-theoretic approach with new ingredients.
result Exact tight bound for generalization error.
New bound on generalization error using mutual information.
problem Generalization error in supervised learning.
method Information-theoretic bound on mutual information between samples and predictions.
result Tighter characterization of generalization error.
Paper improves generalization bounds for noisy stochastic algorithms.
problem Improving generalization bounds for noisy stochastic algorithms.
method Introduces Exponential Family Langevin Dynamics (EFLD) and establishes data-dependent expected stability based generalization bounds.
result Sharp generalization bounds with O(1/n) sample dependence and gradient discrepancy.
New bounds derived using conditional f f f -information for machine learning models.
problem Improving generalization bounds in machine learning.
method Introducing novel information-theoretic generalization bounds via conditional f f f -information. result Derives generalization bounds applicable to both bounded and unbounded loss functions.
New bounds on generalization error using information density moments.
problem Bounding the generalization error of randomized learning algorithms.
method Derives bounds on average and tail probabilities of generalization error using mth central moments of the information density.
result Explicit bounds on generalization error are derived, showing better dependence on confidence level with higher-order information density moments.
The paper offers generalization bounds for Transformers that ignore sequence length.
problem Developing generalization bounds for Transformers that are independent of sequence length.
method Covering number approach to upper bound Rademacher complexity of bounded linear transformations.
result Theoretical bounds for Transformer generalization are independent of sequence length.
New bound relaxes uniform gradient norm assumptions for PAC-Bayesian bounds.
problem Generalization bounds with strict assumptions like uniformly bounded loss.
method Relax uniform bounds assumptions to on-average bounded loss and gradient norm.
result Proposes a new generalization bound with a surrogate of model complexity.
PAC-Bayesian bounds for MLPs with cross entropy loss validated.
problem Generalization bounds for MLPs with cross entropy loss.
method Introduced probabilistic explanations and proved PAC-Bayesian bounds using ELBO.
result MLPs with cross entropy loss inherently guarantee PAC-Bayesian generalization bounds.
Paper improves PAC-Bayes bounds for various loss types.
problem Improving PAC-Bayes bounds for different types of losses.
method Introducing new high-probability PAC-Bayes bounds for bounded and general tail behaviors losses, and extending to anytime-valid bounds.
result New fast-rate and mixed-rate bounds for losses with bounded ranges, and parameter-free bounds for losses with general tail behaviors.
Paper develops a new generalization bound using PAC-Bayes theory and Gibbs distributions.
problem Limits of traditional generalization bounds due to complexity measures.
method Leverages PAC-Bayes bounds with Gibbs distributions to derive a flexible generalization bound.
result Derives a generalization bound that can adapt to both hypothesis class and task complexity.
Improved generalization bounds for uniformly stable algorithms.
problem Deriving meaningful generalization bounds for uniformly stable algorithms in common settings.
method New analysis techniques to improve generalization bounds for uniformly stable algorithms.
result Improved generalization bounds for uniformly stable algorithms, with a bound of O ( ( γ + 1 / n ) log ( 1 / δ ) ) O(\sqrt{(γ+ 1/n) \log(1/δ)}) O ( ( γ + 1/ n ) log ( 1/ δ ) ) with probability at least 1 − δ 1-δ 1 − δ . New PAC bound for meta-learning improves generalization guarantees.
problem Provide strong generalization guarantees in meta-learning.
method PAC-Bayes and uniform stability frameworks applied to gradient-based meta-learning.
result Derives a tighter PAC bound for gradient-based meta-learning.
New bounds improve generalization in learning scenarios.
problem Limitations of existing information-theoretic bounds in SCO problems.
method Sample-conditioned hypothesis stability and neighboring-hypothesis matrix.
result Sharper generalization guarantees in various learning scenarios.
New bounds estimate learning algorithm performance using prediction information.
problem Estimating the performance of black-box learning algorithms.
method Information-theoretic bounds based on prediction information.
result Improved bounds applicable to deterministic algorithms and easier to estimate.
Proves bounds on deep CNN generalization error.
problem Understanding how well deep CNNs perform on unseen data.
method Bounds on generalization error based on training loss, parameters, Lipschitz constant, and weight distance.
result Bounds are independent of input size and feature map dimensions.
New bounds show polyhedral surrogates are optimal for generalization.
problem Proving generalization rates for polyhedral loss functions.
method Developed two general results for polyhedral surrogates.
result Polyhedral surrogates provide linear surrogate regret bounds, translating directly to target rates.
The paper analyzes generalization bounds for NC-SC/NC-C stochastic minimax optimization.
problem Generalization analysis of nonconvex-(strongly)-concave stochastic minimax optimization.
method Established algorithm-agnostic and algorithm-dependent generalization bounds via uniform convergence and stability arguments.
result Sample complexities and generalization bounds for NC-SC and NC-C settings.
New framework connects online learning to statistical learning for better generalization bounds.
problem Deriving generalization bounds for statistical learning algorithms.
method Constructing an online learning game and showing a connection to statistical learning.
result Established a connection between online and statistical learning, leading to new generalization bounds.
Proposes a new generalization bound for Bayesian deep nets without strict assumptions.
problem Lack of generalization bounds for Bayesian deep nets without strict assumptions.
method Exploits contractivity of Log-Sobolev inequalities to add a loss-gradient norm term to the generalization bound.
result Introduces a new generalization bound for Bayesian deep nets that avoids strict assumptions.
Hierarchical Federated Learning bounds generalize using Wasserstein distance.
problem Bounding generalization error in Federated Learning with hierarchical sampling.
method Introduced a hierarchical sampling framework and derived generalization bounds using Wasserstein distance.
result Recover and strictly imply existing CMI bounds for bounded losses.
The study generalizes curvature bounds for manifolds with boundary.
problem Proving curvature bounds for manifolds with boundary.
method Bakry-Émery curvature bounds and splitting theorems.
result Proves curvature bounds for manifolds with boundary.
This paper improves meta-learning by developing new PAC-Bayes bounds.
problem Meta-learning generalization gap across multiple tasks.
method Upper bounding convex functions linking environment and task-level losses.
result New PAC-Bayes bounds for meta-learning with improved algorithms.
New bounds on learning algorithm generalization error derived using information density.
problem Bounding the generalization error of learning algorithms.
method Exponential inequalities and information density/conditional information density.
result Novel bounds on average and tail probability of generalization error.
The paper derives tighter generalization bounds for k k k -dimensional coding schemes in finite-dimensional feature spaces.
problem Previous bounds for k k k -dimensional coding schemes were dimensionality-independent and not suitable for finite-dimensional data. method The paper derives a dimensionality-dependent generalization bound for k k k -dimensional coding schemes by bounding the covering number of the loss function class induced by the reconstruction error. result The derived bound is tighter than previous results and converges faster, especially for finite-dimensional data.
New bounds using samplewise evaluated CMI for deep neural networks.
problem Improving generalization bounds for deep neural networks.
method Introduced a new family of information-theoretic generalization bounds using samplewise evaluated conditional mutual information (CMI).
result The new bounds can be tighter than previous ones for deep neural networks.
New margin-based learning guarantees improve generalization bounds.
problem Improving generalization bounds for machine learning models.
method Relative deviation margin bounds using empirical margin loss and Rademacher complexity.
result Distribution-dependent generalization bounds for unbounded loss functions.
Novel bounds for SGLD show generalization error decreases with more samples.
problem Understanding the generalization error of SGLD in non-convex optimization.
method Information-theoretic approach focusing on Kullback-Leibler divergence and sub-exponential loss function.
result Time-independent generalization bounds for SGLD, independent of step size and number of iterations.