New method prevents 'shattered gradients' in deep networks, improving training of very deep models.
problem Vanishing and exploding gradients in deep learning networks.
method Introducing 'looks linear' (LL) initialization to prevent gradient shattering.
result Gradient shattering is prevented, allowing training of very deep networks without skip-connections.
Improved uniform convergence bound with fat-shattering dimension reduces sample complexity gap.
problem Gap between upper and lower bounds on sample complexity for fat-shattering dimension.
method Provided an improved uniform convergence bound.
result Closed the gap between existing upper and lower bounds on sample complexity.
Estimates fat-shattering dimension of aggregated function classes.
problem Understanding the complexity of aggregated function classes.
method Analyzes fat-shattering dimension of k-fold aggregations of real-valued function classes. result Provides upper and lower bounds on fat-shattering dimension for linear and affine function classes.
Study shows how many domains are needed for generalization, using a new measure called domain shattering dimension.
problem How many domains are needed for domain generalization?
method Introduced a new combinatorial measure called the domain shattering dimension to model domain sample complexity.
result Established a tight quantitative relationship between domain shattering dimension and classic VC dimension.
This paper analyzes the Shattering coefficient for supervised learning algorithms.
problem Ensuring uniform convergence of empirical risk to expected risk.
method Analyzes the Shattering coefficient for Hilbert spaces containing input spaces.
result Proves the Shattering coefficient's polynomial growth for any Hilbert space.
This paper addresses the complexity of labeled datasets using topological methods.
problem Estimating the necessary training sample size for supervised learning.
method Employing equivalence relations from Topology, data separability results, and combinatorics to compute the Shattering coefficient.
result Estimation of required number of hyperplanes and training sample sizes for binary and multi-class datasets.
The paper explores how data geometry influences generalization in neural networks.
problem Understanding generalization in overparameterized neural networks.
method Theoretical exploration of overparametrized two-layer ReLU networks trained below the edge of stability.
result Generalization bounds adapt to the intrinsic dimension of data distributions and deteriorate as data concentrates towards the unit sphere.
New results show flat minima in neural networks suffer from high dimensionality.
problem Flat minima in neural networks generalize poorly in high dimensions.
method Theoretical analysis of two-layer ReLU networks with multivariate inputs.
result Flat minima lead to exponentially slower convergence in high dimensions.
New learning rule for quantum measurement classes overcomes uniform convergence issues.
problem Characterizing learnability of POVM hypothesis classes in quantum settings.
method Introduced a new learning rule called denoised ERM to address uniform convergence issues.
result Characterized learnability conditions and sample complexity bounds for POVM classes.
Study robust regression learning under adversarial attacks.
problem Understanding which function classes are learnable in the presence of adversarial attacks.
method Introduced a novel agnostic sample compression scheme and used fat-shattering dimension to construct adversarially robust sample compression schemes.
result Finite fat-shattering dimension classes are learnable in both realizable and agnostic settings.
New method trains neural networks with threshold activation functions efficiently.
problem Training neural networks with threshold activation functions is challenging due to zero gradients.
method We study weight decay regularized training problems of deep neural networks with threshold activations, showing they can be formulated as convex optimization problems.
result Regularized deep threshold network training problems can be formulated as standard convex optimization problems, paralleling the LASSO method.
The study provides a sample complexity estimate for multi-category classifiers with bounded variation.
problem Controlling the deviation between empirical and generalization performances of multi-category classifiers.
method Using the empirical L1-norm covering number and fat-shattering dimension, the study derives a sample size estimate for classifiers of bounded variation.
result The sample size estimate is sufficient for the performances to be close with high probability, improving the dependency on the number of classes.
Positive results for agnostic regression with various losses.
problem Agnostic regression with bounded sample compression.
method Generic and efficient sample compression schemes for real-valued functions.
result Exact and approximate compression schemes for specific losses.
Paper provides convergence guarantees for rectifier networks using neural Taylor approximations.
problem Smoothness and convexity issues in modern convolutional networks.
method Neural Taylor approximations and Taylor loss for optimization.
result Guarantees match lower bounds for convex nonsmooth functions and accurately capture optimization dynamics.
Investigates scaling deep neural networks to avoid capacity issues.
problem Avoiding the shattering problem in deep neural networks.
method Formulates conjecture and studies various architectures to determine scaling relations.
result Reveals scaling relations for deep residual networks and recurrent networks.
Study online learning with set-valued feedback, showing differences between deterministic and randomized approaches.
problem Online learning with set-valued feedback, where labels are sets rather than single labels.
method Introduced new combinatorial dimensions (Set Littlestone and Measure Shattering) to characterize learnability.
result Characterized deterministic and randomized online learnability, and established bounds for various learning settings.
We obtain a tight distribution-specific characterization of the sample complexity of large-margin classification with L2 regularization: We introduce the margin-adapted dimension, which is a simple function of the second order statistics of the data distribution, and show distribution-specific upper and lower bounds on…
New protocol for online learning with partial feedback, extending classical methods.
problem Learning with partial feedback where only one acceptable label is observed per round.
method Introducing a collection version space to address the lack of direct extension of classical methods.
result Characterization of learnability in set-realizable regime using Partial-Feedback Littlestone dimension and Partial-Feedback Measure Shattering dimension.
New algorithms achieve near-optimal cumulative loss in nonparametric online learning and games.
problem Fast rates of convergence in nonparametric online regression and classification.
method Randomized proper learning algorithms, hierarchical aggregation, multi-scale extension, stability proof.
result Achieved near-optimal cumulative loss bounds for real-valued and binary games.
General lower bounds on neural network approximation in L^p norm.
problem Fundamental limits of neural network expressivity.
method General lower bound proof on approximation in L^p norm, applied to feed-forward neural networks.
result Neural networks can't approximate certain functions as well as previously thought.
Characterizes statistical complexity of realizable regression in PAC and online learning.
problem Understanding the statistical complexity of realizable regression in both PAC and online learning settings.
method Introduces minimax instance optimal learners, novel and combinatorial dimensions to characterize learnability.
result Characterizes which classes of real-valued predictors are learnable and provides necessary conditions for learnability.
New algorithm for learning functions with bounds on error and sample complexity.
problem Learning [0,1]-valued functions in a prediction model. method General-purpose algorithm with upper and lower bounds on expected error and sample complexity.
result Improved bounds on sample complexity and agnostic learning conditions.
New algorithm learns regression models privately under growth condition.
problem Private learning of nonparametric regression models.
method Novel filtering procedure to output stable hypotheses for nonparametric function classes.
result Established first nonparametric private learnability guarantee for diverging fat shattering dimensions.
New findings on neural networks with non-negative weights and low training error.
problem Does a low training error imply a small outer norm for two-layer neural networks?
method Covering number argument and fat-shattering dimension analysis.
result For non-negative output weights, low training error guarantees a well-controlled outer norm.
The paper explores how approximation theory can improve understanding of smooth kernels in machine learning.
problem Understanding the inferential properties of smooth kernels in machine learning.
method Analysis of eigenvalue decay, properties of eigenfunctions/eigenvectors, and fitting capacity of kernels.
result Eigenvalues of kernel matrices show nearly exponential decay, highlighting the 'approximation beats concentration' phenomenon.
Improved robust learning model with tighter generalization bounds.
problem Adversarial robust learning in environments with limited corruptions.
method Model as a zero-sum game, using regret minimization and ERM oracles.
result Improved sample complexity for robust classifiers, handling infinite hypothesis classes.
Characterizes sample complexity for outcome indistinguishability in machine learning.
problem Outcome indistinguishability in machine learning, focusing on distinguishers and predictors.
method Sample complexity characterized by metric entropy of predictor and distinguisher classes, using dual Minkowski norms.
result Equivalence and tightness of sample complexity characterizations in distribution-specific and distribution-free settings.
New algorithm tackles multiclass transductive online learning with unbounded labels.
problem Characterizing optimal mistake bound for unbounded label spaces.
method Introducing new combinatorial dimensions (Level-constrained Littlestone and Branching dimensions) to characterize online learnability.
result Established trichotomy of possible minimax rates for unbounded label spaces: Θ(T), Θ(logT), or Θ(1). Comparative learning combines realizable and agnostic settings for two hypothesis classes, reducing sample complexity.
problem Learning with two hypothesis classes in a more general setting than single hypothesis classes.
method Introduces comparative learning, defines mutual VC dimension and Littlestone dimension, and applies insights to multiaccuracy and multicalibration.
result Sample complexity of comparative learning is characterized by mutual VC dimension and Littlestone dimension.
Characterizes the sample complexity of list regression tasks.
problem Understanding the sample complexity of list learning tasks in regression.
method Introducing two combinatorial dimensions: k-OIG dimension and k-fat-shattering dimension.
result These dimensions characterize realizable and agnostic k-list regression.
Framework for private, noise-tolerant, and efficient learning algorithms.
problem Private and efficient learning of large-margin halfspaces in noisy environments.
method Simple framework using differential privacy and noise tolerance conditions.
result Noise-tolerant and private PAC learners for large-margin halfspaces with sample complexity independent of dimension.
The paper solves open questions in computable PAC learning, providing a complete landscape.
problem Understanding the boundaries and capabilities of computable PAC learning.
method Analyzing and constructing decidable hypothesis classes with different sample complexities and Littlestone dimensions.
result A complete understanding of CPAC learnability, answering open questions and confirming conjectures.
Recent advances in large-margin classification of data residing in general metric spaces (rather than Hilbert spaces) enable classification under various natural metrics, such as string edit and earthmover distance. A general framework developed for this purpose by von Luxburg and Bousquet [JMLR, 2004] left open the qu…
ContraNorm prevents dimensional collapse in GNNs and Transformers.
problem Dimensional collapse in Graph Neural Networks and Transformers.
method Proposes ContraNorm, a novel normalization layer inspired by contrastive learning.
result Proves ContraNorm alleviates both complete and dimensional collapse under certain conditions.
Let F be a family of Borel measurable functions on a complete separable metric space. The gap (or fat-shattering) dimension of F is a combinatorial quantity that measures the extent to which functions f in F can separate finite sets of points at a predefined resolution gamma > 0. We establish a connection between the g…
The article introduces gamma-Psi-dimensions for margin multi-category classifiers.
problem Margin multi-category classifiers' generalization performance under minimal learnability hypotheses.
method Derives gamma-Psi-dimensions, handles capacity measures, and establishes upper bounds on metric entropies and Rademacher complexity.
result Gamma-Psi-dimensions improve over fat-shattering dimension and offer a promising alternative for multi-class to binary transitions.
Extends capacity analysis to neural networks, showing how capacity is distributed across layers.
problem How capacity is distributed in neural networks with non-linear layers.
method Introduces layer decoupling to quantify non-linear activation's impact, and uses a markovian rule for capacity propagation in deep networks.
result Shows that under certain conditions, capacity allocation in neural networks is equivalent to linear capacity allocation in an extended input space.
In this paper we address the problem of understanding the success of algorithms that organize patches according to graph-based metrics. Algorithms that analyze patches extracted from images or time series have led to state-of-the art techniques for classification, denoising, and the study of nonlinear dynamics. The mai…
New algorithm reduces dictionary learning complexity.
problem Efficiently learning dictionaries from high-dimensional data.
method IcTKM algorithm using dimensionality reduction and fast Fourier transform.
result Locally recovers dictionary with high probability.
New algorithm reduces online learning error for unknown feature distributions.
problem Oracle-efficient hybrid online learning with unknown feature and label distributions.
method Computational efficient online predictor using ERM oracle for finite-VC and fat-shattering classes.
result Oracle-efficient sublinear regret bounds for hybrid online learning with unknown feature generation.
The financial crisis offers new business opportunities in heritage management.
problem Financial institutions' weakened financial condition due to fluctuating real estate property prices.
method Proactive management and stakeholder cooperation to stabilize and optimize properties.
result Properties can serve as a solid base for new business and investment opportunities.
In response to a 1997 problem of M. Vidyasagar, we state a criterion for PAC learnability of a concept class C under the family of all non-atomic (diffuse) measures on the domain Ω. The uniform Glivenko--Cantelli property with respect to non-atomic measures is no longer a necessary condition, and consisten…
Linear classifiers in product space forms improve scRNA-seq data classification.
problem Linear classification in products of Euclidean, spherical, and hyperbolic spaces.
method Novel formulations of linear classifiers on Riemannian manifolds, proving expressive power, and formalizing perceptron and SVM classifiers.
result Linear classifiers in product space forms have the same expressive power as in Euclidean space of the same dimension.
Paper combines RL with policy regularization for inventory policies.
problem Optimizing inventory policies using RL and dynamic programming.
method Hybrid approach combining RL with policy regularization.
result Generalization guarantees for inventory policies using VC theory.
The study analyzes deep neural networks for texture classification, deriving upper bounds and intrinsic dimension insights.
problem Classifying image datasets with texture features using deep neural networks.
method Theoretical analysis using Vapnik-Chervonenkis dimension, Convolutional Neural Networks, Dropout, and Dropconnect networks.
result Upper bounds on the VC dimension of Convolutional Neural Networks and Dropout/Dropconnect networks are derived.
Accuracy on in-distribution data correlates with out-of-distribution data when data is noisy or contains nuisance features.
problem Correlation between in-distribution and out-of-distribution accuracy in noisy or feature-rich data.
method Analyzes the impact of noise and nuisance features on model performance.
result Accuracy on in-distribution and out-of-distribution data can become negatively correlated in noisy or feature-rich data.
CMTRF improves recommendation accuracy by transforming rating scales.
problem Non-linear transformation of rating scales disrupts low-rank structure in rating matrices.
method CMTRF performs regression up to unknown monotonic transforms over user segments, coupled with matrix factorization.
result CMTRF outperforms other baselines in synthetic and real-world datasets.
Quantum machine learning has received significant attention in recent years, and promising progress has been made in the development of quantum algorithms to speed up traditional machine learning tasks. In this work, however, we focus on investigating the information-theoretic upper bounds of sample complexity - how ma…