Framework reduces simplicity bias in NNs, improving OOD generalization and robustness.
problem Simplicity bias in deep learning models leads to biased predictions and poor OOD generalization.
method Proposes a framework that regularizes conditional mutual information to encourage use of diverse features.
result Demonstrates effectiveness in various settings, enhancing OOD generalization and robustness.
Research reveals simplicity bias in random logistic map, impacting data analysis and forecasting.
problem Simplicity bias in dynamical systems and its impact on data analysis and prediction.
method Examined the logistic map and random logistic map, focusing on simplicity bias and noise effects.
result Simplicity bias is observable in the random logistic map, persisting even with small noise levels.
Adam avoids simplicity bias in neural networks, leading to better generalization.
problem Simplicity bias in neural networks trained with SGD.
method Comparison of Adam and GD on binary classification tasks with Gaussian data.
result Adam leads to richer and more diverse features, improving generalization.
The study reveals simplicity bias in neural networks leading to better compositional mappings.
problem Understanding when and how to encourage neural networks to learn compositional mappings.
method Examined compositional mappings through coding length and gradient descent dynamics.
result Neural networks tend to learn the simplest bijections, explaining their good generalization.
Two-layer networks favor simple features, especially in complex datasets.
problem Simplicity bias in neural networks over-reliing on simple features.
method Characterization of two-layer neural networks with small weights and gradient flow.
result Features learned in middle training stages are more useful for out-of-distribution transfer.
Neural nets learn simple distributions first, then more complex ones.
problem Understanding how neural networks generalize from simple to complex functions.
method Stochastic gradient descent training, synthetic data, CIFAR10, ImageNet pre-training.
result Neural networks initially use lower-order statistics, then higher-order ones.
Spectral regularization simplifies sequence models by focusing on grammatical simplicity.
problem Sequence modeling challenges in learning tasks.
method Introduces spectral regularization based on Hankel matrices and trace norm, addressing bi-infinite matrices with an unbiased estimator.
result Demonstrates spectral regularization's potential benefits on Tomita grammars.
New datasets reveal neural networks can rely on simple features, leading to poor generalization.
problem Neural networks' reliance on simple features can lead to poor generalization and robustness.
method Designing datasets with varying levels of simplicity and incorporating non-robustness.
result Neural networks can exclusively rely on the simplest feature, leading to poor performance on complex data.
Neural networks favor simple features over complex ones, even when complex features are available.
problem Neural networks exhibit a bias towards simple features over complex ones, even when complex features are present.
method Rigorously defined simplicity bias, theoretical and empirical demonstrations, ensemble approach to improve robustness.
result One hidden layer neural networks favor simple features over complex ones, even in the presence of more robust features.
Label noise SGD converges to a simple model with a single linear feature.
problem Understanding the simplicity bias in neural network training.
method Analyzing the convergence of label noise SGD on two-layer neural networks.
result Label noise SGD converges to a model with a single linear feature.
Simplifies RL training with fewer techniques, reducing bias and instability.
problem Training instabilities and high sample complexity in RL.
method Introduced a simple deterministic policy gradient, used propensity estimation, and delayed policy updates.
result Improved performance and reduced sample complexity through these techniques.
A new reinforcement learning method reduces action complexity for robust control.
problem Deep reinforcement learning's susceptibility to spurious correlations.
method Minimizing trajectory entropy to encourage simple, predictable actions.
result Trajectory Entropy Reinforcement Learning achieves superior performance and robustness.
The Nadaraya-Watson kernel estimator is among the most popular nonparameteric regression technique thanks to its simplicity. Its asymptotic bias has been studied by Rosenblatt in 1969 and has been reported in a number of related literature. However, Rosenblatt's analysis is only valid for infinitesimal bandwidth. In co…
Two-layer ReLU networks often converge to simpler solutions, improving generalization.
problem Understanding generalization in overparametrized neural networks, especially for complex tasks.
method Theoretical analysis of two-layer ReLU networks, focusing on the early alignment phase.
result Two-layer ReLU networks often converge to simpler solutions rather than interpolating the training data, leading to better generalization.
Random SNNs are stable and simple, with low-frequency Fourier spectra.
problem Stability and robustness of spiking neural networks.
method Boolean function analysis and Fourier spectrum concentration.
result Random LIF-SNNs are stable and biased towards simple functions.
Bayesian approaches have become increasingly popular in causal inference problems due to their conceptual simplicity, excellent performance and in-built uncertainty quantification ('posterior credible sets'). We investigate Bayesian inference for average treatment effects from observational data, which is a challenging…
The study identifies spurious correlations in high-dimensional regression and quantifies their impact.
problem Spurious correlations in high-dimensional regression models.
method Statistical characterization of spurious correlations, quantifying their amount via ridge regularization.
result The value of regularization strength that minimizes test loss is in an interval where spurious correlations increase.
Exactly solvable model reveals how data geometry influences ML bias.
problem How data geometry affects machine learning bias.
method High-dimensional data imbalance model, statistical physics tools.
result Exact predictions for fairness metrics and mitigation strategies.
Decomposes bias in linear models under demographic parity constraints.
problem Understanding and quantifying bias in linear models under fairness constraints.
method Post-processing framework to decompose bias into direct and indirect components.
result Analytical characterization of how demographic parity reshapes model coefficients.
TopoNTK kernel captures higher-order interactions in simplicial complexes.
problem Graph neural networks miss higher-order interactions in relational systems.
method Introduces TopoNTK, an infinite-width kernel for simplicial message passing.
result TopoNTK captures topology invisible to graph kernels, improving expressivity and interpretability.
Contrastive learning struggles with class collapse and feature suppression, revealing bias towards simpler solutions.
problem Contrastive learning struggles with class collapse and feature suppression, especially in supervised and unsupervised settings.
method Unified theoretical framework to determine which features are learnt by CL, revealing bias towards simpler solutions.
result Bias towards simpler solutions is a key factor in class collapse and feature suppression.
Algorithms are increasingly used to aid, or in some cases supplant, human decision-making, particularly for decisions that hinge on predictions. As a result, two additional features in addition to prediction quality have generated interest: (i) to facilitate human interaction and understanding with these algorithms, we…
Deep neural networks (DNNs) generalize remarkably well without explicit regularization even in the strongly over-parametrized regime where classical learning theory would instead predict that they would severely overfit. While many proposals for some kind of implicit regularization have been made to rationalise this su…
Combining explicit and implicit regularization improves deep learning performance without needing depth.
problem Improving deep learning performance without increasing model complexity.
method Proposes an explicit penalty to mirror implicit regularization bias in adaptive gradient optimizers.
result Single-layer networks can achieve low-rank approximations with similar performance to deep linear networks.
The paper develops formulas for hyperbolic simplices based on edge lengths.
problem Understanding the geometry of hyperbolic simplices using only edge lengths.
method Develops geometric formulas for hyperbolic simplices based on edge lengths.
result Distance and projection formulas in hyperbolic simplices.
Neural networks learn simpler features first, then more complex ones; Fourier analysis reveals this pattern.
problem Understanding the learning dynamics of neural networks, especially with natural image data.
method Fourier analysis of translation-invariant and power-law spectra to study feature learning.
result Simple neural networks first rely on amplitude information, then phase information, and power-law spectra can accelerate learning phase information.
NoisyDARTS injects random noise to improve neural architecture search.
problem Performance collapse in Differentiable Architecture Search (DARTS).
method Inject unbiased random noise to skip connections to impede gradient flow.
result NoisyDARTS achieves state-of-the-art results across various tasks.
Estimates dimensions of maximal simplices for rational and irrational trees in Outer space.
problem Understanding the structure of trees in Outer space.
method Associate simplices to R-trees and estimate their dimensions. result Estimates the dimensions of maximal simplices for both rational and irrational trees.
The paper establishes conditions for Riemannian connections and semi-simplicity of Lie algebras using spray structures.
problem Conditions for Riemannian connections and semi-simplicity of Lie algebras.
method Using almost product structures and spray, the paper provides necessary and sufficient conditions for these properties.
result Equivalence of semi-simplicity of Lie algebras to derived ideal coincidence, interiority of derivations, and adjoint representation semi-simplicity.
It is proved that the volume of spherical or hyperbolic simplices, when considered as a function of the dihedral angles, can be extended continuously to degenerated simplices.
Develops SCMs for latent selection to simplify causal analysis.
problem Latent selection complicates causal analysis.
method Introduces a conditioning operation for SCMs to encode latent selection.
result Conditioning operation preserves simplicity, acyclicity, and linearity of SCMs.
Geodesic simplices in pseudo-hyperbolic space get a cohomological treatment.
problem Understanding geodesic simplices in pseudo-hyperbolic space.
method Cohomological interpretation and necessary/sufficient condition formulation.
result Every ideal geodesic polytope in (2,2) pseudo-hyperbolic space has finite volume. We consider the problem of estimating the parameters of a d-dimensional rectified Gaussian distribution from i.i.d. samples. A rectified Gaussian distribution is defined by passing a standard Gaussian distribution through a one-layer ReLU neural network. We give a simple algorithm to estimate the parameters (i.e., th…
Paper proposes a robust LPR method using similarity kernels.
problem Outliers and high-leverage points affect traditional LPR's accuracy.
method Integrates predictor and response variables in weighting mechanism using a conditional density kernel.
result Lower empirical bias compared to iterative robust LOWESS.
A new CA-GAN architecture improves minority class data generation in health datasets.
problem Algorithmic bias due to health data poverty and underrepresentation of minority groups.
method Proposes CA-GAN architecture to address shortcomings of resampling and GAN-based approaches.
result CA-GAN outperforms SMOTE and WGAN-GP* in generating authentic minority class data and maintaining original distribution.
Study PL bordism theories with quantitative bounds on filling simplices.
problem Understanding PL bordism theories with geometric constraints.
method Quantitative analysis of PL manifolds and exotic theories.
result Bounding the number of simplices in fillings of cycles.
A sparse modeling is a major topic in machine learning and statistics. LASSO (Least Absolute Shrinkage and Selection Operator) is a popular sparse modeling method while it has been known to yield unexpected large bias especially at a sparse representation. There have been several studies for improving this problem such…
The translation equivariance of convolutional layers enables convolutional neural networks to generalize well on image problems. While translation equivariance provides a powerful inductive bias for images, we often additionally desire equivariance to other transformations, such as rotations, especially for non-image d…
New framework shows C∗-simplicity for groups without certain subalgebras.
problem Characterizing C∗-simplicity of groups. method Introducing confined subalgebras and Uniformly Recurrent States.
result A countable discrete group is C∗-simple if it has no non-trivial amenable confined subalgebras. In this article, we prove a theorem comparing the dihedral angles of simplices in the hyperbolic, spherical and Euclidean geometries.
The paper explains generalization in kernel regression and deep neural networks using spectral bias and task-model alignment.
problem Understanding generalization in machine learning models, especially deep neural networks.
method Analytical expression for generalization error derived from statistical mechanics, applied to various kernels and data distributions.
result Spectral bias and task-model alignment explain generalization in kernel regression and deep neural networks.
We study a natural intrinsic definition of geometric simplices in Riemannian manifolds of arbitrary dimension n, and exploit these simplices to obtain criteria for triangulating compact Riemannian manifolds. These geometric simplices are defined using Karcher means. Given a finite set of vertices in a convex set on t…
The Apollonius theorem is generalized for m-simplices, with applications in geometry and optimization.
problem Generalizing the Apollonius theorem for m-simplices.
method Direct generalization of the theorem to m-simplices in n-dimensional space.
result Applications in geometry and optimization, including minimal surface enclosures, simplex thickness, and root-finding methods.
Similar simplices can be inscribed in most smoothly embedded spheres.
problem Inscribing families of similar simplices in spheres.
method Diffeomorphic mapping and techniques from previous work on inscribing triangles.
result A dense family of spheres allows inscribing similar simplices of every pose.
Graph-based kernels improve GP performance on graph data.
problem Improving Gaussian process performance on graph-structured data.
method Introduced graph neural network-inspired kernels into Gaussian processes.
result Graph convolutional networks are equivalent to certain GP kernels when infinitely wide.
We study prismatics sets analogously to simplical sets except that realization involves prisms, i.e., products of simplices rather than just simplices. Particular examples are the prismatic subdivision of a simplicial set S and the prismatic star of S. Both have the same homotopy type as S and in particular the latter …
Simplicial sets deformation retract onto transverse simplices.
problem Deformation retraction of simplicial sets.
method Showed deformation retraction of singular simplicial set onto transverse simplices.
result Singular simplicial set deformation retracts onto transverse simplices.
We generalize the very well known boundary operator of the ordinary singular homology theory, defined in many books about algebraic topology. We describe a variant of this ordinary simplicial boundary operator where the usual boundary (n-1)-simplices of each n-simplex are replaced by combinations of internal (n-1)- sim…