ALPCAH improves PCA for noisy data by estimating sample-wise noise variances.
problem Noisy data with varying noise levels in different samples.
method Sample-wise heteroscedastic PCA with tail singular value regularization.
result Improves subspace basis estimation for low-rank data.
Localized Lasso improves interpretable models for high-dimensional data.
problem High-dimensional regression with small sample size and interpretability.
method Sample-wise network regularization and exclusive group sparsity.
result Localized Lasso outperforms alternatives in simulated and genomic data.
New bounds show limitations of sample-wise information-theoretic generalization.
problem Limitations of sample-wise information-theoretic generalization bounds.
method Analysis of existing bounds and derivation of new bounds.
result No sample-wise information-theoretic bounds exist for expected squared generalization gap.
New method enhances neural network robustness against adversarial attacks.
problem Enhancing neural network robustness against adversarial attacks.
method Variational framework with per-sample noise level selector.
result Enhanced empirical robustness and certified robustness.
Estimates interactions between modalities for multimodal data.
problem Accurately quantifying interactions between different data types.
method Developed Lightweight Sample-wise Multimodal Interaction (LSMI) estimator using pointwise information theory.
result LSMI reveals fine-grained dynamics in multimodal data.
Measures sample learnability across DNNs, showing consistency.
problem Estimating the learnability of each sample in a training set.
method Train DNN on training set, aggregate hits and misses over epochs.
result Sample-wise learnability measure is highly correlated across DNN models.
Paper proposes a method to adapt domains without target data using attribute information.
problem Domain adaptation without target data when prior attribute changes exist.
method Reweight source data with estimated sample-wise weights based on attribute prior.
result Method provides more precise transferability estimation than attribute-based reweighting.
AdaTrans adapts to feature and sample transfer in high-dimensional regression.
problem High-dimensional linear regression with more features than samples.
method F-AdaTrans and S-AdaTrans methods using fused-penalties and adaptive weights.
result AdaTrans achieves convergence rates close to oracle estimators and near-minimax optimal rates.
More data can actually hurt linear regression performance in certain conditions.
problem The test risk of linear regression estimators increases with additional samples in overparameterized settings.
method An analysis of linear regression with isotropic Gaussian covariates using gradient descent.
result The bias decreases with more samples, but variance increases, leading to a surprising increase in test risk.
The paper analyzes the risk of bagging regularized M-estimators under proportional asymptotics.
problem Characterizing the risk of ensemble estimators trained with subsamples and regularizers.
method Developed a consistent estimator for the risk of ensemble estimators under proportional asymptotics.
result Optimal subsample size k⋆ tends to be in the overparameterized regime for the full-ensemble estimator. ALPCAH improves PCA for noisy data samples.
problem Heteroscedastic data with varying noise levels.
method Subspace learning method estimating sample-wise noise variances.
result Improves subspace basis for low-rank data.
Method removes misleading data to improve ML model accuracy.
problem Unhealthy fear of missing out on data leads to model instability and poor performance.
method Bayesian sequential selection method that identifies and selects critical information.
result Improves sample-wise error convergence and eliminates model instabilities.
The paper studies how more data affects prediction risk in high-dimensional models.
problem The impact of increasing data on prediction risk in high-dimensional models.
method Derives central limit theorem and provides finite-sample distribution and confidence interval for prediction risk.
result Demonstrates 'more data hurt' phenomenon in high-dimensional least squares estimation.
The paper explains two distinct peaks in generalization error for neural networks and simpler models, each governed by different factors.
problem Understanding the peaks in generalization error for neural networks and simpler models.
method Analysis of random feature models and comparison with numerical experiments involving deep neural networks.
result The peaks at N=P and N=D are distinct and governed by different factors (noise sensitivity vs. initialization noise). FCDD improves image anomaly detection without post-hoc explainers.
problem Image anomaly detection, especially pixel-wise.
method Fully Convolutional Data Description (FCDD) directly addresses anomaly detection without post-hoc methods.
result FCDD achieves state-of-the-art results on pixel-wise AD tasks.
CDSSL improves representation quality by integrating linear and nonlinear dependencies.
problem Scarcity of labeled data and neglect of nonlinear dependencies in SSL.
method CDSSL combines linear correlations and nonlinear dependencies using HSIC in RKHS.
result CDSSL enhances representation quality on diverse benchmarks.
Clapping reduces memory usage in distributed optimization by reusing data samples.
problem Significant communication overhead and impractical memory overhead in pipeline-parallel distributed optimization.
method Lazy sampling strategy to reuse data samples across steps, supporting convergence without unbiased gradient assumptions.
result Clapping achieves convergence in few-epoch or online training regimes without sample-size memory overhead.
Riemannian metric matching learns the geometry of high-dimensional datasets using neural networks.
problem Estimating the geometry of high-dimensional datasets from samples
method Riemannian metric matching using neural networks
result Riemannian metric matching rivals or improves k-NN-based diffusion geometry estimators Study shows double and triple descent in unsupervised autoencoders, improving performance in various tasks.
problem Exploring the phenomenon of double descent in unsupervised learning.
method Analytical demonstration and extensive experiments on synthetic and real datasets.
result Over-parameterized unsupervised autoencoders exhibit double and triple descent, enhancing performance in downstream tasks.
The paper analyzes how the one-dimensional Wasserstein distance captures pointwise density differences in finite samples.
problem Uncertainty in identifying density differences when supports overlap and densities have substantial pointwise differences.
method Analysis using the Poisson process and neural spike train decoding.
result The one-dimensional Wasserstein distance highlights meaningful density differences related to both rate and support.
ALPCAHUS clusters data from multiple subspaces with varying noise.
problem Heteroscedastic data with varying noise levels.
method Develops a heteroscedastic PCA method for subspace clustering.
result Improves subspace clustering by accounting for sample-wise noise variances.
Noise makes data unlearnable by tricking models.
problem Unauthorized exploitation of personal data by deep learning models.
method Error-minimizing noise to reduce training examples' learnability.
result Error-minimizing noise can make training examples unlearnable by deep learning models.
Study compares random and learned features in deep Bayesian linear models.
problem Understanding how feature learning affects generalization in deep learning.
method Comparing deep random feature models to deep networks with trained layers.
result Random feature models can display double-descent behavior, while deep networks do not.
RID-Noise improves robust design under noisy conditions using neural networks.
problem Design robustness under noisy environments.
method Robust Inverse Design under Noise (RID-Noise) using conditional invertible neural networks (cINNs).
result RID-Noise achieves more effective robust design compared to state-of-the-art methods.
The paper analyzes learning curves for kernel ridge regression with dot-product kernels.
problem Understanding the learning curves for different scaling regimes of data and model.
method Precise formulas for mean test error, bias, and variance in the mo∞ with m/dr constant regime. result A peak in the learning curve at m≈dr/r! for any integer r. The paper extends PAC-learning to handle evasion adversaries, finding limits on what can be learned.
problem Evasion attacks on machine learning models during testing.
method Extending PAC-learning framework to include evasion adversaries, defining corrupted hypothesis classes, and deriving adversarial VC-dimension.
result The adversarial VC-dimension can be larger or smaller than the standard VC-dimension, offering new insights.
AIR-Net adapts low-rank regularization dynamically for better image completion.
problem Fixed low-rank regularization limits adaptability to different images.
method AIR-Net uses adaptive and implicit regularization parameterized by a dynamic Laplacian matrix.
result AIR-Net enhances implicit regularization and outperforms fixed methods in non-uniform missing data scenarios.
A 6-regular triangulation for hyperbolic plane created.
problem Creating a 6-regular triangulation for hyperbolic plane.
method Constructed a 6-regular geodesic triangulation.
result A 6-regular geodesic triangulation of the hyperbolic plane was successfully created.
The article explores toric spaces of regular polyhedra, highlighting rational and non-rational cases.
problem Exploring toric spaces associated with regular convex polyhedra.
method Symplectic and complex toric spaces associated with five regular convex polyhedra.
result The regular dodecahedron and icosahedron cannot be treated via standard toric geometry.
Gradient descent implicitly regularizes neural networks by penalizing large loss gradients.
problem How to optimize deep neural networks without explicit regularization.
method Backward error analysis to calculate implicit gradient regularization and demonstrate its effectiveness empirically.
result Implicit gradient regularization biases gradient descent toward flat minima, improving model robustness and test errors.
Regularized deep networks improve generalization and robustness.
problem Improving generalization and robustness of deep neural networks.
method Input gradient regularization combined with Lipschitz and adversarial robustness.
result Regularized models show improved adversarial robustness and generalization.
Choquet regularization improves exploration in RL.
problem Improving exploration in reinforcement learning.
method Introducing Choquet regularizers to measure and manage exploration, reformulating RL problems and deriving explicit solutions.
result Explicit optimal distributions and Choquet regularizers for various exploratory samplers.
Gradient-coherent strong regularization improves deep neural networks' generalization.
problem Deep neural networks overfit with strong L1/L2 regularization.
method Imposes regularization only when gradients are coherent, using stochastic gradient descent.
result Significantly improves accuracy and compression (up to 9.9x).
The paper explores optimal regularizers for data sources, linking them to star bodies.
problem Understanding optimal regularizers for data sources.
method Investigates optimal regularizers for data distributions using star bodies and dual Brunn-Minkowski theory.
result Identifies optimal regularizers and assesses amenability to convex regularization.
A triangulation of a connected closed surface is called weakly regular if the action of its automorphism group on its vertices is transitive. A triangulation of a connected closed surface is called degree-regular if each of its vertices have the same degree. Clearly, a weakly regular triangulation is degree-regular. In…
The paper proves existence and multiplicity of affine connections on regular manifolds.
problem Existence and multiplicity of affine connections on regular manifolds.
method Regularity theory and properties of the structural presheaf.
result The space of regular affine connections is an affine space of the space of regular End(TM)-valued 1-forms. Paper develops a new theory on Wasserstein DRO's variation regularization effect.
problem Developing a new theory for Wasserstein DRO's regularization effect.
method General theory on variation regularization effect of Wasserstein DRO.
result New generalization guarantees for adversarial robust learning.
The paper studies convergence rates of Tsallis entropic regularization in optimal transport.
problem Optimal transport with regularization.
method Γ-convergence and quantization/shadow arguments.
result Derives convergence rate of Tsallis entropic regularization.
PL Morse theory proves strong regularity in low dimensions.
problem Understanding regular and critical points in PL manifolds.
method Introducing homologically and strongly regular points, presenting criteria, and constructing examples.
result In low dimensions d≤4, homologically regular points are always strongly regular. Study uses property elicitation to understand how fairness regularizers affect optimal decisions.
problem Understanding how fairness regularizers change the optimal decision in predictive algorithms.
method Property elicitation to analyze the relationship between loss, regularization, and optimal decision.
result Necessary and sufficient condition for when a property changes with the addition of a regularizer.
Fiedler regularization uses spectral graph theory to improve neural network performance.
problem Improving neural network performance by penalizing weights based on connectivity.
method Uses the Fiedler value of the neural network's graph as a regularization tool, providing theoretical and computational methods.
result Demonstrates Fiedler regularization's effectiveness in improving neural network performance.
New input gradient regularization improves adversarial robustness efficiently.
problem Improving adversarial robustness in machine learning models.
method Derive robustness bounds, implement scaleable input gradient regularization, avoid double backpropagation.
result Input gradient regularization is competitive with adversarial training and avoids gradient obfuscation.
Study on convergence rates for optimal transport with regularization.
problem Convergence analysis of divergence-regularized optimal transport.
method Novel methodology using quantization and martingale couplings.
result Sharp rates for various divergences and transport costs.
Dropout is a simple but effective technique for learning in neural networks and other settings. A sound theoretical understanding of dropout is needed to determine when dropout should be applied and how to use it most effectively. In this paper we continue the exploration of dropout as a regularizer pioneered by Wager,…
In this paper, we give a new generalization error bound of Multiple Kernel Learning (MKL) for a general class of regularizations, and discuss what kind of regularization gives a favorable predictive accuracy. Our main target in this paper is dense type regularizations including \ellp-MKL. According to the recent numeri…
Improved optimal regularity for harmonic almost complex structures.
problem Establishing optimal regularity for harmonic almost complex structures.
method Quantitative stratification method and rectifiability of singular strata.
result Optimal regularity theory for energy minimizing harmonic almost complex structures.
We establish continuous maximal regularity results for parabolic differential operators acting on sections of tensor bundles on Riemannian manifolds. As an application, we show that solutions to the Yamabe flow instantaneously regularize and become real analytic in space and time. The regularity result is obtained by i…
Study on the regularity of p-Gauss curvature flow near flat interfaces.
problem Regularity of p-Gauss curvature flow near flat interfaces. method Analysis of convex hypersurface near the interface.
result Regularity of the convex hypersurface near the interface.