We propose a novel SPARsity and Clustering (SPARC) regularizer, which is a modified version of the previous octagonal shrinkage and clustering algorithm for regression (OSCAR), where, the proposed regularizer consists of a K-sparse constraint and a pair-wise ℓ∞ norm restricted on the K largest componen…
Heavy-tailed regularization improves deep neural network performance.
problem Improving generalization of deep neural networks.
method Introducing Heavy-Tailed Regularization, using differentiable penalty terms and Bayesian statistics.
result Heavy-tailed regularization outperforms conventional regularization techniques.
Study on convergence rates for optimal transport with regularization.
problem Convergence analysis of divergence-regularized optimal transport.
method Novel methodology using quantization and martingale couplings.
result Sharp rates for various divergences and transport costs.
Novel regularization for Vision Transformers improves model generalization and sparsity.
problem Improving generalization and sparsity in Vision Transformers.
method Likelihood-guided variational Ising-based regularization.
result Improved generalization and sparsity in Vision Transformers.
Develops a new OT framework for class-based data with improved robustness.
problem Understand and recover class structure in optimal transport schemes.
method Proposes a convex OT program with sum-of-norms regularization and an accelerated proximal algorithm.
result The new regularizer preserves class structure better and is more robust to data geometry.
Novel singularity models for 4D harmonic forms and spinors from polytopes.
problem Understanding harmonic forms and spinors in 4D.
method Homogeneous singularity models based on regular 4-polytopes.
result Models describe cones on the 1-skeletal of polytopes.
New method targets sparsity to prevent overfitting in deep nets.
problem Overfitting in deep neural networks with small datasets.
method Targeted sparsity regularization to visualize and counteract overfitting.
result Significant increase in image classification performance without overfitting.
A novel approach to quantizing neural networks using periodic functions as regularizers.
problem Quantization of neural network parameters to reduce memory usage and computational cost.
method Using periodic functions (sine, cosine, hat) as regularizers during training to push weights into discrete points.
result Quantized models achieve the same accuracy as original models on CIFAR-10 and ImageNet datasets.
We stabilize MI-based losses by adding a regularization term, improving their performance and stability.
problem Instability of MI-based losses in machine learning.
method Added a novel regularization term to stabilize MI-based losses.
result Regularization stabilizes training and improves the performance of MI-based losses.
Structured gradient regularizer boosts neural nets' resistance to adversarial attacks.
problem Improving neural networks' robustness against adversarial perturbations.
method Structured gradient regularizer derived from training with noise.
result Structured gradient regularization acts as an effective defense against low-level signal corruption attacks.
Paper analyzes and improves KL-regularized RL for LLMs with logarithmic regret.
problem Improving efficiency of RL fine-tuning for large language models.
method Optimism-based KL-regularized online contextual bandit algorithm with novel regret analysis.
result Achieves an O(ηlog(NRT)⋅dR) logarithmic regret bound. Path regularization improves GFlowNets exploration and generalization.
problem Improving GFlowNets exploration and generalization.
method Path regularization based on optimal transport theory.
result Path regularization enhances GFlowNets to generate more diverse and novel candidates.
A novel graph-regularized CCA approach for datasets with a common source graph.
problem Discovering hidden sources in datasets with common geometry.
method Graph regularizer to encode common sources' geometry in CCA.
result Improved classification performance over competing methods.
Paper introduces DropFilter and DropFilter-PLUS for CNN regularization.
problem Overfitting in Convolutional Neural Networks (CNNs).
method Randomly modifies convolution filters in CNNs.
result Improves performance on image classification tasks.
Study examines stability of image-reconstruction algorithms using variational regularization.
problem Stability and robustness of image-reconstruction algorithms in medical imaging.
method Review and novel stability results for ℓp-regularized linear inverse problems, focusing on p∈(1,∞). result Guarantees Lipschitz continuity for small p and Hölder continuity for larger p in Lp(Ω) function spaces. Sharp bounds on diameter and eigenvalues for amply regular graphs.
problem Finding bounds for amply regular graphs' diameter and eigenvalues.
method New ideas relating discrete Ricci curvature to local matching properties, including a novel construction of a regular bipartite graph.
result Sharp diameter and eigenvalue bounds for amply regular graphs.
Learning the "blocking" structure is a central challenge for high dimensional data (e.g., gene expression data). Recently, a sparse singular value decomposition (SVD) has been used as a biclustering tool to achieve this goal. However, this model ignores the structural information between variables (e.g., gene interacti…
The paper introduces a novel method for training neural network Stein critics with staged L2-regularization.
problem Learning to differentiate model distributions from observed data in high-dimensional settings.
method Developed a novel staging procedure for L2 regularization over training time, leveraging the advantages of highly-regularized training at early times. result Theoretical guarantees and empirical validation show that the method improves the approximation of the training dynamic by the kernel optimization, leading to faster convergence and better performance.
A novel BMC model with nonconvex regularizers and accelerated proximal algorithm for binary matrix completion.
problem Recovering a binary matrix from partial observed positive elements.
method Proposes a novel BMC model with nonconvex regularizers and accelerates proximal algorithm for solving the nonconvex optimization problem.
result The proposed model and algorithm outperform other methods in both synthetic and real-world data sets.
This paper presents a bias-variance tradeoff of graph Laplacian regularizer, which is widely used in graph signal processing and semi-supervised learning tasks. The scaling law of the optimal regularization parameter is specified in terms of the spectral graph properties and a novel signal-to-noise ratio parameter, whi…
Unified regularization framework for visualizing CNNs.
problem Visualizing concepts learned by convolutional neural networks.
method Mathematical framework unifying regularization methods, Sobolev gradients.
result Sobolev filters provide sharper reconstructions and better control over scales.
Novel kernelized LSTD method improves Q-function approximation in RL.
problem Improving policy evaluation in reinforcement learning.
method Manifold regularization applied to kernelized LSTD.
result Superior performance in Q-function approximation compared to existing methods.
Novel optimization method detects change points in Gaussian data.
problem Detecting change points in univariate Gaussian data sequences.
method Continuous optimization for best subset selection (COMBSS) applied to a reformulated statistical inverse problem.
result Adaptation and evaluation of COMBSS for offline normal mean multiple change-point detection.
Gradient-coherent strong regularization improves deep neural networks' generalization.
problem Deep neural networks overfit with strong L1/L2 regularization.
method Imposes regularization only when gradients are coherent, using stochastic gradient descent.
result Significantly improves accuracy and compression (up to 9.9x).
Novel model captures high-dimensional copulas with spectral dynamics and regularization.
problem Modeling time-varying, asymmetric, tail-dependent copulas in high dimensions.
method Score-driven dynamics for eigenvalues, non-linear shrinkage for biases, parsimonious and scalable.
result Model outperforms recent alternatives in capturing co-movements and diversification potential.
Paper tackles incremental few-shot learning with novel classes.
problem Learning new classes with limited data and without re-training.
method Attention Attractor Network (AAN) for incremental few-shot learning.
result AAN helps recognize new classes without forgetting old ones.
SGD without replacement decouples into curvature-following and flatness-regularizing steps.
problem Theoretical analysis of SGD without replacement for large-scale neural networks.
method Analysis of SGD without replacement in a realistic regime, considering high curvature and flatness.
result Optimizing with SGD without replacement is locally equivalent to an additional regularizer step.
Novel method learns memory kernels in Langevin equations.
problem Estimating memory kernels in Langevin equations.
method Regularized Prony method for correlation functions, followed by regression over Sobolev norm-based loss function with RKHS regularization.
result Method outperforms other regression estimators in exponentially weighted L^2 space.
Paper optimizes TSK fuzzy systems for large datasets with MBGD and novel regularization.
problem Optimizing TSK fuzzy systems for large datasets with high dimensionality.
method Proposes MBGD with UR and BN for TSK fuzzy classifiers.
result UR and BN improve classification performance on various UCI datasets.
Inserts proximal mapping into deep networks for better regularization.
problem Effective regularization of deep learning models to handle adversarial perturbations and correlations between modalities.
method Proposes a new layer that directly produces regularized hidden layer outputs using proximal mapping.
result Outperforms state-of-the-art methods in robust temporal learning and multiview modeling.
Bayesian regularizations are explicitly implemented in CNNs, improving deep learning generalization.
problem Improving generalization in deep learning models.
method Introduced a novel probabilistic representation for CNN hidden layers and demonstrated their Bayesian nature.
result CNNs have explicitly Bayesian regularizations based on Bayesian regularization theory.
A novel density regularizer improves data interpolation on non-simply-connected manifolds.
problem Topological differences between model-defined simply-connected manifolds and dataset's non-simply-connected regions.
method Density regularizer to circumvent low-probability-density regions (holes).
result Consistently better interpolation results on real-world image datasets.
Efficiently preconditions machine learning problems with adaptive regularization.
problem Prohibitively expensive full-matrix adaptive regularization for large parameter problems.
method Modified full-matrix adaptive regularization with efficient inverse square root computation.
result Improved convergence rates and better solutions through careful preconditioning.
A new method improves graph-based learning for high-dimensional data.
problem Inconsistent high-dimensional learning efficiency of semi-supervised graph regularization.
method Introducing a novel regularization approach involving centering operation.
result Empirical results show improved performance over spectral clustering.
DRIVE improves IV estimation by accounting for distributional uncertainties.
problem Challenges in IV estimation due to untestable model assumptions and poor finite sample properties.
method DRIVE is a distributionally robust IV estimation method that minimizes a square root TSLS objective with a Wasserstein ambiguity set.
result DRIVE achieves consistency without requiring regularization parameter to vanish, ensuring robustness to distributional uncertainties.
Regularization can improve machine learning models' robustness against poisoning attacks.
problem Poisoning attacks degrade machine learning models' performance by manipulating a fraction of the training data.
method Proposes a novel optimal attack formulation considering the effect of hyperparameters on regularization, leading to better evaluation of robustness.
result Demonstrates that L2 regularization can help mitigate the impact of poisoning attacks. Proposes MR-GAN to improve GAN training by respecting real data manifold geometry.
problem Challenges in training GANs, especially mode collapse and poor generalization.
method Introduces manifold regularizer to regularize GAN training.
result Improves GAN performance in terms of generalization, equilibrium, and stability.
Data augmentation implicitly regularizes deep networks by penalizing rugosity.
problem Understanding generalization in overparameterized deep networks.
method Data augmentation introduces an implicit regularization penalty based on rugosity.
result Data augmentation penalizes rugosity, leading to better generalization.
Fiedler regularization uses spectral graph theory to improve neural network performance.
problem Improving neural network performance by penalizing weights based on connectivity.
method Uses the Fiedler value of the neural network's graph as a regularization tool, providing theoretical and computational methods.
result Demonstrates Fiedler regularization's effectiveness in improving neural network performance.
Paper develops a method to learn optimal sparsity-promoting regularizers for linear inverse problems.
problem Solving linear inverse problems with sparse solutions.
method Bilevel optimization framework to select an optimal synthesis operator B. result Established well-posedness and theoretical guarantees for the learning process.
Tikhonov regularization is robust under specific martingale constraints in distributionally robust optimization.
problem Distributionally robust optimization and regularization of learning models.
method Optimal transport approach with martingale constraints.
result Tikhonov regularization is optimal transport robust under specified martingale constraints.
Fiedler regularization uses graph sparsity to improve neural network training.
problem Improving neural network training by respecting graph structure.
method Using the Fiedler value of the neural network's graph as a regularization tool.
result Fiedler regularization outperforms traditional methods like dropout and weight decay.
We present a data dependent generalization bound for a large class of regularized algorithms which implement structured sparsity constraints. The bound can be applied to standard squared-norm regularization, the Lasso, the group Lasso, some versions of the group Lasso with overlapping groups, multiple kernel learning a…
Paper tackles selfless sequential learning with neural inhibition to improve future task capacity.
problem Learning tasks in sequence with limited model capacity.
method Study regularization strategies and activation functions, proposing a novel representation sparsity regularizer.
result Representation sparsity regularizer improves performance over alternative regularizers.
RotationOut rotates input vectors to regularize neural networks.
problem Reduction of co-adaptation in neural networks.
method Randomly rotates input vectors of the input layer.
result RotationOut reduces co-adaptation better than Dropout.
New method uses deep learning to solve inverse problems with provable guarantees.
problem Solving inverse problems with high-quality results and provable guarantees.
method Convex-Nonconvex (CNC) framework with input weakly convex neural network (IWCNN).
result The method provides provably convergent regularization for inverse problems.
Adversarial training linked to operator norm regularization, proving network sensitivity to attacks.
problem Robustifying neural networks against adversarial attacks.
method Theoretical link established between adversarial training and operator norm regularization.
result Adversarial training is equivalent to data-dependent operator norm regularization.
Optimizes quantile and semi-adversarial regret with novel root-logarithmic regularizers.
problem Minimizes regret in adversarial and semi-adversarial online learning.
method FTRL with root-logarithmic regularizers for quantile and semi-adversarial settings.
result Achieves minimax optimal regret bounds in both paradigms.