Cross-regularization adapts model complexity during training.
problem Manual tuning of model complexity for overfitting prevention.
method Directly adapts regularization parameters through validation gradients during training.
result Organic emergence of architecture-specific regularization during training.
This paper proves geodesic curvature measures are bounded for curves near cross cap singularities.
problem Boundedness of geodesic curvature measures near cross cap singularities.
method Analyzes intrinsic cross cap singularities and extends Gauss-Bonnet formula.
result Proves boundedness of geodesic curvature measures for curves near cross cap singularities.
The paper improves ALO for ℓ1-regularized models.
problem Estimating out-of-sample error for ℓ1-regularized models. method Developed a novel theory for ℓ1-regularized problems, bounding ALO error. result For ℓ1-regularized problems, ALO error goes to zero as p goes to infinity. Optimizes hyperparameter tuning for models using approximate leave-one-out cross-validation.
problem Finding optimal hyperparameters for regularized models using approximate leave-one-out cross-validation.
method Derive efficient formulas for gradient and hessian of approximate leave-one-out cross-validation, apply second-order optimization.
result Demonstrates the effectiveness of the approach on real-world data sets.
Tutorial defines and explains overfitting, cross validation, regularization, bagging, and boosting.
problem Understanding and preventing overfitting in machine learning models.
method Defining and explaining mean squared error, variance, covariance, and bias; using Stein's Unbiased Risk Estimator (SURE); introducing cross validation, regularization, bagging, and boosting.
result Boosting prevents overfitting by reducing variance and improving generalization.
We develop an approximate formula for evaluating a cross-validation estimator of predictive likelihood for multinomial logistic regression regularized by an ℓ1-norm. This allows us to avoid repeated optimizations required for literally conducting cross-validation; hence, the computational time can be significantl…
The key condition A3w of Ma, Trudinger and Wang for regularity of optimal transportation maps is implied by the nonnegativity of a pseudo-Riemannian curvature -- which we call cross-curvature -- induced by the transportation cost. For the Riemannian distance squared cost, it is shown that (1) cross-curvature nonnegativ…
DNN-based cross-modal retrieval has become a research hotspot, by which users can search results across various modalities like image and text. However, existing methods mainly focus on the pairwise correlation and reconstruction error of labeled data. They ignore the semantically similar and dissimilar constraints bet…
Proposes a new method for manifold alignment using geometry-regularized twin autoencoders.
problem Traditional MA methods lack out-of-sample extension and real-world applicability.
method Guided representation learning with geometry-regularized twin autoencoders.
result Improves cross-domain generalization and robustness while maintaining alignment fidelity.
Improved Kriging model reduces prediction errors.
problem Improving prediction accuracy in Kriging models.
method Theta-regularized Kriging model with Lasso, Ridge, and Elastic-net penalties.
result The Theta-regularized Kriging model outperforms other penalized Kriging models in accuracy and stability.
An increasing sequence of integers is said to be universal for knots if every knot has a reduced regular projection on the sphere such that the number of edges of each complementary face of the projection comes from the given sequence. Adams, Shinjo, and Tanaka have, in a work, shown that (2,4,5) and (3,4,n) (where n i…
Unified study of ridge regression structure, cross-validation, and acceleration.
problem Understanding and optimizing ridge regression in large-data settings.
method Unified large-data linear model analysis, cross-validation bias correction, sketching accuracy study.
result Unified understanding and improved methods for ridge regression.
Recent GAN-based architectures have been able to deliver impressive performance on the general task of image-to-image translation. In particular, it was shown that a wide variety of image translation operators may be learned from two image sets, containing images from two different domains, without establishing an expl…
Unified framework for estimating high-dimensional conditional factor models.
problem Estimating high-dimensional conditional latent factor models with practical limitations.
method Constrained nuclear norm regularization and cross-validation for parameter selection.
result Imposing homogeneity improves model predictability, with new method outperforming alternatives.
Study Dirac operators on finite warped cylinders with gauge fields.
problem Characterize spectral flow on finite warped cylinders with gauge fields.
method Identify endpoint operators, derive determinant characterization, introduce regularized APS conditions.
result Regularized APS conditions admit a spectral-flow framework, matching zero-mode sets.
The paper analyzes trace regression with low-rank matrices under various regularization methods.
problem Estimating low-rank matrices with near-optimal error bounds under unknown regularization parameters.
method General spikiness notion, restricted strong convexity of sampling operator, cross-validation for parameter selection.
result Cross-validated estimators select near-optimal penalty parameters and outperform theory-inspired approaches.
Study evaluates various regularization methods for electricity price forecasting.
problem Improving accuracy of electricity price predictions.
method Applied ten different penalty functions to two model structures in two electricity markets.
result LQ and elastic net consistently produce more accurate forecasts than other regularization types.
Researchers analyze inverse optimal transport, deriving theoretical and empirical insights.
problem Understanding the inverse problem of inferring cost matrices from optimal couplings.
method Formalized and analyzed using entropy-regularized optimal transport, with theoretical and empirical contributions.
result Characterization of the manifold of cross-ratio equivalent costs and derivation of an MCMC sampler.
New method estimates spatial weights matrix for lattice data, improving prediction accuracy.
problem Estimating spatial dependence structure for regular lattice data.
method Adaptive lasso with cross-sectional resampling to estimate sparse spatial weights matrix.
result Improves prediction accuracy of nitrogen dioxide concentrations.
Careful tuning of a regularization parameter is indispensable in many machine learning tasks because it has a significant impact on generalization performances. Nevertheless, current practice of regularization parameter tuning is more of an art than a science, e.g., it is hard to tell how many grid-points would be need…
CLSB models system dynamics from cross-sectional data with population-level regularization.
problem Challenges in modeling system dynamics from limited cross-sectional samples and heterogeneous individual behaviors.
method Introduces CLSB framework for learning dynamics, regularized for population-level temporal variations.
result Empirically superior in single-cell sequencing data analyses, e.g., simulating cell development and drug response.
Regularization improves stability and consistency of sparse autoencoders.
problem Varying features across random seeds and training choices in SAEs.
method Added L1 or L2 penalties on encoder and decoder weights.
result L2 regularization increases cross-seed feature consistency.
In this paper, we extend the existence and regularity theorems for Kähler-Einstein metrics having conic singularities along a simple normal crossing divisor to the case of normal crossing divisor, i.e. when components of the divisor are allowed to intersect themselves transversely.
New method for cross-validation in high-dimensional data with dependent or heavy-tailed covariates.
problem Inconsistent cross-validation in high-dimensional settings with dependent or heavy-tailed covariates.
method ROTI-GCV framework for cross-validation under proportional asymptotics regime.
result Demonstrated accuracy of ROTI-GCV in synthetic and semi-synthetic settings.
We investigate polyhedral 2k-manifolds as subcomplexes of the boundary complex of a regular polytope. We call such a subcomplex {\it k-Hamiltonian} if it contains the full k-skeleton of the polytope. Since the case of the cube is well known and since the case of a simplex was also previously studied (these are so…
Proposes a new method for selecting regularization parameters in sparse precision matrix estimation.
problem Selecting an appropriate regularization parameter for sparse precision matrix estimation.
method Developed a closed-form matrix-valued regularization parameter based on the sampling distribution of optimality conditions.
result The proposed method achieves comparable estimation accuracy and superior support recovery to cross-validation, with significant runtime improvements.
LPL optimizes embeddings to align local neighborhoods, improving cross-lingual word alignment.
problem Aligning embeddings across different datasets and languages.
method Locality Preserving Loss (LPL) optimizes model to project embeddings while maintaining local neighborhoods and aligning them.
result LPL-based alignment leads to better and consistent accuracy, especially in small training set settings.
The paper analyzes the risk of CV-tuned regularized estimators and connects it to SURE.
problem Understanding the risk of CV-tuned regularized estimators.
method Derives asymptotic risk function of CV-tuned estimators and connects it to SURE.
result The risk function provides a more detailed picture of predictive performance than uniform bounds.
Manifold regularization, such as laplacian regularized least squares (LapRLS) and laplacian support vector machine (LapSVM), has been widely used in semi-supervised learning, and its performance greatly depends on the choice of some hyper-parameters. Cross-validation (CV) is the most popular approach for selecting the …
New method classifies C-boundaries up to 6 crossings.
problem Classifying C-boundaries with up to 6 crossings. method Proposed a new construction method.
result Extended classification of C-boundaries up to 6 crossings. This work improves deep learning from noisy crowdsourced labels.
problem Learning label correction and neural classifier from noisy crowdsourced data.
method Coupled Cross-Entropy Minimization (CCEM) with identifiability and regularization.
result The CCEM criterion correctly identifies annotators' confusion and neural classifier under realistic conditions.
In high-dimensional data analysis, regularization methods pursuing sparsity and/or low rank have received a lot of attention recently. To provide a proper amount of shrinkage, it is typical to use a grid search and a model comparison criterion to find the optimal regularization parameters. However, we show that fixing …
Optimal model improves AUC, recall, and F1 score for class-imbalanced business risk.
problem Improving prediction of class-imbalanced business risk.
method Resampling, regularization, and model ensembling techniques.
result Boosting on DT with SMOTE oversampling achieves AUC, recall, and F1 score of 0.8633, 0.9260, and 0.8907, respectively.
Many classical objects on a surface S can be interpreted as cross-ratio functions on the circle at infinity of the universal covering. This includes closed curves considered up to homotopy, metrics of negative curvature considered up to isotopy and, in the case of interest here, tangent vectors to the Teichmüller space…
In this note we prove convexity, in the sense of Colding-Naber, of the regular set of solutions to some complex Monge-Ampere equations with conical singularities along simple normal crossing divisors. In particular, any two points in the regular set can be joined by a smooth minimal geodesic lying entirely in the regul…
Regularization-induced exploration improves contextual bandit performance.
problem Complex reward models in real-world contextual bandits are hard to explore effectively.
method Regularization-induced exploration using stochasticity in cross-validation.
result Regularization-induced exploration leads to reliable exploration in large-scale business environments.
Many information retrieval algorithms rely on the notion of a good distance that allows to efficiently compare objects of different nature. Recently, a new promising metric called Word Mover's Distance was proposed to measure the divergence between text passages. In this paper, we demonstrate that this metric can be ex…
Real analytic functions can be extended on manifolds with normal crossings.
problem Extending continuous functions to Cω functions on manifolds with normal crossings. method Employing Cartan Theorems A and B from real analytic geometry.
result Continuous functions on the union of submanifolds with normal crossings can be extended to Cω functions on the entire manifold. The ratio of volume to crossing number of a hyperbolic knot is known to be bounded above by the volume of a regular ideal octahedron, and a similar bound is conjectured for the knot determinant per crossing. We investigate a natural question motivated by these bounds: For which knots are these ratios nearly maximal? We…
We consider the parametric learning problem, where the objective of the learner is determined by a parametric loss function. Employing empirical risk minimization with possibly regularization, the inferred parameter vector will be biased toward the training samples. Such bias is measured by the cross validation procedu…
Researchers create spectral triples for twisted crossed products using Kasparov's external product.
problem Constructing spectral triples for twisted crossed products.
method Using Kasparov's external product, the construction of spectral triples for twisted crossed products is achieved.
result The construction of spectral triples for twisted crossed products is possible under suitable assumptions.
Sparse coding has shown its power as an effective data representation method. However, up to now, all the sparse coding approaches are limited within the single domain learning problem. In this paper, we extend the sparse coding to cross domain learning problem, which tries to learn from a source domain to a target dom…
Cross-validation is one of the most popular model selection methods in statistics and machine learning. Despite its wide applicability, traditional cross validation methods tend to select overfitting models, due to the ignorance of the uncertainty in the testing sample. We develop a new, statistically principled infere…
This paper addresses image classification through learning a compact and discriminative dictionary efficiently. Given a structured dictionary with each atom (columns in the dictionary matrix) related to some label, we propose cross-label suppression constraint to enlarge the difference among representations for differe…
Regularization and data augmentation can be class-dependent, leading to poor performance on some classes.
problem Class-dependent effects of regularization and data augmentation.
method Evaluation of regularization and data augmentation techniques on Imagenet and INaturalist datasets.
result Regularization and data augmentation can lead to significant performance drops on some classes.
ALO-CV approximates leave-one-out error in proportional regime.
problem Estimating generalization error in high-dimensional settings.
method Developed new analysis for ALO-CV, showed consistency under strong convexity.
result ALO-CV approximates leave-one-out error up to negligible error.
Study spectral flow on a warped cylinder with special boundary conditions.
problem Analyzing spectral flow on a warped cylinder with specific boundary conditions.
method Complexifying the twisting bundle, diagonalizing the orthogonal twist, and regrouping conjugate and reflection-paired blocks.
result Explicit formula for RO(O(2))-valued spectral flow, refining ordinary spectral flow. This work shows that supervised contrastive learning achieves similar results to cross-entropy but requires more iterations.
problem The question of whether there are fundamental differences in representation geometry between supervised contrastive learning and cross-entropy.
method The authors prove that both losses attain their minimum when representations of each class collapse to the vertices of a regular simplex, and they empirically validate this finding.
result Supervised contrastive learning requires more iterations to reach a close-to-optimal state compared to cross-entropy, indicating different optimization behavior.