In this paper, we develop a relative error bound for nuclear norm regularized matrix completion, with the focus on the completion of full-rank matrices. Under the assumption that the top eigenspaces of the target matrix are incoherent, we derive a relative upper bound for recovering the best low-rank approximation of t…
This work tightens generalization error bounds using Wasserstein distance.
problem Improving expected generalization error bounds in machine learning.
method Introduces bounds based on Wasserstein distance for various settings.
result New, tighter bounds based on relative entropy and other information measures.
The CUR matrix decomposition is an important extension of Nyström approximation to a general matrix. It approximates any data matrix in terms of a small number of its columns and rows. In this paper we propose a novel randomized CUR algorithm with an expected relative-error bound. The proposed algorithm has the advanta…
This work bounds the generalization error of private algorithms for discrete data.
problem Bounding the generalization error of private algorithms for discrete data.
method Information-theoretic approach using relative entropy and the method of types.
result Explicit upper bounds on the generalization error of stable private algorithms for discrete data.
The paper proposes an efficient NMF algorithm using geometric assumptions and rank-one NMFs.
problem Nonnegative matrix factorization (NMF) for clustering and factorization.
method Geometric assumption on data matrices, rank-one NMF initialization, and clustering.
result The proposed algorithm provides faster speeds and comparable relative errors to classical NMF algorithms.
Kernel quadrature uses DPPs for sampling with tight error bounds.
problem Efficiently sampling nodes for quadrature rules in RKHS.
method Nodes sampled from a truncated and saturated DPP kernel.
result Tighter quadrature error bounds using DPPs.
New method preserves unitarity for Schrödinger equation learning, reducing errors and improving time generalization.
problem Learning the evolution operator for time-dependent Schrödinger equation with varying Hamiltonians.
method Linear estimator preserving weak unitarity, with theoretical error bounds and time generalization.
result Achieves up to two orders of magnitude smaller relative errors than existing methods.
A new framework evaluates HTE estimators using relative error.
problem Lack of robust evaluation methods for HTE estimators.
method Proposes a relative error-based evaluation framework and neural network architecture to estimate nuisance parameters and robustly compare HTE estimators.
result Demonstrates reliable comparisons and improved HTE estimation through the proposed framework and learning algorithm.
This monograph deals with adaptive supervised classification, using tools borrowed from statistical mechanics and information theory, stemming from the PACBayesian approach pioneered by David McAllester and applied to a conception of statistical learning theory forged by Vladimir Vapnik. Using convex analysis on the se…
Develops a method to estimate rare-event probabilities under distributional uncertainty.
problem Distributional uncertainty limits the effectiveness of rare-event simulation techniques.
method Wasserstein distributionally robust rare-event simulation (DRIS) framework.
result DRIS achieves vanishing relative error in estimating rare-event probabilities.
Improved kernel k-means clustering for large datasets with reduced computational cost.
problem High computational cost of kernel k-means clustering for large datasets.
method Applying linear k-means clustering to a subset of features constructed using rank-restricted Nyström approximation.
result Achieves a 1+ε approximation ratio for kernel k-means cost function.
The study optimizes distribution estimation from samples with relative entropy error, adapting to sparse distributions.
problem Estimating discrete distributions with high-probability accuracy in relative entropy.
method Analysis of Laplace estimator and confidence-dependent smoothing techniques, including data-dependent smoothing.
result Optimal high-probability risk bounds for various estimators, including a new data-dependent smoothing method.
Ensembling improves performance when classifiers disagree more than average.
problem When do ensembles provide significant performance improvements in classification tasks?
method Theoretical and empirical analysis of ensemble improvement rate and disagreement-error ratio.
result Ensembling improves performance significantly when the disagreement rate is large relative to the average error rate.
The study bounds generalization error using mutual information and proposes methods to control it.
problem Understanding and controlling the generalization capability of learning algorithms.
method Derives upper bounds on generalization error using mutual information, proposes regularization methods.
result Provides theoretical guidelines for balancing data fit and generalization.
The study analyzes numerical stability in large language models using mixed-precision arithmetic.
problem Numerical stability of large language models using low-precision arithmetic.
method Developed a mixed-precision analysis of transformer inference, deriving bounds for condition numbers and forward error.
result Established that numerical stability is determined by the interplay between weight magnitude and the growth of the residual stream.
This paper uses normalizing flows to approximate transport maps between densities.
problem Approximating transport maps between given densities.
method Construct time-dependent controls using normalizing flows.
result Provides bounds on the number of switches for piecewise constant approximations.
This paper tightens information-theoretic bounds on generalization errors.
problem Understanding the discrepancy between training and testing data losses.
method Investigates the tightness of information-theoretic bounds on generalization error.
result The individual sample mutual information bound can be asymptotically tight under specific assumptions.
Study error bounds and optimal schedules for Masked Diffusions with factorized approximations.
problem Analyzing trade-offs between computation and accuracy in Masked Diffusion Models.
method Provided general error bounds and identified optimal schedules based on data distribution information profiles.
result Identified optimal schedule sizes for Masked Diffusion Models.
Optimal transport bounds improve generalization in learning algorithms.
problem Understanding and improving generalization in machine learning.
method Using algorithmic transport cost and Wasserstein distance to derive upper bounds on generalization error.
result Generalization error decreases exponentially with the number of layers in deep neural networks.
The belief propagation (BP) algorithm is widely applied to perform approximate inference on arbitrary graphical models, in part due to its excellent empirical properties and performance. However, little is known theoretically about when this algorithm will perform well. Using recent analysis of convergence and stabilit…
This paper improves entropy bounds for ranking time-series complexity.
problem Ranking the complexity of time series processes.
method Building on information theoretic bounds, the paper improves the upper bound of conditional differential entropy using Hadamard's inequality and covariance matrix properties.
result The improved bounds can be used to rank the complexity of time series processes.
We propose a voted dual averaging method for online classification problems with explicit regularization. This method employs the update rule of the regularized dual averaging (RDA) method, but only on the subsequence of training examples where a classification error is made. We derive a bound on the number of mistakes…
The paper explores learning metrics in low dimensions with bounds and complexities.
problem Learning metrics in low dimensions with bounds and complexities.
method Develops upper and lower bounds on generalization error, quantifies sample complexity, and bounds accuracy relative to the true metric.
result Novel mathematical approaches to metric learning and insights into ordinal embedding.
This handbook translates lead time analysis into R code.
problem Tracking divergence in booking lead times over time.
method Translated original article's methodology into R code.
result Demonstrated reproducibility and error bounds in lead time forecasts.
Estimates inner products between nonparametric distributions using Fourier basis.
problem Estimating inner products between two nonparametric distributions.
method Proposes estimators for inner products and induced norms, proves mean squared error bounds and minimax lower bounds.
result Proposed estimators are rate-optimal over Fourier ellipsoids.
Kernel density estimation (KDE) is a popular statistical technique for estimating the underlying density distribution with minimal assumptions. Although they can be shown to achieve asymptotic estimation optimality for any input distribution, cross-validating for an optimal parameter requires significant computation do…
Paper proposes a fast stochastic algorithm for neural network quantization with error bounds.
problem Error analysis for quantized neural networks with non-convex loss functions and nonlinear activations.
method Greedy path-following mechanism combined with stochastic quantizer.
result Established full-network error bounds for quantized neural networks.
New GaussianSketch approximates kernel distances with almost relative error and small additive term.
problem Approximating kernel distances between point sets efficiently.
method Truncating Gaussian kernel expansions and using RecursiveTensorSketch.
result Approximates kernel distance with almost (1+ε)-relative error and small additive α term. We analyze how errors in interbank liabilities affect the clearing vector in financial systems.
problem Estimation errors in interbank liabilities can lead to inaccuracies in the clearing vector, impacting risk assessments.
method We quantify the sensitivity of the clearing vector to estimation errors in the interbank liabilities matrix using a basis for permissible perturbations.
result We derive analytical solutions for the maximal deviations of the clearing vector and compute upper bounds for worst-case perturbations.
Efficiently estimates graphlet statistics in large networks.
problem Limited ability to compute graphlets in massive networks.
method Unbiased estimation framework for graphlets, parallel and scalable.
result Accurate and fast estimation of graphlet statistics in billions of edges networks.
AGCA approximates angular variation on the unit sphere, reducing extremal dependence problems to eigenanalysis.
problem Approximating angular variation in multivariate extremes.
method Anchored geodesic component analysis (AGCA) approximates angular variation by great subspheres constrained to pass through a chosen reference direction.
result AGCA finds concentrated tail directions in daily equity-portfolio losses, explaining about 91% of anchored variation.
Study on H-consistency bounds for machine learning surrogates.
problem Estimating target loss error relative to surrogate loss error in machine learning.
method Developed H-consistency bounds for various surrogates and loss functions. result Stronger guarantees than existing methods, offering distribution-dependent and -independent bounds.
Meta-learning improves relative density-ratio estimation from limited data.
problem Estimating relative density-ratios from few instances.
method Meta-learning using neural networks to extract and embed dataset information for relative DRE.
result Meta-learning enables efficient and effective adaptation to few instances for relative DRE.
The study improves representation learning bounds using data-dependent Gaussian mixtures.
problem Improving generalization in representation learning.
method Established bounds using relative entropy and MDL of latent variables.
result The approach significantly improves generalization over existing methods.
The paper tackles learning from non-irreducible Markov chains, proving learnability and generalization bounds.
problem Learning from temporal dependent data with non-irreducible Markov chains.
method Uniform convergence and generalization bounds for sample error under uniform ergodicity.
result Learnability and generalization bounds for approximate sample error minimization algorithm.
Paper develops estimators for unbounded density ratios with applications in error control.
problem Estimating density ratios with unbounded domains and ranges.
method Least squares and logistic regression loss functions for density ratio estimation.
result Established upper bounds on estimation errors with optimal rates for unbounded density ratios.
Unified framework for understanding GRPO as U-statistic.
problem Theoretical properties of GRPO remain less studied.
method Unified framework through classical U-statistics.
result GRPO is asymptotically equivalent to an oracle policy gradient algorithm.
Sharp ℓ∞-bounds for Q-learning derived using cone-contractive operators.
problem Deriving optimal sample complexity for Q-learning algorithms. method Stochastic approximation with cone-contractive operators, deriving non-asymptotic ℓ∞-norm bounds. result Derives the sharpest known ℓ∞-norm bounds for Q-learning. Active learning method improves local model validity estimation.
problem Ensuring local model validity in machine learning applications.
method Learning model error to estimate local validity using active learning.
result The proposed method can estimate local validity with a small amount of data.
The density matrices are positively semi-definite Hermitian matrices of unit trace that describe the state of a quantum system. The goal of the paper is to develop minimax lower bounds on error rates of estimation of low rank density matrices in trace regression models used in quantum state tomography (in particular, i…
Meta-learning bounds derived using PAC-Bayes theory for improved generalization.
problem Uncertainty in generalization performance for meta-learning with new tasks.
method PAC-Bayes relative entropy bounds and empirical risk minimization (ERM) method.
result Competitive generalization performance and rapid convergence with data-dependent prior.
Entropy-SGD optimizes a PAC-Bayes bound, leading to improved generalization.
problem Improving generalization in machine learning models.
method Entropy-SGD optimizes a PAC-Bayes bound by adjusting the prior, which is typically chosen independently of the data.
result Entropy-SGD can yield relatively tight generalization bounds and still fit real labels.
Study on multicalibration for multiple properties, establishing sample complexity bounds.
problem Ensuring unbiasedness across multiple related properties in predictions.
method Establishing upper and lower bounds on sample complexity for multicalibration of multiple properties.
result Matching upper and lower bounds on sample complexity for multicalibration of k properties. New approach for reward-free exploration reduces estimation error.
problem Reward-free exploration in reinforcement learning.
method Adaptive approach reducing MDP estimation error.
result Reward-free UCRL algorithm improves sample complexity.
We consider the question of efficient estimation in the tails of Gaussian copulas. Our special focus is estimating expectations over multi-dimensional constrained sets that have a small implied measure under the Gaussian copula. We propose three estimators, all of which rely on a simple idea: identify certain \emph{dom…
Bayesian framework improves robustness in nonlinear regression models.
problem Measurement error, model misspecification, and distributional misspecification in regression analyses.
method Joint Dirichlet process prior on latent covariate-response distribution, updating with posterior pseudo-samples.
result Improved stability and consistency in estimators under increasing measurement error.
Paper improves distributed mean estimation and variance reduction without relying on input norm.
problem Distributed mean estimation and variance reduction with large input norms.
method Quantization and lattice theory connection for improved error bounds.
result Output error bounds depend only on input distance, not norm.
Researchers create integral representations for two-layer ReLU networks with quantitative bounds.
problem Approximating functions with two-layer ReLU networks using explicit integral representations.
method Developed integral representations involving harmonic extension and projection, providing L2 bounds. result Functions can be approximated with L2 errors independent of dimension or degree, depending on coefficients and distribution.