Paper shows how gradient concentration helps in learning from inexact data.
problem Learning from inexact and stochastic training data.
method Combines probabilistic gradient concentration with inexact optimization techniques.
result Derives sharp test error guarantees for learning.
In several real-world applications involving decision making under uncertainty, the traditional expected value objective may not be suitable, as it may be necessary to control losses in the case of a rare but extreme event. Conditional Value-at-Risk (CVaR) is a popular risk measure for modeling the aforementioned objec…
New methods improve accuracy in detecting concentric objects.
problem Detecting concentric geometric objects in noisy data.
method Developed new estimators and compared performance of existing methods.
result New methods outperform existing non-iterative methods and are robust to noise.
Paper analyzes sample complexity for offline f-divergence-regularized contextual bandits.
problem Lack of tight analyses for sample complexity in offline reinforcement learning.
method Novel pessimism-based analysis for reverse KL divergence, establishing ildeO(ε−1) sample complexity. result Achieves ildeO(ε−1) sample complexity for reverse KL divergence, surpassing existing bounds. New inequalities for unbounded functions improve denoising score matching.
problem Statistical error bounds for denoising score matching with unbounded objective functions.
method Derive new concentration inequalities using McDiarmid's inequality and Rademacher complexity bounds.
result Improved statistical error bounds for denoising score matching.
New method learns shared structures in non-linear tasks.
problem Learning shared linear representations in non-linear tasks.
method Convex optimization with structural assumptions.
result Rank and clustered estimators recover shared structures under certain conditions.
The paper offers a framework to analyze machine learning problems using concentration of measure.
problem Analyzing machine learning algorithms defined by implicit equations.
method Develops a concentration of measure framework to solve convex problems and implicit formulations.
result Provides precise estimations for the first moments of the solution, describing the behavior and performance of machine learning classifiers.
Method captures shared information across many views robustly.
problem Modeling hundreds of views per event and learning robust embeddings without view knowledge.
method View bootstrapping using multi-view correlation and matrix concentration theory.
result View bootstrapping captures shared information across many views robustly.
Paper develops privacy-preserving federated learning for nonsmooth objectives.
problem Solving nonsmooth objective functions in a privacy-preserving manner.
method Zero-concentrated differential privacy (zCDP) with Gaussian noise, distributed ADMM, and approximation of augmented Lagrangian.
result The algorithm achieves a competitive privacy-accuracy trade-off and converges to the exact solution.
Expanding on techniques of concentration of measure, we develop a quantitative framework for modeling liquidity risk using convex risk measures. The fundamental objects of study are curves of the form (ρ(λX))λ≥0, where ρ is a convex risk measure and X a random variable, and we call such a curve a \emph{liqu…
Stochastic approximation algorithms show exponential progress bounds.
problem Analyzing the convergence of stochastic approximation algorithms.
method Developed geometric ergodicity proofs to establish exponential concentration bounds.
result Proved faster convergence rates for specific algorithms.
Proposes a new latent variable model for hyperspherical latent spaces.
problem Efficiently modeling heavy-tailed distributions in hyperspherical latent spaces.
method Introduces spherical Cauchy (spCauchy) latent variables and applies Möbius transformations.
result Shows spCauchy recovers vMF geometry in high-concentration limits and avoids complex evaluations.
Optimal transport theory has recently found many applications in machine learning thanks to its capacity for comparing various machine learning objects considered as distributions. The Kantorovitch formulation, leading to the Wasserstein distance, focuses on the features of the elements of the objects but treat them in…
Firms with different ownership structures could be argued to have different levels of efficiency.Highly concentrated firms are expected to be more efficient as this type of ownership structure may alleviate the conflict of interest between managers and shareholders.In Malaysia, public-listed firms have been found to ha…
New algorithm finds global maxima in multi-modal functions.
problem Finding global maxima in multi-modal functions with saddle points.
method G-PFSO algorithm using particle filter and averaging.
result G-PFSO efficiently finds global maximizer at optimal rate.
Paper introduces risk assessment for contextual bandits without experiments.
problem Evaluate policies using logged data in context bandits.
method Lipschitz risk functionals and Off-Policy Risk Assessment (OPRA) framework.
result OPRA provides finite sample guarantees for various risk estimates.
Optimizes non-linear outcomes from summed contributions.
problem Maximizing a non-linear function of summed small contributions.
method Derives a scalable descent algorithm leveraging concentration properties.
result Directly optimizes for stated objective, e.g., A/B test success criterion.
Develops privacy-preserving multivariate median estimation methods.
problem Lack of rigorous privacy guarantees for robust multivariate location estimation.
method Novel finite-sample performance guarantees for differentially private multivariate depth-based medians.
result Sharp performance guarantees for multivariate depth-based medians under differential privacy.
SGD converges to an invariant distribution with sub-Gaussian or sub-exponential properties.
problem Optimizing smooth and strongly convex objectives using SGD.
method Analysis through Markov chains, focusing on convergence and concentration properties.
result SGD iterates and their invariant limit distribution inherit sub-Gaussian or sub-exponential concentration properties.
Bayesian neural networks explore rare fluctuations for better feature learning.
problem Understanding rare but dominant fluctuations in Bayesian neural networks.
method Large-deviation theory and joint optimization over predictors and internal kernels.
result Posterior rate function optimization reveals data-dependent kernel selection.
In this paper we study the concentration properties for the eigenvalues of kernel matrices, which are central objects in a wide range of kernel methods and, more recently, in network analysis. We present a set of concentration inequalities tailored for each individual eigenvalue of the kernel matrix with respect to its…
New bounds for convex clustering under graph connectivity.
problem Understanding clustering performance under different graph connectivity structures.
method Random walks and concentration inequalities for random graph models.
result Improved rates of convergence for centroid recovery.
The paper geometrically characterizes graded manifolds and proves the Frobenius theorem.
problem Understanding and characterizing graded manifolds.
method Geometric characterization and Frobenius theorem proof.
result Frobenius theorem proven for graded distributions.
The paper clusters PK curves using ML, finding it useful for identifying similar patterns.
problem Improving drug development and patient outcomes through ML in pharmacogenomics.
method Unsupervised clustering of PK curves using various dissimilarity measures.
result Euclidean distance is most suitable for clustering PK curves, and clustering can validate pharmacogenomic results.
Many machine learning algorithms are based on the assumption that training examples are drawn independently. However, this assumption does not hold anymore when learning from a networked sample because two or more training examples may share some common objects, and hence share the features of these shared objects. We …
Theory for algebraic data on categories via concentration structures.
problem Defining algebraic structures on categories.
method Introducing concentration structures and concentration monoids.
result Every group can be represented as a concentration monoid of a trivial category.
Paper addresses concentration of distances for fractional quasi p-norms, identifying conditions for concentration and anti-concentration.
problem Understanding concentration of distances for fractional quasi p-norms in high dimensions.
method Analyzes conditions for concentration and anti-concentration of distances for fractional quasi p-norms.
result Identifies conditions for concentration and anti-concentration of fractional quasi p-norms, ruling out some approaches and specifying conditions for control.
Study Finsler metric measure manifolds' concentration properties.
problem Understanding concentration properties in Finsler metric measure manifolds.
method Established relationships with observable diameter, isoperimetric inequalities, and first eigenvalue.
result Derived a Cheng type upper bound estimate for the first closed eigenvalue.
We analyze critical points of the Sliced Wasserstein Distance for optimization stability.
problem Understanding the behavior of optimization algorithms for models trained with the Sliced Wasserstein Distance.
method Explicit perturbations and critical point analysis of the SW objective.
result Stable critical points of SW cannot concentrate on segments, providing optimization stability.
New method improves missing mass concentration bounds.
problem Missing mass concentration problem
method New method of estimating concentration of heterogenic sums
result Slightly improved state-of-the-art bounds
MAGT generates data efficiently by aligning to manifold structure.
problem Efficiently generating data near a low-dimensional structure embedded in high-dimensional space.
method MAGT is a flow-like generator that learns a one-shot, manifold-aligned transport from a low-dimensional base distribution to the data space, using a fixed Gaussian smoothing level and self-normalized importance sampling.
result MAGT samples in a single forward pass, concentrates probability near the learned support, and induces an intrinsic density with respect to the manifold volume measure, enabling principled likelihood evaluation for generated samples.
Reducing barriers to entry in large-scale ML markets, study shows multi-objective learning can lower data requirements.
problem Barriers to entry in emerging markets for large-scale machine learning models.
method Defined a multi-objective high-dimensional regression framework to study reputational damage and data requirements.
result The number of data points needed for a new company to enter the market can be significantly smaller than the incumbent company's dataset size.
Improved fast rates for decision making with forward-KL regularization in contextual bandits.
problem Improving fast rates for decision making with forward-KL regularization in contextual bandits.
method Streamlined analysis of forward-KL-regularized offline CBs, exploiting the pessimism principle and convex-analytical pipeline.
result First ildeO(ε−1) upper bounds in tabular and general function approximation settings. Paper develops sparse learning for heavy-tailed time series with locally stationary dynamics.
problem Sparse learning for high-dimensional heavy-tailed locally stationary time series.
method Additive modeling with kernel smoothing, sparsity-inducing penalized estimation.
result Prediction-error bounds and convergence rates for different sparsity structures.
Surfaces in 3-manifolds concentrate at curvature critical points.
problem Understanding concentration of surfaces in 3-manifolds.
method Proving surfaces concentrate at critical points of scalar curvature.
result Simply connected H-surfaces concentrate at curvature critical points.
Sharp concentration bounds for i.i.d. variables.
problem Controlling the tail probabilities of independent variables.
method Extension of Sanov's theorem using large deviations and information theory.
result Matching concentration and anti-concentration bounds for i.i.d. samples of any size.
We survey recent results related to the concentration of eigenfunctions. We also prove some new results concerning ball-concentration, as well as showing that eigenfunctions saturating lower bounds for L1-norms must also, in a measure theoretical sense, have extreme concentration near a geodesic.
A new model CDTM improves text classification by concentrating document topics.
problem Unsupervised text classification with diverse topic distributions.
method Imposes an exponential entropy penalty on document topic distribution to encourage concentration.
result More coherent topics and concentrated, sparse document-topic distributions.
Simplified proof of Gaussian concentration inequality using covariance.
problem Gaussian concentration inequality proof
method Covariance representation based on characteristic functions
result Elementary proof of Gaussian concentration inequality
Study provides bounds for estimating intrinsic dimension using Gaussian kernels.
problem Estimating intrinsic dimension from data.
method Finite-sample concentration and anti-concentration bounds for Gaussian kernel sums.
result Explicit dependence on sample size, bandwidth, and geometric parameters.
Developed concentrated liquidity in n-dimensional AMM with polar coordinates in Rust.
problem Risk of stacking too many stablecoin pools.
method Building concentrated liquidity positions with ticks in polar coordinates in Rust.
result Hedging risk of stacking stablecoin pools.
In this paper, we consider a concentration of measure problem on Riemannian manifolds with boundary. We study concentration phenomena of non-negative 1-Lipschitz functions with Dirichlet boundary condition around zero, which is called boundary concentration phenomena. We first examine relation between boundary concen…
Study on volume of tubes and concentration in Riemannian geometry.
problem Understanding concentration loci in Riemannian manifolds and their relation to tube volumes.
method Provided a general formula for tube volumes, specialized to totally geodesic submanifolds, and investigated concentration loci.
result Explicitly proved concentration for codimension one cases and explored characterizations in Wasserstein and Box distances.
Polluting fine dusts in South Korea which are mainly consisted of biomass burning and fugitive dust blown from dust belt is significant problem these days. Predicting concentrations of fine dust particles in Seoul is challenging because they are product of complicate chemical reactions among gaseous pollutants and also…
This study assesses risk concentration in MDB portfolios using Monte Carlo simulations.
problem Risk concentration in MDB portfolios of a few borrowers.
method Realistic MDB portfolio simulations and Monte Carlo analysis.
result Current risk adjustments may be overly conservative.
This paper continues study, both theoretical and empirical, of the method of Venn prediction, concentrating on binary prediction problems. Venn predictors produce probability-type predictions for the labels of test objects which are guaranteed to be well calibrated under the standard assumption that the observations ar…
A new algorithm, Regular Tree Search, tackles non-convex simulation optimization problems.
problem Non-convex objective functions in simulation optimization.
method Integrates adaptive sampling with recursive partitioning of the search space.
result Proves global convergence and reliably identifies the global optimum.
While the objective in traditional multi-armed bandit problems is to find the arm with the highest mean, in many settings, finding an arm that best captures information about other arms is of interest. This objective, however, requires learning the underlying correlation structure and not just the means of the arms. Se…