New concentration inequalities for tensors with heavy-tailed coefficients.
problem Developing bounds for Euclidean functions of tensors with sub-Weibull distributions.
method Extending concentration inequalities to sub-Weibull random tensors, using new inequalities for heavy-tailed random variables and martingale analysis.
result Established a phase transition between sub-gaussian and heavy-tailed regimes for Euclidean functions of tensors.
We prove that an approximated version of the Brunn--Minkowski inequality with volume distortion coefficient implies a Gaussian concentration-of-measure phenomenon. Our main theorem is applicable to discrete spaces.
New insights on offline RL with state aggregation and trajectory data.
problem Understanding sample complexity in offline policy evaluation.
method Analyzing concentrability coefficient in aggregated Markov Transition Model.
result Sample complexity depends on concentrability coefficient in aggregated model.
Paper tackles offline preference-based RL with human feedback.
problem Offline Preference-based Reinforcement Learning with preference feedback.
method Two-step approach: MLE for reward estimation and distributionally robust planning.
result First guarantee for learning any target policy with polynomial samples.
In the celebrated book entitled Metric Structures for Riemannian and Non-Riemannian Spaces, so-called Green Book, Gromov presented a problem regarding a metric measure space. Gromov posed the question Bound the expansion coefficient from below in terms of the observable diameter. The overall aim of the current study is…
Paper tackles robust offline RL for non-Markovian processes, improving efficiency and applicability.
problem Learning robust policies for non-Markovian decision processes with limited offline data.
method Proposes a novel algorithm with dataset distillation and LCB design for robust values, derived new dual forms, and introduces concentrability coefficients.
result Proves polynomial sample efficiency for finding ε-optimal robust policies.
The paper analyzes sparse high-dimensional linear regression with random design and unknown error variance, providing adaptiveness and concentration rates.
problem Sparse high-dimensional linear regression with random design and unknown error variance.
method Analysis of posterior concentration rates, employing techniques to address model misspecification.
result Adaptiveness and concentration rates of the posterior for sparse high-dimensional linear regression.
New models reduce regional inequality by adjusting exchange range and asset distribution bias.
problem Reduction of regional inequality in economic systems.
method Proposed new asset exchange models with spatial exchange range and local support bias to adjust asset distribution and circulation rates.
result Achieved asset distribution from over-concentration to exponential and eventually normal, reducing Gini coefficient.
Uncertainty principles such as Heisenberg's provide limits on the time-frequency concentration of a signal, and constitute an important theoretical tool for designing and evaluating linear signal transforms. Generalizations of such principles to the graph setting can inform dictionary design for graph signals, lead to …
Study improves robustness and sparsity in linear regression with adversarial outliers and heavy-tailed noise.
problem Outliers and heavy-tailed noise in linear regression coefficients.
method Sharp concentration inequalities and generic chaining.
result Sharper error bounds under weaker assumptions.
The Fokker-Planck equation with diffusion coefficient quadratic in space variable, linear drift coefficient, and nonlocal nonlinearity term is considered in the framework of a model of analysis of asset returns at financial markets. For special cases of such a Fokker-Planck equation we describe a construction of exact …
A concentration graph associated with a random vector is an undirected graph where each vertex corresponds to one random variable in the vector. The absence of an edge between any pair of vertices (or variables) is equivalent to full conditional independence between these two variables given all the other variables. In…
Improved concentration inequalities for sub-Weibull variables enhance statistical and machine learning applications.
problem Improving concentration inequalities for sub-Weibull random variables.
method Developed new concentration inequalities for sums of independent sub-Weibull random variables, including a new sub-Weibull parameter.
result New concentration inequalities with sharper constants and a mixture of sub-Gaussian and sub-Weibull tails.
New Gini indices capture more nuanced income inequality.
problem Measuring joint dispersion across multiple observations.
method Axiomatic approach to define and characterize n-th order Gini deviations.
result Higher-order Gini coefficients reveal more extreme income disparities.
High-dimensional, large-sample astrophysical databases of galaxy clusters, such as the Chandra Deep Field South COMBO-17 database, provide measurements on many variables for thousands of galaxies and a range of redshifts. Current understanding of galaxy formation and evolution rests sensitively on relationships between…
This paper proposes a new methodology to compute Value at Risk (VaR) for quantifying losses in credit portfolios. We approximate the cumulative distribution of the loss function by a finite combination of Haar wavelets basis functions and calculate the coefficients of the approximation by inverting its Laplace transfor…
Study on Fox's trapezoidal conjecture for specific alternating links.
problem Investigating Fox's trapezoidal conjecture for alternating links.
method Diagrammatic Murasugi sums, Alexander polynomial, and concordance analysis.
result Established inequalities and conditions for the Alexander polynomial and trapezoidal conjecture.
Using geometric quantization, we represent curve operators in the TQFT of Witten-Reshetikhin-Turaev with jauge group SU_2 as Toeplitz operators with symbols corresponding to trace functions. As an application, we show that eigenvectors of these operators are concentrated near the level sets of these trace functions, an…
Paper tackles offline CMDP problems with near-optimal algorithm and sample complexity bound.
problem Offline CMDP problems with only offline data available.
method DPDL algorithm using single-policy concentrability coefficient C∗ and deviation control mechanism. result DPDL algorithm matches sample complexity lower bound with ildeO((1−γ)−1) factor. Recently the interest of researchers has shifted from the analysis of synchronous relationships of financial instruments to the analysis of more meaningful asynchronous relationships. Both of those analyses are concentrated only on Pearson's correlation coefficient and thus intraday lead-lag relationships associated wi…
SCOPE estimator improves covariance and precision matrix estimation.
problem Estimating covariance and precision matrices accurately.
method Distributionally robust optimization with convex spectral divergence.
result SCOPE estimator reduces spectral bias and improves condition number.
The paper develops bounds for predictive values in binary classification.
problem Lack of confidence intervals for positive and negative predictive values.
method Bi-criterion framework and distribution-free large deviation and uniform convergence bounds.
result New bounds for predictive values without relying on concentration inequalities.
Unified hybrid RL algorithm improves online RL performance with offline data.
problem Improving reinforcement learning performance with limited online data.
method A unified hybrid RL algorithm combining offline and online data.
result Unified algorithm achieves state-of-the-art results in sub-optimality gap and online learning regret.
In recent work, Boltzmann and Fokker-Planck equations were derived for the "Yard-Sale Model" of asset exchange. For the version of the model without redistribution, it was conjectured, based on numerical evidence, that the time-asymptotic state of the model was oligarchy -- complete concentration of wealth by a single …
X in R^D has mean zero and finite second moments. We show that there is a precise sense in which almost all linear projections of X into R^d (for d < D) look like a scale-mixture of spherical Gaussians -- specifically, a mixture of distributions N(0, sigma^2 I_d) where the weight of the particular sigma component is P …
Proposes HDBEN for heteroscedastic regression with improved sparsity and variance modeling.
problem Violation of constant error variance in high-dimensional regression.
method HDBEN framework using hierarchical Bayesian priors with ℓ1 and ℓ2 penalties. result Achieves posterior concentration, variable selection consistency, and asymptotic normality.
We consider linear models where d potential causes X1,...,Xd are correlated with one target quantity Y and propose a method to infer whether the association is causal or whether it is an artifact caused by overfitting or hidden common causes. We employ the idea that in the former case the vector of regression c…
Paper develops sparse learning for heavy-tailed time series with locally stationary dynamics.
problem Sparse learning for high-dimensional heavy-tailed locally stationary time series.
method Additive modeling with kernel smoothing, sparsity-inducing penalized estimation.
result Prediction-error bounds and convergence rates for different sparsity structures.
Positive definite kernels and their associated Reproducing Kernel Hilbert Spaces provide a mathematically compelling and practically competitive framework for learning from data. In this paper we take the approximation theory point of view to explore various aspects of smooth kernels related to their inferential proper…
Study on hemisphere threshold for Escobar functional on Riemannian manifolds, revealing mass and boundary invariant behaviors.
problem Analyzing the hemisphere threshold for the Escobar functional on compact Riemannian manifolds.
method Near-threshold landscape organization by boundary invariants, exact evaluation of weighted profile moments, Lyapunov-Schmidt correction, and blow-up analysis.
result At threshold, blow-ups concentrate at umbilic points with vanishing mass and gradient, leading to compactness and hemispherical rigidity.
We consider strictly stationary heavy tailed time series whose finite-dimensional exponent measures are concentrated on axes, and hence their extremal properties cannot be tackled using classical multivariate regular variation that is suitable for time series with extremal dependence. We recover relevant information ab…
This paper analyzes SHAP values using Fourier expansions for model interpretability.
problem Understanding and interpreting SHAP values in complex models.
method Developed a spectral framework using Fourier expansions for SHAP values in various model regimes.
result SHAP values are Lipschitz continuous in the deterministic regime and converge to Gaussian process values in the probabilistic regime.
The paper improves count data regression models for overdispersed data.
problem Improving regression models for overdispersed count data.
method Double ℓ1-regularized negative binomial regressions. result Oracle inequalities and consistency for Lasso estimators of partial regression coefficients.
Automatically differentiable estimation for BLP model reduces bias in demand estimation.
problem Estimating the BLP model with reduced bias and improved performance.
method Phrasing BLP as an automatically differentiable moment function, using CUE for estimation, and incorporating MCMC credible intervals.
result CUE estimation shows lower bias but higher MAE compared to 2S-GMM, with MCMC providing closest empirical coverage.
Subspace recovery from corrupted and missing data is crucial for various applications in signal processing and information theory. To complete missing values and detect column corruptions, existing robust Matrix Completion (MC) methods mostly concentrate on recovering a low-rank matrix from few corrupted coefficients w…
Study analyzes financial distributions and inequality in professional cycling teams.
problem Financial inequality and concentration among cycling teams.
method Rank-size law and various inequality indices applied to Tour de France data.
result Financial gains distribution is hyperbolic with a decay exponent of about -1, contrary to Pareto principle.
A new inference method using regression and batched discrepancies.
problem Simulating parameters from simulator outputs.
method Regression-based projection and batched discrepancy weighting.
result Method produces a self-normalized pseudo-posterior.
Theory for algebraic data on categories via concentration structures.
problem Defining algebraic structures on categories.
method Introducing concentration structures and concentration monoids.
result Every group can be represented as a concentration monoid of a trivial category.
The paper introduces a new method for tail bounds of random vectors and matrices.
problem Estimating norms of random vectors and matrices under moment assumptions.
method Variational tail bounds for norms of random vectors and matrices.
result Dimension-free concentration inequalities for various norms of random vectors and matrices.
We investigate second order quasilinear equations of the form f_{ij} u_{x_ix_j}=0 where u is a function of n independent variables x_1, ..., x_n, and the coefficients f_{ij} are functions of the first order derivatives p^1=u_{x_1}, >..., p^n=u_{x_n} only. We demonstrate that the natural equivalence group of the problem…
This study assesses the reproducibility of 1H-MRS scans across different vendors and sessions.
problem Lack of harmonization in magnetic resonance spectroscopy protocols among vendors.
method Analysis of CV and ICC for within- and between-sessions, and correlation coefficients for across machines.
result Metabolite concentrations are highly reproducible across different vendors and sessions.
Paper addresses concentration of distances for fractional quasi p-norms, identifying conditions for concentration and anti-concentration.
problem Understanding concentration of distances for fractional quasi p-norms in high dimensions.
method Analyzes conditions for concentration and anti-concentration of distances for fractional quasi p-norms.
result Identifies conditions for concentration and anti-concentration of fractional quasi p-norms, ruling out some approaches and specifying conditions for control.
Study Finsler metric measure manifolds' concentration properties.
problem Understanding concentration properties in Finsler metric measure manifolds.
method Established relationships with observable diameter, isoperimetric inequalities, and first eigenvalue.
result Derived a Cheng type upper bound estimate for the first closed eigenvalue.
New method improves missing mass concentration bounds.
problem Missing mass concentration problem
method New method of estimating concentration of heterogenic sums
result Slightly improved state-of-the-art bounds
Surfaces in 3-manifolds concentrate at curvature critical points.
problem Understanding concentration of surfaces in 3-manifolds.
method Proving surfaces concentrate at critical points of scalar curvature.
result Simply connected H-surfaces concentrate at curvature critical points.
Sharp concentration bounds for i.i.d. variables.
problem Controlling the tail probabilities of independent variables.
method Extension of Sanov's theorem using large deviations and information theory.
result Matching concentration and anti-concentration bounds for i.i.d. samples of any size.
SLT reveals how grokking occurs via basin selection in training.
problem Understanding grokking in machine learning models.
method Singular Learning Theory (SLT) to analyze the loss landscape and local learning coefficient (LLC).
result LLC ranks basins by statistical preference, leading to grokking.
We survey recent results related to the concentration of eigenfunctions. We also prove some new results concerning ball-concentration, as well as showing that eigenfunctions saturating lower bounds for L1-norms must also, in a measure theoretical sense, have extreme concentration near a geodesic.