New method estimates and samples high-dimensional probability distributions avoiding optimization and approximation curse.
problem Estimating high-dimensional probability distributions from data samples.
method Hierarchic probability flow from coarse to fine scales, defined by conditional probabilities across scales.
result Sampling hierarchic models avoids critical slowing down at phase transitions and generates turbulence and dark matter images.
New method approximates high-dimensional probability densities efficiently.
problem Approximating high-dimensional probability densities accurately and efficiently.
method Hierarchical tensor-network approach using randomized SVD and linear equations.
result The method effectively approximates high-dimensional densities with linear complexity.
Matrix completion is a modern missing data problem where both the missing structure and the underlying parameter are high dimensional. Although missing structure is a key component to any missing data problems, existing matrix completion methods often assume a simple uniform missing mechanism. In this work, we study ma…
Paper proposes new density estimators for high-dimensional data.
problem Prohibitive computational cost and slow convergence rate in high-dimensional density estimation.
method Adaptive hyperbolic cross density estimators in mixed smooth Sobolev spaces.
result Proposed estimators do not suffer curse of dimensionality under Integral Probability Metrics.
Transformer model for probabilistic dynamical systems.
problem Modeling high-dimensional dynamical systems from noisy observations.
method Parallel between dynamical systems and language modeling; transformer-based model with geometrical properties; iterative training algorithm.
result Fine-grid approximation of conditional probabilities for high-dimensional systems.
A new method optimizes slicing directions for SW distances to improve high-dimensional probability measure comparison.
problem Challenging identification of informative slicing directions for SW distances.
method Constrained learning approach to optimize slicing directions, using continuous relaxations and gradient-based primal-dual approach.
result Demonstrated efficacy in learning more informative slicing directions on various high-dimensional data.
Efficient learning of minimax risk classifiers in high dimensions.
problem Efficient learning of classifiers in high-dimensional data.
method Iterative algorithm leveraging constraint generation methods for minimax risk classifiers.
result The algorithm provides efficient learning and feature selection in high-dimensional scenarios.
We study high-dimensional asymptotic performance limits of binary supervised classification problems where the class conditional densities are Gaussian with unknown means and covariances and the number of signal dimensions scales faster than the number of labeled training samples. We show that the Bayes error, namely t…
New method improves sampling from high-dimensional target densities.
problem Sampling from high-dimensional target densities using Monte Carlo algorithms.
method Extends Metropolis-Adjusted Langevin Diffusion algorithm with random precondition matrix modeling.
result Significantly improves performance and computational efficiency over standard MCMC methods.
A new IPM uses ReLU networks to measure probability discrepancies.
problem Measuring the difference between two probability distributions in high dimensions.
method Proposes a new parametric IPM using ReLU neural networks to optimize and distinguish between distributions.
result The proposed IPM has good convergence rates and can be used as a surrogate for other IPMs.
Quantum probability metrics improve distribution comparison in high dimensions.
problem Challenges in comparing probability distributions, especially in high-dimensional and non-compact domains.
method Quantum probability metrics (QPMs) derived from quantum state spaces, overcoming limitations of MMD.
result QPMs offer enhanced sensitivity to subtle distributional differences in high dimensions and improve performance in generative modeling.
Two-parameter models can learn high-dimensional targets via gradient flow.
problem Learning high-dimensional targets with limited parameters.
method Gradient flow approach for W<d models. result Two-parameter models can learn targets with arbitrarily high success probability.
This paper rigorously establishes that the existence of the maximum likelihood estimate (MLE) in high-dimensional logistic regression models with Gaussian covariates undergoes a sharp `phase transition'. We introduce an explicit boundary curve hMLE, parameterized by two scalars measuring the overall magnitu…
Model selection is indispensable to high-dimensional sparse modeling in selecting the best set of covariates among a sequence of candidate models. Most existing work assumes implicitly that the model is correctly specified or of fixed dimensions. Yet model misspecification and high dimensionality are common in real app…
The problem of categorical data analysis in high dimensions is considered. A discussion of the fundamental difficulties of probability modeling is provided, and a solution to the derivation of high dimensional probability distributions based on Bayesian learning of clique tree decomposition is presented. The main contr…
Develops a measure-theoretic framework for complex co-occurrence data.
problem Modeling and interpreting complex co-occurrences in high-dimensional data.
method Introduces measure-theoretic probability and conditional probability, investigates E-integrals.
result Establishes a rigorous measure-theoretic foundation for co-occurrence modeling.
A new method estimates rare events using tensor trains.
problem Estimating rare event probabilities in high-dimensional problems.
method Approximating optimal importance distribution via tensor-train decompositions and compositions.
result Better variance reduction and efficient computation of rare event probabilities.
Model selection is crucial to high-dimensional learning and inference for contemporary big data applications in pinpointing the best set of covariates among a sequence of candidate interpretable models. Most existing work assumes implicitly that the models are correctly specified or have fixed dimensionality. Yet both …
The paper solves optimal bounds for separating data points in high dimensions.
problem Correcting AI errors and analyzing vulnerabilities in high-dimensional data.
method General stochastic separation theorems with optimal probability estimates.
result Explicit and optimal estimates of separation probabilities for important classes of distributions.
This work improves deep neural network probability estimation methods.
problem Estimating probabilities from high-dimensional data with inherent uncertainty.
method Investigates and compares methods for probability estimation using deep neural networks, proposing a new method that promotes consistent probabilities.
result The new method outperforms existing approaches on most metrics on simulated and real-world data.
Some high-dimensional data.sets can be modelled by assuming that there are many different linear constraints, each of which is Frequently Approximately Satisfied (FAS) by the data. The probability of a data vector under the model is then proportional to the product of the probabilities of its constraint violations. We …
We consider a high dimensional binary classification problem and construct a classification procedure by minimizing the empirical misclassification risk with a penalty on the number of selected features. We derive non-asymptotic probability bounds on the estimated sparsity as well as on the excess misclassification ris…
Improved location estimation for high-dimensional data with finite sample size.
problem Estimating the shift in high-dimensional data with limited samples.
method Smoothed estimators and bounds on subgamma vectors.
result Convergence to Cramér-Rao bound for finite sample sizes.
A new method for estimating density ratios in high dimensions.
problem Difficulty in accurately comparing probability distributions in high-dimensional settings.
method Divide-and-conquer approach via an infinite continuum of bridge distributions and time score matching.
result The proposed method effectively estimates density ratios and performs well on complex datasets.
One of the fundamental problems in machine learning is the estimation of a probability distribution from data. Many techniques have been proposed to study the structure of data, most often building around the assumption that observations lie on a lower-dimensional manifold of high probability. It has been more difficul…
We introduce a framework using Generative Adversarial Networks (GANs) for likelihood--free inference (LFI) and Approximate Bayesian Computation (ABC) where we replace the black-box simulator model with an approximator network and generate a rich set of summary features in a data driven fashion. On benchmark data sets, …
Estimates high-dimensional posterior densities by marginal distributions and neural networks.
problem High-dimensional probability density estimation for inference is difficult.
method Direct estimation of lower-dimensional marginal distributions, using Moment Networks for fast computation of moments.
result Demonstrates estimation of gravitational wave time series and applications in cosmology.
In general, the clustering problem is NP-hard, and global optimality cannot be established for non-trivial instances. For high-dimensional data, distance-based methods for clustering or classification face an additional difficulty, the unreliability of distances in very high-dimensional spaces. We propose a distance-ba…
fastHDMI improves neuroimaging variable selection in high-dimensional data.
problem Efficient variable screening in high-dimensional neuroimaging datasets.
method Three mutual information estimation methods implemented in fastHDMI.
result FFTKDE-based method superior for continuous nonlinear outcomes.
New test for comparing high-dimensional text data.
problem Testing equality of multinomial distributions in high dimensions.
method Proposed a test statistic with asymptotic normality under null.
result Achieves optimal detection boundary across parameter space.
Our paper characterizes how ReLU affects GD's implicit bias in high-dimensional neural networks.
problem Understanding the implicit bias of gradient descent on neural networks.
method Novel primal-dual analysis tracking predictions and coefficients.
result The implicit bias approximates the minimum-ℓ2-norm solution with high probability. Deep FPF approximates gain function for high-dimensional particle filtering.
problem Approximating the exact gain function in high-dimensional settings.
method Represent the gain function as a neural network gradient and solve a variational Poisson equation via optimization.
result The approach allows parallel processing of particles and is applicable to high-dimensional problems.
We formulate and analyze a graphical model selection method for inferring the conditional independence graph of a high-dimensional nonstationary Gaussian random process (time series) from a finite-length observation. The observed process samples are assumed uncorrelated over time and having a time-varying marginal dist…
The paper analyzes the robustness of a minimum ℓ2 interpolator in high-dimensional linear regression.
problem Analyzing the robustness of a minimum ℓ2 interpolator in high-dimensional linear regression. method The paper analyzes the interpolator with minimal ℓ2-norm in a general high-dimensional linear regression framework, proving bounds on prediction loss. result The paper shows that the prediction loss of the interpolator is bounded by (∥β∗∥22rcn(Σ)∨∥ξ∥2)/n with high probability, revealing a transition in rates. Paper extends tail bounds to high-dimensional random objects on Riemannian manifolds.
problem Need for tail bounds in high-dimensional data.
method Random walks on graph approximating the manifold, ensuring spectral similarity.
result Derived tensor Chernoff bound for Riemannian manifolds.
In this paper, we consider the problem of classification of M high dimensional queries y1,⋯,yM∈BS to N high dimensional classes x1,⋯,xN∈AS where A and B are discrete alphabets and the probabilistic model that relates data to the classes P(x,y) is known. This problem has applications …
Unified framework for semi-supervised regression with misspecified models.
problem Estimating regression coefficients in conditional mean models with unlabeled data.
method Developed an augmented inverse probability weighted (AIPW) method using regularized calibrated estimators for PS and OR nuisance models.
result The proposed estimator is consistent, asymptotically normal, and provides valid confidence intervals even with misspecified OR models and high-dimensional data.
The paper provides statistical guarantees for SGD and ASGD in high-dimensional settings.
problem Theoretical understanding of SGD and ASGD in high-dimensional settings.
method Transfer of tools from high-dimensional time series to online learning, using coupling techniques.
result Established geometric-moment contraction and q-th moment convergence of SGD and ASGD. NNLMs optimize poorly for word probabilities due to embedding space structure.
problem NNLMs assign suboptimal probabilities to some words.
method Analyzed the inductive bias of NNLMs and the structure of word embeddings.
result Words on the convex hull have bounded probability, affecting others.
Classification is an important statistical learning tool. In real application, besides high prediction accuracy, it is often desirable to estimate class conditional probabilities for new observations. For traditional problems where the number of observations is large, there exist many well developed approaches. Recentl…
Proposes a group-splicing algorithm for efficient BSGS in high-dimensional settings.
problem Efficiently selecting a small part of non-overlapping groups for best interpretability in high-dimensional settings.
method Iteratively detects relevant groups and excludes irrelevant ones using a novel group information criterion.
result Certifiable polynomial-time algorithm for identifying the optimal subset of groups with high probability.
Expands learning paradigm to stochastic orders using Choquet-Toland distance and Variational Dominance Criterion.
problem Learning high-dimensional distributions with stochastic orders.
method Introduces Choquet-Toland distance and Variational Dominance Criterion, uses input convex maxout networks (ICMNs).
result Proposes surrogates for Choquet-Toland distance and Variational Dominance Criterion with parametric rates.
We consider the problem of the estimation of a high-dimensional probability distribution from i.i.d. samples of the distribution using model classes of functions in tree-based tensor formats, a particular case of tensor networks associated with a dimension partition tree. The distribution is assumed to admit a density …
Robust variable selection for high-dimensional data with missing and measurement errors.
problem Missing data and measurement errors confound data distribution.
method Exponential loss function with inverse probability weighting and additive error models.
result The Atan punishment method improves robust variable selection.
CCs learn high-dimensional distributions from heterogeneous data.
problem Learning high-dimensional distributions from heterogeneous data.
method Introducing characteristic circuits (CCs) that learn from data and use spectral domain.
result CCs outperform state-of-the-art density estimators on common benchmark data sets.
Ridge regression reveals surprising high-dimensional behaviors via random matrix theory.
problem Understanding power-law scalings in high-dimensional regression models.
method Random matrix theory and free probability.
result Analytic formulas for training and generalization errors derived from S-transform. New CLT for AIPW estimator in high-dimensional settings.
problem Estimating ATE in high-dimensional covariate scenarios.
method Cross-fitting AIPW estimator with well-specified models.
result Established a new CLT for the scaled cross-fit AIPW.
Method identifies low-dimensional structure in high-dimensional probability measures.
problem Identifying low-dimensional structure in high-dimensional probability measures.
method Extends prior work on minimizing majorizations of the Kullback-Leibler divergence to identify optimal approximations within a specific class of measures.
result Connection between dimensional logarithmic Sobolev inequality and approximations with the ansatz.