In this paper, we first study the Poisson reductions of controlled Hamiltonian (CH) system and symmetric CH system by controllability distributions. These reductions are the extension of Poisson reductions by distribution for Poisson manifolds to that for phase spaces of CH systems with external force and control. We g…
We realise the first and second Grushin distributions as symmetry reductions of the 3-dimensional Heisenberg distribution and 4-dimensional Engel distribution respectively. Similarly, we realise the Martinet distribution as an alternative symmetry reduction of the Engel distribution. These reductions allow us to derive…
The version of Marsden-Ratiu reduction theorem for Nambu-Poisson manifolds by a regular distribution has been studied by Ibaˊn~ez et al. In this paper we show that the reduction is always ensured unless the distribution is zero. Next we extend the more general Falceto-Zambon Poisson reduct…
We show that the distribution of symmetry of a naturally reductive nilpotent Lie group coincides with the invariant distribution induced by the set of fixed vectors of the isotropy. This extends a known result on compact naturally reductive spaces. We also address the study of the quotient by the foliation of symmetry.
Unified framework for DR and clustering using Gromov-Wasserstein.
problem Capturing structure in high-dimensional datasets.
method Distributional reduction framework using Gromov-Wasserstein.
result Unified approach recovers DR and clustering as special cases.
Study of symmetry distributions in Lorentzian naturally reductive nilmanifolds.
problem Understanding symmetry in Lorentzian naturally reductive nilmanifolds.
method Analysis of 2-step nilpotent Lorentzian Lie groups with transitive isometry subgroups.
result Fixed points of isotropy representation indicate the distribution of symmetry.
This work introduces a unified approach to the reduction of Poisson manifolds using their description by graded symplectic manifolds. This yields a generalization of the classical Poisson reduction by distributions (Marsden-Ratiu reduction). Further it allows one to construct actions of strict Lie 2-groups and to descr…
Efficiently transforms samples from various statistical models.
problem Approximately transforming samples from one statistical model to another without knowing the source model's parameters.
method Constructs computationally efficient procedures to reduce uniform, Erlang, and Laplace models to general target families.
result Establishes nonasymptotic reductions between canonical high-dimensional problems, such as mixtures of experts, phase retrieval, and signal denoising.
Efficiently transforms Gaussian data to simulate various target distributions.
problem Generating observations from different target distributions given a single Gaussian observation.
method Designs computationally efficient procedures to approximate target distributions.
result Establishes reduction-based computational lower bounds for high-dimensional statistical models.
Normal distributions ensure asymptotic variance reduction in moment matching Monte Carlo.
problem Asymptotic variance reduction in general integration problems.
method Characterization of conditions for asymptotic variance reduction using normal distributions.
result Asymptotic variance reduction is guaranteed for normal distributions in moment matching Monte Carlo.
New algorithms reduce computational burden for principal support vector machines.
problem High computational cost of principal support vector machines for large datasets.
method Two distributed estimation algorithms for principal support vector machines.
result Statistical efficiency is maintained with distributed algorithms.
New Lie systems derived from Goursat distributions with applications to differential equations.
problem Analyzing Lie systems associated with Goursat distributions and their applications.
method Analyzing bracket-generating distributions and their relation to Lie systems, focusing on reductions and reconstructions.
result Lie systems associated with Goursat distributions can be reduced and solutions reconstructed from reduced systems.
New method for reducing dimensions of distributional data.
problem Nonlinear sufficient dimension reduction for distribution-on-distribution regression.
method Building universal kernels on metric spaces to characterize conditional independence.
result Method outperforms competing methods in synthetic and real data applications.
This paper analyses the parabolic geometries generated by a free n-distribution in the tangent space of a manifold. It shows that certain holonomy reductions of the associated normal Tractor connections, imply preferred connections with special properties, along with Riemannian or sub-Riemannian structures on the man…
Study star products on Poisson manifolds compatible with reduction.
problem Finding star products compatible with coisotropic reduction.
method Compute second constraint Hochschild cohomology of constraint algebra.
result Determine infinitesimal star products on Poisson manifolds.
Adaptive importance sampling for stochastic optimization is a promising approach that offers improved convergence through variance reduction. In this work, we propose a new framework for variance reduction that enables the use of mixtures over predefined sampling distributions, which can naturally encode prior knowledg…
SQFA learns features maximizing Fisher-Rao distance for better classification.
problem Improving classification accuracy through feature learning.
method SQFA learns linear features maximizing Fisher-Rao distance between class-conditional distributions.
result SQFA-H features achieve the best classification accuracy.
Unified neural network for linear and nonlinear dimension reduction.
problem Efficiently perform linear and nonlinear sufficient dimension reduction.
method Belted and Ensembled Neural Network (BENN) framework.
result Unified framework for both linear and nonlinear dimension reduction.
Variance reduction (VR) methods boost the performance of stochastic gradient descent (SGD) by enabling the use of larger, constant stepsizes and preserving linear convergence rates. However, current variance reduced SGD methods require either high memory usage or an exact gradient computation (using the entire dataset)…
Pairs (Hamiltonian system, Lagrangian distribution), called dynamical Lagrangian distributions, appear naturally in Differential Geometry, Calculus of Variations and Rational Mechanics. The basic differential invariants of a dynamical Lagrangian distribution w.r.t. the action of the group of symplectomorphisms of the a…
Improves gradient estimation for discrete distributions with variance reduction techniques.
problem Excessive variance in gradient estimation for discrete distributions.
method Stein operators for discrete distributions and control variates.
result Substantially lower variance in gradient estimation.
Proposes variance reduction for optimizing permutation models.
problem High variance in gradient estimates for discrete latent variables.
method Control variates for the Plackett-Luce distribution.
result Optimization of black-box functions over permutations using SGD.
In the covariate shift learning scenario, the training and test covariate distributions differ, so that a predictor's average loss over the training and test distributions also differ. In this work, we explore the potential of extreme dimension reduction, i.e. to very low dimensions, in improving the performance of imp…
Paper improves SDR estimation speed and conditions.
problem Improving sufficient dimension reduction for multi-index models.
method Estimating expected smoothed gradient outer product.
result Achieves fast parametric convergence rate of Cd⋅n−1/2. Paper improves distributed mean estimation and variance reduction without relying on input norm.
problem Distributed mean estimation and variance reduction with large input norms.
method Quantization and lattice theory connection for improved error bounds.
result Output error bounds depend only on input distance, not norm.
Reducing communication in training large-scale machine learning applications on distributed platform is still a big challenge. To address this issue, we propose a distributed hierarchical averaging stochastic gradient descent (Hier-AVG) algorithm with infrequent global reduction by introducing local reduction. As a gen…
PGPCA improves PCA for nonlinear data in neuroscience.
problem Nonlinear data distribution in neuroscience.
method Developed PGPCA for nonlinear manifolds, incorporating EM algorithm.
result PGPCA outperforms PPCA in modeling data around nonlinear manifolds.
Paper solves NGCA for discrete distributions using LLL method.
problem Learning hidden non-Gaussian components in discrete distributions.
method Utilizes LLL lattice basis reduction method.
result Sample and computationally efficient algorithm for NGCA in discrete distributions.
We consider supervised dimension reduction problems, namely to identify a low dimensional projection of the predictors $\-x$ which can retain the statistical relationship between $\-x$ and the response variable y. We follow the idea of the sliced inverse regression (SIR) and the sliced average variance estimation (SA…
Bayesian model fuses multiple classifiers with explicit correlation modeling.
problem Combining outputs of multiple classifiers with explicit correlation.
method Hierarchical Bayesian model with correlated Dirichlet distribution.
result Fused classifier performance can be Bayes optimal even for highly correlated base classifiers.
Paper interprets UMAP and t-SNE as probabilistic MAP inference.
problem Understanding and interpreting UMAP and t-SNE.
method Interprets UMAP and t-SNE as MAP inference methods corresponding to a probabilistic model of the graph Laplacian.
result Shows UMAP and t-SNE can be understood as probabilistic inference methods.
This paper simplifies finding least favorable priors by reducing dimensionality.
problem Finding least favorable priors is challenging due to infinite-dimensional optimization.
method Develops a dimensionality reduction method using Bregman divergences.
result Allows use of gradient ascent algorithms for finding least favorable priors.
Standard methods for anomaly detection assume that all features are observed at both learning time and prediction time. Such methods cannot process data containing missing values. This paper studies five strategies for handling missing values in test queries: (a) mean imputation, (b) MAP imputation, (c) reduction (redu…
This report concerns the problem of dimensionality reduction through information geometric methods on statistical manifolds. While there has been considerable work recently presented regarding dimensionality reduction for the purposes of learning tasks such as classification, clustering, and visualization, these method…
MCE reduces embedding instability in nonlinear dimensionality reduction.
problem Embedding instability caused by random initialization.
method Median of multiple embeddings (MCE) based on large deviation theory.
result MCE achieves consistency at an exponential rate and effectively mitigates instability.
Ensemble learning has had many successes in supervised learning, but it has been rare in unsupervised learning and dimensionality reduction. This study explores dimensionality reduction ensembles, using principal component analysis and manifold learning techniques to capture linear, nonlinear, local, and global feature…
Unified framework for decentralized optimization combining gradient tracking and variance reduction.
problem Solving finite-sum minimization problems in distributed systems with privacy and resource constraints.
method Unified algorithmic framework combining variance-reduction and gradient tracking.
result Unified methods achieve robust performance and fast convergence for smooth and strongly-convex objectives, and are applicable to non-convex problems.
The paper presents a probabilistic framework for SPD matrices in machine learning.
problem Machine learning on SPD matrices is fragmented; this paper aims to unify it.
method Unified probabilistic framework using Gaussian distributions and Bayes classifiers.
result Different SPD machine learning tools can be reinterpreted and extended using Gaussian distributions.
DMT enhances deep neural networks to better preserve data structures.
problem Preserving geometric, topological, and distributional structures of data in NLDR.
method Deep manifold transformation (DMT) using cross-layer LGP constraints.
result DMT networks outperform existing NLDR methods in preserving data structures.
Paper optimizes classification of distributions using Wasserstein metric.
problem Classifying instances represented by distributions on a vector space.
method Maximizing Fisher's ratio in the Wasserstein metric space through iterative algorithm.
result The method enhances classification performance and is robust to variations in distribution summaries.
Explains SNE, t-SNE, and their variants for manifold learning.
problem Dimensionality reduction and manifold learning.
method Probabilistic approach using Gaussian and Student-t distributions.
result Out-of-sample extension and acceleration methods for t-SNE.
AB-SAGA optimizes distributed optimization over directed graphs using variance reduction and stochastic weights.
problem Optimizing distributed stochastic optimization over directed graphs with stochastic weights.
method AB-SAGA combines variance reduction and network-level gradient tracking, using both row and column stochastic weights.
result AB-SAGA converges linearly to the global optimal with a constant step-size and achieves a linear speed-up over centralized methods.
A new method using energy distance for ensemble and scenario reduction.
problem Solving complex dynamic and stochastic programs, especially in energy systems.
method Proposes a new method based on energy distance for ensemble and scenario reduction.
result Reduced scenario sets exhibit better statistical properties for energy distance than Wasserstein distance.
Sequential or online dimensional reduction is of interests due to the explosion of streaming data based applications and the requirement of adaptive statistical modeling, in many emerging fields, such as the modeling of energy end-use profile. Principal Component Analysis (PCA), is the classical way of dimensional redu…
Paper reduces vocabulary losslessly for language model cooperation.
problem Language models struggle to cooperate with different tokenizations.
method Established a theoretical framework for lossless vocabulary reduction.
result Efficiently converts models with different tokenizations to cooperate with maximal common vocabulary.
Reduces data leakage in distributed deep learning models.
problem Prevents reconstruction of sensitive raw data patterns during client communications.
method Reduces distance correlation between raw data and learned representations.
result Resilient to reconstruction attacks while maintaining model accuracy.
New algorithms improve distributional TD learning with linear approximations.
problem Estimating return distributions in reinforcement learning.
method Fine-grained analysis of linear-categorical Bellman equation, variance reduction techniques.
result Tight sample complexity bounds for distributional TD learning with linear approximations.
Survey on geometric foundations of data reduction methods.
problem High-dimensional data with intrinsic nonlinear structure.
method Spectral manifold learning methods.
result Derivation and convergence analysis of spectral manifold learning.