New algorithm learns two-layer neural networks under symmetric input distributions.
problem Learning two-layer neural networks with symmetric inputs.
method Method-of-moments framework, spectral algorithms.
result Guaranteed to recover parameters of ground-truth network under certain conditions.
New neural network models learn symmetric functions of varying input sizes.
problem Learning symmetric functions with varying input sizes.
method Functional perspective on neural networks, treating symmetric functions as functions over probability measures.
result Established approximation and generalization bounds for shallow architectures that extend across input sizes.
Transformationally invariant processors constructed by transformed input vectors or operators have been suggested and applied to many applications. In this study, transformationally identical processing based on combining results of all sub-processes with corresponding transformations at one of the processing steps or …
New neural networks respect symmetries in symmetric tensors, improving efficiency and generalization.
problem Learning from symmetric tensors efficiently and respecting their inherent symmetries.
method Developed two characterizations of linear permutation equivariant functions between symmetric power spaces of R^n.
result These functions are highly data efficient compared to standard MLPs and generalize well to different sizes of symmetric tensors.
Non-symmetric rectangular correlation matrices occur in many problems in economics. We test the method of extracting statistically meaningful correlations between input and output variables of large dimensionality and build a toy model for artificially included correlations in large random time series.The results are t…
Given a symmetric nonnegative matrix A, symmetric nonnegative matrix factorization (symNMF) is the problem of finding a nonnegative matrix H, usually with much fewer columns than A, such that A≈HHT. SymNMF can be used for data analysis and in particular for various clustering tasks. In this paper, we p…
Kernel embeddings of distributions and the Maximum Mean Discrepancy (MMD), the resulting distance between distributions, are useful tools for fully nonparametric two-sample testing and learning on distributions. However, it is rarely that all possible differences between samples are of interest -- discovered difference…
Deep single-index Fréchet regression for metric space-valued outputs
problem Predicting outputs in non-Euclidean spaces
method DeSI (Deep Single-Index Fréchet Regression)
result Interpretable index direction for inputs
New bounds tighten the generalization error of Gibbs algorithm.
problem Bounding the generalization error of Gibbs algorithm.
method Characterization of generalization error in terms of symmetrized KL information.
result Exact characterization of Gibbs algorithm's expected generalization error.
Theorems and techniques to form different types of transformationally invariant processing and to produce the same output quantitatively based on either transformationally invariant operators or symmetric operations have recently been introduced by the authors. In this study, we further propose to compose a geared rota…
We present a systematic analysis on the performance of a phonetic recogniser when the window of input features is not symmetric with respect to the current frame. The recogniser is based on Context Dependent Deep Neural Networks (CD-DNNs) and Hidden Markov Models (HMMs). The objective is to reduce the latency of the sy…
We propose a new input perturbation mechanism for publishing a covariance matrix to achieve (ε,0)-differential privacy. Our mechanism uses a Wishart distribution to generate matrix noise. In particular, We apply this mechanism to principal component analysis. Our mechanism is able to keep the positive semi-definitene…
The study predicts Kronecker coefficients using interpretable machine learning models.
problem Predicting Kronecker coefficients of the symmetric group.
method Employed interpretable machine learning models with input features of triples of partitions and b-loadings.
result Achieved an accuracy of approximately 83% and over 99% with transformer-based models.
Reservoir computing's success depends on mapping different input time series to separable states.
problem Quantifying the ability of random linear reservoirs to map different input time series.
method Mathematical framework using spectral properties of the connectivity matrix.
result Separation capacity is fully characterized by the spectral properties of the connectivity matrix.
New derivation shows how a three-factor learning rule is derived from Oja's rule.
problem Deriving a three-factor learning rule from Oja's rule.
method Using frame theory to systematically derive EGHR-PCA from Oja's rule.
result A principled derivation of a biologically plausible learning rule.
Capsule networks can only represent symmetric functions due to routing limitations.
problem Capsule networks' expressivity is limited to symmetric functions.
method Proved and empirically demonstrated that EM-routing and routing-by-agreement prevent capsule networks from distinguishing inputs and their negative counterpart.
result Capsule networks are not universal approximators due to the limitation of expressivity.
Prediction and explanation are key objects in supervised machine learning, where predictive models are known as black boxes and explanatory models are known as glass boxes. Explanation provides the necessary and sufficient information to interpret the model output in terms of the model input. It includes assessments of…
New method improves mean estimation for heavy-tailed data.
problem Estimating mean of heavy-tailed distributions.
method Median-of-Means (MoM) with symmetrization technique.
result Improved sample complexity bound for mean estimation.
A new approach to sensitivity analysis without the Sobol decomposition.
problem Traditional sensitivity indices like Sobol indices have limitations.
method Introducing sensitivity measures that generalize existing indices and define interaction effects.
result Sensitivity measures can create new indices and define interaction effects.
Equivariant CNNs improve RL performance in symmetric environments.
problem Learning equivariant representations for RL in symmetric environments.
method Proposed and studied equivariant CNNs for RL.
result Equivariant CNNs enhance RL performance and sample efficiency.
New method uses spherical harmonics to simplify learning single-index models.
problem Learning single-index models with unknown one-dimensional projections.
method Proposes using spherical harmonics instead of Hermite polynomials to capture rotational symmetry.
result Characterizes the complexity of learning single-index models under arbitrary spherically symmetric input distributions.
Olshausen and Field (OF) proposed that neural computations in the primary visual cortex (V1) can be partially modeled by sparse dictionary learning. By minimizing the regularized representation error they derived an online algorithm, which learns Gabor-filter receptive fields from a natural image ensemble in agreement …
New conditions ensure deep neural networks can approximate any function on non-Euclidean spaces.
problem Understanding how to modify neural network architectures to approximate functions on non-Euclidean spaces.
method Developed conditions for feature and readout maps that preserve universal approximation capabilities.
result Modified architectures can deterministically approximate any classifier on non-Euclidean spaces.
Neuc-MDS extends MDS for non-Euclidean data.
problem Limitations of classical MDS with non-Euclidean data.
method Generalizes inner product to symmetric bilinear forms, optimizes eigenvalues of dissimilarity Gram matrix.
result Optimizes STRESS for non-Euclidean data.
We introduce SARR for symmetric object pose estimation, improving CNN performance.
problem Ambiguities in symmetric object orientations hinder deep learning pose estimation.
method Numeric rotation representation using symmetry-derived trigonometric identities.
result SARR enables standard CNNs to achieve state-of-the-art performance.
New algorithms learn multi-index models via harmonic analysis, achieving statistical and computational trade-offs.
problem Learning multi-index models with unknown projections of input data.
method Exploiting the equivariance of the problem under the orthogonal group, we derive lower bounds and construct spectral algorithms based on harmonic tensor unfolding.
result Achieve statistical and computational trade-offs between sample and runtime complexity.
Deep neural networks favor symmetric structures, enabling multilevel symmetries.
problem Understanding and optimizing deep neural networks.
method Formulating DNN training as convex Lasso problems with geometric algebra.
result Deep networks inherently favor symmetric structures, enabling multilevel symmetries.
A fast, robust AMP algorithm for quadratic optimization problems.
problem Implementing robust approximate-message passing algorithms for quadratic optimization problems.
method Spectral pre-processing and mild modification of AMP algorithm iterates.
result Output solution close to AMP algorithm output for perturbed inputs.
Efficient neural network invariant to symmetry subgroups.
problem Designing neural networks invariant to symmetry subgroups for computational efficiency.
method A new G-invariant transformation module and multi-layer perceptron. result The proposed architecture is computationally and memory efficient, and universal.
Convolutional neural networks handle rotated image symmetries without dimensionality issues.
problem Binary image classification with rotational symmetry.
method Least squares plug-in classifiers based on convolutional neural networks under rotationally symmetric assumptions.
result Convolutional neural networks can circumvent the curse of dimensionality in binary image classification with rotational symmetry.
Transformer's attention mechanism is re-examined using kernel smoothing.
problem Understanding and optimizing the Transformer's attention mechanism.
method Presented a new kernel-based formulation of Transformer's attention mechanism.
result The new kernel-based formulation provides a better understanding of Transformer's attention components and introduces a new variant achieving competitive performance.
E2M predicts metric space outputs using deep learning.
problem Predicting non-Euclidean outputs like distributions and matrices.
method Weighted Fréchet means over learned weights.
result E2M achieves state-of-the-art performance across various outputs.
We demonstrate how the novel approach to the local geometry of structures of nonholonomic nature, originated by Andrei Agrachev, works in the following two situations: rank 2 distributions of maximal class in R^n with non-zero generalized Wilczynski invariants and rank 2 distributions of maximal class in R^n with addit…
Proposes a new neural head for asymmetric representation learning.
problem Asymmetric representation learning in directed relations.
method Role-aware neural convex divergence head.
result Role-aware projections improve directional accuracy over plain ICNN-Bregman heads.
New basis for permutation equivariant layers reduces computation costs.
problem Efficiently computing permutation equivariant layers in neural networks.
method Generalized partition algebra basis with low-rank tensors.
result Low-rank tensors enable faster computation compared to orbit basis.
Abstraction and realization are bilateral processes that are key in deriving intelligence and creativity. In many domains, the two processes are approached through rules: high-level principles that reveal invariances within similar yet diverse examples. Under a probabilistic setting for discrete input spaces, we focus …
Optimal Gaussian noise mechanisms achieve nearly optimal error in unbiased mean estimation.
problem Efficiently estimating the mean of high-dimensional data while preserving privacy.
method Differential privacy mechanisms with Gaussian noise, focusing on optimal covariance.
result Gaussian noise mechanisms achieve nearly optimal error among all private unbiased mean estimation mechanisms.
Study evaluates quantum and classical conditional Boltzmann machines for time-series forecasting.
problem Time-series forecasting using quantum and classical conditional Boltzmann machines.
method Developed and compared four conditional energy-based forecasting architectures: Gaussian-Bernoulli CRBM, QCRBM, QQRBM, and QFeatureQRBM. Evaluated using symmetric hyperparameter optimisation.
result No systematic evidence of a quantum advantage in time-series forecasting at the available sample size.
The paper improves confidence ellipsoids for ridge regression with PAC bounds.
problem Uncertainty quantification in ridge regression for insufficiently exciting inputs.
method Extension of SPS EOA algorithm to ridge regression with PAC bounds.
result Explicitly shows how regularization parameter affects region sizes and provides tighter bounds.
We propose Power Slow Feature Analysis, a gradient-based method to extract temporally slow features from a high-dimensional input stream that varies on a faster time-scale, as a variant of Slow Feature Analysis (SFA) that allows end-to-end training of arbitrary differentiable architectures and thereby significantly ext…
Stable unactivated neurons reduce expressiveness in ReLU networks.
problem Reducing expressiveness in ReLU neural networks due to stably unactivated neurons.
method Investigated the probability of neurons being stably unactivated in ReLU networks with symmetric weight and bias distributions.
result Proved the probability of a neuron being stably unactivated in the second hidden layer of a ReLU network.
A new neural network model identifies hysteresis universally.
problem Inability of existing models to simulate hysteresis universally.
method Inspired by the Preisach model, an Extended Preisach Neural Network (EPNN) is introduced with two hidden layers and a hybrid training algorithm.
result EPNN successfully identifies various hysteresis phenomena from different fields.
The paper defines symmetric brackets for skew-symmetric algebroids with totally skew-symmetric torsion.
problem Defining symmetric brackets for skew-symmetric algebroids.
method Using connections with totally skew-symmetric torsion and pseudo-Riemannian metrics.
result Explicit formula for the Levi-Civita connection and symmetric brackets on almost Hermitian manifolds.
The paper classifies symmetric triads with multiplicities and their applications.
problem Classifying symmetric triads with multiplicities and their applications.
method Developed the theory of symmetric triads with multiplicities, classified abstract triads, and determined corresponding triads for commutative compact triads.
result Classified symmetric triads with multiplicities and their applications.
We find all Ricci semi-symmetric as well as all conformally semi-symmetric spacetimes. Neither of these properties implies the other. We verify that only conformally flat spacetimes can be Ricci semi-symmetric without being conformally semi-symmetric and show that only vacuum spacetimes and spacetimes with just a Λ-t…
We establish a new symmetrization procedure for the isoperimetric problem in symmetric spaces of noncompact type. This symmetrization generalizes the well known Steiner symmetrization in euclidean space. In contrast to the classical construction the symmetrized domain is obtained by solving a nonlinear elliptic equatio…
Study on totally symmetric sets with group applications.
problem Understanding totally symmetric sets and their group applications.
method Survey of existing theory and applications to various groups.
result Exploration of totally symmetric sets in multiple group contexts.
We give convergence guarantees for estimating the coefficients of a symmetric mixture of two linear regressions by expectation maximization (EM). In particular, we show that the empirical EM iterates converge to the target parameter vector at the parametric rate, provided the algorithm is initialized in an unbounded co…