Deep neural networks approximate functions in shift-invariant spaces with controlled error.
problem Approximating functions in shift-invariant spaces with neural networks.
method Using deep ReLU neural networks, estimating approximation error bounds based on network width and depth.
result Deep neural networks achieve optimal approximation rates for Sobolev spaces up to a logarithmic factor.
New adaptive signal denoising method mimics oracle with better statistical properties.
problem Adaptive discrete-time signal denoising with linear oracle structure.
method Minimizes the ℓ2-norm of the estimation residual, proving oracle inequalities for ℓ2-loss. result Improved statistical properties over ℓ∞-fit estimators, especially in ℓ2- and pointwise losses. The paper addresses instability in CNNs' first layer by proving max pooling's shift invariance.
problem Instability in CNNs' first layer, leading to sensitivity to small input shifts.
method Establishing conditions for max pooling's shift invariance and deriving a measure of stability.
result Max pooling approximates a nearly shift-invariant complex modulus under certain conditions.
New forms of symmetric shift-invariant subspaces found for harmonic maps.
problem Understanding harmonic maps into symmetric and k-symmetric spaces. method Imposing a symmetry condition on shift-invariant subspaces of a Hilbert space.
result Obtained new general forms for symmetric shift-invariant subspaces and extended solutions.
We consider the problem of improving the efficiency of randomized Fourier feature maps to accelerate training and testing speed of kernel methods on large datasets. These approximate feature maps arise as Monte Carlo approximations to integral representations of shift-invariant kernel functions (e.g., Gaussian kernel).…
Study links harmonic maps to shift-invariant subspaces in complex function spaces.
problem Understanding the relationship between harmonic maps and shift-invariant subspaces.
method Operator-theoretic methods to derive a criterion for the finiteness of the uniton number.
result Derives a criterion for the finiteness of the uniton number in harmonic maps.
Nonlinear kernel regression models are often used in statistics and machine learning because they are more accurate than linear models. Variable selection for kernel regression models is a challenge partly because, unlike the linear regression setting, there is no clear concept of an effect size for regression coeffici…
New algorithms for learning shift-invariant components and aligning signals.
problem Learning shift-invariant components and aligning signals.
method Formulated optimization problems using circulant and convolutional matrices, proposed efficient solutions.
result Effective algorithms for learning shift-invariant components and aligning signals.
Efficiently estimates densities of multidimensional shift-invariant distributions.
problem Density estimation for shift-invariant multidimensional distributions.
method Efficient algorithms for learning any distribution in the class from samples, using total variation distance.
result Shift-invariant distributions can be learned efficiently with a number of samples and time proportional to 1/εd+2 and 1/ε2d+2 respectively. Proposes a new method to improve CNNs' shift invariance and accuracy.
problem Improving CNNs' shift invariance and prediction accuracy.
method Replaces RMax with CMod, a Gabor-like structure, to increase shift invariance and accuracy.
result Achieves superior accuracy on ImageNet and CIFAR-10 classification tasks.
FNNs detect EEG signals without position dependence.
problem Detecting EEG signals without position dependence.
method Shift invariant functional neural networks (FNNs) using FDA methods.
result FNNs outperform FDA benchmarks in EEG classification.
Estimating signals with linear recurrence relations under Gaussian noise is nearly as hard as sparse signals.
problem Estimating discrete-time signals with unknown linear recurrence relations in Gaussian noise.
method Analyzing shift-invariant subspaces and their Fourier coefficients as reproducing filters.
result The statistical complexity is nearly the same as for s-sparse signals, and the estimator is tractable. New quantization methods improve accuracy of Random Fourier Features.
problem Improving accuracy of Random Fourier Features for machine learning.
method Sigma-Delta and distributed noise-shaping quantization methods for 1-bit and low bit-depth quantization.
result Quantized RFFs allow high accuracy approximation of underlying kernels with polynomial error decay.
New deep network derived from rate reduction principles, explaining features and efficiency.
problem Understanding and optimizing deep learning architectures.
method Gradient ascent scheme for rate reduction leading to multi-layer deep network.
result Explicitly constructed multi-layer network with precise optimization and interpretation.
New Fourier features improve high-precision approximation in large-scale problems.
problem Designing scalable, high-precision Fourier features for large-scale kernel methods.
method Introducing a new family of quadrature rules that accurately approximate the Gaussian measure in higher dimensions.
result Improved approximation bounds with new Fourier features.
Kernel methods represent one of the most powerful tools in machine learning to tackle problems expressed in terms of function values and derivatives due to their capability to represent and model complex relations. While these methods show good versatility, they are computationally intensive and have poor scalability t…
Study approximates probability measures using structured classes of functions.
problem Approximating probability measures in Wasserstein-p distance. method Structured classes of approximators for functions in Lp(Ω), transferring to measures in Wp(Ω). result Linear rate approximation for measures with densities bounded away from zero.
MCLNN improves sound event recognition with fewer parameters.
problem Improving sound event recognition with deep neural networks.
method Developed MCLNN to enforce sparseness and frequency shift invariance.
result MCLNN achieved competitive performance with 12% fewer parameters.
New method learns compressed transforms with flexible displacement operators.
problem Efficiently representing and learning shift-invariant patterns in neural networks.
method Explicitly learns over displacement operators and low-rank components in LDR matrices.
result Reduces sample complexity and improves model accuracy with fewer parameters.
New framework explains leading digit patterns without probabilistic assumptions.
problem Explaining leading digit distributions without relying on probabilistic models.
method Shift-invariant functional equation and affine-plus-periodic formulas.
result Unified mathematical foundation for understanding digit distributions.
Neural time-series data contain a wide variety of prototypical signal waveforms (atoms) that are of significant importance in clinical and cognitive research. One of the goals for analyzing such data is hence to extract such 'shift-invariant' atoms. Even though some success has been reported with existing algorithms, t…
This paper tackles spatio-temporal information preservation in machine learning.
problem Conventional machine learning assumes orthogonal data attributes, disrupting spatio-temporal information.
method Shift-invariant k-means, convolutional dictionary learning, and spatio-temporal hypercomplex encoding schemes are proposed.
result Gabor feature extraction outperforms convolutional dictionary learning in spatio-temporal information preservation.
The maximum mean discrepancy (MMD) is a recently proposed test statistic for two-sample test. Its quadratic time complexity, however, greatly hampers its availability to large-scale applications. To accelerate the MMD calculation, in this study we propose an efficient method called FastMMD. The core idea of FastMMD is …
We connect shift-invariant characteristic kernels to infinitely divisible distributions on Rd. Characteristic kernels play an important role in machine learning applications with their kernel means to distinguish any two probability measures. The contribution of this paper is two-fold. First, we show, usi…
The paper explores high-dimensional learning in finance, proving key aspects and setting lower bounds.
problem Understanding when and how large, over-parameterized models achieve predictive success in finance.
method Theoretical foundations and empirical validation of two key aspects: standardization and information-theoretic lower bounds.
result Empirical validation shows that high-dimensional learning in finance often relies on lower-complexity artefacts rather than the intended mechanism.
Proposes a new CSC model for handling unknown noise.
problem Existing CSC methods can only model Gaussian noise, which is restrictive.
method Uses Gaussian mixture model for unknown noise and EM algorithm for optimization.
result Effective modeling of complicated unknown noise with high-quality filters and representation.
In this paper, we introduce DICOD, a convolutional sparse coding algorithm which builds shift invariant representations for long signals. This algorithm is designed to run in a distributed setting, with local message passing, making it communication efficient. It is based on coordinate descent and uses locally greedy u…
Bird sounds possess distinctive spectral structure which may exhibit small shifts in spectrum depending on the bird species and environmental conditions. In this paper, we propose using convolutional recurrent neural networks on the task of automated bird audio detection in real-life environments. In the proposed metho…
Sparse coding is an unsupervised learning algorithm that learns a succinct high-level representation of the inputs given only unlabeled data; it represents each input as a sparse linear combination of a set of basis functions. Originally applied to modeling the human visual cortex, sparse coding has also been shown to …
Single model estimates uncertainty via biased data shifts.
problem Estimating uncertainties in deep neural networks.
method Trivial input transformation to approximate ensemble behavior.
result Single model uncertainty estimates are superior to current methods.
We analyze the structure of the \emph{frequency space} Q(F) of a nonabelian free group F=F(a1,...,ak) consisting of all shift-invariant Borel probability measures on ∂F and construct a natural action of Out(F) on Q(F). In particular we prove that for any outer automorphism φ of F the \emph{conju…
New spectral mixture representation for isotropic kernels simplifies random Fourier features.
problem Applying Random Fourier Features to complex kernels.
method Decompose isotropic kernels into scale mixtures of α-stable random vectors.
result Constructive spectral sampling formula for various kernels.
Quantum kernels can be efficiently embedded into classical feature spaces.
problem Can all quantum kernels be efficiently embedded into classical feature spaces?
method Invoking computational universality and using techniques like random Fourier features, the authors show that certain classes of quantum kernels can be efficiently embedded.
result For shift-invariant and composition kernels, embedding quantum kernels are universal and efficient.
Scalable methods integrate multiview data for clinical outcomes.
problem Jointly associate and predict outcomes from multiple data sources.
method Randomized Fourier bases for nonlinear mappings, view-independent low-dimensional representations.
result Identified molecular signatures for COVID-19 status and severity.
Tensor methods have emerged as a powerful paradigm for consistent learning of many latent variable models such as topic models, independent component analysis and dictionary learning. Model parameters are estimated via CP decomposition of the observed higher order input moments. However, in many domains, additional inv…
This article addresses the issue of representing electroencephalographic (EEG) signals in an efficient way. While classical approaches use a fixed Gabor dictionary to analyze EEG signals, this article proposes a data-driven method to obtain an adapted dictionary. To reach an efficient dictionary learning, appropriate s…
Median activation functions improve GNNs by capturing local graph signal behavior.
problem Lack of local nonlinear graph signal encoding in GNNs.
method Proposed median activation functions with support on graph neighborhoods.
result Median activation functions improve GNN capacity with minimal complexity increase.
Optimizes asset allocation for risk measures in a Lévy market.
problem Maximizing time-consistent mean-risk reward with general risk measures.
method Uses a generalized Lévy market model and Hamilton-Jacobi-Bellman equation.
result Deterministic optimal solution under certain conditions.
A distributed algorithm learns patterns in large images and signals.
problem High-dimensional optimization in large images and signals.
method Distributed asynchronous algorithm with locally greedy coordinate descent.
result Patterns can be learned on large scales images from the Hubble Space Telescope.
ForecastNet uses a time-variant deep feed-forward neural network for better multi-step-ahead time series forecasting.
problem Time-invariant architectures limit multi-step-ahead forecasting.
method ForecastNet employs a deep feed-forward architecture with time-variant parameters and interleaved outputs.
result ForecastNet outperforms other models on multi-step-ahead time series forecasting tasks.
Designs CNNs for better image reconstruction.
problem Image reconstruction from limited data.
method Parseval convolution operators and chaining of elementary modules.
result CNN-based algorithm yields better results than sparsity-based methods.
This paper develops methods for obtaining distribution-free prediction regions for invariant representations.
problem Distributional shifts in machine learning models.
method Invariant risk minimization and weighted conformity scores.
result Proves the effectiveness of adaptive conformal intervals for uncertainty estimation.
Many modern tools in machine learning and signal processing, such as sparse dictionary learning, principal component analysis (PCA), non-negative matrix factorization (NMF), K-means clustering, etc., rely on the factorization of a matrix obtained by concatenating high-dimensional vectors from a training collection. W…
Multi-task learning leverages shared information among data sets to improve the learning performance of individual tasks. The paper applies this framework for data where each task is a phase-shifted periodic time series. In particular, we develop a novel Bayesian nonparametric model capturing a mixture of Gaussian proc…
New DP algorithms with margin guarantees for various hypothesis sets.
problem Differential privacy in machine learning with margin guarantees.
method Developed pure and efficient DP learning algorithms for linear, kernel-based, and neural network hypotheses.
result Margin guarantees are independent of input dimension and hypothesis type.
New framework learns sufficient invariant features robustly across distribution shifts.
problem Learning robust models under distribution shifts between training and test datasets.
method Sufficient Invariant Learning (SIL) framework and Adaptive Sharpness-aware Group Distributionally Robust Optimization (ASGDRO) algorithm.
result Empirical evaluations confirm ASGDRO's robustness against distribution shifts.
Paper simplifies CNNs for irregular data using MIMO graph filters.
problem Challenges in applying CNNs to irregularly structured data.
method Introduces MIMO graph filters to CNNs, simplifying architectures.
result Proposed architectures reduce model complexity and computational cost.
New method uses neural networks to solve complex PDEs from optimal control theory.
problem Solving high-dimensional Hamilton-Jacobi-Bellman PDEs.
method Iterative diffusion optimization techniques, focusing on path measures and divergences.
result Favourable properties of log-variance divergence for Monte Carlo estimators.