Efficiently estimates densities of multidimensional shift-invariant distributions.
problem Density estimation for shift-invariant multidimensional distributions.
method Efficient algorithms for learning any distribution in the class from samples, using total variation distance.
result Shift-invariant distributions can be learned efficiently with a number of samples and time proportional to 1/εd+2 and 1/ε2d+2 respectively. New forms of symmetric shift-invariant subspaces found for harmonic maps.
problem Understanding harmonic maps into symmetric and k-symmetric spaces. method Imposing a symmetry condition on shift-invariant subspaces of a Hilbert space.
result Obtained new general forms for symmetric shift-invariant subspaces and extended solutions.
Study links harmonic maps to shift-invariant subspaces in complex function spaces.
problem Understanding the relationship between harmonic maps and shift-invariant subspaces.
method Operator-theoretic methods to derive a criterion for the finiteness of the uniton number.
result Derives a criterion for the finiteness of the uniton number in harmonic maps.
New algorithms for learning shift-invariant components and aligning signals.
problem Learning shift-invariant components and aligning signals.
method Formulated optimization problems using circulant and convolutional matrices, proposed efficient solutions.
result Effective algorithms for learning shift-invariant components and aligning signals.
The paper addresses instability in CNNs' first layer by proving max pooling's shift invariance.
problem Instability in CNNs' first layer, leading to sensitivity to small input shifts.
method Establishing conditions for max pooling's shift invariance and deriving a measure of stability.
result Max pooling approximates a nearly shift-invariant complex modulus under certain conditions.
Deep neural networks approximate functions in shift-invariant spaces with controlled error.
problem Approximating functions in shift-invariant spaces with neural networks.
method Using deep ReLU neural networks, estimating approximation error bounds based on network width and depth.
result Deep neural networks achieve optimal approximation rates for Sobolev spaces up to a logarithmic factor.
Proposes a new method to improve CNNs' shift invariance and accuracy.
problem Improving CNNs' shift invariance and prediction accuracy.
method Replaces RMax with CMod, a Gabor-like structure, to increase shift invariance and accuracy.
result Achieves superior accuracy on ImageNet and CIFAR-10 classification tasks.
New framework explains leading digit patterns without probabilistic assumptions.
problem Explaining leading digit distributions without relying on probabilistic models.
method Shift-invariant functional equation and affine-plus-periodic formulas.
result Unified mathematical foundation for understanding digit distributions.
FNNs detect EEG signals without position dependence.
problem Detecting EEG signals without position dependence.
method Shift invariant functional neural networks (FNNs) using FDA methods.
result FNNs outperform FDA benchmarks in EEG classification.
We connect shift-invariant characteristic kernels to infinitely divisible distributions on Rd. Characteristic kernels play an important role in machine learning applications with their kernel means to distinguish any two probability measures. The contribution of this paper is two-fold. First, we show, usi…
We discuss the problem of adaptive discrete-time signal denoising in the situation where the signal to be recovered admits a "linear oracle" -- an unknown linear estimate that takes the form of convolution of observations with a time-invariant filter. It was shown by Juditsky and Nemirovski (2009) that when the $\ell_2…
New spectral mixture representation for isotropic kernels simplifies random Fourier features.
problem Applying Random Fourier Features to complex kernels.
method Decompose isotropic kernels into scale mixtures of α-stable random vectors.
result Constructive spectral sampling formula for various kernels.
Estimating signals with linear recurrence relations under Gaussian noise is nearly as hard as sparse signals.
problem Estimating discrete-time signals with unknown linear recurrence relations in Gaussian noise.
method Analyzing shift-invariant subspaces and their Fourier coefficients as reproducing filters.
result The statistical complexity is nearly the same as for s-sparse signals, and the estimator is tractable. In this paper, we introduce DICOD, a convolutional sparse coding algorithm which builds shift invariant representations for long signals. This algorithm is designed to run in a distributed setting, with local message passing, making it communication efficient. It is based on coordinate descent and uses locally greedy u…
Neural time-series data contain a wide variety of prototypical signal waveforms (atoms) that are of significant importance in clinical and cognitive research. One of the goals for analyzing such data is hence to extract such 'shift-invariant' atoms. Even though some success has been reported with existing algorithms, t…
New deep network derived from rate reduction principles, explaining features and efficiency.
problem Understanding and optimizing deep learning architectures.
method Gradient ascent scheme for rate reduction leading to multi-layer deep network.
result Explicitly constructed multi-layer network with precise optimization and interpretation.
Nonlinear kernel regression models are often used in statistics and machine learning because they are more accurate than linear models. Variable selection for kernel regression models is a challenge partly because, unlike the linear regression setting, there is no clear concept of an effect size for regression coeffici…
We consider the problem of improving the efficiency of randomized Fourier feature maps to accelerate training and testing speed of kernel methods on large datasets. These approximate feature maps arise as Monte Carlo approximations to integral representations of shift-invariant kernel functions (e.g., Gaussian kernel).…
A distributed algorithm learns patterns in large images and signals.
problem High-dimensional optimization in large images and signals.
method Distributed asynchronous algorithm with locally greedy coordinate descent.
result Patterns can be learned on large scales images from the Hubble Space Telescope.
This paper develops methods for obtaining distribution-free prediction regions for invariant representations.
problem Distributional shifts in machine learning models.
method Invariant risk minimization and weighted conformity scores.
result Proves the effectiveness of adaptive conformal intervals for uncertainty estimation.
New quantization methods improve accuracy of Random Fourier Features.
problem Improving accuracy of Random Fourier Features for machine learning.
method Sigma-Delta and distributed noise-shaping quantization methods for 1-bit and low bit-depth quantization.
result Quantized RFFs allow high accuracy approximation of underlying kernels with polynomial error decay.
New method learns compressed transforms with flexible displacement operators.
problem Efficiently representing and learning shift-invariant patterns in neural networks.
method Explicitly learns over displacement operators and low-rank components in LDR matrices.
result Reduces sample complexity and improves model accuracy with fewer parameters.
New framework learns sufficient invariant features robustly across distribution shifts.
problem Learning robust models under distribution shifts between training and test datasets.
method Sufficient Invariant Learning (SIL) framework and Adaptive Sharpness-aware Group Distributionally Robust Optimization (ASGDRO) algorithm.
result Empirical evaluations confirm ASGDRO's robustness against distribution shifts.
Boosted Control Functions improve prediction under distributional shifts.
problem Prediction under distributional shifts in the presence of hidden confounding.
method Boosted Control Function (BCF) and ControlTwicing algorithm.
result BCF allows for distribution generalization and invariance under nonlinear, non-identifiable structural functions.
Proposes DRIG for robust predictions using noise interventions.
problem Developing robust prediction models against distribution shifts.
method Distributional Robustness via Invariant Gradients (DRIG) exploiting general noise interventions.
result DRIG yields robust predictions among a data-dependent class of distribution shifts.
This paper tackles spatio-temporal information preservation in machine learning.
problem Conventional machine learning assumes orthogonal data attributes, disrupting spatio-temporal information.
method Shift-invariant k-means, convolutional dictionary learning, and spatio-temporal hypercomplex encoding schemes are proposed.
result Gabor feature extraction outperforms convolutional dictionary learning in spatio-temporal information preservation.
Novel framework improves graph learning for out-of-distribution generalization.
problem Graph out-of-distribution generalization challenges in neural networks.
method Invariant Graph Learning based on Information bottleneck theory (InfoIGL).
result Achieves state-of-the-art performance in graph classification tasks under OOD generalization.
Bird sounds possess distinctive spectral structure which may exhibit small shifts in spectrum depending on the bird species and environmental conditions. In this paper, we propose using convolutional recurrent neural networks on the task of automated bird audio detection in real-life environments. In the proposed metho…
Sparse coding is an unsupervised learning algorithm that learns a succinct high-level representation of the inputs given only unlabeled data; it represents each input as a sparse linear combination of a set of basis functions. Originally applied to modeling the human visual cortex, sparse coding has also been shown to …
Many modern tools in machine learning and signal processing, such as sparse dictionary learning, principal component analysis (PCA), non-negative matrix factorization (NMF), K-means clustering, etc., rely on the factorization of a matrix obtained by concatenating high-dimensional vectors from a training collection. W…
MIP framework improves urban flow prediction by adapting to distribution shifts.
problem Distribution shifts in urban flow data make prediction models unreliable.
method Memory-enhanced Invariant Prompt learning with learnable memory bank.
result MIP ensures robust predictions by focusing on invariant features.
We analyze the structure of the \emph{frequency space} Q(F) of a nonabelian free group F=F(a1,...,ak) consisting of all shift-invariant Borel probability measures on ∂F and construct a natural action of Out(F) on Q(F). In particular we prove that for any outer automorphism φ of F the \emph{conju…
Quantum kernels can be efficiently embedded into classical feature spaces.
problem Can all quantum kernels be efficiently embedded into classical feature spaces?
method Invoking computational universality and using techniques like random Fourier features, the authors show that certain classes of quantum kernels can be efficiently embedded.
result For shift-invariant and composition kernels, embedding quantum kernels are universal and efficient.
Proposes a new CSC model for handling unknown noise.
problem Existing CSC methods can only model Gaussian noise, which is restrictive.
method Uses Gaussian mixture model for unknown noise and EM algorithm for optimization.
result Effective modeling of complicated unknown noise with high-quality filters and representation.
Single model estimates uncertainty via biased data shifts.
problem Estimating uncertainties in deep neural networks.
method Trivial input transformation to approximate ensemble behavior.
result Single model uncertainty estimates are superior to current methods.
Tensor methods have emerged as a powerful paradigm for consistent learning of many latent variable models such as topic models, independent component analysis and dictionary learning. Model parameters are estimated via CP decomposition of the observed higher order input moments. However, in many domains, additional inv…
The paper studies Lipschitz bounds for integral kernels under differentiability assumptions.
problem Understanding the Lipschitz continuity of feature maps associated with integral kernels.
method Analyzes differentiability assumptions to derive explicit formulas for Lipschitz constants and conditions for non-Lipschitz continuity.
result Explicit formulas and conditions for Lipschitz continuity of feature maps associated with various kernels.
This article addresses the issue of representing electroencephalographic (EEG) signals in an efficient way. While classical approaches use a fixed Gabor dictionary to analyze EEG signals, this article proposes a data-driven method to obtain an adapted dictionary. To reach an efficient dictionary learning, appropriate s…
Bayesian Empirical Bayes extends EB to complex structures using probabilistic symmetry.
problem Improving simultaneous inference in complex settings like arrays and graphs.
method Generalized empirical Bayes approach based on probabilistic symmetry.
result BEB outperforms existing methods in denoising arrays and spatial data.
Deep neural network architectures designed for application domains other than sound, especially image recognition, may not optimally harness the time-frequency representation when adapted to the sound recognition problem. In this work, we explore the ConditionaL Neural Network (CLNN) and the Masked ConditionaL Neural N…
Median activation functions improve GNNs by capturing local graph signal behavior.
problem Lack of local nonlinear graph signal encoding in GNNs.
method Proposed median activation functions with support on graph neighborhoods.
result Median activation functions improve GNN capacity with minimal complexity increase.
New research shows input-gradients can be manipulated without changing model's core function, challenging their use for model interpretation.
problem Current methods for model interpretability using input-gradients are flawed due to their arbitrary manipulability.
method Investigated by reinterpreting logits as unnormalized log-densities, proposing novel approximations for score-matching.
result Improving alignment between implicit density model and data distribution enhances gradient structure and explanatory power.
Optimizes asset allocation for risk measures in a Lévy market.
problem Maximizing time-consistent mean-risk reward with general risk measures.
method Uses a generalized Lévy market model and Hamilton-Jacobi-Bellman equation.
result Deterministic optimal solution under certain conditions.
ForecastNet uses a time-variant deep feed-forward neural network for better multi-step-ahead time series forecasting.
problem Time-invariant architectures limit multi-step-ahead forecasting.
method ForecastNet employs a deep feed-forward architecture with time-variant parameters and interleaved outputs.
result ForecastNet outperforms other models on multi-step-ahead time series forecasting tasks.
Superior performance and ease of implementation have fostered the adoption of Convolutional Neural Networks (CNNs) for a wide array of inference and reconstruction tasks. CNNs implement three basic blocks: convolution, pooling and pointwise nonlinearity. Since the two first operations are well-defined only on regular-s…
Designs CNNs for better image reconstruction.
problem Image reconstruction from limited data.
method Parseval convolution operators and chaining of elementary modules.
result CNN-based algorithm yields better results than sparsity-based methods.
The paper explores high-dimensional learning in finance, proving key aspects and setting lower bounds.
problem Understanding when and how large, over-parameterized models achieve predictive success in finance.
method Theoretical foundations and empirical validation of two key aspects: standardization and information-theoretic lower bounds.
result Empirical validation shows that high-dimensional learning in finance often relies on lower-complexity artefacts rather than the intended mechanism.
Kernel methods represent one of the most powerful tools in machine learning to tackle problems expressed in terms of function values and derivatives due to their capability to represent and model complex relations. While these methods show good versatility, they are computationally intensive and have poor scalability t…