We study the problem of structured output learning from a regression perspective. We first provide a general formulation of the kernel dependency estimation (KDE) problem using operator-valued kernels. We show that some of the existing formulations of this problem are special cases of our framework. We then propose a c…
A new kernel for multi-output Gaussian processes reduces undesirable scale effects.
problem Predicting multiple output variables simultaneously with Gaussian processes.
method Design a new kernel (MOCSM) using convolution in the spectral domain to model cross channel dependencies.
result MOCSM kernel reduces undesirable scale effects compared to the Multi-Output Spectral Mixture kernel.
The paper explores the identifiability and interpretability of Gaussian process models using different kernel structures.
problem Identifiability and interpretability issues in Gaussian process models.
method The paper examines both single-output and multi-output Gaussian process models using additive and multiplicative mixtures of Matérn kernels.
result The smoothness of a mixture of Matérn kernels is determined by the least smooth component, and none of the mixing weights or parameters are identifiable.
Deep neural networks for structured prediction using kernel-induced losses.
problem Structured prediction tasks for images and texts.
method Designing a novel family of deep neural architectures that predict in a finite-dimensional subspace derived from the kernel-induced loss.
result Gradient descent algorithms can be used for structured prediction with deep neural networks.
New method learns multiple tasks efficiently by relaxing positive semidefinite constraint.
problem Efficiently learn multiple tasks with shared structure.
method Relax the positive semidefinite constraint on output kernel, solve unconstrained dual problem.
result Efficiently solve multi-task learning problems without eigendecomposition.
New kernel models multi-output Gaussian processes accurately.
problem Challenges in modelling cross-covariances for multiple-output Gaussian processes.
method Replaced Gaussian components with block components of finite bandwidth in spectral mixture kernel.
result First multi-output generalization of spectral mixture kernel that can approximate any stationary multi-output kernel to arbitrary precision.
Bayesian kernel regression improves functional output prediction.
problem Functional output regression in supervised learning.
method Kernel methods, leveraging covariance structure within function values.
result Enhanced prediction accuracy and handling of high-dimensional nonlinearity.
Paper proposes a new method for learning kernels that depend on both inputs and outputs.
problem Common kernels are limited in their ability to handle complex tasks.
method Developed a spectral kernel learning framework that uses non-stationary kernels and learns from data.
result Derived a data-dependent generalization error bound and suggested regularization terms.
Sketching accelerates structured prediction methods for large datasets.
problem Scaling surrogate kernel methods for large datasets.
method Sketching approximations applied to input and output feature maps.
result Achieves close-to-optimal rates with reduced sketch size.
Interest in multioutput kernel methods is increasing, whether under the guise of multitask learning, multisensor networks or structured output data. From the Gaussian process perspective a multioutput Mercer kernel is a covariance function over correlated output functions. One way of constructing such kernels is based …
Random Fourier Features adapted for operator-valued kernels to scale multi-task and structured output learning.
problem Scaling operator-valued kernels for multi-task and structured output learning.
method Adapted Random Fourier Features for operator-valued kernels, using a generalization of Bochner's theorem.
result Uniform convergence of kernel approximation for operator-valued Random Fourier Features.
The goal of supervised feature selection is to find a subset of input features that are responsible for predicting output values. The least absolute shrinkage and selection operator (Lasso) allows computationally efficient feature selection based on linear dependency between input features and output values. In this pa…
Kernel methods are among the most popular techniques in machine learning. From a frequentist/discriminative perspective they play a central role in regularization theory as they provide a natural choice for the hypotheses space and the regularization functional through the notion of reproducing kernel Hilbert spaces. F…
Develops nonstationary MOGP kernels for better performance.
problem Limited applicability of existing MOGP kernels for nonstationary data.
method Harmonizable spectral mixture kernels for nonstationary MOGP.
result Automatic identification of nonstationary behavior in data.
New method learns output embeddings for structured prediction.
problem Structured prediction with output embeddings.
method Jointly learns output embedding and regression function.
result Structured predictor is a consistent estimator with smaller complexity.
We study the stability properties of nonlinear multi-task regression in reproducing Hilbert spaces with operator-valued kernels. Such kernels, a.k.a. multi-task kernels, are appropriate for learning prob- lems with nonscalar outputs like multi-task learning and structured out- put prediction. We show that multi-task ke…
Paper develops a duality approach for robust loss functions in infinite-dimensional RKHSs.
problem Robustness issues in infinite-dimensional RKHSs with operator-valued kernels.
method Develops a duality approach to solve OVK machines for various loss functions.
result Empirical improvements and theoretical stability analysis for robust structured data applications.
Improved sample efficiency with normalized RBF kernels in neural networks.
problem Learning more with less data in deep learning models.
method Two-phase method to train neural networks with normalized RBF kernels as output layer.
result Normalized RBF kernel networks achieve higher sample efficiency, compactness, and separability.
New kernels allow learning from non-separable data.
problem Learning from non-separable data.
method Introducing entangled kernels and a two-step algorithm.
result Efficient algorithm for learning entangled kernels.
Paper studies t-SNE convergence with generalized kernels.
problem Understanding convergence of t-SNE with generalized kernels.
method Concrete formulation of generalized kernels, proving convergence to an equilibrium distribution.
result t-SNE converges to an equilibrium distribution under certain conditions for generalized kernels.
The study addresses negative transfer in multi-output Gaussian processes by proposing latent structures.
problem Negative transfer in multi-output Gaussian processes leading to decreased performance.
method Defining negative transfer, deriving conditions for avoiding it, proposing latent structures.
result Latent structures can avoid negative transfer and scale to large datasets.
New recursive algorithm estimates conditional kernel mean embeddings in Hilbert space.
problem Estimating conditional distributions in RKHS for supervised learning.
method Recursive algorithm in L2 space for conditional kernel mean map. result Strong L2 consistency of recursive estimator proved. New spectral mixture kernels improve MOGP cross-covariance interpretation.
problem Limited parametric interpretation of cross-covariances in MOGPs.
method Complex-valued cross-spectral densities, Cramér's Theorem, phase shifts, delays.
result Improved expressive and interpretable multivariate covariance functions.
The paper introduces a method for learning nonparametric Volterra kernels using Gaussian processes.
problem Learning nonparametric nonlinear operators from data.
method NVKM model using Volterra series and Gaussian processes for unobserved and observed input functions.
result The NVKM model can perform both single and multiple output regression and system identification.
Extends OC-KSR for multi-task one-class classification.
problem Improving one-class classification performance with shared information.
method Linear and non-linear structure learning mechanisms for multi-task one-class classification.
result Improved performance on multiple one-class problems.
Paper shows robustness of kernel-based pairwise learning without strict assumptions.
problem Statistical robustness of kernel-based pairwise learning under minimal conditions.
method No assumptions on input and output spaces; derives influence function and robustness.
result Qualitative robustness of kernel-based estimator established.
Kernel regression predicts graph signals in noisy environments.
problem Predicting smooth graph signals in the presence of sparse noise.
method Kernel regression with ℓ1-norm and ℓ2-norm optimization using IRLS. result Efficacy demonstrated on real-world temperature data.
Study confirms learning rates for vector-valued spectral algorithms, proving consistency.
problem Theoretical confirmation of learning rates for vector-valued spectral algorithms.
method Rigorous analysis of learning rates for various vector-valued spectral algorithms, including kernel ridge regression and gradient descent.
result Upper and lower bounds on learning rates for vector-valued spectral algorithms, proving minimax optimality in various scenarios.
Paper converts deep networks to flat, equivalent kernel machines.
problem Capacity control and uniform convergence in deep learning.
method Push-forward transformation from deep networks to indefinite kernel machines.
result Flat network weights are Lp-norm regularized (0<p<1).
A new model for estimating multivariate densities efficiently.
problem Estimating complex multivariate densities efficiently.
method CDO model based on kernel mean embeddings and RKHS.
result Competitive performance with neural models and Gaussian processes.
A new kernel improves Volterra series model selection and prediction.
problem Hard identification of Volterra series from limited data.
method Proposes a novel regularization network using a multiplicative polynomial kernel.
result Better selection of influential monomials improves model prediction.
In this paper we introduce a novel method for linear system identification with quantized output data. We model the impulse response as a zero-mean Gaussian process whose covariance (kernel) is given by the recently proposed stable spline kernel, which encodes information on regularity and exponential stability. This s…
The study explains how neural networks align their kernels to target functions during training.
problem Understanding how neural networks align their kernels to target functions during training.
method Theoretical analysis of kernel evolution in toy models and deep networks.
result Kernel alignment naturally emerges during training to accelerate convergence and improve generalization.
We consider the problem of learning a vector-valued function f in an online learning setting. The function f is assumed to lie in a reproducing Hilbert space of operator-valued kernels. We describe two online algorithms for learning f while taking into account the output structure. A first contribution is an algorithm,…
New method for identifying systems with quantized data using Gaussian process and stable spline kernel.
problem Identifying linear systems with quantized output data.
method Bayesian framework with Markov Chain Monte Carlo and Gibbs sampler.
result Effectiveness of the proposed scheme compared to state-of-the-art methods.
Kernel methods summarize and integrate posterior similarity matrices from Bayesian clustering.
problem Summarizing and integrating posterior similarity matrices from Bayesian clustering.
method Positive semi-definite PSMs, kernel matrices, kernel methods, combining kernels.
result Kernel methods effectively summarize and integrate posterior similarity matrices.
The report analyzes infinite-dimensional output space regression.
problem Learning theory in vector-valued RKHS regression.
method Integral operator technique with spectral theory for non-compact operators.
result Results with minimal assumptions using Chebyshev's inequality.
Novel kernel-based PSI algorithm handles non-linearity and structured data.
problem Non-linearity and structured data in independence measures.
method Develops a PSI algorithm using HSIC, capable of handling non-linearity and structured data.
result Successfully identifies important features in real-world data.
This work challenges the Neural Tangent Kernel's role in overparameterized neural networks, especially with large width and depth.
problem The Neural Tangent Kernel's behavior in overparameterized neural networks with large width and depth is unclear.
method Experimental and theoretical analysis of ReLU networks with large width and depth.
result The aggregate norm of hidden neuron deviations does not vanish in infinitely-wide ReLU networks, indicating non-trivial behavior.
Recent developments in system identification have brought attention to regularized kernel-based methods, where, adopting the recently introduced stable spline kernel, prior information on the unknown process is enforced. This reduces the variance of the estimates and thus makes kernel-based methods particularly attract…
Study on Bayesian deep linear networks with multiple outputs and convolutional layers.
problem Characterize feature learning in finite-width Bayesian deep linear networks.
method Exact and analytical formulas for joint and posterior distributions, using large deviation theory.
result Quantitative description of feature learning in infinite-width regime.
A new model family of zero-inflated Gaussian processes improves prediction and interpretability of rare event data.
problem Poor performance of conventional machine learning on zero-inflated datasets.
method Sparse kernels and latent probit Gaussian processes to zero out kernel rows and columns.
result Improves prediction of zero-inflated data and interpretability of latent mixing models.
This paper uses random Fourier features to simplify latent force models and convolved Gaussian processes.
problem Expensive covariance matrix calculation in latent force models due to double integrals.
method Approximates double integrals using random Fourier features to obtain simpler analytical expressions.
result Simplified analytical expressions for covariance functions, leading to faster computation.
New method uses Kernel Flows to improve neural network training without changing structure.
problem Improving neural network training without altering structure or output classifier.
method Combines KFs with a classical output loss to aggregate a subset of KF losses.
result Reduced test errors, decreased generalization gaps, increased robustness to distribution shift.
Deep neural-kernel models combine neural networks and kernel machines for scalable large datasets.
problem Combining neural networks and kernel machines for efficient large-scale learning.
method Hybrid neural-kernel architecture using explicit feature mapping and pooling layers.
result The deep neural-kernel models are effective and scalable on benchmark datasets.
Distributed learning with least squares regularization achieves good performance without eigenfunction assumptions.
problem Efficiently learning from large datasets distributed across multiple machines.
method Divide-and-conquer approach, least squares regularization, RKHS, error bounds in expectation.
result The global estimator is a good approximation to the full data estimator, with sharp error bounds.
A novel dictionary-based approach for predicting functions.
problem Functional-output regression with non-orthogonal dictionaries.
method Projection learning (PL) with reproducing kernel Hilbert spaces (KPL).
result KPL offers a flexible and computationally efficient solution.
This work simplifies Gaussian process regression for multiple outputs.
problem Exponential computational complexity in Gaussian process regression.
method Approximating the covariance kernel using eigenvalues and functions.
result Significant reduction in training and regression complexity.