Develops a new model for deep structured prediction with non-linear output transformations.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper converts deep networks to flat, equivalent kernel machines.
Paper introduces non-linear process convolutions for multi-output Gaussian processes.
We consider the problem of training input-output recurrent neural networks (RNN) for sequence labeling tasks. We propose a novel spectral approach for learning the network parameters. It is based on decomposition of the cross-moment tensor between the output and a non-linear transformation of the input, based on score …
ALAMO is a computational methodology for leaning algebraic functions from data. Given a data set, the approach begins by building a low-complexity, linear model composed of explicit non-linear transformations of the independent variables. Linear combinations of these non-linear transformations allow a linear model to b…
Neural networks use their hidden layers to transform input data into linearly separable data clusters, with a linear or a perceptron type output layer making the final projection on the line perpendicular to the discriminating hyperplane. For complex data with multimodal distributions this transformation is difficult t…
Different neural networks learn similar mappings with different weights.
Linear Memory Network separates memory and function in RNNs.
Transformer attention layers solve single-location regression tasks.
Veronese webs are rich geometric structures with deep relationships to various domains of mathematics. The PDEs which determine the Veronese web are overdetermined if dim >3, but in the case dim =3 they reduce to a special flavor of a non-linear wave equation. The symmetries embedded in the definition of a Veronese web…
Paper learns dynamic generator models for video sequences.
This paper tackles efficient and scalable estimation of a complex model involving stochastic linear combinations of non-linear regressions.
Unified framework for non-linear attention using modern Hopfield networks.
Paper tackles conditional learning between different domains.
HFNO enhances interpretability of turbulent flows through parallel wavenumber bin processing.
The dynamic emulation of non-linear deterministic computer codes where the output is a time series, possibly multivariate, is examined. Such computer models simulate the evolution of some real-world phenomenon over time, for example models of the climate or the functioning of the human brain. The models we are interest…
FFCP improves FCP's speed without sacrificing accuracy.
NDMs enable non-linear transformations in diffusion models for better generative tasks.
Extends neural network verification to non-linear specifications.
Gaussian processes improve system identification models.
When approximating a black-box function, sampling with active learning focussing on regions with non-linear responses tends to improve accuracy. We present the FLOLA-Voronoi method introduced previously for deterministic responses, and theoretically derive the impact of output uncertainty. The algorithm automatically p…
A new algorithm estimates output ranges for deep neural networks efficiently.
We show how the tangent functor extends from ordinary smooth maps to "microformal morphisms" (also called "thick morphisms") of supermanifolds. Microformal morphisms generalize ordinary maps and correspond to formal canonical relations between the cotangent bundles specified by generating functions depending on positio…
Variational auto-encoder frameworks have demonstrated success in reducing complex nonlinear dynamics in molecular simulation to a single non-linear embedding. In this work, we illustrate how this non-linear latent embedding can be used as a collective variable for enhanced sampling, and present a simple modification th…
Transformer improves sequence generation with insertion and deletion phases.
The price of financial assets are, since Bachelier, considered to be described by a (discrete or continuous) time sequence of random variables, i.e a stochastic process. Sharp scaling exponents or unifractal behavior of such processes has been reported in several works. In this letter we investigate the question of sca…
Paper uses black-box inference to estimate non-linear latent force models.
Recently, we proposed to transform the outputs of each hidden neuron in a multi-layer perceptron network to have zero output and zero slope on average, and use separate shortcut connections to model the linear dependencies instead. We continue the work by firstly introducing a third transformation to normalize the scal…
Variational inference is a powerful tool for approximate inference, and it has been recently applied for representation learning with deep generative models. We develop the variational Gaussian process (VGP), a Bayesian nonparametric variational family, which adapts its shape to match complex posterior distributions. T…
Deep learning improves sensor performance optimization.
We consider the boundary rigidity problem for asymptotically hyperbolic manifolds. We show injectivity of the X-ray transform in several cases and consider the non-linear inverse problem which consists of recovering a metric from boundary measurements for the geodesic flow.
New algorithm predicts multiple types of outputs with dependencies.
Regularizes GAMs to improve interpretability by reducing concurvity.
We propose a novel kernel based post selection inference (PSI) algorithm, which can not only handle non-linearity in data but also structured output such as multi-dimensional and multi-label outputs. Specifically, we develop a PSI algorithm for independence measures, and propose the Hilbert-Schmidt Independence Criteri…
This paper recovers input data from transformer models using attention weights.
Extends OC-KSR for multi-task one-class classification.
In this paper, we generalize the famous Hasimoto's transformation by showing that the dynamics of a closed unidimensional vortex filament embedded in a three-dimensional manifold of constant curvature gives rise under Hasimoto's transformation to the non-linear Schrodinger equation. We also give a natural interpretatio…
Traditional linear methods for forecasting multivariate time series are not able to satisfactorily model the non-linear dependencies that may exist in non-Gaussian series. We build on the theory of learning vector-valued functions in the reproducing kernel Hilbert space and develop a method for learning prediction func…
We introduce the Randomized Dependence Coefficient (RDC), a measure of non-linear dependence between random variables of arbitrary dimension based on the Hirschfeld-Gebelein-Rényi Maximum Correlation Coefficient. RDC is defined in terms of correlation of random non-linear copula projections; it is invariant with respec…
CMTRF improves recommendation accuracy by transforming rating scales.
In recent work on both generative and discriminative score to log-likelihood-ratio calibration, it was shown that linear transforms give good accuracy only for a limited range of operating points. Moreover, these methods required tailoring of the calibration training objective functions in order to target the desired r…
Transformer with denoising diffusion improves probabilistic density estimation.
Transformers use a unique Hessian structure that differs from classical networks, affecting optimization.
This paper explains a mechanism called phase collapse that improves image classification accuracy.
The goal of supervised feature selection is to find a subset of input features that are responsible for predicting output values. The least absolute shrinkage and selection operator (Lasso) allows computationally efficient feature selection based on linear dependency between input features and output values. In this pa…
Solves a PDE for Landsberg surfaces using new Finsler surface insights.
A new method quantizes output space for multi-target regression.
Paper investigates Lipschitz constants of self-attention modules in neural networks.