New tensor formulation reveals gradient flow's bias in linear neural networks.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New dual formulation reduces generalization error for ERM-fDR.
A notion of implicit difference equation on a Lie groupoid is introduced and an algorithm for extracting the integrable part (backward or/and forward) is formulated. As an application, we prove that discrete Lagrangian dynamics on a Lie groupoid may be described in terms of Lagrangian implicit difference equations …
Semi-Implicit Variational Inference (SIVI) is improved with SIVI-SM using score matching.
Paper analyzes dynamics of nonholonomic systems with collisions using variational techniques.
Implicit generative models are difficult to train as no explicit density functions are defined. Generative adversarial nets (GANs) present a minimax framework to train such models, which however can suffer from mode collapse due to the nature of the JS-divergence. This paper presents a learning by teaching (LBT) approa…
Proposes efficient, modular method for implicit differentiation.
Solves inverse problem for Maxwell equations using vector fields.
Kernel semi-implicit variational inference improves variational inference without additional optimization.
A method to visualize multidimensional local subspaces using implicit differentiation.
A method for fair representation learning through bi-level optimization and implicit differentiation.
Market makers optimize trading with a new implicit scheme for complex inequalities.
New algorithm improves dataset distillation with 108% improvement on ImageNet.
The paper offers a framework to analyze machine learning problems using concentration of measure.
Kernel SIVI improves variational inference by avoiding lower-level optimization.
This work explores non-negative low-rank matrix factorization based on regularized Poisson models (PF or "Poisson factorization" for short) for recommender systems with implicit-feedback data. The properties of Poisson likelihood allow a shortcut for very fast computations over zero-valued inputs, and oftentimes result…
A new method for optimizing non-decomposable metrics with constraints.
The paper challenges the belief that more inner iterations at test time improve performance in implicit deep learning.
Although conservative Hamiltonian systems with constraints can be formulated in terms of Dirac structures, a more general framework is necessary to cover also dissipative systems such as gradient and metriplectic systems with constraints. We define Leibniz-Dirac structures which lead to a natural generalization of Dira…
In science and especially in economics, agent-based modeling has become a widely used modeling approach. These models are often formulated as a large system of difference equations. In this study, we discuss two aspects, numerical modeling and the probabilistic description for two agent-based computational economic mar…
Unified framework approximates gradient descent's implicit bias in high dimensions.
New method trains generative models without discriminators, improving stability and accuracy.
A core capability of intelligent systems is the ability to quickly learn new tasks by drawing on prior experience. Gradient (or optimization) based meta-learning has recently emerged as an effective approach for few-shot learning. In this formulation, meta-parameters are learned in the outer loop, while task-specific m…
Although many convex relaxations of clustering have been proposed in the past decade, current formulations remain restricted to spherical Gaussian or discriminative models and are susceptible to imbalanced clusters. To address these shortcomings, we propose a new class of convex relaxations that can be flexibly applied…
We propose and analyze a novel theoretical and algorithmic framework for structured prediction. While so far the term has referred to discrete output spaces, here we consider more general settings, such as manifolds or spaces of probability measures. We define structured prediction as a problem where the output space l…
Dual optimization connects ERM-fDR to normalization function.
We present a continuous formulation of machine learning, as a problem in the calculus of variations and differential-integral equations, in the spirit of classical numerical analysis. We demonstrate that conventional machine learning models and algorithms, such as the random feature model, the two-layer neural network …
HomoODE connects DEQs and Neural ODEs via homotopy continuation, improving accuracy and memory efficiency.
This paper presents a brief introduction to the key points of the Grey Machine Learning (GML) based on the kernels. The general formulation of the grey system models have been firstly summarized, and then the nonlinear extension of the grey models have been developed also with general formulations. The kernel implicit …
REALFIN benchmarks financial reasoning by removing implicit assumptions, revealing model weaknesses.
Adaptive sampler improves recommendation for implicit feedback data.
Gradient Descent with small random initialization solves rank-1 matrix completion efficiently.
Adam's bias shifts from full-batch to max-margin of different norms for separable data.
The mathematical problem concerning intrinsic storage optimisation is formulated and solved by means of variational analysis. The solution, though obtained in implicit form, still sheds light on many important features of the optimal exercise strategy. It is shown how the solution depends on different constraint types …
It remains a puzzle that why deep neural networks (DNNs), with more parameters than samples, often generalize well. An attempt of understanding this puzzle is to discover implicit biases underlying the training process of DNNs, such as the Frequency Principle (F-Principle), i.e., DNNs often fit target functions from lo…
We present a novel approximate inference method for diffusion processes, based on the Wasserstein gradient flow formulation of the diffusion. In this formulation, the time-dependent density of the diffusion is derived as the limit of implicit Euler steps that follow the gradients of a particular free energy functional.…
Images seen during test time are often not from the same distribution as images used for learning. This problem, known as domain shift, occurs when training classifiers from object-centric internet image databases and trying to apply them directly to scene understanding tasks. The consequence is often severe performanc…
The application of the Legendre transformation to a hyperregular Lagrangian system results in a Hamiltonian vector field generated by a Hamiltonian defined on the phase space of the mechanical system. The Legendre transformation in its usual interpretation can not be applied to homogeneous Lagrangians found in relativi…
We introduce the implicitly constrained least squares (ICLS) classifier, a novel semi-supervised version of the least squares classifier. This classifier minimizes the squared loss on the labeled data among the set of parameters implied by all possible labelings of the unlabeled data. Unlike other discriminative semi-s…
A number of techniques have been proposed to explain a machine learning model's prediction by attributing it to the corresponding input features. Popular among these are techniques that apply the Shapley value method from cooperative game theory. While existing papers focus on the axiomatic motivation of Shapley values…
New method tunes prior IP to data for flexible predictive distributions.
This paper shows how to train only the implicit layer of overparameterized implicit neural networks.
Paper studies the theoretical equivalence between implicit and explicit neural networks in high dimensions.
Gradient matching method estimates implicit regularization in complex deep learning systems.
This research uses DPPs to improve semi-parametric regression models.
Polynomial-time convex optimization for CNNs with ReLU activations.
Continuous semi-implicit models enable faster training and better performance in generative modeling.
We propose a \textbf{uni}fied \textbf{f}ramework for \textbf{i}mplicit \textbf{ge}nerative \textbf{m}odeling (UnifiGem) with theoretical guarantees by integrating approaches from optimal transport, numerical ODE, density-ratio (density-difference) estimation and deep neural networks. First, the problem of implicit gene…