This research explores inductive biases for deep learning to improve AI's higher-level cognition.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Extends positive mass theorem to arbitrary dimensions using a new inductive scheme.
This research formalizes inductive generalization and proposes a new learning paradigm called Inductive Learning.
Extract symbolic models from deep learning with inductive biases.
In this paper, we investigate minimizing properties of the map from the Euclidean unit ball to its boundary , for the weighted energy functionals . We establish the following induction principle: if the map $\fra…
The paper explores how equivariant models' biases affect latent representations for better performance.
Deep learning's anomalous generalization explained by standard frameworks.
In the first part of this series, we defined an equivariant index without assuming the group acting or the orbit space of the action to be compact. This allowed us to generalise an index of deformed Dirac operators, defined for compact groups by Braverman. In this paper, we investigate properties and applications of th…
Looped transformers outperform standard transformers in complex reasoning tasks due to a specific loss landscape geometry.
We introduce an approach for imposing physically motivated inductive biases on graph networks to learn interpretable representations and improved zero-shot generalization. Our experiments show that our graph network models, which implement this inductive bias, can learn message representations equivalent to the true fo…
The article consists of the Russian and English variants of Ph.D. Thesis in which the answers is given on the following questions: 1. how to construct the spinor formalism for n=6; 2. how to construct the spinor formalism for n=8; 3. how to prolong the Riemannian connection from the tangent bundle into the spinor one w…
Conditional Random Fields (CRFs) are undirected graphical models, a special case of which correspond to conditionally-trained finite state machines. A key advantage of these models is their great flexibility to include a wide array of overlapping, multi-granularity, non-independent features of the input. In face of thi…
The ability of deep neural networks to generalize well in the overparameterized regime has become a subject of significant research interest. We show that overparameterized autoencoders exhibit memorization, a form of inductive bias that constrains the functions learned through the optimization process to concentrate a…
Clinical diagnostic decision making and population-based studies often rely on multi-modal data which is noisy and incomplete. Recently, several works proposed geometric deep learning approaches to solve disease classification, by modeling patients as nodes in a graph, along with graph signal processing of multi-modal …
This paper improves sample efficiency in noisy inductive matrix completion with side-information.
In this paper we formulate in general terms an approach to prove strong consistency of the Empirical Risk Minimisation inductive principle applied to the prototype or distance based clustering. This approach was motivated by the Divisive Information-Theoretic Feature Clustering model in probabilistic space with Kullbac…
SGD-trained deep nets often generalize well due to a strong inductive bias towards low-error, low-complexity functions.
Concerns about interpretability, computational resources, and principled inductive priors have motivated efforts to engineer sparse neural models for NLP tasks. If sparsity is important for NLP, might well-trained neural models naturally become roughly sparse? Using the Taxi-Euclidean norm to measure sparsity, we find …
Develops neural networks for learning physics of complex systems by enforcing thermodynamics principles.
Paper explores how knowledge distillation transfers inductive biases between models.
Interpolated-MLPs control inductive bias for better performance in low-compute tasks.
New method quantifies inductive bias for machine learning tasks.
One-layer transformers can't solve induction heads task efficiently.
IMA improves representation learning even when assumptions are violated.
OTI extends OTP for inductive semi-supervised learning.
Transformers with multiple layers learn to estimate bigram distributions, while single-layer models often get stuck in unigram local minima.
Physics-informed GCRL tackles sparse feedback learning with hybrid dynamics.
Strong inductive biases prevent harmless interpolation in overparameterized models.
Deep ResNets favor low bottleneck rank with proper hyperparameters.
The paper explores fundamental limits of learning non-hallucinating generative models.
We introduce several methods to define the self-inductance of a single loop as the regularization of divergent integrals which we obtain by applying Neumann (or Weber) formula for the mutual inductance of a pair of loops to the case when two loops are identical.
This paper applies machine learning techniques to student modeling. It presents a method for discovering high-level student behaviors from a very large set of low-level traces corresponding to problem-solving actions in a learning environment. Basic actions are encoded into sets of domain-dependent attribute-value patt…
We introduce the notion of large scale inductive dimension for asymptotic resemblance spaces. We prove that the large scale inductive dimension and the asymptotic dimensiongrad are equal in the class of r-convex metric spaces. This class contains the class of all geodesic metric spaces and all finitely generated groups…
If is an unramified covering map between two compact oriented surfaces of genus at least two, then it is proved that the embedding map, corresponding to , from the Teichmüller space , for , to actually extends to an embedding between the Thurston compactification of the tw…
The dominant paradigm for relation prediction in knowledge graphs involves learning and operating on latent representations (i.e., embeddings) of entities and relations. However, these embedding-based methods do not explicitly capture the compositional logical rules underlying the knowledge graph, and they are limited …
Novel framework for Bayesian reinforcement learning infers value function distributions.
Adversarial training is a principled approach for training robust neural networks. Despite of tremendous successes in practice, its theoretical properties still remain largely unexplored. In this paper, we provide new theoretical insights of gradient descent based adversarial training by studying its computational prop…
The problem of adaptive learning from evolving and possibly non-stationary data streams has attracted a lot of interest in machine learning in the recent past, and also stimulated research in related fields, such as computational intelligence and fuzzy systems. In particular, several rule-based methods for the incremen…
Noise affects the effectiveness of interpolating models, especially those with strong inductive biases.
Unsupervised machine translation---i.e., not assuming any cross-lingual supervision signal, whether a dictionary, translations, or comparable corpora---seems impossible, but nevertheless, Lample et al. (2018) recently proposed a fully unsupervised machine translation (MT) model. The model relies heavily on an adversari…
PACOH improves meta-learning with theoretical guarantees and practical efficiency.
Autoencoders are popular among neural-network-based matrix completion models due to their ability to retrieve potential latent factors from the partially observed matrices. Nevertheless, when training data is scarce their performance is significantly degraded due to overfitting. In this paper, we mit- igate overfitting…
We propose a neural network approach to price EU call options that significantly outperforms some existing pricing models and comes with guarantees that its predictions are economically reasonable. To achieve this, we introduce a class of gated neural networks that automatically learn to divide-and-conquer the problem …
New approach relaxes inductive biases of physics-inspired NNs for better performance.
Study links neural network inductive bias, feature learning, and generalization on Boolean functions.
Novel approach trains LLMs for inductive reasoning using probabilistic programs.
We prove addition and subspace theorems for asymptotic large inductive dimension. We investigate a transfinite extension of this dimension and show that it is trivial.
In this work, we develop a novel regularizer to improve the learning of long-range dependency of sequence data. Applied on language modelling, our regularizer expresses the inductive bias that sequence variables should have high mutual information even though the model might not see abundant observations for complex lo…