This research formalizes inductive generalization and proposes a new learning paradigm called Inductive Learning.
problem Generalization from easy to hard tasks, especially out-of-domain generalization.
method Formalizes inductive generalization, introduces Inductive Learning, and outlines steps to adapt techniques for learning model successors.
result A new learning paradigm (Inductive Learning) that emphasizes induction and universal properties of learning and computation.
Study the inductive bias of neural networks using neural tangent kernels.
problem Understanding the generalization properties of over-parameterized neural networks.
method Analysis of the neural tangent kernel and its corresponding function space (RKHS).
result Stability properties of functions with finite norm, including stability to image deformations in convolutional networks.
Integrates inductive biases into VAEs using intermediary latent variables.
problem Ineffective mechanisms for incorporating inductive biases into VAEs.
method InteL-VAEs use an intermediary latent space to control encoding, with a parametric function to enforce desired properties.
result InteL-VAEs lead to better generative models and representations.
Survey of latent factor models for relational learning to improve their inductive abilities.
problem Understanding and improving latent factor models for multi-relational knowledge graphs.
method Experimental survey of state-of-the-art models, creating synthetic genealogies to assess strengths and weaknesses.
result Proposed new research directions to improve latent factor models.
Deep learning's success is puzzling from a statistical perspective.
problem Deep learning's success is puzzling from a statistical perspective.
method Physics-informed investigation of deep learning features and surprises.
result Neural scaling laws and their interplay with inductive biases.
New model discovers complex structures without explicit forms.
problem Understanding how humans and children discover organizing structures.
method Broad hypothesis space with preference for sparse connectivity.
result Model can learn complex structures and human property induction judgments.
Defines an equivariant index for proper actions, proving properties and applications.
problem Generalizing an index for non-compact groups and orbit spaces.
method Investigates properties and applications of the equivariant index, proving induction and relating to assembly maps.
result Shows the equivariant index has an induction property and is related to the Baum-Connes conjecture.
Quantum learning mimics classical learning by splitting training and testing.
problem How to apply classical learning theory to quantum systems.
method Proved a quantum de Finetti theorem for quantum channels, showing equivalence with classical inductive learning.
result Quantum and classical inductive learning coincide in the asymptotic setting.
A new geometric method for clustering SPD data improves upon Euclidean and Riemannian approaches.
problem Skewed interpretations of SPD data in Euclidean analysis and computational inefficiency of Riemannian methods.
method Proposes a geometric method based on the Thompson metric for unsupervised clustering of SPD data.
result Demonstrates improved clustering results using inductive midrange centroid computation.
Improved stability and generalization for blackbox learned optimizers.
problem Stability and generalization issues in blackbox learned optimizers.
method Investigation using dynamical systems, modifications to optimizer architecture and meta-training procedure.
result Improved stability and generalization of learned optimizers.
NLM combines neural networks and logic programming for complex reasoning.
problem Complex reasoning tasks involving logic and properties.
method Neural-symbolic architecture combining neural networks and logic programming.
result NLM achieves perfect generalization on various tasks.
This paper tackles data-efficient nonlinear control in Hamiltonian systems using symplectic geometry.
problem Data-efficient nonlinear control in Hamiltonian systems.
method Combines symplectic geometry, recurrence on energy level sets, and chain policies to solve target reachability problems.
result Data requirements depend on geometric and recurrence properties of the Hamiltonian, not the state dimension.
New methods to define self-inductance by regularizing divergent integrals.
problem Defining self-inductance for identical loops.
method Regularization of divergent integrals using Neumann/Weber formula.
result Established new methods to calculate self-inductance.
Neural networks struggle with periodic functions, a new activation fixes this.
problem Neural networks fail to learn simple periodic functions.
method Proposed a new activation function, x+sin2(x), to learn periodic functions. result The new activation function successfully learns and predicts periodic functions.
Paper explores how knowledge distillation transfers inductive biases between models.
problem Transferring inductive biases between models for tasks with limited data.
method Knowledge distillation applied to models with different inductive biases (LSTMs vs. Transformers, CNNs vs. MLPs).
result Effect of inductive biases is transferred through knowledge distillation, impacting both performance and solution characteristics.
Interpolated-MLPs control inductive bias for better performance in low-compute tasks.
problem Low-compute performance gap between MLPs and CNNs.
method Introduced Interpolated MLP (I-MLP) approach to control inductive bias incrementally.
result Continuous logarithmic relationship between inductive bias and performance in low-compute tasks.
New method quantifies inductive bias for machine learning tasks.
problem Quantifying the amount of inductive bias in machine learning models.
method Estimates inductive bias by modeling loss distribution of random hypotheses.
result Higher dimensional tasks require greater inductive bias.
The paper investigates how multi-label evaluation metrics can prune rule search space.
problem Challenges in inducing rules with multiple labels in multi-label classification.
method Examines anti-monotonicity and decomposability properties of multi-label evaluation metrics.
result Commonly used multi-label evaluation metrics exhibit anti-monotonicity, aiding rule search space pruning.
A new method for multi-task learning improves performance without weakening inductive bias.
problem Joint optimization of parameters for multiple tasks remains challenging.
method Maximum Roaming, a novel parameter partitioning method inspired by dropout.
result Maximum Roaming improves performance compared to recent multi-task learning formulations.
Gradient descent-based adversarial training converges to robust classifiers on linearly separable data.
problem Understanding the inductive bias of adversarial training for robustness.
method Gradient descent on binary classification tasks with linearly separable data, focusing on inductive bias and convergence rates.
result Gradient descent-based adversarial training converges to the maximum margin classifier at a faster rate than clean data training.
Recent unsupervised representation learning methods maximize mutual information, but their success depends on architecture and estimator inductive biases.
problem Estimating mutual information is hard, and MI maximization can lead to entangled representations.
method The paper argues that the success of mutual information maximization methods depends on the choice of feature extractor architectures and the parametrization of MI estimators.
result The paper provides empirical evidence that the success of mutual information maximization methods is not solely due to MI properties.
One-layer transformers can't solve induction heads task efficiently.
problem Solving the induction heads task efficiently with one-layer transformers.
method Communication complexity argument showing exponential size requirement.
result No one-layer transformer can solve the induction heads task efficiently.
OTI extends OTP for inductive semi-supervised learning.
problem Inductive semi-supervised learning for out-of-sample data.
method Optimal transport-based approach extended to inductive tasks.
result OTI outperforms state-of-the-art methods in experiments.
Natural gradient descent avoids the magic of model parametrization, leading to different optimization outcomes.
problem Understanding the impact of model parametrization on optimization and generalization in deep learning.
method Characterization of natural gradient flow in deep linear networks and nonlinear neural networks.
result Natural gradient descent fails to generalize in some cases, while gradient descent with the right architecture performs well.
Proposes a framework for deep learning on hypergraphs.
problem Lack of effective, unified framework for hypergraph learning.
method Jointly uses vertex and hyperedge embeddings for transductive and inductive learning.
result Achieves state-of-the-art performance on benchmark datasets.
Enhances graph neural networks with structural message-passing for better generalization.
problem Limited representation power and inability to learn basic graph topological properties.
method Proposes a framework that includes a one-hot encoding of nodes and parametrized message and update functions ensuring permutation equivariance.
result Achieves state-of-the-art results on molecular graph regression on the ZINC dataset.
Study generalizes matrix completion with side info in low noise settings.
problem Matrix completion with side information in low noise conditions.
method Inductive matrix completion with i.i.d. subgaussian noise, uniform sampling, and side information.
result Generalization bounds with noise scaling, convergence to zero, and logarithmic dependence on matrix size.
Strong inductive biases prevent harmless interpolation in overparameterized models.
problem Understanding the conditions under which overparameterized models can interpolate noise without overfitting.
method Theoretical analysis of high-dimensional kernel regression and deep neural networks, focusing on the role of inductive biases.
result The strength of an estimator's inductive bias determines whether interpolation is harmless or requires fitting noise for good generalization.
We consider the problem of learning a binary classifier from a training set of positive and unlabeled examples, both in the inductive and in the transductive setting. This problem, often referred to as \emph{PU learning}, differs from the standard supervised classification problem by the lack of negative examples in th…
Deep ResNets favor low bottleneck rank with proper hyperparameters.
problem Understanding the inductive bias of deep neural networks.
method Computed minimum-norm weights of a deep linear ResNet.
result Deep nonlinear ResNets have an inductive bias towards minimizing bottleneck rank.
The study refines algebraic domains with specific boundary conditions.
problem Understanding shapes of regions bounded by real algebraic curves.
method Inductive definition of regions, respecting characteristic finite sets.
result Generalized Poincar'e-Reeb Graphs for new types of regions.
Graph contrastive learning reveals unique inductive biases.
problem Understanding and optimizing graph contrastive learning methods.
method Systematic study of various GCL methods and their properties.
result GCL methods can work without positive or negative samples, and data augmentations have less impact.
The article consists of the Russian and English variants of Ph.D. Thesis in which the answers is given on the following questions: 1. how to construct the spinor formalism for n=6; 2. how to construct the spinor formalism for n=8; 3. how to prolong the Riemannian connection from the tangent bundle into the spinor one w…
This paper introduces hierarchical Gaussian process priors for neural networks to capture weight correlations and inductive biases.
problem Capturing weight correlations and inductive biases in neural networks.
method Hierarchical Gaussian process priors with unit embeddings and input-dependent kernels.
result Hierarchical Gaussian process priors provide competitive predictive performance and desirable uncertainty estimates.
We introduce the notion of large scale inductive dimension for asymptotic resemblance spaces. We prove that the large scale inductive dimension and the asymptotic dimensiongrad are equal in the class of r-convex metric spaces. This class contains the class of all geodesic metric spaces and all finitely generated groups…
Unsupervised MT struggles with morphologically rich languages.
problem Limitations of unsupervised machine translation on morphologically rich languages.
method Adversarial unsupervised alignment of word embedding spaces for bilingual dictionary induction.
result A simple trick exploiting weak supervision from identical words improves unsupervised bilingual dictionary induction performance.
New framework verifies reinforcement learning systems without altering neural networks.
problem Lack of assurance guarantees in reinforcement learning applications.
method Repurposes formal verification techniques for reinforcement learning, synthesizing simpler programs that preserve safety.
result Synthesized programs ensure safety of reinforcement learning systems without modifying neural networks.
SGD-trained deep nets often generalize well due to a strong inductive bias towards low-error, low-complexity functions.
problem Understanding why overparameterized deep nets generalize well despite fitting training data perfectly.
method Empirical investigation of PSGD(f∣S) and PB(f∣S) for various architectures and datasets. result The probability of SGD-converging on a function consistent with training data correlates well with the Bayesian posterior probability of expressing that function.
If p:Y→X is an unramified covering map between two compact oriented surfaces of genus at least two, then it is proved that the embedding map, corresponding to p, from the Teichmüller space T(X), for X, to T(Y) actually extends to an embedding between the Thurston compactification of the tw…
The paper challenges the unsupervised learning of disentangled representations, showing it's fundamentally impossible without biases.
problem The unsupervised learning of disentangled representations is fundamentally impossible without inductive biases.
method Theoretical analysis and a large-scale experimental study on seven different datasets.
result Well-disentangled models cannot be identified without supervision, and increased disentanglement does not decrease sample complexity.
Novel framework for Bayesian reinforcement learning infers value function distributions.
problem Bayesian reinforcement learning's challenges in inferring value function distributions.
method Inferential Induction framework for Bayesian reinforcement learning, developing Bayesian Backwards Induction algorithm.
result Proposed algorithm is competitive with state-of-the-art methods.
Algorithm finds minimal colorings of tree structures.
problem Finding minimal unbounded factor complexity colorings of trees.
method Induction algorithm using colored balls.
result Characterization of Sturmian colorings.
Noise affects the effectiveness of interpolating models, especially those with strong inductive biases.
problem The impact of noise on interpolating models with strong inductive biases.
method Analyzing linear and classification models with sparse ground truths, proving fast rates for interpolators.
result Strong inductive biases can lead to faster but noisier interpolators, contrary to intuition.
Study shows NCCP can replace CP for ACI in non-exchangeable data.
problem Ensuring reliable prediction under non-exchangeability.
method Demonstrates NCCP as a valid alternative to CP for ACI.
result NCCP offers computational advantages and comparable predictive efficiency.
GraIL predicts relations by reasoning over subgraphs, outperforming embeddings.
problem Relation prediction in knowledge graphs using latent representations is limited.
method Graph neural network with inductive bias to learn entity-independent relational semantics.
result GraIL outperforms existing rule-induction baselines in the inductive setting.
New approach relaxes inductive biases of physics-inspired NNs for better performance.
problem Challenges in applying physics-inspired NNs to real-world systems.
method Examined and relaxed inductive biases of Hamiltonian NNs, improving performance on non-conservative systems.
result Improved performance on practical, non-conservative systems by relaxing inductive biases.
Quantum kernels offer potential speed-ups but require encoding problem-specific knowledge.
problem Generalization difficulty in high-dimensional feature spaces.
method Analysis of spectral properties of quantum kernels and their RKHS.
result Quantum advantage is expected if RKHS is low-dimensional and contains hard-to-compute functions.
IGMC learns inductive matrix completion without side info.
problem Inductive matrix completion without side information.
method Graph Neural Network (GNN) trained on 1-hop subgraphs of the rating matrix.
result Achieves competitive performance with state-of-the-art transductive baselines.