Machine learning needs stronger mathematical foundations for scientific applications.
problem Lack of rigorous mathematical foundations for machine learning in scientific and engineering contexts.
method Further mathematical developments and incorporation of prior knowledge and inductive biases.
result Stronger mathematical rigor is essential for reliable and interpretable machine learning results in scientific fields.
Artin groups have a special structure that helps prove a complex mathematical conjecture.
problem Proving the Farrell-Jones isomorphism conjecture for Artin groups.
method Identifying an inductive structure in Artin groups and applying it to the conjecture.
result The Farrell-Jones isomorphism conjecture is proven for certain Artin groups.
Study reveals biases in gradient descent for GLNs, improving neural network performance.
problem Understanding and improving the inductive biases of deep neural networks.
method Derive infinite-time training limit of gated linear networks and generalize to other networks.
result Theoretical framework captures key inductive biases of ReLU networks.
New method uses reinforced regression for solving optimal stopping problems.
problem Solving optimal stopping problems in mathematical finance.
method Reinforced regression based on previously estimated continuation values.
result Illustrated by a numerical example from mathematical finance.
We use mathematical induction to prove that the horizontal composition in the class of coherently diagonal complexes is indeed a binary operation. That is to say, the embedding of two coherently diagonal complexes in an alternating planar diagram produces a coherently diagonal complex.
Novel approach trains LLMs for inductive reasoning using probabilistic programs.
problem Training LLMs for inductive reasoning with sparse, ambiguous data.
method Program-based Posterior Training (PPT) using probabilistic inference.
result Significant improvement in estimation accuracy and alignment with human judgments.
Abstractor enhances Transformers for relational reasoning, improving sample efficiency and performance.
problem Improving sample efficiency and performance in relational tasks.
method Introduces Abstractor module with relational cross-attention to enable explicit relational reasoning.
result Dramatic improvements in sample efficiency and performance on various relational tasks.
The paper proves a link concordance implies homotopy theorem in high codimensions.
problem Link concordance and homotopy equivalence in high-dimensional spaces.
method Analyzes link maps and spherical link maps, proving theorems through mathematical induction and other methods.
result Proves a link concordance implies homotopy theorem in codimension ≥3. A new decision tree induction method using MIP for faster optimization.
problem Optimizing decision tree split rules for better classification performance.
method Mixed-integer programming (MIP) for Gini reduction maximization, efficient search algorithm.
result bsnsing trees outperform other tree models in new case discrimination.
Mathematical theory explains neural network semantic development.
problem Understanding how neural networks acquire and organize abstract knowledge.
method Mathematical analysis of deep linear networks.
result Exact solutions reveal principles of semantic development.
This paper constructs cohomological Hall algebras for 3-Calabi-Yau categories.
problem Mathematical definition of the algebra of BPS states.
method Construction of cohomological Hall algebras for 3-Calabi-Yau categories.
result Construction of cohomological Hall algebras and proof of Joyce's conjecture.
Mathematical analysis improves SGMs, resolving memorization issues.
problem Improving performance and avoiding memorization in SGMs.
method Formulated SGMs using Wasserstein proximal operators and mean-field games.
result Improved SGM performance in terms of training samples and time.
New methods to define self-inductance by regularizing divergent integrals.
problem Defining self-inductance for identical loops.
method Regularization of divergent integrals using Neumann/Weber formula.
result Established new methods to calculate self-inductance.
Paper develops a framework to identify latent dynamics from high-dimensional data.
problem Identifying latent dynamics from high-dimensional time-series data.
method Combines physics inductive bias and learn-to-identify strategy.
result Meta-HyLaD framework effectively identifies hybrid latent dynamics.
Paper explores how knowledge distillation transfers inductive biases between models.
problem Transferring inductive biases between models for tasks with limited data.
method Knowledge distillation applied to models with different inductive biases (LSTMs vs. Transformers, CNNs vs. MLPs).
result Effect of inductive biases is transferred through knowledge distillation, impacting both performance and solution characteristics.
Interpolated-MLPs control inductive bias for better performance in low-compute tasks.
problem Low-compute performance gap between MLPs and CNNs.
method Introduced Interpolated MLP (I-MLP) approach to control inductive bias incrementally.
result Continuous logarithmic relationship between inductive bias and performance in low-compute tasks.
New method quantifies inductive bias for machine learning tasks.
problem Quantifying the amount of inductive bias in machine learning models.
method Estimates inductive bias by modeling loss distribution of random hypotheses.
result Higher dimensional tasks require greater inductive bias.
MFGs explain and enhance generative models, revealing new model types.
problem Understanding and improving generative models.
method Mean-field games (MFGs) as a framework to explain and enhance generative models.
result Established connections between MFGs and generative flows, diffusions, and gradient flows.
One-layer transformers can't solve induction heads task efficiently.
problem Solving the induction heads task efficiently with one-layer transformers.
method Communication complexity argument showing exponential size requirement.
result No one-layer transformer can solve the induction heads task efficiently.
OTI extends OTP for inductive semi-supervised learning.
problem Inductive semi-supervised learning for out-of-sample data.
method Optimal transport-based approach extended to inductive tasks.
result OTI outperforms state-of-the-art methods in experiments.
Strong inductive biases prevent harmless interpolation in overparameterized models.
problem Understanding the conditions under which overparameterized models can interpolate noise without overfitting.
method Theoretical analysis of high-dimensional kernel regression and deep neural networks, focusing on the role of inductive biases.
result The strength of an estimator's inductive bias determines whether interpolation is harmless or requires fitting noise for good generalization.
Deep ResNets favor low bottleneck rank with proper hyperparameters.
problem Understanding the inductive bias of deep neural networks.
method Computed minimum-norm weights of a deep linear ResNet.
result Deep nonlinear ResNets have an inductive bias towards minimizing bottleneck rank.
This research formalizes inductive generalization and proposes a new learning paradigm called Inductive Learning.
problem Generalization from easy to hard tasks, especially out-of-domain generalization.
method Formalizes inductive generalization, introduces Inductive Learning, and outlines steps to adapt techniques for learning model successors.
result A new learning paradigm (Inductive Learning) that emphasizes induction and universal properties of learning and computation.
We introduce the notion of large scale inductive dimension for asymptotic resemblance spaces. We prove that the large scale inductive dimension and the asymptotic dimensiongrad are equal in the class of r-convex metric spaces. This class contains the class of all geodesic metric spaces and all finitely generated groups…
Enhances neural models with simple functions to improve language modeling.
problem Neural models struggle with certain spatial, temporal, or quantitative relationships.
method Integrates simple functions into neural architecture to form a hierarchical NSLM.
result NSLMs significantly reduce perplexity in small-corpus language modeling.
Unsupervised MT struggles with morphologically rich languages.
problem Limitations of unsupervised machine translation on morphologically rich languages.
method Adversarial unsupervised alignment of word embedding spaces for bilingual dictionary induction.
result A simple trick exploiting weak supervision from identical words improves unsupervised bilingual dictionary induction performance.
The paper explains how language models acquire complex skills through scaling laws and statistical analysis.
problem Understanding how language models acquire new skills as their size and training data increase.
method Statistical framework and mathematical analysis of scaling laws.
result Language models can learn complex skills efficiently due to a strong inductive bias.
If p:Y→X is an unramified covering map between two compact oriented surfaces of genus at least two, then it is proved that the embedding map, corresponding to p, from the Teichmüller space T(X), for X, to T(Y) actually extends to an embedding between the Thurston compactification of the tw…
Novel framework for Bayesian reinforcement learning infers value function distributions.
problem Bayesian reinforcement learning's challenges in inferring value function distributions.
method Inferential Induction framework for Bayesian reinforcement learning, developing Bayesian Backwards Induction algorithm.
result Proposed algorithm is competitive with state-of-the-art methods.
Algorithm finds minimal colorings of tree structures.
problem Finding minimal unbounded factor complexity colorings of trees.
method Induction algorithm using colored balls.
result Characterization of Sturmian colorings.
Noise affects the effectiveness of interpolating models, especially those with strong inductive biases.
problem The impact of noise on interpolating models with strong inductive biases.
method Analyzing linear and classification models with sparse ground truths, proving fast rates for interpolators.
result Strong inductive biases can lead to faster but noisier interpolators, contrary to intuition.
GraIL predicts relations by reasoning over subgraphs, outperforming embeddings.
problem Relation prediction in knowledge graphs using latent representations is limited.
method Graph neural network with inductive bias to learn entity-independent relational semantics.
result GraIL outperforms existing rule-induction baselines in the inductive setting.
New approach relaxes inductive biases of physics-inspired NNs for better performance.
problem Challenges in applying physics-inspired NNs to real-world systems.
method Examined and relaxed inductive biases of Hamiltonian NNs, improving performance on non-conservative systems.
result Improved performance on practical, non-conservative systems by relaxing inductive biases.
IGMC learns inductive matrix completion without side info.
problem Inductive matrix completion without side information.
method Graph Neural Network (GNN) trained on 1-hop subgraphs of the rating matrix.
result Achieves competitive performance with state-of-the-art transductive baselines.
Study links neural network inductive bias, feature learning, and generalization on Boolean functions.
problem Understanding how neural networks learn and generalize on Boolean data.
method End-to-end analysis of depth-2 discrete fully connected networks and DNF formulas, using Monte Carlo learning.
result Predictable training dynamics and interpretable features emerge, linking inductive bias and generalization.
Extends program induction for probabilistic programming.
problem Automatic probabilistic program synthesis for diverse data types.
method Further steps to extend previous work on program induction.
result Generalization over various data types (text, image, video).
Extends positive mass theorem to arbitrary dimensions using a new inductive scheme.
problem Overcoming singularities in the Schoen-Yau proof for arbitrary dimensions.
method Inductive scheme combining shielding principle, conformal blow-up, and Cheeger-Naber bound.
result Proof of positive mass theorem in arbitrary dimensions.
We prove addition and subspace theorems for asymptotic large inductive dimension. We investigate a transfinite extension of this dimension and show that it is trivial.
I-BERT extends Transformer's self-attention to arbitrary input lengths.
problem Transformer models struggle with inductive generalization to unseen input lengths.
method Replaces positional encodings with a recurrent layer.
result I-BERT achieves state-of-the-art results on algorithmic tasks.
Transformers learn rich in-context dependencies efficiently.
problem Understanding how transformers learn long-range dependencies efficiently.
method Approximation and dynamics analysis of induction head mechanisms.
result Abrupt transition from lazy to rich mechanisms during training.
Novel approach uses inductive biases for semiconductor etching.
problem Significant violations of physics in etching process predictions.
method Introduced deep learning model with inductive biases.
result Fits measurements faster and follows physical behavior.
Integrates inductive biases into VAEs using intermediary latent variables.
problem Ineffective mechanisms for incorporating inductive biases into VAEs.
method InteL-VAEs use an intermediary latent space to control encoding, with a parametric function to enforce desired properties.
result InteL-VAEs lead to better generative models and representations.
This paper explains the theoretical inductive bias of Isolation Forest.
problem Lack of theoretical foundation explaining Isolation Forest's success.
method Formulated the growth process of iForest as a random walk, derived expected depth function using transition probabilities.
result Established a theoretical understanding of iForest's effectiveness and parameter adaptability.
Feed-forward nets fail to learn equality relations, but adding DR units helps.
problem Feed-forward neural networks struggle to learn equality relations reliably.
method Introduced differential rectifier (DR) units to create an inductive bias.
result DR units enable feed-forward nets to learn equality relations reliably.
RuleKit aids in creating interpretable models for various data types.
problem Creating interpretable models for different data types.
method Sequential covering induction algorithm for classification, regression, and survival problems.
result Facilitates verification of hypotheses about data dependencies.
R2N learns interpretable rules and literals from numerical features.
problem Lack of expressive vocabulary in rule-based decision models.
method Relational Rule Network (R2N) learns literals and rules end-to-end.
result Learned literals improve prediction accuracy and rule conciseness.
SSINNs learn Hamiltonian systems from data with interpretable, low-memory models.
problem Learning Hamiltonian dynamical systems from data efficiently and accurately.
method Combines fourth-order symplectic integration with sparse regression for a learned Hamiltonian.
result Outperforms state-of-the-art techniques in system prediction and energy conservation.
Machine learning refactors knowledge to improve learning efficiency.
problem Inductive program synthesis efficiency through knowledge restructuring.
method Introduces Knorf, a system that refactors knowledge bases using constraint optimization.
result Learning from refactored knowledge improves predictive accuracy fourfold and reduces learning time by half.