Transformers can generalize to a large task family with only a few demonstrations.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Transformers can learn new tasks from diverse pretraining data but struggle with out-of-domain tasks.
NeuTSFlow models continuous functions behind time series forecasting.
We consider a problem of learning kernels for use in SVM classification in the multi-task and lifelong scenarios and provide generalization bounds on the error of a large margin classifier. Our results show that, under mild conditions on the family of kernels used for learning, solving several related tasks simultaneou…
NatPN provides fast, accurate uncertainty estimation for exponential family distributions.
Representing networks in a low dimensional latent space is a crucial task with many interesting applications in graph learning problems, such as link prediction and node classification. A widely applied network representation learning paradigm is based on the combination of random walks for sampling context nodes and t…
This work improves AI's ability to solve physical tasks by optimizing world models in abstracted spaces.
New tractable density models from squaring neural networks.
Adaptive methods learn from multiple datasets, leveraging similarities and robust to outliers.
AdMRL improves meta-reinforcement learning by minimizing worst-case sub-optimality gap.
We study the problem of recovering the subspace spanned by the first principal components of -dimensional data under the streaming setting, with a memory bound of . Two families of algorithms are known for this problem. The first family is based on the framework of stochastic gradient descent. Nevertheles…
Meta-learn sparse Gaussian process inference for faster predictions.
We propose a robust estimator to improve maximum likelihood in probabilistic models.
Meta-Dynamic models learn shared neural dynamics across tasks.
TOPPO improves PPO for MTRL by balancing critic gradients, outperforming SAC.
New variational flows improve Monte Carlo and normalization tasks.
Study compares adaptive vs fixed query learning methods.
While Reinforcement Learning (RL) approaches lead to significant achievements in a variety of areas in recent history, natural language tasks remained mostly unaffected, due to the compositional and combinatorial nature that makes them notoriously hard to optimize. With the emerging field of Text-Based Games (TBGs), re…
Training-free source selection for LLM families with shared vocabularies
We focus on two supervised visual reasoning tasks whose labels encode a semantic relational rule between two or more objects in an image: the MNIST Parity task and the colorized Pentomino task. The objects in the images undergo random translation, scaling, rotation and coloring transformations. Thus these tasks involve…
New framework assesses neural sensitivity to small perturbations.
We propose the Neural Logic Machine (NLM), a neural-symbolic architecture for both inductive learning and logic reasoning. NLMs exploit the power of both neural networks---as function approximators, and logic programming---as a symbolic processor for objects with properties, relations, logic connectives, and quantifier…
Sloth predicts LLM performance using latent skills across families.
This paper compares log-likelihood and BLEU scores for sequence generation tasks.
Boosted decision trees enjoy popularity in a variety of applications; however, for large-scale datasets, the cost of training a decision tree in each round can be prohibitively expensive. Inspired by ideas from the multi-arm bandit literature, we develop a highly efficient algorithm for computing exact greedy-optimal d…
Non-coding RNA (ncRNA) are RNA sequences which don't code for a gene but instead carry important biological functions. The task of ncRNA classification consists in classifying a given ncRNA sequence into its family. While it has been shown that the graph structure of an ncRNA sequence folding is of great importance for…
Sparse Gaussian processes with compact kernels for faster inference.
Neural networks improve cancer risk prediction from family history data.
Semi-implicit variational inference (SIVI) is introduced to expand the commonly used analytic variational distribution family, by mixing the variational parameter with a flexible distribution. This mixing distribution can assume any density function, explicit or not, as long as independent random samples can be generat…
We consider the task of fitting a regression model involving interactions among a potentially large set of covariates, in which we wish to enforce strong heredity. We propose FAMILY, a very general framework for this task. Our proposal is a generalization of several existing methods, such as VANISH [Radchenko and James…
Multi-task learning (MTL) aims to improve generalization performance by learning multiple related tasks simultaneously. While sometimes the underlying task relationship structure is known, often the structure needs to be estimated from data at hand. In this paper, we present a novel family of models for MTL, applicable…
We study the problem of designing models for machine learning tasks defined on \emph{sets}. In contrast to traditional approach of operating on fixed dimensional vectors, we consider objective functions defined on sets that are invariant to permutations. Such problems are widespread, ranging from estimation of populati…
New method for fast inference in diffusion models.
New method infers unknown parameters in quantum sensing with high probability.
Exponential families and mixture families are parametric probability models that can be geometrically studied as smooth statistical manifolds with respect to any statistical divergence like the Kullback-Leibler (KL) divergence or the Hellinger divergence. When equipping a statistical manifold with the KL divergence, th…
Single neural network learns multiple tasks from combined data.
We investigate the non-identifiability issues associated with bidirectional adversarial training for joint distribution matching. Within a framework of conditional entropy, we propose both adversarial and non-adversarial approaches to learn desirable matched joint distributions for unsupervised and supervised tasks. We…
Multi-task learning aims to learn multiple tasks jointly by exploiting their relatedness to improve the generalization performance for each task. Traditionally, to perform multi-task learning, one needs to centralize data from all the tasks to a single machine. However, in many real-world applications, data of differen…
A new framework integrates classification and regression tasks in multi-task learning.
With a view to bridging the gap between deep learning and symbolic AI, we present a novel end-to-end neural network architecture that learns to form propositional representations with an explicitly relational structure from raw pixel data. In order to evaluate and analyse the architecture, we introduce a family of simp…
The paper extends optimal transport for linear separability of sheared distributions in supervised learning.
The paper argues that uncertainty quantification in ML is application-specific and proposes a flexible family of measures.
DORO improves DRO's performance and stability in tasks with subpopulation shift.
Deep kernel learning provides an elegant and principled framework for combining the structural properties of deep learning algorithms with the flexibility of kernel methods. By means of a deep neural network, we learn a parametrized kernel operator that can be combined with a differentiable kernel algorithm during infe…
Informative and discriminative feature descriptors play a fundamental role in deformable shape analysis. For example, they have been successfully employed in correspondence, registration, and retrieval tasks. In the recent years, significant attention has been devoted to descriptors obtained from the spectral decomposi…
-algebra consists of expressions constructed with four kinds operations, the minimum, maximum, difference and additively homogeneous generalized means. Five families of -classifiers are investigated on binary classification tasks between English phonemes. It is shown that the classifiers are able to reflect well…
Variational inference for latent variable models is prevalent in various machine learning problems, typically solved by maximizing the Evidence Lower Bound (ELBO) of the true data likelihood with respect to a variational distribution. However, freely enriching the family of variational distribution is challenging since…
Proposes semi-supervised feature ranking for handling high-dimensional, unlabeled data.