We introduce a conceptually simple and scalable framework for continual learning domains where tasks are learned sequentially. Our method is constant in the number of parameters and is designed to preserve performance on previously encountered tasks while accelerating learning progress on subsequent problems. This is a…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
PS-KD distills a model's own knowledge to soften hard targets during training.
Extracting a curriculum from a teacher network improves distillation efficiency.
The concept of progress has characterized human society from millennia. However, this concept is elusive and too often given for certain. The goal of this paper is to suggest a general definition of human progress that satisfies, whenever possible the conditions of independence, generality, epistemological applicabilit…
This paper improves disentanglement in VAEs by progressively learning hierarchical representations.
LHM integrates expert ODEs with neural ODEs for disease progression prediction.
New equations reveal how cylinder power in progressive lenses depends on geodesic curvature.
PAS method improves UDA by progressively refining subspaces for reliable pseudo-labels.
New embedding method in function spaces improves expressiveness.
KT models improved slightly with synthetic student data.
AlphaZero reveals new chess concepts learnable by top experts.
Survey explores how transfer learning improves deep reinforcement learning.
Recent progress in artificial intelligence (AI) has renewed interest in building systems that learn and think like people. Many advances have come from using deep neural networks trained end-to-end in tasks such as object recognition, video games, and board games, achieving performance that equals or even beats humans …
We define general linguistic intelligence as the ability to reuse previously acquired knowledge about a language's lexicon, syntax, semantics, and pragmatic conventions to adapt to new tasks quickly. Using this definition, we analyze state-of-the-art natural language understanding models and conduct an extensive empiri…
Bernstein processes are Brownian diffusions that appear in Euclidean Quantum Mechanics. Knowledge of the symmetries of the Hamilton-Jacobi-Bellman equation associated with these processes allows one to obtain relations between stochastic processes (Lescot-Zambrini, Progress in Probability, vols 58 and 59). More recentl…
HOPE uses Hilbert space to deconstruct deep network representations.
PSI-KT improves KT accuracy and interpretability in learning materials.
Knowledge Matters: Importance of Prior Information for Optimization [7], by Gulcehre et. al., sought to establish the limits of current black-box, deep learning techniques by posing problems which are difficult to learn without engineering knowledge into the model or training procedure. In our work, we completely solve…
This paper introduces self-paced task selection to multitask learning, where instances from more closely related tasks are selected in a progression of easier-to-harder tasks, to emulate an effective human education strategy, but applied to multitask machine learning. We develop the mathematical foundation for the appr…
OpenHAIV integrates OOD detection and incremental learning for open-world models.
Machine reading using differentiable reasoning models has recently shown remarkable progress. In this context, End-to-End trainable Memory Networks, MemN2N, have demonstrated promising performance on simple natural language based reasoning tasks such as factual reasoning and basic deduction. However, other tasks, namel…
Purpose: Machine learning is broadly used for clinical data analysis. Before training a model, a machine learning algorithm must be selected. Also, the values of one or more model parameters termed hyper-parameters must be set. Selecting algorithms and hyper-parameter values requires advanced machine learning knowledge…
Despite the advancement of supervised image recognition algorithms, their dependence on the availability of labeled data and the rapid expansion of image categories raise the significant challenge of zero-shot learning. Zero-shot learning (ZSL) aims to transfer knowledge from labeled classes into unlabeled classes to r…
DBULL learns new clusters without forgetting past knowledge in streaming unlabelled data.
Framework predicts patient risk progression over time.
Proposes a method to compress deep networks using adversarial training and attention mechanisms.
ProSelfLC improves robustness of deep neural networks by automatically deciding trust in predictions.
AI uses KGs to assess economic impact of selective lockdowns on Italian companies.
GENE tackles sparse reward in RL by generating states to explore and exploit.
Learning from many real-world datasets is limited by a problem called the class imbalance problem. A dataset is imbalanced when one class (the majority class) has significantly more samples than the other class (the minority class). Such datasets cause typical machine learning algorithms to perform poorly on the classi…
QuantAgent learns trading signals through self-improvement.
While recent progress has spawned very powerful machine learning systems, those agents remain extremely specialized and fail to transfer the knowledge they gain to similar yet unseen tasks. In this paper, we study a simple reinforcement learning problem and focus on learning policies that encode the proper invariances …
Deep Neural Networks have achieved huge success at a wide spectrum of applications from language modeling, computer vision to speech recognition. However, nowadays, good performance alone is not sufficient to satisfy the needs of practical deployment where interpretability is demanded for cases involving ethics and mis…
We compare the acquisition of knowledge in humans and machines. Research from the field of developmental psychology indicates, that human-employed hypothesis are initially guided by simple rules, before evolving into more complex theories. This observation is shared across many tasks and domains. We investigate whether…
Meta-learning shows negative transfer between tasks, which MetaCRL addresses.
This paper proposes a method to compress and adapt CNNs for real-world applications.
Galactica learns from scientific literature to help researchers.
Recent progress in deep learning is revolutionizing the healthcare domain including providing solutions to medication recommendations, especially recommending medication combination for patients with complex health conditions. Existing approaches either do not customize based on patient health history, or ignore existi…
A Convolutional Neural Network was used to predict kidney function in patients with chronic kidney disease from high-resolution digital pathology scans of their kidney biopsies. Kidney biopsies were taken from participants of the NEPTUNE study, a longitudinal cohort study whose goal is to set up infrastructure for obse…
AxCell automates results extraction from machine learning papers.
Survey of LLMs in finance tasks, highlighting progress and challenges.
The abstract warns against flawed empirical research in machine learning.
Proposes a method to derive knowledge graphs from EHR data.
AutoKE automates embedding physical knowledge into neural networks for complex engineering problems.
The paper characterizes the efficiency of transferring knowledge from a teacher to a student classifier over finite domains.
SpanishTinyRoBERTa distills large Spanish models into efficient question-answering models.
In this article, we investigate when the set of primitive geodesic lengths on a Riemannian manifold have arbitrarily long arithmetic progressions. We prove that in the space of negatively curved metrics, a metric having such arithmetic progressions is quite rare. We introduce almost arithmetic progressions, a coarsific…
Study evaluates large language models' ability to understand probabilistic real-world distributions.