DeepSeekMoE improves language model efficiency with shared experts and normalized gating.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
PIMA autoencoders discover shared features in multimodal scientific data.
PCGS-TF uses a Transformer to adaptively control expert switching in non-stationary environments.
Algorithm detects changes online using expert tracking.
When dealing with time series with complex non-stationarities, low retrospective regret on individual realizations is a more appropriate goal than low prospective risk in expectation. Online learning algorithms provide powerful guarantees of this form, and have often been proposed for use with non-stationary processes …
We develop the setting of sequential prediction based on shifting experts and on a "smooth" version of the method of specialized experts. To aggregate experts predictions, we use the AdaHedge algorithm, which is a version of the Hedge algorithm with adaptive learning rate, and extend it by the meta-algorithm Fixed Shar…
STAR improves equivariant and invariant representation learning by routing projection heads.
The article is devoted to investigating the application of hedging strategies to online expert weight allocation under delayed feedback. As the main result, we develop the General Hedging algorithm based on the exponential reweighing of experts' losses. We build the artificial probabilistic framework and …
A new method for multi-expert learning-to-defer avoids optimization issues.
We consider the forecast aggregation problem in repeated settings, where the forecasts are done on a binary event. At each period multiple experts provide forecasts about an event. The goal of the aggregator is to aggregate those forecasts into a subjective accurate forecast. We assume that experts are Bayesian; namely…
ICIL learns policies invariant to multiple environments, improving generalization.
Unified scoring model improves efficiency and performance across multiple tasks.
Bayesian Experience Reuse improves learning from multiple experts.
Paper proposes a reinforcement learning method for trading using expert trajectories.
Deep learning models for semantic segmentation of images require large amounts of data. In the medical imaging domain, acquiring sufficient data is a significant challenge. Labeling medical image data requires expert knowledge. Collaboration between institutions could address this challenge, but sharing medical data to…
The step of expert taxa recognition currently slows down the response time of many bioassessments. Shifting to quicker and cheaper state-of-the-art machine learning approaches is still met with expert scepticism towards the ability and logic of machines. In our study, we investigate both the differences in accuracy and…
This paper considers the challenge of evaluating a set of classifiers, as done in shared task evaluations like the KDD Cup or NIST TREC, without expert labels. While expert labels provide the traditional cornerstone for evaluating statistical learners, limited or expensive access to experts represents a practical bottl…
DSelect-k improves MoE models for multi-task learning with better performance and smoother training.
Adversarial MoE learns category-specific models for product search.
Current imitation learning techniques are too restrictive because they require the agent and expert to share the same action space. However, oftentimes agents that act differently from the expert can solve the task just as good. For example, a person lifting a box can be imitated by a ceiling mounted robot or a desktop…
Online Apprenticeship Learning aims to match expert performance without access to cost functions.
Learning generative models that span multiple data modalities, such as vision and language, is often motivated by the desire to learn more useful, generalisable representations that faithfully capture common underlying factors between the modalities. In this work, we characterise successful learning of such models as t…
Automatic methods for generating state-of-the-art neural network architectures without human experts have generated significant attention recently. This is because of the potential to remove human experts from the design loop which can reduce costs and decrease time to model deployment. Neural architecture search (NAS)…
Data extracted from software repositories is used intensively in Software Engineering research, for example, to predict defects in source code. In our research in this area, with data from open source projects as well as an industrial partner, we noticed several shortcomings of conventional data mining approaches for c…
This work models the interconnection of company's investment managers' representations and the market attraction of its shares. The models that reflect the connection of the company's market effectiveness indices and parameters of its economic activity are created on the basis of the Mean-Variance Analysis and Regressi…
We consider the setting of sequential prediction of arbitrary sequences based on specialized experts. We first provide a review of the relevant literature and present two theoretical contributions: a general analysis of the specialist aggregation rule of Freund et al. (1997) and an adaptation of fixed-share rules of He…
In several natural language tasks, labeled sequences are available in separate domains (say, languages), but the goal is to label sequences with mixed domain (such as code-switched text). Or, we may have available models for labeling whole passages (say, with sentiments), which we would like to exploit toward better po…
Proposes a FoE prior for improving CNN performance in distribution shifts.
In this paper, we present a data science automation system called Prediction Factory. The system uses several key automation algorithms to enable data scientists to rapidly develop predictive models and share them with domain experts. To assess the system's impact, we implemented 3 different interfaces for creating pre…
Limited data access is a longstanding barrier to data-driven research and development in the networked systems community. In this work, we explore if and how generative adversarial networks (GANs) can be used to incentivize data sharing by enabling a generic framework for sharing synthetic datasets with minimal expert …
New model improves multimodal autoencoders by learning joint and conditional distributions.
Automates neural network design for diverse tasks.
Accessibility is a major challenge of machine learning (ML). Typical ML models are built by specialists and require specialized hardware/software as well as ML experience to validate. This makes it challenging for non-technical collaborators and endpoint users (e.g. physicians) to easily provide feedback on model devel…
Multi-expert L2D underfits more severely, requiring new methods.
Neural architecture search (NAS) is a promising research direction that has the potential to replace expert-designed networks with learned, task-specific architectures. In this work, in order to help ground the empirical results in this field, we propose new NAS baselines that build off the following observations: (i) …
In real-world machine learning applications, data subsets correspond to especially critical outcomes: vulnerable cyclist detections are safety-critical in an autonomous driving task, and "question" sentences might be important to a dialogue agent's language understanding for product purposes. While machine learning mod…
TENP prunes experts and neurons in Mixture-of-Experts models for efficient deployment.
A method to select important experts for Gaussian processes to balance computational efficiency and uncertainty quantification.
We consider the problem of contextual bandits with stochastic experts, which is a variation of the traditional stochastic contextual bandit with experts problem. In our problem setting, we assume access to a class of stochastic experts, where each expert is a conditional distribution over the arms given a context. We p…
Improved time series forecasting with expert loss integration.
FL improves insurance claims loss prediction without sharing data.
HS-MoE selects sparse experts using adaptive priors and data-adaptive gating.
NAMEx merges experts using Nash bargaining for improved performance.
Improved sales forecasting for new products using transfer learning.
Expert augmentation improves hybrid model generalization.
Unified model for prediction and deferral selects top-k entities efficiently.
New method calibrates Gaussian product experts for better predictions.
Sym-NCO leverages symmetricities to improve DRL-NCO performance.