When constructing a classifier ensemble, diversity among the base classifiers is one of the important characteristics. Several studies have been made in the context of standard static data, in particular, when analyzing the relationship between a high ensemble predictive performance and the diversity of its components.…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study shows diverse data sources improve cryptocurrency forecasting models.
New framework shows diverse training data improves subgroup and overall performance.
D-CBRS manages memory for continual learning by accounting for intra-class diversity.
DiwE uses regional distribution changes to create diverse ensemble classifiers for concept drift.
Ensembles depend on diversity for improved performance. Many ensemble training methods, therefore, attempt to optimize for diversity, which they almost always define in terms of differences in training set predictions. In this paper, however, we demonstrate the diversity of predictions on the training set does not nece…
The paper improves experimental design by weighting diversity metrics with quality, leading to more diverse and effective discoveries.
SharpBalance improves deep ensemble performance by balancing sharpness and diversity.
MO-PaDGAN generates diverse, high-performance designs with multiple metrics.
Proposes Ada-Sit method for mortality prediction of rare diseases.
METASET selects diverse unit cells for efficient data-driven metamaterial design.
Diversity or complementarity of experts in ensemble pattern recognition and information processing systems is widely-observed by researchers to be crucial for achieving performance improvement upon fusion. Understanding this link between ensemble diversity and fusion performance is thus an important research question. …
Proposes Vendi Score for evaluating diversity in ML models.
Domain adaptation approaches seek to learn from a source domain and generalize it to an unseen target domain. At present, the state-of-the-art unsupervised domain adaptation approaches for subjective text classification problems leverage unlabeled target data along with labeled source data. In this paper, we propose a …
This work tackles semi-supervised federated learning by reducing model gradient diversity.
This paper improves ensemble learning for vision tasks by encouraging diversity in predictions.
Sampling methods that choose a subset of the data proportional to its diversity in the feature space are popular for data summarization. However, recent studies have noted the occurrence of bias (under- or over-representation of a certain gender or race) in such data summarization methods. In this paper we initiate a s…
Though data augmentation has become a standard component of deep neural network training, the underlying mechanism behind the effectiveness of these techniques remains poorly understood. In practice, augmentation policies are often chosen using heuristics of either distribution shift or augmentation diversity. Inspired…
Pantypes improve prototypical models by capturing diverse input distributions.
Exploration is a key problem in reinforcement learning, since agents can only learn from data they acquire in the environment. With that in mind, maintaining a population of agents is an attractive method, as it allows data be collected with a diverse set of behaviors. This behavioral diversity is often boosted via mul…
New method recovers diverse policies from expert data using state-action pair weighting.
Unified framework for portfolio optimization using multiple hypotheses.
Mixreg improves RL generalization by mixing diverse training environments.
Data Science is currently a popular field of science attracting expertise from very diverse backgrounds. Current learning practices need to acknowledge this and adapt to it. This paper summarises some experiences relating to such learning approaches from teaching a postgraduate Data Science module, and draws some learn…
We consider the problem of diversity enhancing clustering, i.e, developing clustering methods which produce clusters that favour diversity with respect to a set of protected attributes such as race, sex, age, etc. In the context of fair clustering, diversity plays a major role when fairness is understood as demographic…
New measures quantify diversity of latent representations using metric space magnitude.
SplitNN-driven Vertical Partitioning enables distributed learning from diverse data sources.
Database activity monitoring (DAM) systems are commonly used by organizations to protect the organizational data, knowledge and intellectual properties. In order to protect organizations database DAM systems have two main roles, monitoring (documenting activity) and alerting to anomalous activity. Due to high-velocity …
Product diversity of large US firms has declined steadily since 1997.
The paper clusters hypergraphs to find diverse and experienced groups based on past experiences.
Task-agnostic data valuation without validation requirements.
Deep generative models are proven to be a useful tool for automatic design synthesis and design space exploration. When applied in engineering design, existing generative models face three challenges: 1) generated designs lack diversity and do not cover all areas of the design space, 2) it is difficult to explicitly im…
The explosion of time series data in recent years has brought a flourish of new time series analysis methods, for forecasting, clustering, classification and other tasks. The evaluation of these new methods requires either collecting or simulating a diverse set of time series benchmarking data to enable reliable compar…
It is widely known in the machine learning community that class noise can be (and often is) detrimental to inducing a model of the data. Many current approaches use a single, often biased, measurement to determine if an instance is noisy. A biased measure may work well on certain data sets, but it can also be less effe…
Meta-learning improves few-shot land cover classification across diverse regions.
DivDis learns diverse hypotheses from underspecified data to improve robustness.
Identifies latent actions and dynamics from offline data with diverse demonstrators.
Proposes a Bayesian federated learning method for diverse tasks.
A new confidence measure improves self-training in biased data.
We address the problem of partial index tracking, replicating a benchmark index using a small number of assets. Accurate tracking with a sparse portfolio is extensively studied as a classic finance problem. However in practice, a tracking portfolio must also be diverse in order to minimise risk -- a requirement which h…
A new LLM-based method enhances diversity in oversampling for imbalanced classification.
This paper introduces new invariants for time series analysis.
New metric evaluates generative models across domains, diagnosing fidelity, diversity, and generalization.
Improved uncertainty estimation through diverse sampling in neural networks.
Ensembles, as a widely used and effective technique in the machine learning community, succeed within a key element -- "diversity." The relationship between diversity and generalization, unfortunately, is not entirely understood and remains an open research issue. To reveal the effect of diversity on the generalization…
Generative models have proven to be an outstanding tool for representing high-dimensional probability distributions and generating realistic-looking images. An essential characteristic of generative models is their ability to produce multi-modal outputs. However, while training, they are often susceptible to mode colla…
The highly detailed international trade data among all countries in the world during 1971-2000 shows that the kinds of export goods and the logarithmic GDP (gross domestic production) of a country has an S-shaped relationship. This indicates all countries can be divided into three stages accordingly. First, the poor co…
We propose a method to efficiently learn diverse strategies in reinforcement learning for query reformulation in the tasks of document retrieval and question answering. In the proposed framework an agent consists of multiple specialized sub-agents and a meta-agent that learns to aggregate the answers from sub-agents to…