New method learns diverse solutions in reinforcement learning without gradient bias.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
CPRA efficiently finds diverse solutions in CO problems using UL and parallelization.
EDU method finds diverse optimal solutions for expensive simulators.
Proposes method to discover diverse near-optimal policies in reinforcement learning.
Proposes MOGFNs for generating diverse Pareto optimal solutions in multi-objective optimization.
To cope with the high level of ambiguity faced in domains such as Computer Vision or Natural Language processing, robust prediction methods often search for a diverse set of high-quality candidate solutions or proposals. In structured prediction problems, this becomes a daunting task, as the solution space (image label…
Entropy minimization has been widely used in unsupervised domain adaptation (UDA). However, existing works reveal that entropy minimization only may result into collapsed trivial solutions. In this paper, we propose to avoid trivial solutions by further introducing diversity maximization. In order to achieve the possib…
A new method for diverse Pareto solutions in multi-objective learning.
RL enhances LLM planning but introduces spurious solutions and diversity collapse.
Quality-Diversity algorithms explore multiple high-performing solutions in a search space.
BOP-Elites uses Bayesian Optimisation for QD search, improving efficiency and insight.
In an ordinary feature selection procedure, a set of important features is obtained by solving an optimization problem such as the Lasso regression problem, and we expect that the obtained features explain the data well. In this study, instead of the single optimal solution, we consider finding a set of diverse yet nea…
Transformers learn to generalize out-of-distribution with diverse pretraining tasks.
UCPO improves diversity in reinforcement learning models, maintaining high accuracy.
Paper tackles diversity in Airbnb search results.
Multi-view clustering aims at integrating complementary information from multiple heterogeneous views to improve clustering results. Existing multi-view clustering solutions can only output a single clustering of the data. Due to their multiplicity, multi-view data, can have different groupings that are reasonable and …
We focus on the challenge of finding a diverse collection of quality solutions on complex continuous domains. While quality diver-sity (QD) algorithms like Novelty Search with Local Competition (NSLC) and MAP-Elites are designed to generate a diverse range of solutions, these algorithms require a large number of evalua…
The need for diversification of recommendation lists manifests in a number of recommender systems use cases. However, an increase in diversity may undermine the utility of the recommendations, as relevant items in the list may be replaced by more diverse ones. In this work we propose a novel method for maximizing the u…
Developing a range-aware Bayesian optimization framework for discovering diverse designs within target property windows.
Proposes a meta-learning method for robust portfolio optimization.
DMNL bandits optimize assortment choices balancing relevance and diversity.
We address the problem of partial index tracking, replicating a benchmark index using a small number of assets. Accurate tracking with a sparse portfolio is extensively studied as a classic finance problem. However in practice, a tracking portfolio must also be diverse in order to minimise risk -- a requirement which h…
New method improves neural architecture search by optimizing for both performance and diversity.
Hierarchical reinforcement learning (HRL) has recently shown promising advances on speeding up learning, improving the exploration, and discovering intertask transferable skills. Most recent works focus on HRL with two levels, i.e., a master policy manipulates subpolicies, which in turn manipulate primitive actions. Ho…
The MAP-Elites algorithm produces a set of high-performing solutions that vary according to features defined by the user. This technique has the potential to be a powerful tool for design space exploration, but is limited by the need for numerous evaluations. The Surrogate-Assisted Illumination algorithm (SAIL), introd…
Standard reinforcement learning methods aim to master one way of solving a task whereas there may exist multiple near-optimal policies. Being able to identify this collection of near-optimal policies can allow a domain expert to efficiently explore the space of reasonable solutions. Unfortunately, existing approaches t…
We address the challenge of effective exploration while maintaining good performance in policy gradient methods. As a solution, we propose diverse exploration (DE) via conjugate policies. DE learns and deploys a set of conjugate policies which can be conveniently generated as a byproduct of conjugate gradient descent. …
This study presents an ANWSER model (asset network systemic risk model) to quantify the risk of financial contagion which manifests itself in a financial crisis. The transmission of financial distress is governed by a heterogeneous bank credit network and an investment portfolio of banks. Bankruptcy reproductive ratio …
ePF improves PF for ITS by balancing exploration and exploitation, outperforming baselines.
We propose to tackle the mode collapse problem in generative adversarial networks (GANs) by using multiple discriminators and assigning a different portion of each minibatch, called microbatch, to each discriminator. We gradually change each discriminator's task from distinguishing between real and fake samples to disc…
Automatic program repair holds the potential of dramatically improving the productivity of programmers during the software development process and correctness of software in general. Recent advances in machine learning, deep learning, and NLP have rekindled the hope to eventually fully automate the process of repairing…
A new method improves recommendation accuracy by learning from multiple networks and time-dependent user preferences.
Database activity monitoring (DAM) systems are commonly used by organizations to protect the organizational data, knowledge and intellectual properties. In order to protect organizations database DAM systems have two main roles, monitoring (documenting activity) and alerting to anomalous activity. Due to high-velocity …
New framework for optimal transport with jumps over intermediate spaces.
Bayesian nonparametrics adapt model complexity to diverse datasets.
This paper solves the multiple reference model problem in RLHF with exact solutions and sample complexity guarantees.
NHC learns scalable algorithmic solutions from diverse tasks.
New method learns diverse protein scaffolds for motif design.
DivDis learns diverse hypotheses from underspecified data to improve robustness.
Deep generative models are proven to be a useful tool for automatic design synthesis and design space exploration. When applied in engineering design, existing generative models face three challenges: 1) generated designs lack diversity and do not cover all areas of the design space, 2) it is difficult to explicitly im…
MOBO-OSD optimizes multi-objective functions using orthogonal search directions.
GEMSS discovers multiple sparse solutions in high-dimensional data.
In this work, we provide an efficient and realistic data-driven approach to simulate astronomical images using deep generative models from machine learning. Our solution is based on a variant of the generative adversarial network (GAN) with progressive training methodology and Wasserstein cost function. The proposed so…
Standard acquisition functions are sufficient for asynchronous Bayesian optimization.
This paper proposes new search algorithms for counterfactual explanations based upon mixed integer programming. We are concerned with complex data in which variables may take any value from a contiguous range or an additional set of discrete states. We propose a novel set of constraints that we refer to as a "mixed pol…
We establish Schauder a priori estimates and regularity for solutions to a class of boundary-degenerate elliptic linear second-order partial differential equations. Furthermore, given a smooth source function, we prove regularity of solutions up to the portion of the boundary where the operator is degenerate. Degenerat…
The instance segmentation problem intends to precisely detect and delineate objects in images. Most of the current solutions rely on deep convolutional neural networks but despite this fact proposed solutions are very diverse. Some solutions approach the problem as a network problem, where they use several networks or …
Machine learning (ML) needs industry-standard performance benchmarks to support design and competitive evaluation of the many emerging software and hardware solutions for ML. But ML training presents three unique benchmarking challenges absent from other domains: optimizations that improve training throughput can incre…