Safe RL with binary feedback using SABRE algorithm.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We compute the spectral action of with the trivial spin structure and the round metric and find it in each case to be equal to . We do this by explicitly computing the spectrum of the Dirac operator for equipped with the trivial …
A simple DQN-based multi-agent RL system for binary actions.
Action detection and recognition tasks have been the target of much focus in the computer vision community due to their many applications, namely, security, robotics and recommendation systems. Recently, datasets like AVA, provide multi-person, multi-label, spatiotemporal action detection and recognition challenges. Be…
Smooth and symplectic symmetries of an infinite family of distinct exotic surfaces are studied, and comparison with the corresponding symmetries of the standard is made. The action on the lattice induced by a smooth finite group action is shown to be strongly restricted, and as a result, nonsmoothability…
Due to the high variance of policy gradients, on-policy optimization algorithms are plagued with low sample efficiency. In this work, we propose Augment-Reinforce-Merge (ARM) policy gradient estimator as an unbiased low-variance alternative to previous baseline estimators on tasks with binary action space, inspired by …
Low precision networks in the reinforcement learning (RL) setting are relatively unexplored because of the limitations of binary activations for function approximation. Here, in the discrete action ATARI domain, we demonstrate, for the first time, that low precision policy distillation from a high precision network pro…
Sequence generation models are commonly refined with reinforcement learning over user-defined metrics. However, high gradient variance hinders the practical use of this method. To stabilize this method, we adapt to contextual generation of categorical sequences a policy gradient estimator, which evaluates a set of corr…
Estimation of importance sampling weights for off-policy evaluation of contextual bandits often results in imbalance - a mismatch between the desired and the actual distribution of state-action pairs after weighting. In this work we present balanced off-policy evaluation (B-OPE), a generic method for estimating weights…
We consider sequential decision making problems for binary classification scenario in which the learner takes an active role in repeatedly selecting samples from the action pool and receives the binary label of the selected alternatives. Our problem is motivated by applications where observations are time consuming and…
A new calibration metric bridges testability and actionability.
We determine the homotopy type of quotients of by free actions of where . Much like free actions, they can be classified via the first -localized -invariant, but there are restrictions on the possibilities, and these restrictions are…
A new uplift modeling approach uses binary treatment indicators more efficiently.
Study extends binary omniprediction to multiclass setting with improved sample complexity.
Policy learning can be used to extract individualized treatment regimes from observational data in healthcare, civics, e-commerce, and beyond. One big hurdle to policy learning is a commonplace lack of overlap in the data for different actions, which can lead to unwieldy policy evaluation and poorly performing learned …
The UNKNOT problem solved using natural language processing and machine learning.
New method minimizes decision errors in large treatment spaces.
Neural Index Policy for multi-action bandits with heterogeneous budgets.
In this paper, we present simple algorithms for Dueling Bandits. We prove that the algorithms have regret bounds for time horizon T of order O(T^rho ) with 1/2 <= rho <= 3/4, which importantly do not depend on any preference gap between actions, Delta. Dueling Bandits is an important extension of the Multi-Armed Bandit…
New method constructs multiple group racks, differing from known constructions.
Research on the multi-armed bandit problem has studied the trade-off of exploration and exploitation in depth. However, there are numerous applications where the cardinal absolute-valued feedback model (e.g. ratings from one to five) is not suitable. This has motivated the formulation of the duelling bandits problem, w…
We study the logistic bandit, in which rewards are binary with success probability and actions and coefficients are within the -dimensional unit ball. While prior regret bounds for algorithms that address the logistic bandit exhibit exponential dependence on the slop…
Study evaluates machine learning methods for large-scale network reliability, revealing ANN's and PR's performance.
Loss-calibrated EP improves Bayesian decision-making by focusing on utility-sensitive posterior approximations.
Study of zero-sum games with noisy observations and commitments.
We address the problem of regret minimization in logistic contextual bandits, where a learner decides among sequential actions or arms given their respective contexts to maximize binary rewards. Using a fast inference procedure with Polya-Gamma distributed augmentation variables, we propose an improved version of Thomp…
Algorithm BGLM-OFU minimizes regret in combinatorial causal bandits with binary models.
Improved Thompson Sampling for logistic bandits with information-theoretic analysis.
New method for evaluating and learning in complex decision-making scenarios.
Existing algorithms aiming to learn a binary classifier from positive (P) and unlabeled (U) data generally require estimating the class prior or label noises ahead of building a classification model. However, the estimation and classifier learning are normally conducted in a pipeline instead of being jointly optimized.…
Binary PheNorm extends phenotype labeling for EHRs using binary silver labels.
Study detects and explains positional bias in financial LLMs.
We address online linear optimization problems when the possible actions of the decision maker are represented by binary vectors. The regret of the decision maker is the difference between her realized loss and the best loss she would have achieved by picking, in hindsight, the best possible action. Our goal is to unde…
Quantum circuits represent binary classification trees with binary features.
The Bouncy Particle Sampler is a novel rejection-free non-reversible sampler for differentiable probability distributions over continuous variables. We generalize the algorithm to piecewise differentiable distributions and apply it to generic binary distributions using a piecewise differentiable augmentation. We illust…
An attractive approach for fast search in image databases is binary hashing, where each high-dimensional, real-valued image is mapped onto a low-dimensional, binary vector and the search is done in this binary space. Finding the optimal hash function is difficult because it involves binary constraints, and most approac…
Extracting actionable insight from Electronic Health Records (EHRs) poses several challenges for traditional machine learning approaches. Patients are often missing data relative to each other; the data comes in a variety of modalities, such as multivariate time series, free text, and categorical demographic informatio…
We present a comprehensive study of multilayer neural networks with binary activation, relying on the PAC-Bayesian theory. Our contributions are twofold: (i) we develop an end-to-end framework to train a binary activated deep neural network, (ii) we provide nonvacuous PAC-Bayesian generalization bounds for binary activ…
Probabilistic learning for binary classification with categorical variables.
G-Net constructs binary neural networks with high accuracy using randomized binary embeddings.
QNNs can't distinguish binary signals from their negations, revealing a new symmetry.
Study identifies conditions for proxy adjustment in confounded binary treatment outcomes.
Machine learning struggles to predict binary options movements due to randomness.
Artificial Neural Networks (ANNs) are currently being used as function approximators in many state-of-the-art Reinforcement Learning (RL) algorithms. Spiking Neural Networks (SNNs) have been shown to drastically reduce the energy consumption of ANNs by encoding information in sparse temporal binary spike streams, hence…
We introduce a novel scheme to train binary convolutional neural networks (CNNs) -- CNNs with weights and activations constrained to {-1,+1} at run-time. It has been known that using binary weights and activations drastically reduce memory size and accesses, and can replace arithmetic operations with more efficient bit…
Paper resolves open problems on sample complexity in binary hypothesis testing.
BIND removes background noise from binary matrices, improving detection accuracy and fairness.
Study shows how adjusting for a binary proxy can bound causal effects.