Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

4896143191 · Jun 202019922001200920182026
48 results for Programmatic Policies

PROPEL learns interpretable programmatic policies using imitation and projection.

problem Learning interpretable programmatic policies in reinforcement learning.
method PROPEL is a meta-algorithm based on three insights: optimization in policy space, neural-program mixing, and imitation synthesis.
result PROPEL significantly outperforms state-of-the-art approaches in learning programmatic policies.

The paper tackles calibrating long-term behaviors with multiple styles using programmatic style-consistency.

problem Generating long-term sequential behaviors with multiple styles simultaneously.
method Leverage programmatic labeling functions to specify controllable styles and derive style-consistency as a learning objective.
result Learned policies can be calibrated for up to 1024 distinct style combinations.

Paper introduces methods for more reliable probabilistic predictions with confidence intervals.

problem Inaccurate labeling of datasets due to unreliable probabilistic predictions from weak labeling functions.
method Proposes a methodology to provide confidence intervals for label probabilities using uncertainty sets of distributions.
result Improves reliability of probabilistic predictions and provides confidence intervals for label probabilities.

MORL uses program synthesis to improve reinforcement learning policies.

problem Difficult to interpret and impose constraints on learned policies from black-box neural networks.
method Iterative framework combining program synthesis and behavior cloning.
result Programmatic representation allows for high-level modifications leading to improved learning.

FABLE incorporates instance features into PWS label models for improved performance.

problem Lack of instance features in existing label models limits their performance.
method FABLE uses a mixture of Bayesian label models and a Gaussian Process classifier to incorporate instance features.
result FABLE achieves the highest averaged performance across nine baselines on benchmark datasets.

Adapting neural networks to guide program optimization for better classifiers.

problem Learning differentiable programs with complex architectures.
method Formulating program optimization as a graph search problem, using neural networks as heuristic relaxations.
result Trained neural networks can guide combinatorial search for programmatic classifiers, improving accuracy and interpretability.

Improves understanding of PWS by calculating influence of sources and data.

problem Understanding the influence of each component in PWS.
method Proposes source-aware Influence Function (IF) to decompose and calculate influence.
result Improves end model's generalization performance and identifies mislabeling.

A hierarchical model learns multi-agent trajectories from weak labels.

problem Training sequential generative models for coordinated multi-agent behavior.
method Hierarchical framework with programmatically produced weak labels.
result Effective learning of long-term coordination and high-level semantics.

Improved learning algorithm for first-price auctions reduces regret significantly.

problem Challenges in learning optimal bidding strategies for first-price auctions.
method Introduced novel ideas to achieve lower regret in sequential learning.
result Achieved log2(T)log^2(T) regret when opponents' bid distribution is known, and T1/3+εT^{1/3+ ε} regret in learning case.

Stablecoin system improves resilience to extreme market events.

problem Vulnerability of stablecoins to extreme volatility and adversarial attacks.
method MVF-Composer uses multi-agent simulations to stress-test and down-weight manipulative signals.
result Reduces peak peg deviation by 57% and mean recovery time by 3.1x under adversarial conditions.

Generative AI improves stock selection by synthesizing features from diverse data sources.

problem Automating feature discovery in stock market data.
method Used large language models with retrieval-augmented generation and structured prompting to synthesize features from various data sources.
result AI-generated features consistently outperform baselines, with Sharpe improvements ranging from 14% to 91%.

The paper introduces metrics to objectively evaluate interpretability methods.

problem Lack of objective evaluation metrics for interpretability methods.
method Proposes a set of metrics to evaluate interpretability methods along simplicity and broadness.
result Validated metrics on different benchmark tasks and showed their utility in method selection.

New method estimates model performance bounds without ground truth labels.

problem Evaluation of weakly supervised models without direct access to ground truth labels.
method Formulates model evaluation as a partial identification problem and uses Fréchet bounds for performance estimation.
result Derives accurate and computationally efficient bounds for key metrics like accuracy, precision, recall, and F1-score.

Paper presents a model for identifying informative COVID-19 tweets.

problem Identifying informative COVID-19 tweets on Twitter.
method Leveraged transformers (RoBERTa, XLNet, BERTweet) trained in Semi-Supervised Learning (SSL) setting.
result Achieved F1 score of 0.9011 on test set, ranking 7th on leaderboard.

New method uses multi-task learning to improve molecule representations.

problem Cost, bias, and data requirements in chemical representation generation.
method Intelligent task selection in deep multitask networks with transfer learning.
result Deep representations capture more expressive task-based information.

This paper examines interest rates and market efficiency in DeFi loanable funds protocols.

problem Equilibrium of supply and demand for loanable funds in DeFi protocols.
method Review of interest rate mechanisms in Compound, Aave, and dYdX; empirical analysis of market efficiency and inter-connectedness.
result Interest rate rules in DeFi protocols do not always equilibrate supply and demand.

Cross-modal data programming speeds medical machine learning.

problem Labeling medical datasets is time-consuming and requires expert knowledge.
method Generates training labels by writing rules over auxiliary modalities, estimating accuracies and correlations.
result Matches or exceeds hand-labeling with statistical significance, faster and more flexible.

DPBD simplifies labeling functions through interactive demonstrations.

problem Difficulty in writing labeling functions for large-scale labeled training data.
method Data Programming by Demonstration (DPBD) framework using interactive demonstrations.
result Ruler system generates labeling rules more easily and with higher user satisfaction.

Large labeled training sets are the critical building blocks of supervised learning methods and are key enablers of deep learning techniques. For some applications, creating labeled training sets is the most time-consuming and expensive part of applying machine learning. We therefore propose a paradigm for the programm…

2016-05-25abs ↗pdf ↗

The paper shows how to improve policies on- and off-policy using bounds.

problem Improving reinforcement learning policies on- and off-policy.
method Lower bounding the performance difference of two policies to ensure monotonic improvement from mixture samples.
result An optimization procedure that applies the proposed bound can be seen as an off-policy natural policy gradient method.

Paper tackles efficient evaluation of natural stochastic policies in offline RL.

problem Efficiency issues in evaluating natural stochastic policies due to unknown evaluation policy.
method Derive efficiency bounds for tilting and modified treatment policies, propose nonparametric estimators.
result Proposed estimators attain efficiency bounds under lax conditions and enjoy partial double robustness.

Study shows accurate OPE depends on calibrated behaviour policy models.

problem Estimating a behaviour policy for OPE when true policy is unknown.
method Empirical studies comparing parametric vs non-parametric models.
result Simple non-parametric k-nearest neighbors model produces better calibrated behaviour policy estimates.

Stabilizes policy optimization with off-policy data using divergence augmentation.

problem Premature convergence and instability in policy optimization with off-policy data.
method Incorporates Bregman divergence between behavior and current policies to ensure safe policy updates.
result Empirically shows better performance in data-scarce scenarios compared to other algorithms.

New method estimates state-action stationary distribution for better off-policy policy evaluation.

problem Accurately estimating state-action stationary distribution for off-policy policy evaluation.
method Estimated Mixture Policy (EMP) for state and state-action stationary distribution corrections.
result Empirical validation shows improved accuracy over state-of-the-art methods.

New methods estimate policy value and gradients for deterministic policies from off-policy data.

problem Estimating policy value and gradients for deterministic policies from off-policy data.
method Proposed new doubly robust estimators based on kernelization approaches.
result Demonstrated a rate independent of horizon length for policy value and gradient estimation.