Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

81162242323 · Jun 202019922001200920182026
48 results for agent selection

Simplified feature selection using a single agent with restructured choice strategy.

problem Efficiency and cost issues in multi-agent reinforced feature selection.
method Single-agent approach with restructured choice strategy, including scanning method, feature prioritization, state representation, and reward scheme.
result Improved efficiency and effectiveness of feature selection.

A bandit framework selects the best reinforcement learning agent for costly interactions.

problem Optimal reinforcement learning agent selection in environments without simulators.
method Multi-arm bandit framework with surrogate rewards to select and learn from reinforcement learning agents.
result The framework consistently selects the optimal agent with more cumulative reward.

Selection mechanisms impact market volatility in evolving markets.

problem Determining how selection mechanisms affect market volatility in evolving markets.
method Used a population of evolving zero-intelligence agents and a frequent batch auction price-discovery mechanism to analyze the role of selection mechanisms.
result Local fitness-proportionate selection mechanisms correlate with high correlation between risk-aversion and volatility, while quantile-based selection mechanisms show less correlation.

AI traders learn to exploit meta-orders from slower traders, increasing their profits.

problem Adverse selection of medium-frequency traders by high-frequency AI agents.
method Reinforcement learning in a Hawkes LOB model, with impulse control and PPO.
result AI agents can learn to capitalize on meta-orders, increasing their profits.

Study on investment strategy for agents with periodic preferences and discounting.

problem Investment decisions by agents with periodic S-shaped preferences and present bias.
method Infinite-horizon, continuous-time portfolio selection problem with quasi-hyperbolic discounting.
result Time-consistent planning strategy can be formulated as an equilibrium to a static mean field game.

CausalGame benchmarks LLM agents' causal thinking in games.

problem Evaluating causal thinking in AI Scientists with LLMs.
method Interactive games with 14 scenarios incorporating selection bias, measurement error, and hidden confounders.
result None of the 30 LLM agents demonstrated reliable causal thinking, with the best model achieving only 68.0% survival.

AMSAs adaptively manage crypto-currency trading by selecting multiple strategies based on market conditions.

problem Maximizing gains in volatile crypto-currency markets with high uncertainty.
method AMSAs use multiple sub-agents with different strategies, dynamically selecting them based on market conditions.
result AMSAs can achieve high positive alpha in long-term crypto-currency trading.

Developing an Agent-Based Model to Mitigate Adverse Selection in Uniswap v3 Liquidity Providers

problem Adverse selection in Uniswap v3 liquidity providers
method Agent-Based Model incorporating blockchain microstructure and volatility dynamics
result Dynamic fee schedules improve hedged Profit and Loss for liquidity providers

The paper examines how slightly biasing towards under-represented groups in sequential selection processes can lead to long-term fairness.

problem Designing fair sequential decision-making processes for long-term social fairness.
method Proposes Multi-agent Fair-Greedy policy to balance score maximization and fairness.
result Proves convergence to long-term fairness target set by agents when score distributions are identical.

AutoFS combines trainers to improve feature selection efficiency and effectiveness.

problem Balancing feature selection efficiency and effectiveness.
method Interactive Reinforced Feature Selection (IRFS) framework with diverse trainers.
result Improved feature selection efficiency and effectiveness compared to existing methods.

Benchmark evaluates LLM trading agents by masking identifiers to prevent memory leaks.

problem Evaluate LLM trading agents without relying on market memory or noise.
method Data-side masking protocol, Barra-style performance attribution framework.
result LLM agents' returns are largely explained by market and style exposure, not stock selection.

Investigates portfolio selection among competitive agents with mean-variance preferences.

problem Optimizing portfolios with multi-agent competition and relative wealth comparison.
method Reformulated as a constrained, non-homogeneous stochastic linear-quadratic control problem; derived optimal feedback strategies; used decoupling techniques and fixed-point theory to solve nonlinear BSDEs.
result Characterized three scenarios based on market and competition parameters: unique Nash equilibrium, no Nash equilibrium, or infinitely many Nash equilibria.

MaxMax Q-Learning improves coordination in multi-agent reinforcement learning by refining action selection.

problem Relative over-generalization in decentralized multi-agent reinforcement learning.
method MaxMax Q-Learning employs iterative sampling and evaluation of potential next states to refine approximations of ideal state transitions.
result MaxMax Q-Learning frequently outperforms existing baselines, demonstrating enhanced convergence and sample efficiency.

The paper introduces a method for multi-agent reinforcement learning to coordinate exploration.

problem Sparse rewards in multi-agent settings lead to independent exploration.
method Designing intrinsic rewards that encourage coordination and developing a hierarchical policy.
result The approach accelerates and improves exploration in cooperative multi-agent settings.

MarketSenseAI system outperforms passive benchmarks by 25.2% on S&P 500, adding value over random selection.

problem Identifying alpha in stock recommendations from multi-agent LLM systems.
method Deployed multi-agent LLM equity system generating live signals, combining four specialist agents into a synthesis agent.
result Strong-buy equal-weight portfolio on S&P 500 earns +2.18%/month, significantly outperforming passive benchmarks.

HabitatAgent offers a multi-agent system for transparent housing consultation.

problem Opaque reasoning and brittle multi-constraint handling in housing recommendation systems.
method HabitatAgent is a multi-agent architecture with specialized roles for memory, retrieval, generation, and validation.
result HabitatAgent achieves 95% accuracy in real user consultation scenarios, significantly outperforming a strong baseline.

Agent learns from an expert, adapting to constraints in concept learning.

problem Insufficient query selection in active learning for realistic human domains.
method Imitation learning to reason about both internal goals and external constraints.
result Agent outperforms other active learners under most constrained conditions.

Study shows informed traders harm market makers but price discovery benefits outweigh costs.

problem Informed traders' impact on market makers' profitability.
method Agent-based model with heterogeneous learning agents, multi-agent reinforcement learning.
result Informed market order flow is harmful when aggregate informedness is low but beneficial as it increases.

Study counterfactuals in combinatorial choice using a representative agent model.

problem Analyzing decision-making from aggregated binary polytope data.
method Nonparametric approach based on a representative agent model, solving polynomial and mixed-integer convex programs.
result Developed a method for counterfactual prediction that works even under model misspecification.

DARL framework tackles partial domain adaptation by selecting source instances for positive transfer.

problem Tackles the challenge of selecting source instances for positive transfer in partial domain adaptation.
method Proposes a Domain Adversarial Reinforcement Learning (DARL) framework that uses deep Q-learning and domain adversarial learning to select source instances and learn domain-invariant features.
result Demonstrates superior performance over existing methods for partial domain adaptation on several benchmark datasets.

Study on decision-making cascades with agents having varying beliefs and noise levels.

problem Optimizing decision-making in a cascade of agents with heterogeneous beliefs and noise.
method Recursive belief update and analysis of optimal decision rules, predecessor selection problem characterization.
result Optimal decisions can deviate from true prior beliefs in certain conditions, highlighting the importance of social learning.

This paper addresses reward estimation and incentive design for agents with hidden rewards.

problem Estimating and incentivizing agents with unknown rewards in a learning setting.
method Repeated adverse selection game with a self-interested learning agent and a learning principal. Introduces an estimator for consistent reward estimation and a data-driven incentive policy.
result Finite-sample consistency of the estimator and a rigorous regret bound for the principal.

We study the market selection hypothesis in complete financial markets, populated by heterogeneous agents. We allow for a rich structure of heterogeneity: individuals may differ in their beliefs concerning the economy, information and learning mechanism, risk aversion, impatience and 'catching up with Joneses' preferen…

2011-06-15abs ↗pdf ↗

We develop a model to study the role of rationality in economics and biology. The model's agents differ continuously in their ability to make rational choices. The agents' objective is to ensure their individual survival over time or, equivalently, to maximize profits. In equilibrium, however, rational agents who maxim…

2015-07-14abs ↗pdf ↗

Datasets with hundreds to tens of thousands features is the new norm. Feature selection constitutes a central problem in machine learning, where the aim is to derive a representative set of features from which to construct a classification (or prediction) model for a specific task. Our experimental study involves micro…

2016-03-16abs ↗pdf ↗

New theorems show agents need specific internal structures to perform well under uncertainty.

problem How do agents need to be structured to perform well under uncertainty?
method Proved selection theorems showing strong task performance forces specific internal structures.
result Strong task performance forces world models, belief-like memory, and persistent regime-tracking variables.

A novel method for efficient CDRL over wireless networks.

problem Challenges in collaborative deep reinforcement learning over wireless networks.
method Semantic-aware heterogeneous federated deep reinforcement learning (HFDRL) algorithm.
result Superior performance compared to state-of-the-art baselines.

New algorithm for multi-agent reinforcement learning with attention mechanism.

problem Challenges in training decentralized policies in multi-agent settings.
method Actor-attention-critic algorithm with centrally computed critics and attention mechanism.
result More effective and scalable learning in complex multi-agent environments.

Study optimal portfolios for many players in a market model with random coefficients.

problem Optimal portfolio selection for many players under relative performance criteria in a market model with random coefficients.
method Game theory and stochastic optimal control, focusing on CARA and CRRA risk preferences, and extending to continuum of players.
result Existence of forward Nash equilibrium and mean field equilibrium for the n-agent game and corresponding mean field stochastic optimal control problem.

This work tackles uncertainty in multi-agent multi-modal trajectory forecasting.

problem Measuring and ranking uncertainty in multi-agent multi-modal trajectory forecasting.
method Proposes collaborative uncertainty (CU) and a CU-aware regression framework.
result The CU-aware regression framework improves SOTA systems' performances.

Paper presents a method to efficiently learn ordered representations of multi-agent data.

problem Challenges in learning consistent representations of multi-agent interactions.
method Dynamic alignment method to order multi-agent data for faster representation learning.
result Representation learning of multi-agent data is significantly accelerated.

A RL framework selects features to balance bias and accuracy dynamically.

problem Bias in automated feature selection when predictors are correlated.
method Multi-component reward function with policy gradient for dynamic regularization and bias mitigation.
result Model balances fairness and accuracy during training.

A new decentralized policy for multi-agent MAB problems outperforms random communication.

problem Solving multi-agent multi-armed bandit problems efficiently.
method UCB strategy for individual option selection and communication strategy based on neighbor exploration potential.
result The proposed policy significantly outperforms random communication strategies.

In this paper, we investigate a new form of automated curriculum learning based on adaptive selection of accuracy requirements, called accuracy-based curriculum learning. Using a reinforcement learning agent based on the Deep Deterministic Policy Gradient algorithm and addressing the Reacher environment, we first show …

2018-06-25abs ↗pdf ↗

Enhances index selection for databases with task-specific inductive biases.

problem Challenges in traditional and automatic tuning strategies for database index set selection.
method Applies deep RL with task-specific inductive biases to index set selection, reformulating the problem as permutation learning.
result Improves index selection, achieving up to 40% smaller configurations with similar latency.

New neural policies learn multi-agent relationships directly, improving coordination in dynamic environments.

problem Training coordination among varying numbers of agents in reinforcement learning.
method Attentional architecture for shared policies that adapt to each agent's context.
result Superior performance on multi-agent vehicle coordination problem, especially with many agents.

Policy-gradient method controls multiple non-cohesive targets.

problem Controlling multiple non-cohesive targets in a decentralized manner.
method Proximal Policy Optimization for target selection and driving.
result Effective control of non-cohesive targets without prior dynamics knowledge.

A meta-learning approach for efficient algorithm selection in budget-limited scenarios.

problem Efficiently selecting the best-performing machine learning algorithm with limited computational resources.
method A Markov Decision Process framework where an agent decides whether to train, wake up, or start new algorithms based on partial learning curves.
result Meta-learning from learning curves improves algorithm selection, especially when learning curves do not intersect frequently.

We present a detailed numerical analysis of the modified version of a conservative self-organized extremal model introduced by Pianegonda et. al. for the distribution of wealth of the people in a society. Here the trading process has been modified by the stochastic bipartite trading rule. More specifically in a trade o…

2011-09-30abs ↗pdf ↗

In our simplified description `wealth' is money (mm). A kinetic theory of gas like model of money is investigated where two agents interact (trade) selectively and exchange some amount of money between them so that sum of their money is unchanged and thus total money of all the agents remains conserved. The probabilit…

2005-09-21abs ↗pdf ↗

We study ranking quantilized mean-field games to select top-performing agents.

problem Selecting top-performing agents in competitive scenarios.
method Developed two formulations: target-based and threshold-based, and provided analytic and semi-explicit solutions.
result Analytic and semi-explicit solutions for quantilized mean-field consistency conditions.