Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

114228341455 · Jun 202019922001200920182026
48 results for debate complexity

PEAR dynamically reconfigures agent roles to prevent persistent biases in multi-agent debates.

problem Persistent positional biases and sensitivity to role assignments in fixed topologies.
method Dynamic reconfiguration of agent roles and sparse topologies based on evolving agent states.
result Significantly improves average accuracy over debate baselines across multiple reasoning benchmarks.

AI systems learn complex goals via debate with human judges.

problem Complex human goals and preferences for AI systems.
method Training agents via self-play on a debate game, where humans judge which agent gives more true, useful information.
result Boosted classifier accuracy from 48.2% to 85.2% given 4 pixels, and from 59.4% to 88.9% given 6 pixels.

The UN General Debate Corpus analyzes speeches from UN member states to reveal their political positions.

problem Lack of data on state preferences in international politics.
method Text analysis of over 7,700 speeches from 1970-2016.
result Demonstrates how the UN General Debate Corpus can reveal country positions on various policy dimensions.

A novel fact-checking method using debate dynamics on knowledge graphs.

problem Fact-checking on knowledge graphs with user comprehension and interactive reasoning.
method Reinforcement learning agents debate on paths in the graph to classify facts as true or false.
result Interactive reasoning and user understanding of AI decisions on knowledge graphs.

Paper proposes using word embeddings to detect trolls in social media debates.

problem Preventing online harassment through rapid detection of offensive posts.
method Word embedding models for identifying fast-changing topics and negative content.
result GloVe model helps in discovering new keywords for trolling detection.

Simplified LSTM models improve sentiment analysis on Twitter debate data.

problem Performing sentiment analysis on long sequence data from Twitter debates.
method Developed six parameter-reduced LSTM models (slim LSTM) for faster training and reduced computational cost.
result Slim LSTM models outperform standard LSTM model in sentiment analysis of GOP Debate Twitter dataset.

Crowd opinions in microblogs can predict event outcomes, matching with expert opinions.

problem Utilizing crowd wisdom for event outcome prediction in microblogs.
method Multi-label sentiment classification of tweets to gauge crowd opinion and compare with expert predictions.
result Crowd opinions in microblogs often match with expert opinions, especially in non-debate events.

Paper uses neural word embeddings to analyze UN speeches for policy preferences and voting behavior.

problem Analyzing policy preferences and paradigm shifts in international politics.
method Applied neural word embeddings (Word2vec) to UN General Debate speeches.
result Found statistical relation between speech semantic content and voting behavior, contrary to hypothesis.

Study uses few-shot learning to analyze claims and arguments in German debate on arms deliveries.

problem Limited data and computational resources for automated content analysis.
method Multilingual transformer model with adapter extension and few-shot learning.
result Parameter-efficient approach performs well on varying training set sizes.

This study models FOMC policy decisions using debate-based LLMs.

problem Accurately predicting central bank policy decisions, especially FOMC's, is challenging.
method A novel framework that simulates FOMC's collective decision-making process through iterative rounds of LLMs interacting as agents.
result The debate-based approach significantly outperforms standard LLMs in prediction accuracy.

New theory of co~events resolves debates between Bayesianists and frequentists.

problem Fierce debates between Bayesianists and frequentists over Bayesian scheme.
method Developed new co~event axiomatics and theory to study experience and chance as a single co~event.
result Demonstrated effectiveness of new theory in resolving debates over Bayesian scheme.

Paper resolves the debate on process vs. outcome supervision in reinforcement learning.

problem Distinguishing between process and outcome supervision in reinforcement learning.
method Developed a technical tool (Change of Trajectory Measure Lemma) to show equivalence between outcome and process supervision under standard data coverage assumptions.
result Reinforcement learning through outcome supervision is statistically equivalent to process supervision, up to polynomial factors in horizon.

Investment strategy depends on many factors for venture capital funds.

problem Finding the optimal portfolio size for venture capital funds.
method Analyzes various factors affecting fund returns and optimal portfolio size, starting with basic assumptions and increasing complexity.
result Investment strategy depends on many factors, not a one-size-fits-all formula.

This work addresses fairness constraints for multiple subpopulations in machine learning models.

problem Fairness constraints for multiple subpopulations in machine learning models.
method Constraining the expected outcome of subpopulations in kernel regression and decision tree regression, specifically random forests and boosted trees.
result The proposed solution does not affect the computational or memory complexity of decision trees and can be easily integrated post training.

Potential Future Exposure (PFE) is a standard risk metric for managing business unit counterparty credit risk but there is debate on how it should be calculated. The debate has been whether to use one of many historical ("physical") measures (one per calibration setup), or one of many risk-neutral measures (one per num…

2015-12-19abs ↗pdf ↗

Crowdsourced investigation shows differing results for technical analysis strategies.

problem Differing results in technical analysis strategies due to lack of method.
method Collaborative scientific computational framework using Monte Carlo simulations and historical back testing.
result Results are not repeatable by other researchers, highlighting the need for transparency and robustness.

MakerDAO's governance is centralized despite its decentralized claim.

problem Decentralization illusion in Decentralized Finance (DeFi) governance.
method Empirical analysis using financial, transaction, network, and sentiment indicators.
result Centralized governance impacts Maker protocol and voting power distribution.

This paper reviews bank performance determinants, highlighting future research areas.

problem Understanding bank performance factors to improve sector efficiency and knowledge.
method Analysis of 54 studies in peer-reviewed journals.
result Bank performance factors remain largely unexplored, especially post-COVID-19.

Topological parallax assesses AI models' geometric similarity to datasets for safety.

problem Ensuring AI models' robustness and safety in deep learning applications.
method Topological parallax compares a trained model to a reference dataset using Rips complexes and geodesic distortions.
result Topological parallax indicates whether a model shares similar multiscale geometric features with the dataset.

The paper analyzes tech specialization and diversification at various scales.

problem Trade-offs between specialization and diversification in economic development.
method Patent data and Economic Complexity framework.
result Technological Coherence positively impacts growth at metropolitan areas but negatively at larger scales.

The Information Plane theory predicts autoencoders do not compress input information.

problem Understanding the training dynamics of hidden layers in autoencoders.
method Derive a theoretical convergence for the Information Plane of autoencoders using a Gram-matrix based mutual information estimator.
result Ideal autoencoders with a large bottleneck layer size do not compress input information, while a small size causes compression only in the encoder layers.

In the current environment of financial distress, many governments are likely to soon become major holders of financial assets, but the policy debate focuses only on the likelihood and extent of short-term market stabilization. This paper shows that government intervention and propping up are likely to lead to long-ter…

2010-02-11abs ↗pdf ↗

Bayesian neural networks integrate uncertainty into neural networks for improved performance.

problem Overconfidence, lack of interpretability, and adversarial attacks in neural networks.
method Integrates Bayesian inference into neural networks to address limitations.
result BNNs improve model performance and provide uncertainty estimates.

This work investigates how multi-round reasoning improves LLM performance.

problem Improving problem-solving abilities in complex tasks with LLMs.
method Investigates approximation, learnability, and generalization properties of multi-round auto-regressive models.
result Transformers with finite context windows are universal approximators for Turing-computable functions and can approximate any Turing-computable sequence-to-sequence function through multi-round reasoning.

Geospatial ML models need special evaluation methods due to their unique challenges.

problem Evaluating geospatial machine learning models is challenging due to their specific characteristics.
method Delineated unique challenges and proposed concrete takeaways for improving geospatial model evaluations.
result Concrete takeaways for improving evaluations of geospatial model performance.