Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

246493739985 · Jun 202019922001200920182026
48 results for AI performance

Experiment shows cognitive biases impact human-AI collaboration, highlighting the need for diverse evaluator samples.

problem Cognitive biases affect human-AI collaboration, leading to suboptimal outcomes.
method Randomized experiment with 2,784 participants, manipulating AI suggestion quality, task burden, and financial incentives.
result Individual attitudes toward AI are the strongest predictor of performance, influencing accuracy and overcorrection.

The paper highlights AI brittleness and the need for robust testing out-of-distribution performance.

problem The brittleness of AI systems, especially Deep Neural Networks, limits their reliability and certification.
method Analysis of AI brittleness and OOD performance, emphasizing the need for resilience and improved evaluation methods.
result AI systems are more failure-prone than certified in critical systems, and OOD performance falls off gradually.

FST.ai 2.0 improves Taekwondo decision-making with AI, reducing review time and increasing trust.

problem Fair, transparent, and explainable decision-making in Taekwondo.
method Pose-based action recognition, epistemic uncertainty modeling, interactive dashboards.
result 85% reduction in decision review time, 93% referee trust in AI-assisted decisions.

Gen AI improves document understanding but not data analysis in public sector tasks.

problem Understanding the impact of Gen AI on public sector tasks.
method Pre-registered field experiment comparing Gen AI to control group performance.
result Mixed results: Gen AI improves document understanding but not data analysis.

This paper compares AI performance in native vs. browser-based implementations.

problem Performance demands in AI applications, especially client-side.
method Comparison study between native code and browser-based (JS, ASM.js, WebAssembly) implementations.
result Current runtime optimizations push browser performance close to native binary.

AI threatens financial stability through misuse and stealth adoption.

problem Misuse and stealth adoption of AI in financial regulations.
method Analysis of AI's potential risks and criteria for AI suitability.
result AI will likely become widely used by stealth, affecting high-level financial functions.

The study evaluates AI model performance measures for medical use.

problem Selecting appropriate performance measures for AI models in medical practice.
method Assessed 32 performance measures across five domains for binary outcomes.
result 17 measures are both proper and reflect decision-analytic performance.

Optimal allocation of human effort to correct AI assessments in decision-making.

problem How to allocate costly human effort to correct noisy or biased AI-generated assessments.
method Decision-theoretic framework treating AI assessments as signals and human judgments as costly information. Developed estimation procedures under nonparametric and linear models.
result Our approach substantially outperforms LLM-only predictions and achieves performance comparable to full human review while using only 20-30% of the human information.

AlphaX uses AI to outperform Brazilian stock market benchmarks.

problem AI strategies often overperform in backtests but underperform in real markets due to lookahead bias.
method Controlled simulations to mitigate lookahead bias, using Value Investing principles.
result AlphaX strategy outperforms major benchmarks and technical indicators.

Optimizes AIS hyperparameters for efficient marginal likelihood estimation.

problem Limited computation budget affects AIS performance.
method Flexible intermediary distributions defined by residual density, parameter sharing, and fix linear schedule.
result Optimized-Path AIS reduces sampling iterations and improves performance.

Ray is a distributed system for AI applications that learn from continuous interactions.

problem Demanding systems requirements for next-gen AI applications.
method Unified interface for task-parallel and actor-based computations, distributed scheduler, fault-tolerant store.
result Demonstrated scaling beyond 1.8 million tasks per second and better performance for reinforcement learning.

A machine learning environment for detecting autonomous vehicle corner cases.

problem Testing autonomous driving software in the real world is difficult.
method Connecting CARLA simulation software to TensorFlow and custom AI client software.
result The system can identify situations where AI software fails to understand the scenario.

Paper proposes a new AI design philosophy for distributed learning.

problem Designing AI systems that mimic human cognitive capabilities.
method Democratized learning (Dem-AI) with self-organized hierarchical agents.
result Self-organized learning systems can perform complex tasks more efficiently.

AI-driven investment strategies self-defeat at scale due to signal crowding and erosion.

problem Excess returns from AI-driven investment strategies diminish at scale due to signal crowding and erosion.
method Theoretical model and empirical validation using SEC Form 13F filings and hedge fund return dynamics.
result The alpha half-life of signals decreases significantly with AI adoption, leading to diminishing returns.

AI systems need reliable testing to ensure safety and trustworthiness.

problem Current AI Act lacks functional trustworthiness for AI systems.
method Define technical application distribution, set risk-based performance, and conduct statistically valid testing.
result Reliable functional trustworthiness is essential for AI systems.

Do-AIQ framework evaluates AI algorithms' quality using DOE.

problem Quality evaluation of AI mislabel detection algorithms.
method Design-of-experiment approach with high-dimensional constraint space design and surrogate modeling.
result Established framework for evaluating AI algorithm quality robustly.

AI triage and diagnostic system performs similarly to human doctors, with safer triage advice.

problem Improving patient care through more reliable symptom checkers.
method Prospective validation study comparing AI system to human doctors.
result AI system's accuracy in identifying conditions comparable to human doctors, with safer triage advice.

AI methods often fail to outperform classical CPU-based solvers on Maximum Independent Set problems.

problem Comparing AI methods with classical CPU-based solvers on Maximum Independent Set problems.
method Comparison of AI methods (e.g., generative models, reinforcement learning) with classical CPU-based solvers (e.g., KaMIS) on Maximum Independent Set problem.
result AI-inspired methods are often outperformed by classical CPU-based solvers, even with post-processing techniques.

AI pipeline simplifies AI deployment on embedded devices.

problem Training and deployment of custom AI solutions on embedded devices requires expertise and integration barriers.
method Modular AI pipeline integrating data, algorithms, and deployment tools.
result LPDNN consistently outperforms other deployment frameworks on embedded platforms.

CAT framework improves AI medical screening fairness and reliability.

problem Imbalanced data, varying performance across cohorts, and patient-level inconsistencies in traditional metrics.
method CAT framework introduces patient-level assessment, entropy-based distribution weighting, and cohort-weighted sensitivity and specificity.
result Enhanced predictive reliability, fairness, and interpretability of AI-driven medical screening models.

AI learns to adapt driving behavior based on human interest.

problem Adapting AI behavior to individual human preferences for trust and comfort.
method Hybrid brain-computer interface (hBCI) detects human interest, which is used to adapt AI behavior in a virtual vehicle.
result AI agent slows down when passengers encounter objects of interest, increasing viewing time for subjects.

AI-driven tax policies improve economic equality and productivity.

problem Lack of appropriate economic data and limited opportunity to experiment.
method Two-level deep reinforcement learning approach to learn dynamic tax policies from observational data.
result AI-driven tax policies improve the trade-off between equality and productivity by 16%.

The paper tackles AI advice giving by considering adherence levels and defer options.

problem Inadequate consideration of human adherence to AI recommendations.
method Sequential decision-making model that considers adherence levels and incorporates a defer option.
result Specialized learning algorithms provide better convergence and empirical performance.

The paper investigates AI robustness through experiments and statistical analysis.

problem Inaccurate AI predictions can lead to safety and adoption issues.
method Design of experiments framework to study AI classification robustness.
result AI algorithms' robustness is influenced by various factors.

AI agents learn to cooperate with users of unknown type.

problem Designing AI agents that can cooperate with new users effectively.
method Modeling user behavior as parameters, observing user actions to infer type, and adapting policies.
result Adaptive AI agents perform significantly better than non-adaptive ones in real scenarios.

The paper integrates AI and expert knowledge to optimize radiotherapy decisions.

problem Optimizing radiation dose planning considering patient-specific information.
method Integrating Gaussian process models with deep neural networks to quantify uncertainty.
result Improves AI model performance and guides clinical decision making.

ProEval efficiently estimates AI performance and discovers failures using pre-trained Gaussian Processes.

problem Resource-intensive evaluation of generative AI models.
method ProEval uses pre-trained Gaussian Processes and Bayesian quadrature to estimate performance and discover failures.
result ProEval requires significantly fewer samples to achieve accurate performance estimates and reveals more diverse failure cases.

Optimizes AI learning with limited human feedback budgets.

problem Optimizing allocation of a fixed annotation budget for AI learning.
method Preference-Calibrated Active Learning (PCAL) using semi-parametric inference.
result Proves asymptotic optimality and robustness of the PCAL estimator.