The abstract warns against flawed empirical research in machine learning.
problem Flawed empirical research in machine learning leading to unreliable results.
method Call for more awareness of experimental knowledge plurality and epistemic limitations.
result Current empirical machine learning research should be exploratory, not confirmatory.
QRAFTI uses multi-agent framework to improve equity factor research.
problem Replicating and developing new equity factors in large financial datasets.
method Integrates a research toolkit with MCP servers for data access and custom coding operations.
result Improves performance and explainability in multi-step empirical tasks.
Cryptocurrencies use blockchain tech for secure transactions, offering new research opportunities.
problem Misunderstanding of cryptocurrency technology and lack of empirical data.
method Analyzing detailed transaction data and summarizing statistics.
result Opportunity for academic research in financial economics.
Study examines challenges and applications of machine learning in finance.
problem Challenges in applying machine learning to financial research due to market idiosyncrasies and methodological differences.
method Discussion of adjustments needed to conventional machine learning methodology to account for financial market peculiarities.
result Machine learning can be unified with financial research as a robust complement to econometric methods.
AI agents improve forecast combination in empirical economics.
problem Hidden researcher degrees of freedom in AI-generated code.
method Adapted agent-loop architecture to empirical economics, added holdout evaluation.
result Independent agent searches find better forecast methods than benchmarks.
AI agents improve forecast combination but require transparency.
problem AI coding agents increase flexibility in empirical economics, leading to hidden degrees of freedom.
method Adapted open-source agent-loop architecture to empirical economics workflow, adding post-search holdout evaluation.
result Multiple agent runs outperform standard benchmarks in rolling evaluation but not all on post-search holdout.
Recently, a unified model for image-to-image translation tasks within adversarial learning framework has aroused widespread research interests in computer vision practitioners. Their reported empirical success however lacks solid theoretical interpretations for its inherent mechanism. In this paper, we reformulate thei…
Publication bias skews asset pricing research findings.
problem Bias in sharing and publishing research findings.
method Meta-studies and empirical Bayes corrections.
result Publication bias effects are minimal and not dominant.
Paper uses AI methods to forecast Bitcoin prices.
problem Inaccurate Bitcoin price predictions in previous studies.
method Combines EEMD and LSTM for next-day price forecast.
result Improves Bitcoin price prediction accuracy.
Peer-reviewed research and mined data predict stock returns similarly.
problem Predicting stock returns using research quality.
method Cross-sectional analysis of 29,000 accounting ratios with t-statistics > 2.0.
result Post-sample performance is largely independent of whether the predictor is peer-reviewed or mined.
New research shows common ID estimators in neural representations are inaccurate.
problem Inaccurate estimation of intrinsic dimensions in neural representations.
method Theoretical and empirical investigation of ID estimators in neural representations.
result Common ID estimators do not accurately reflect the true underlying ID of neural representations.
New theory explains how self-supervised learning converges, advancing AI research.
problem Lack of precise theoretical explanation for self-supervised learning convergence.
method Synthesized Identifiability Theory with empirical evidence to propose Singular Identifiability Theory (SITh).
result SITh provides deeper insights into SSL's implicit data assumptions and advances representation learning.
Introduces foundation priors for using model-generated data in empirical research.
problem Using model-generated data as real observations in empirical research.
method Introduces foundation priors as an exponential-tilted, generalized Bayesian update of the user's primitive prior.
result Synthetic data reflects both model patterns and user's priors, enabling principled use in empirical work.
New datasets improve fairness research by revealing UCI Adult's limitations.
problem Limitations of UCI Adult dataset in fairness research.
method Reconstructed a superset of UCI Adult data from US Census sources.
result New datasets reveal trade-offs between fairness criteria and performance.
Study finds LLMs hallucinate in finance tasks, needing research.
problem Hallucination in LLMs in finance.
method Empirical investigation of four methods to mitigate hallucination.
result LLMs hallucinate in financial tasks.
New research suggests privileged information doesn't improve model performance.
problem Challenges in transferring knowledge using privileged information in machine learning.
method Critical examination of existing theoretical and empirical analyses of LUPI methods.
result LUPI methods often fail to effectively transfer knowledge from privileged information.
A major line of contemporary research on complex networks is based on the development of statistical models that specify the local motifs associated with macro-structural properties observed in actual networks. This statistical approach becomes increasingly problematic as network size increases. In the context of curre…
Empirical evidence shows that ensembles, such as bagging, boosting, random and rotation forests, generally perform better in terms of their generalization error than individual classifiers. To explain this performance, Schapire et al. (1998) developed an upper bound on the generalization error of an ensemble based on t…
Fundamental portfolio beats market portfolio under certain conditions.
problem Empirical evidence of fundamental portfolio outperformance.
method Theoretical foundation based on stock price reversion to fundamental values.
result Fundamental portfolio outperforms market portfolio under strong reversion conditions.
Empirical law predicts accuracy of Google Translate's translation chains.
problem Predicting accuracy in machine translation with multiple hops.
method Empirical testing of Google Translate's sequential translation.
result Accuracy decreases with the number of translating hops, following a power law.
Low-precision training reduces computational cost and produces efficient models. Recent research in developing new low-precision training algorithms often relies on simulation to empirically evaluate the statistical effects of quantization while avoiding the substantial overhead of building specific hardware. To suppor…
learn2learn simplifies meta-learning research by providing a library and standardized interfaces.
problem Prototyping and reproducibility issues in meta-learning.
method Developed a library (learn2learn) with common routines and standardized interfaces.
result Fosters a community around standardized software for meta-learning research.
This study analyzes financial equity research reports to identify frequently asked questions and automates 80% of them.
problem Insufficient empirical analysis of questions answered in financial equity research reports.
method Analyzed 72 financial equity research reports, classifying sentences into 169 unique question archetypes. Used public corporate reports to classify questions' potential for automation.
result Approximately 80% of financial equity research reports can be automated, with 78.7% of questions automatable.
Theory and methods to mitigate omitted variable bias in causal machine learning.
problem Mitigating omitted variable bias in causal machine learning models.
method Developed a general theory and flexible statistical inference methods for bounding and testing the magnitude of omitted variable bias.
result Simple plausibility judgments can bound the magnitude of omitted variable bias in complex, nonlinear models.
This research analyzes deep PDE solvers for option pricing accuracy.
problem Understanding the accuracy of deep learning methods for solving PDEs in option pricing.
method Comparative experiments with two neural network algorithms in Black--Scholes and Heston models.
result Empirical convergence rates and training times of TDGF method determined.
Examines optimal risk sharing with realistic risk attitudes, finding risk seeking in certain subdomains.
problem Optimal risk sharing with empirically realistic risk attitudes.
method Allows for risk-seeking agents, generalizes expected utility, and uses counter-monotonic improvement theorem.
result First empirical results on optimal risk sharing with realistic risk attitudes.
The paper examines domain generalization algorithms and finds empirical risk minimization performs well.
problem Comparing domain generalization algorithms is difficult due to inconsistent experimental conditions.
method Implemented DomainBed, a testbed for domain generalization with seven datasets and model selection criteria.
result Empirical risk minimization shows state-of-the-art performance across all datasets.
Simulation-based inference methods can produce unreliable posterior approximations.
problem Reliability of simulation-based inference methods for scientific use cases.
method Benchmarked algorithms including Neural Posterior Estimation, Neural Ratio Estimation, Sequential Neural Likelihood, and Approximate Bayesian Computation.
result Ensembling posterior surrogates provides more reliable approximations.
Plotting a learner's average performance against the number of training samples results in a learning curve. Studying such curves on one or more data sets is a way to get to a better understanding of the generalization properties of this learner. The behavior of learning curves is, however, not very well understood and…
Research finds investors may lose from more diverse workplaces.
problem Investors' returns may be lower in more diverse companies.
method Examined D&I scores by Refinitiv in US and European markets.
result Investors may suffer lower returns for investing in more diverse companies.
LLMs help less-resourced researchers access costly data.
problem Unequal access to costly datasets limits research contributions.
method RAG framework with GPT-4o-mini for automated data collection.
result LLMs can collect CEO pay ratios and CAMs from corporate disclosures with high accuracy and low cost.
This research uses empirical copulas to price quanto options, showing significant differences from traditional models.
problem The dependence relation between currency and asset prices affects quanto option pricing.
method Empirical copulas are used to model the dependence between currency and asset prices.
result Empirical copulas provide non-negligible pricing differences compared to traditional models.
We conduct an extensive evaluation of price jump tests based on high-frequency financial data. After providing a concise review of multiple alternative tests, we document the size and power of all tests in a range of empirically relevant scenarios. Particular focus is given to the robustness of test performance to the …
Finding a well-performing architecture is often tedious for both DL practitioners and researchers, leading to tremendous interest in the automation of this task by means of neural architecture search (NAS). Although the community has made major strides in developing better NAS methods, the quality of scientific empiric…
Survey of large language models in financial prediction and trading.
problem Improving predictability and robustness of financial predictions and trading decisions.
method Task-centered taxonomy, review of empirical evidence, design patterns, benchmarks, and challenges analysis.
result Improved predictability and robustness of financial predictions and trading decisions through large language models.
Why do nations produce scientific research? This is a fundamental problem in the field of social studies of science. The paper confronts this question here by showing vital determinants of science to explain the sources of social power and wealth creation by nations. Firstly, this study suggests a new general definitio…
Significant advances have been made in artificial systems by using biological systems as a guide. However, there is often little interaction between computational models for emergent communication and biological models of the emergence of language. Many researchers in language origins and emergent communication take co…
Recent advances in Reinforcement Learning, grounded on combining classical theoretical results with Deep Learning paradigm, led to breakthroughs in many artificial intelligence tasks and gave birth to Deep Reinforcement Learning (DRL) as a field of research. In this work latest DRL algorithms are reviewed with a focus …
Researchers have constantly asked whether stock returns can be predicted by some macroeconomic data. However, it is known that macroeconomic data may exhibit nonstationarity and/or heavy tails, which complicates existing testing procedures for predictability. In this paper we propose novel empirical likelihood methods …
The present research aims to highlight the main factors influencing the development of entrepreneurial innovation in a rural environment and to perform an empirical study with the purpose of assessing the main problems in rural development. The research performed is mostly of a quantitative nature, being based on the u…
Researchers adaptively analyze market regimes to reveal investor behavior shifts.
problem Market relationships shift across different regimes, affecting investor behavior.
method Combining Kalman filtering, Markov-switching, and asymmetric response estimation.
result Foreign investors' predictive power increases during crises, while individual investors react more strongly to positive shocks.
Causal inference is central to many areas of artificial intelligence, including complex reasoning, planning, knowledge-base construction, robotics, explanation, and fairness. An active community of researchers develops and enhances algorithms that learn causal models from data, and this work has produced a series of im…
News might trigger jump arrivals in financial time series. The "bad" and "good" news seems to have distinct impact. In the research, a double exponential jump distribution is applied to model downward and upward jumps. Bayesian double exponential jump-diffusion model is proposed. Theorems stated in the paper enable est…
Paper proposes a data augmentation method for LLM-generated data in market research.
problem Bias in LLM-generated data in market research.
method Statistical data augmentation approach integrating LLM-generated and real data.
result Statistically robust estimators with reduced bias and cost savings.
We offer an experimental benchmark and empirical study for off-policy policy evaluation (OPE) in reinforcement learning, which is a key problem in many safety critical applications. Given the increasing interest in deploying learning-based methods, there has been a flurry of recent proposals for OPE method, leading to …
New research suggests continual learning should focus on both optimization objective and optimization trajectory.
problem Even with perfect joint loss approximation, continual learning still suffers from forgetting when starting a new task.
method Proposes focusing on both optimization objective and optimization trajectory, combining replay-approximated joint objectives with gradient projection-based optimization routines.
result Combining replay-approximated joint objectives with gradient projection-based optimization routines did not show clear benefits in initial experiments.
As researchers and practitioners of applied machine learning, we are given a set of requirements on the problem to be solved, the plausibly obtainable data, and the computational resources available. We aim to find (within those bounds) reliably useful combinations of problem, data, and algorithm. An emphasis on algori…
Recent years have seen adversarial losses been applied to many fields. Their applications extend beyond the originally proposed generative modeling to conditional generative and discriminative settings. While prior work has proposed various output activation functions and regularization approaches, some open questions …