The abstract warns against flawed empirical research in machine learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
QRAFTI uses multi-agent framework to improve equity factor research.
Cryptocurrencies use blockchain tech for secure transactions, offering new research opportunities.
Study examines challenges and applications of machine learning in finance.
AI agents improve forecast combination in empirical economics.
AI agents improve forecast combination but require transparency.
Recently, a unified model for image-to-image translation tasks within adversarial learning framework has aroused widespread research interests in computer vision practitioners. Their reported empirical success however lacks solid theoretical interpretations for its inherent mechanism. In this paper, we reformulate thei…
Publication bias skews asset pricing research findings.
Paper uses AI methods to forecast Bitcoin prices.
Peer-reviewed research and mined data predict stock returns similarly.
New research shows common ID estimators in neural representations are inaccurate.
New theory explains how self-supervised learning converges, advancing AI research.
Introduces foundation priors for using model-generated data in empirical research.
New datasets improve fairness research by revealing UCI Adult's limitations.
Study finds LLMs hallucinate in finance tasks, needing research.
New research suggests privileged information doesn't improve model performance.
A major line of contemporary research on complex networks is based on the development of statistical models that specify the local motifs associated with macro-structural properties observed in actual networks. This statistical approach becomes increasingly problematic as network size increases. In the context of curre…
Empirical evidence shows that ensembles, such as bagging, boosting, random and rotation forests, generally perform better in terms of their generalization error than individual classifiers. To explain this performance, Schapire et al. (1998) developed an upper bound on the generalization error of an ensemble based on t…
Fundamental portfolio beats market portfolio under certain conditions.
Low-precision training reduces computational cost and produces efficient models. Recent research in developing new low-precision training algorithms often relies on simulation to empirically evaluate the statistical effects of quantization while avoiding the substantial overhead of building specific hardware. To suppor…
learn2learn simplifies meta-learning research by providing a library and standardized interfaces.
This study analyzes financial equity research reports to identify frequently asked questions and automates 80% of them.
Theory and methods to mitigate omitted variable bias in causal machine learning.
This research analyzes deep PDE solvers for option pricing accuracy.
Examines optimal risk sharing with realistic risk attitudes, finding risk seeking in certain subdomains.
In this research, we have established, through empirical testing, a law that relates the number of translating hops to translation accuracy in sequential machine translation in Google Translate. Both accuracy and size decrease with the number of hops; the former displays a decrease closely following a power law. Such a…
The paper examines domain generalization algorithms and finds empirical risk minimization performs well.
Simulation-based inference methods can produce unreliable posterior approximations.
Plotting a learner's average performance against the number of training samples results in a learning curve. Studying such curves on one or more data sets is a way to get to a better understanding of the generalization properties of this learner. The behavior of learning curves is, however, not very well understood and…
Research finds investors may lose from more diverse workplaces.
LLMs help less-resourced researchers access costly data.
This research uses empirical copulas to price quanto options, showing significant differences from traditional models.
We conduct an extensive evaluation of price jump tests based on high-frequency financial data. After providing a concise review of multiple alternative tests, we document the size and power of all tests in a range of empirically relevant scenarios. Particular focus is given to the robustness of test performance to the …
Finding a well-performing architecture is often tedious for both DL practitioners and researchers, leading to tremendous interest in the automation of this task by means of neural architecture search (NAS). Although the community has made major strides in developing better NAS methods, the quality of scientific empiric…
Survey of large language models in financial prediction and trading.
Why do nations produce scientific research? This is a fundamental problem in the field of social studies of science. The paper confronts this question here by showing vital determinants of science to explain the sources of social power and wealth creation by nations. Firstly, this study suggests a new general definitio…
Significant advances have been made in artificial systems by using biological systems as a guide. However, there is often little interaction between computational models for emergent communication and biological models of the emergence of language. Many researchers in language origins and emergent communication take co…
Recent advances in Reinforcement Learning, grounded on combining classical theoretical results with Deep Learning paradigm, led to breakthroughs in many artificial intelligence tasks and gave birth to Deep Reinforcement Learning (DRL) as a field of research. In this work latest DRL algorithms are reviewed with a focus …
Researchers have constantly asked whether stock returns can be predicted by some macroeconomic data. However, it is known that macroeconomic data may exhibit nonstationarity and/or heavy tails, which complicates existing testing procedures for predictability. In this paper we propose novel empirical likelihood methods …
The present research aims to highlight the main factors influencing the development of entrepreneurial innovation in a rural environment and to perform an empirical study with the purpose of assessing the main problems in rural development. The research performed is mostly of a quantitative nature, being based on the u…
Researchers adaptively analyze market regimes to reveal investor behavior shifts.
Causal inference is central to many areas of artificial intelligence, including complex reasoning, planning, knowledge-base construction, robotics, explanation, and fairness. An active community of researchers develops and enhances algorithms that learn causal models from data, and this work has produced a series of im…
News might trigger jump arrivals in financial time series. The "bad" and "good" news seems to have distinct impact. In the research, a double exponential jump distribution is applied to model downward and upward jumps. Bayesian double exponential jump-diffusion model is proposed. Theorems stated in the paper enable est…
Paper proposes a data augmentation method for LLM-generated data in market research.
We offer an experimental benchmark and empirical study for off-policy policy evaluation (OPE) in reinforcement learning, which is a key problem in many safety critical applications. Given the increasing interest in deploying learning-based methods, there has been a flurry of recent proposals for OPE method, leading to …
New research suggests continual learning should focus on both optimization objective and optimization trajectory.
As researchers and practitioners of applied machine learning, we are given a set of requirements on the problem to be solved, the plausibly obtainable data, and the computational resources available. We aim to find (within those bounds) reliably useful combinations of problem, data, and algorithm. An emphasis on algori…
Stochastic gradient descent (SGD), which dates back to the 1950s, is one of the most popular and effective approaches for performing stochastic optimization. Research on SGD resurged recently in machine learning for optimizing convex loss functions and training nonconvex deep neural networks. The theory assumes that on…