MIM adds indicator variables to improve model performance on incomplete data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Computer science scans LLMs to understand and manipulate their economic forecasts.
Economies are complex man-made systems where organisms and markets interact according to motivations and principles not entirely understood yet. The increasing dissatisfaction with the postulates of traditional economics i.e. perfectly rational agents, interacting through efficient markets in the search of equilibrium,…
Economics does not need a scientific revolution. Economics needs accurate measurements according to high standards of natural sciences and meticulous work on revealing empirical relationships between measured variables.
This book presents a methodology and philosophy of empirical science based on large scale lossless data compression. In this view a theory is scientific if it can be used to build a data compression program, and it is valuable if it can compress a standard benchmark database to a small size, taking into account the len…
metabeta uses neural networks to speed up Bayesian mixed-effects regression.
The paper proposes using network science to improve portfolio optimization by reducing noise in covariance estimation.
Why do nations produce scientific research? This is a fundamental problem in the field of social studies of science. The paper confronts this question here by showing vital determinants of science to explain the sources of social power and wealth creation by nations. Firstly, this study suggests a new general definitio…
A society or country with income equally distributed among its people is truly a fiction! The phenomena of socioeconomic inequalities have been plaguing mankind from times immemorial. We are interested in gaining an insight about the co-evolution of the countries in the inequality space, from a data science perspective…
We propose a novel approach for sampling realistic financial correlation matrices. This approach is based on generative adversarial networks. Experiments demonstrate that generative adversarial networks are able to recover most of the known stylized facts about empirical correlation matrices estimated on asset returns.…
New research shows graph embeddings fail to capture key network properties.
Physicists use quantum models to describe the behavior of physical systems. Quantum models owe their success to their interpretability, to their relation to probabilistic models (quantization of classical models) and to their high predictive power. Beyond physics, these properties are valuable in general data science. …
Empirical analysis is often the first step towards the birth of a conjecture. This is the case of the Birch-Swinnerton-Dyer (BSD) Conjecture describing the rational points on an elliptic curve, one of the most celebrated unsolved problems in mathematics. Here we extend the original empirical approach, to the analysis o…
BL learns interpretable optimization structures from data.
Machine learning algorithms for prediction are increasingly being used in critical decisions affecting human lives. Various fairness formalizations, with no firm consensus yet, are employed to prevent such algorithms from systematically discriminating against people based on certain attributes protected by law. The aim…
Review of automation's role in chemical discovery, emphasizing future challenges.
New estimator improves statistical validity of synthetic data integration.
New method speeds up solving orthogonality constrained problems.
Researchers analyze and compare nonparametric meta-learners for estimating heterogeneous treatment effects.
P.W. Anderson proposed the concept of complexity in order to describe the emergence and growth of macroscopic collective patterns out of the simple interactions of many microscopic agents. In the physical sciences this paradigm was implemented systematically and confirmed repeatedly by successful confrontation with rea…
New method uses Gaussian processes for solving linear PDEs with boundary conditions.
Data Science is currently a popular field of science attracting expertise from very diverse backgrounds. Current learning practices need to acknowledge this and adapt to it. This paper summarises some experiences relating to such learning approaches from teaching a postgraduate Data Science module, and draws some learn…
ML methods improve planetary science data analysis.
Machine learning improves wildfire science and management, but requires expert knowledge.
Evaluates six ETSC algorithms on various datasets.
Causal inference from observational data is the goal of many data analyses in the health and social sciences. However, academic statistics has often frowned upon data analyses with a causal objective. The introduction of the term "data science" provides a historic opportunity to redefine data analysis in such a way tha…
Develops new techniques for learning from sequential data groups.
A physics-based method improves data interpolators and regression tasks.
Today, the prominence of data science within organizations has given rise to teams of data science workers collaborating on extracting insights from data, as opposed to individual data scientists working alone. However, we still lack a deep understanding of how data science workers collaborate in practice. In this work…
Foundation models alter medical data science workflow, challenging veridical data science principles.
Many scientifically well-motivated statistical models in natural, engineering and environmental sciences are specified through a generative process, but in some cases it may not be possible to write down a likelihood for these models analytically. Approximate Bayesian computation (ABC) methods, which allow Bayesian inf…
Max-sliced Wasserstein metric reduces high-dimensional data to 1D for better estimation.
FCM clustering adapts to persistence diagrams for topological data analysis.
Defines data science as a natural ecosystem with challenges and missions.
Paper tackles circularity issues in machine learning predictions.
A new method uses deep learning to evaluate causal theories without strict assumptions.
Data science enhances knot theory by analyzing invariant relations.
Paper uses SGD for solving linear inverse problems, improving empirical performance.
Donoho's JCGS (in press) paper is a spirited call to action for statisticians, who he points out are losing ground in the field of data science by refusing to accept that data science is its own domain. (Or, at least, a domain that is becoming distinctly defined.) He calls on writings by John Tukey, Bill Cleveland, and…
Bayesian machine scientist uncovers accurate models from data.
Data science models, although successful in a number of commercial domains, have had limited applicability in scientific problems involving complex physical phenomena. Theory-guided data science (TGDS) is an emerging paradigm that aims to leverage the wealth of scientific knowledge for improving the effectiveness of da…
The goal of this article is to inspire data scientists to participate in the debate on the impact that their professional work has on society, and to become active in public debates on the digital world as data science professionals. How do ethical principles (e.g., fairness, justice, beneficence, and non-maleficence) …
This paper explores data science applications in economics using a taxonomy of models and hybrid models showing higher accuracy.
We are in the middle of a complex debate as to whether Economics is really a proper natural science. The 'Discussion & Debate' issue of this Euro. Phys. J. Special Topic volume is: 'Can economics be a Physical Science?' I discuss some aspects here.
Develops c-GNF for personalized social science policy analysis.
Complexity science offers new insights into macroeconomics and finance.
Centroid-Encoder reduces high-dimensional data for better visualization.
Federated learning approach for binary matrix factorization.