Causal inference from observational data is the goal of many data analyses in the health and social sciences. However, academic statistics has often frowned upon data analyses with a causal objective. The introduction of the term "data science" provides a historic opportunity to redefine data analysis in such a way tha…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
ML methods improve planetary science data analysis.
Diffusion models explained via cognitive science.
Bayesian approach groups observations with similar effects for better inference.
This paper explores a real-world fundamental theme under a data science perspective. It specifically discusses whether fraud or manipulation can be observed in and from municipality income tax size distributions, through their aggregation from citizen fiscal reports. The study case pertains to official data obtained fr…
New method shows data-driven causal studies can be misleading.
This paper improves prediction accuracy for multi-input classification tasks using p-value aggregation.
SLdisco uses supervised learning to discover causal models from observational data.
Develops scalable methods to assess sensitivity and uncertainty in continuous treatment effects.
Citizen science projects are successful at gathering rich datasets for various applications. However, the data collected by citizen scientists are often biased --- in particular, aligned more with the citizens' preferences than with scientific objectives. We propose the Shift Compensation Network (SCN), an end-to-end l…
The paper uses a graph autoencoder to learn unbiased plant-pollinator interaction embeddings.
Matrix completion is a classical problem in data science wherein one attempts to reconstruct a low-rank matrix while only observing some subset of the entries. Previous authors have phrased this problem as a nuclear norm minimization problem. Almost all previous work assumes no explicit structure of the matrix and uses…
We present a novel method for obtaining high-quality, domain-targeted multiple choice questions from crowd workers. Generating these questions can be difficult without trading away originality, relevance or diversity in the answer options. Our method addresses these problems by leveraging a large corpus of domain-speci…
Improved meta-learning for dynamics using additional structured knowledge.
With the large-scale penetration of the internet, for the first time, humanity has become linked by a single, open, communications platform. Harnessing this fact, we report insights arising from a unified internet activity and location dataset of an unparalleled scope and accuracy drawn from over a trillion (1.5$\times…
Causal relationships in time series with latent variables are discovered using LPCMCI.
New model estimates species population trends from citizen science data.
AA extracts archetypes from data for clear feature extraction.
Peer effects, in which the behavior of an individual is affected by the behavior of their peers, are posited by multiple theories in the social sciences. Other processes can also produce behaviors that are correlated in networks and groups, thereby generating debate about the credibility of observational (i.e. nonexper…
Automated detection of new, interesting, unusual, or anomalous images within large data sets has great value for applications from surveillance (e.g., airport security) to science (observations that don't fit a given theory can lead to new discoveries). Many image data analysis systems are turning to convolutional neur…
In this paper, we show how simple logistic growth that was studied intensively during the last 200 years in many domains of science could be extended in a rather simple way and with these extensions is capable to produce a collection of behaviors widely observed in an enormous number of real-life systems in Economics, …
LUQ learns QoI from dynamical systems for consistent observation inversion.
AI boosts study of rare weather extremes with lower costs.
Deep Reinforcement Learning (DRL) has emerged as a powerful control technique in robotic science. In contrast to control theory, DRL is more robust in the thorough exploration of the environment. This capability of DRL generates more human-like behaviour and intelligence when applied to the robots. To explore this capa…
Review of clustering methods for functional data across various fields.
S-DIDML integrates structural DID with ML for causal inference in high-dimensional data.
Blockchain technology, and more specifically Bitcoin (one of its foremost applications), have been receiving increasing attention in the scientific community. The first publications with Bitcoin as a topic, can be traced back to 2012. In spite of this short time span, the production magnitude (1162 papers) makes it nec…
StepMix estimates mixture models with covariates for social science applications.
Symmetric observations don't necessarily imply symmetric causal explanations.
metabeta uses neural networks to speed up Bayesian mixed-effects regression.
Data Science is currently a popular field of science attracting expertise from very diverse backgrounds. Current learning practices need to acknowledge this and adapt to it. This paper summarises some experiences relating to such learning approaches from teaching a postgraduate Data Science module, and draws some learn…
Data science teams collaborate extensively, using various tools and stakeholders.
Machine learning methods have been remarkably successful for a wide range of application areas in the extraction of essential information from data. An exciting and relatively recent development is the uptake of machine learning in the natural sciences, where the major goal is to obtain novel scientific insights and di…
Estimates price sensitivity from transaction data using a novel odds ratio method.
Machine learning improves wildfire science and management, but requires expert knowledge.
Python library for causal discovery from observational data.
The paper proposes using network science to improve portfolio optimization by reducing noise in covariance estimation.
BayesFlow learns complex models using neural networks.
The inference of the causal relationship between a pair of observed variables is a fundamental problem in science, and most existing approaches are based on one single causal model. In practice, however, observations are often collected from multiple sources with heterogeneous causal models due to certain uncontrollabl…
Evaluates six ETSC algorithms on various datasets.
Foundation models alter medical data science workflow, challenging veridical data science principles.
A society or country with income equally distributed among its people is truly a fiction! The phenomena of socioeconomic inequalities have been plaguing mankind from times immemorial. We are interested in gaining an insight about the co-evolution of the countries in the inequality space, from a data science perspective…
Gaussian processes model sparse data in astrophysics and chemistry.
Understanding the nature of dark energy, the mysterious force driving the accelerated expansion of the Universe, is a major challenge of modern cosmology. The next generation of cosmological surveys, specifically designed to address this issue, rely on accurate measurements of the apparent shapes of distant galaxies. H…
Study ranks of elliptic curves via prime averages.
Novel method for Bayesian model comparison using deep learning.
We find polynomial-time solutions to the word problem for free-by-cyclic groups, the word problem for automorphism groups of free groups, and the membership problem for the handlebody subgroup of the mapping class group. All of these results follow from observing that automorphisms of the free group strongly resemble s…
Defines data science as a natural ecosystem with challenges and missions.