Donoho's JCGS (in press) paper is a spirited call to action for statisticians, who he points out are losing ground in the field of data science by refusing to accept that data science is its own domain. (Or, at least, a domain that is becoming distinctly defined.) He calls on writings by John Tukey, Bill Cleveland, and…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New geometric methods improve optimization and data science problems.
Data science enhances knot theory by analyzing invariant relations.
Foundation models alter medical data science workflow, challenging veridical data science principles.
Ridge regularization simplifies model complexity in data science.
The goal of this article is to inspire data scientists to participate in the debate on the impact that their professional work has on society, and to become active in public debates on the digital world as data science professionals. How do ethical principles (e.g., fairness, justice, beneficence, and non-maleficence) …
This guide explains statistical distances for evaluating generative models.
A quantum circuit designed for efficient statistical model preparation and training.
Responds to critiques on tests for causal parameter confidence intervals.
New method uses imperfect LLM annotations for valid statistical inference in social science.
The need for new methods to deal with big data is a common theme in most scientific fields, although its definition tends to vary with the context. Statistical ideas are an essential part of this, and as a partial response, a thematic program on statistical inference, learning, and models in big data was held in 2015 i…
Causal inference from observational data is the goal of many data analyses in the health and social sciences. However, academic statistics has often frowned upon data analyses with a causal objective. The introduction of the term "data science" provides a historic opportunity to redefine data analysis in such a way tha…
This paper examines risks and uncertainties of changing data sources in machine learning for official statistics.
This article provides the role of big idea statisticians in future of Big Data Science. We describe the `United Statistical Algorithms' framework for comprehensive unification of traditional and novel statistical methods for modeling Small Data and Big Data, especially mixed data (discrete, continuous).
Recent experimental advances in neuroscience have opened new vistas into the immense complexity of neuronal networks. This proliferation of data challenges us on two parallel fronts. First, how can we form adequate theoretical frameworks for understanding how dynamical network processes cooperate across widely disparat…
This is a review article for Encyclopedia of Complexity and System Science, to be published by Springer http://refworks.springer.com/complexity/. The paper reviews statistical models for money, wealth, and income distributions developed in the econophysics literature since late 1990s.
Maximum likelihood estimation and a test of fit based on the Anderson-Darling statistic is presented for the case of the power law distribution when the parameters are estimated from a left-censored sample. Expressions for the maximum likelihood estimators and tables of asymptotic percentage points for the A^2 statisti…
Teaches deep learning to statisticians.
High-dimensional statistics advances in complex data domains.
We discuss several multi-agent models that have their origin in the kinetic exchange theory of statistical mechanics and have been recently applied to a variety of problems in the social sciences. This class of models can be easily adapted for simulations in areas other than physics, such as the modeling of income and …
Statistical test evaluates if personalizing interventions is cost-effective.
Kan extensions help in data science extrapolation and learning.
Study designs statistical inference for collaborative science teams.
Review of Gerber-Shiu function for practical actuarial science.
New estimator improves statistical validity of synthetic data integration.
In these notes we describe heuristics to predict computational-to-statistical gaps in certain statistical problems. These are regimes in which the underlying statistical problem is information-theoretically possible although no efficient algorithm exists, rendering the problem essentially unsolvable for large instances…
This paper uses information theory to improve risk modeling in big data.
New method shows data-driven causal studies can be misleading.
New method improves active statistical inference by reducing noise.
Bayesian framework explains diverse explanatory values.
Building and expanding on principles of statistics, machine learning, and scientific inquiry, we propose the predictability, computability, and stability (PCS) framework for veridical data science. Our framework, comprised of both a workflow and documentation, aims to provide responsible, reliable, reproducible, and tr…
Spectral methods simplify data analysis, improving accuracy and stability.
Many questions in Data Science are fundamentally causal in that our objective is to learn the effect of some exposure, randomized or not, on an outcome interest. Even studies that are seemingly non-causal, such as those with the goal of prediction or prevalence estimation, have causal elements, including differential c…
In the quest to align deep learning with the sciences to address calls for rigor, safety, and interpretability in machine learning systems, this contribution identifies key missing pieces: the stages of hypothesis formulation and testing, as well as statistical and systematic uncertainty estimation -- core tenets of th…
Statistical Machine Learning (SML) refers to a body of algorithms and methods by which computers are allowed to discover important features of input data sets which are often very large in size. The very task of feature discovery from data is essentially the meaning of the keyword `learning' in SML. Theoretical justifi…
Pipeline for comparing trading algorithms in finance and crypto.
Machine learning's data-centric philosophy conflicts with natural sciences' standards.
We define and study the statistical models in exponential family form whose sufficient statistics are the degree distributions and the bi-degree distributions of undirected labelled simple graphs. Graphs that are constrained by the joint degree distributions are called -graphs in the computer science literature and…
Conversion of raw data into insights and knowledge requires substantial amounts of effort from data scientists. Despite breathtaking advances in Machine Learning (ML) and Artificial Intelligence (AI), data scientists still spend the majority of their effort in understanding and then preparing the raw data for ML/AI. Th…
MIC consistently estimates dependence in large datasets.
AI improves precision health through adaptive interventions.
New method improves local precipitation predictions using video diffusion.
The paper proposes using network science to improve portfolio optimization by reducing noise in covariance estimation.
Networks are ubiquitous in science and have become a focal point for discussion in everyday life. Formal statistical models for the analysis of network data have emerged as a major topic of interest in diverse areas of study, and most of these involve a form of graphical representation. Probability models on graphs dat…
BayesFlow learns complex models using neural networks.
A physics-based method improves data interpolators and regression tasks.
This article is the rejoinder for the paper "Probabilistic Integration: A Role in Statistical Computation?" to appear in Statistical Science with discussion. We would first like to thank the reviewers and many of our colleagues who helped shape this paper, the editor for selecting our paper for discussion, and of cours…
New method for scalable inference in large-scale regression models with complex error structures.