Foundation models alter medical data science workflow, challenging veridical data science principles.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Data science principles enhance AI interpretability for better user control.
A quantum circuit designed for efficient statistical model preparation and training.
Building and expanding on principles of statistics, machine learning, and scientific inquiry, we propose the predictability, computability, and stability (PCS) framework for veridical data science. Our framework, comprised of both a workflow and documentation, aims to provide responsible, reliable, reproducible, and tr…
The goal of this article is to inspire data scientists to participate in the debate on the impact that their professional work has on society, and to become active in public debates on the digital world as data science professionals. How do ethical principles (e.g., fairness, justice, beneficence, and non-maleficence) …
A framework for network embedding using VDS principles.
Machine learning's data-centric philosophy conflicts with natural sciences' standards.
Review of clustering methods for functional data across various fields.
Survey of de Casteljau's algorithm's applications in geometric data analysis.
New estimator improves statistical validity of synthetic data integration.
Economies are complex man-made systems where organisms and markets interact according to motivations and principles not entirely understood yet. The increasing dissatisfaction with the postulates of traditional economics i.e. perfectly rational agents, interacting through efficient markets in the search of equilibrium,…
This book presents a methodology and philosophy of empirical science based on large scale lossless data compression. In this view a theory is scientific if it can be used to build a data compression program, and it is valuable if it can compress a standard benchmark database to a small size, taking into account the len…
Bayesian methods detect clusters in noisy data more reliably.
Federated learning (FL) is a machine learning setting where many clients (e.g. mobile devices or whole organizations) collaboratively train a model under the orchestration of a central server (e.g. service provider), while keeping the training data decentralized. FL embodies the principles of focused data collection an…
A society or country with income equally distributed among its people is truly a fiction! The phenomena of socioeconomic inequalities have been plaguing mankind from times immemorial. We are interested in gaining an insight about the co-evolution of the countries in the inequality space, from a data science perspective…
Quantum memory limits set by relativity theory.
SoftBart improves BART for high-noise modeling in science.
These notes review six lectures given by Prof. Andrea Montanari on the topic of statistical estimation for linear models. The first two lectures cover the principles of signal recovery from linear measurements in terms of minimax risk. Subsequent lectures demonstrate the application of these principles to several pract…
Causal inference from observational data is the goal of many data analyses in the health and social sciences. However, academic statistics has often frowned upon data analyses with a causal objective. The introduction of the term "data science" provides a historic opportunity to redefine data analysis in such a way tha…
Data Science is currently a popular field of science attracting expertise from very diverse backgrounds. Current learning practices need to acknowledge this and adapt to it. This paper summarises some experiences relating to such learning approaches from teaching a postgraduate Data Science module, and draws some learn…
The Widrow-Hoff rule simplifies language data simulation.
New method falsifies causal graphs using outlier events.
Use simplified layerwise linear models to understand neural dynamics.
Today, the prominence of data science within organizations has given rise to teams of data science workers collaborating on extracting insights from data, as opposed to individual data scientists working alone. However, we still lack a deep understanding of how data science workers collaborate in practice. In this work…
This report aims to improve trust in AI by explaining machine learning models.
Defines data science as a natural ecosystem with challenges and missions.
MissBGM uses AI and Bayesian modeling for better missing data imputation.
ML methods improve planetary science data analysis.
SPREV simplifies visualization of complex labeled datasets.
Data science enhances knot theory by analyzing invariant relations.
Quantum dynamics reveals hidden geometric structure in data.
Some optimization problems coming from the Differential Geometry, as for example, the minimal submanifolds problem and the harmonic maps problem are solved here via interior solutions of appropriate multitime optimal control problems. Section 1 underlines some science domains where appear multitime optimal control prob…
Statistical Machine Learning (SML) refers to a body of algorithms and methods by which computers are allowed to discover important features of input data sets which are often very large in size. The very task of feature discovery from data is essentially the meaning of the keyword `learning' in SML. Theoretical justifi…
FLOWGEM generates complete datasets from incomplete data with non-monotone MAR missingness.
Donoho's JCGS (in press) paper is a spirited call to action for statisticians, who he points out are losing ground in the field of data science by refusing to accept that data science is its own domain. (Or, at least, a domain that is becoming distinctly defined.) He calls on writings by John Tukey, Bill Cleveland, and…
We present MRPC, an R package that learns causal graphs with improved accuracy over existing packages, such as pcalg and bnlearn. Our algorithm builds on the powerful PC algorithm, the canonical algorithm in computer science for learning directed acyclic graphs. The improvement in accuracy results from online control o…
We use the language of uninformative Bayesian prior choice to study the selection of appropriately simple effective models. We advocate for the prior which maximizes the mutual information between parameters and predictions, learning as much as possible from limited data. When many parameters are poorly constrained by …
Set-functions appear in many areas of computer science and applied mathematics, such as machine learning, computer vision, operations research or electrical networks. Among these set-functions, submodular functions play an important role, similar to convex functions on vector spaces. In this tutorial, the theory of sub…
Simplified tutorial on doubly robust learning for causal inference.
Data science models, although successful in a number of commercial domains, have had limited applicability in scientific problems involving complex physical phenomena. Theory-guided data science (TGDS) is an emerging paradigm that aims to leverage the wealth of scientific knowledge for improving the effectiveness of da…
Researchers analyze and compare nonparametric meta-learners for estimating heterogeneous treatment effects.
HierGP improves emulator efficiency for sparse, structured data.
This paper explores data science applications in economics using a taxonomy of models and hybrid models showing higher accuracy.
Modern investigation in economics and in other sciences requires the ability to store, share, and replicate results and methods of experiments that are often multidisciplinary and yield a massive amount of data. Given the increasing complexity and growing interaction across diverse bodies of knowledge it is becoming im…
Machine learning algorithms increasingly influence our decisions and interact with us in all parts of our daily lives. Therefore, just as we consider the safety of power plants, highways, and a variety of other engineered socio-technical systems, we must also take into account the safety of systems involving machine le…
This study analyzes data science vocabulary changes over 13 years.
In this paper we focus on the beneficial role of random strategies in social sciences by means of simple mathematical and computational models. We briefly review recent results obtained by two of us in previous contributions for the case of the Peter principle and the efficiency of a Parliament. Then, we develop a new …
Kan extensions help in data science extrapolation and learning.