Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

2855708551,140 · Jun 202019922001200920172026
48 results for Veridical Data Science

Foundation models alter medical data science workflow, challenging veridical data science principles.

problem Foundation models disrupt traditional data science practices in medicine.
method Critically examined the medical foundation model lifecycle and its deviation from veridical data science principles.
result Foundation models challenge veridical data science principles of predictability, computability, and stability.

Building and expanding on principles of statistics, machine learning, and scientific inquiry, we propose the predictability, computability, and stability (PCS) framework for veridical data science. Our framework, comprised of both a workflow and documentation, aims to provide responsible, reliable, reproducible, and tr…

2019-01-23abs ↗pdf ↗

The study examines how language models learn to represent the world, identifying conditions for ecological veridicality.

problem Understanding when language models learn to represent the world accurately and how this learning process can fail.
method Analyzes the Bayes-optimal next-token cross-entropy decomposition and the role of training ecology in shaping model representations.
result The minimum-complexity zero-excess solution is the quotient partition by training equivalence, and this solution is not preserved in in-context learning or per-task adaptation.

PCS-UQ framework improves uncertainty quantification for machine learning models.

problem Ensuring trustworthy uncertainty quantification for machine learning models in high-stakes domains.
method PCS-UQ framework based on Predictability, Computability, and Stability principles, integrating prediction-checking, bootstrap samples, and multiplicative calibration.
result PCS-UQ maintains target coverage while outperforming or matching conformal methods in interval width and subgroup coverage.

Today, the prominence of data science within organizations has given rise to teams of data science workers collaborating on extracting insights from data, as opposed to individual data scientists working alone. However, we still lack a deep understanding of how data science workers collaborate in practice. In this work…

2020-01-18abs ↗pdf ↗

Defines data science as a natural ecosystem with challenges and missions.

problem Challenges and missions in data science due to 5D complexities and data life cycle phases.
method Systemic and data-centric view of data science as a fusion of data universe and its challenges, formalizing a general-purpose architecture.
result Essential data science as a natural ecosystem integrating specific disciplines and high-impact applications.

The goal of this article is to inspire data scientists to participate in the debate on the impact that their professional work has on society, and to become active in public debates on the digital world as data science professionals. How do ethical principles (e.g., fairness, justice, beneficence, and non-maleficence) …

2019-01-14abs ↗pdf ↗

Donoho's JCGS (in press) paper is a spirited call to action for statisticians, who he points out are losing ground in the field of data science by refusing to accept that data science is its own domain. (Or, at least, a domain that is becoming distinctly defined.) He calls on writings by John Tukey, Bill Cleveland, and…

2017-10-24abs ↗pdf ↗

Proposes a methodology to improve data science ROI by addressing key business questions.

problem Companies often fail to maximize data science value, focusing on basic analysis.
method Categorizes and answers 'The Big Three' questions using data science methods.
result Shows how to apply the methodology to real business use cases.

This study analyzes data science vocabulary changes over 13 years.

problem Understanding evolution of data science terms over time.
method Exploratory Data Analysis, Latent Semantic Analysis, Latent Dirichlet Analysis, N-grams Analysis.
result Identified new vocabulary and its incorporation into scientific literature.

Paper relaxes optimal transport using convex functions for data science.

problem Optimal transport problem on finite spaces.
method Relaxation via strictly convex functions (Kullback-Leibler divergence, Bregman divergences). Gradient descent iterative process.
result Mathematical foundations and iterative process for the relaxed optimal transport problem.

New model estimates species population trends from citizen science data.

problem Interannual confounding in citizen science data.
method Double Machine Learning framework to estimate population change and propensity scores for confounding adjustment.
result Spatially detailed trend estimates from citizen science data with low error rates.

Autoencoder learns group representations from actions, improving future prediction accuracy.

problem Learning internal models of interactions with the real world.
method Homomorphism autoencoder with group representation trained on equivariance-derived loss.
result Agents can predict future actions with improved accuracy.

The Prescriptive Canvas improves business outcomes by directly prescribing actions based on predictions.

problem Sub-optimal performance in business projects due to a two-step approach of prediction and decision-making.
method The Prescriptive Canvas methodology for framing and communicating actions directly based on predictions.
result Improves framing and communication across stakeholders for successful business impact.

Machine learning's data-centric philosophy conflicts with natural sciences' standards.

problem Conflict between machine learning's ontology and epistemology and natural sciences' practices.
method Identifying and analyzing contexts where ML can be beneficial or harmful in natural sciences.
result ML can enhance trustworthiness in causal inference but introduces biases in emulation and labeling.

The study explores how machine learning can enhance scientific research.

problem Improving scientific models with machine learning.
method Analysis of data-driven models versus manually added variables in regression.
result Complex models may not always improve over simpler ones in scientific contexts.

SCIENCE improves prediction intervals for individual causal effects.

problem Wide prediction intervals limit practical utility of causal inference.
method Surrogate-assisted conformal inference for efficient individual causal effects.
result SCIENCE produces more efficient prediction intervals for individual causal effects.

The paper uses a graph autoencoder to learn unbiased plant-pollinator interaction embeddings.

problem Sampling bias in citizen science data affects ecological network analysis.
method Bipartite graph variational autoencoder with HSIC for fairness.
result The method mitigates sampling bias and provides unbiased embeddings.