Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

0.3%0.5%0.8%0.9% · Aug 201819922001200920182026
48 results for data-science

Foundation models alter medical data science workflow, challenging veridical data science principles.

problem Foundation models disrupt traditional data science practices in medicine.
method Critically examined the medical foundation model lifecycle and its deviation from veridical data science principles.
result Foundation models challenge veridical data science principles of predictability, computability, and stability.

Data science redefines causal inference from observational data, classifying tasks into description, prediction, and counterfactual prediction.

problem Widespread misunderstandings about data science's role in causal inference from observational data.
method Organizing data science tasks into three classes: Description, prediction, and counterfactual prediction (including causal inference).
result The necessity of subject-matter expert knowledge for causal analyses in data science.

Defines data science as a natural ecosystem with challenges and missions.

problem Challenges and missions in data science due to 5D complexities and data life cycle phases.
method Systemic and data-centric view of data science as a fusion of data universe and its challenges, formalizing a general-purpose architecture.
result Essential data science as a natural ecosystem integrating specific disciplines and high-impact applications.

Proposes a methodology to improve data science ROI by addressing key business questions.

problem Companies often fail to maximize data science value, focusing on basic analysis.
method Categorizes and answers 'The Big Three' questions using data science methods.
result Shows how to apply the methodology to real business use cases.

This study analyzes data science vocabulary changes over 13 years.

problem Understanding evolution of data science terms over time.
method Exploratory Data Analysis, Latent Semantic Analysis, Latent Dirichlet Analysis, N-grams Analysis.
result Identified new vocabulary and its incorporation into scientific literature.

Paper relaxes optimal transport using convex functions for data science.

problem Optimal transport problem on finite spaces.
method Relaxation via strictly convex functions (Kullback-Leibler divergence, Bregman divergences). Gradient descent iterative process.
result Mathematical foundations and iterative process for the relaxed optimal transport problem.

The Prescriptive Canvas improves business outcomes by directly prescribing actions based on predictions.

problem Sub-optimal performance in business projects due to a two-step approach of prediction and decision-making.
method The Prescriptive Canvas methodology for framing and communicating actions directly based on predictions.
result Improves framing and communication across stakeholders for successful business impact.

Data science reveals co-evolution of income inequality and savings across countries.

problem Understanding the co-evolution of income inequality and savings across countries.
method Time series data for Gini indices and Gross Domestic Savings (% of GDP) were used to construct correlation and similarity matrices, and a multi-dimensional scaling technique was applied. Linear regression was used to test the empirical linkage between income inequality and savings.
result The empirical model proposed by Chakraborti-Chakrabarti (2000) holds reasonably true for many economies of the world, showing a moderate relationship between income inequality and savings.

Causal inference is crucial for understanding data in Data Science.

problem Understanding causal effects in data science, even when data is non-causal.
method Review of causal roadmap, including scientific question, causal model, estimands, statistical estimators, and interpretation.
result Using the causal roadmap framework improves statistical analysis and interpretation in Data Science.

Analyzes 6M Python notebooks and 2M enterprise DS pipelines to guide investments in data science.

problem Challenges in following the rapidly evolving landscape of data science technologies and applications.
method Downloaded and analyzed over 6M Python notebooks and 2M enterprise DS pipelines, performing statistical and comparative analyses.
result Identifies actionable conclusions for system builders and technology bets for practitioners based on current trends.

Machine learning applied to particle physics, focusing on future research and development.

problem Improving particle physics analysis and event identification.
method Roadmap for machine learning applications, resource requirements, and collaborative initiatives.
result Identification of resource needs and areas for external collaboration.

This article guides data scientists on avoiding discrimination in machine learning.

problem Machine learning systems can create or exacerbate societal disparities.
method Provides a taxonomy of practices and measures to mitigate discrimination.
result Data scientists should be intentional about modeling and reducing discriminatory outcomes.

This paper introduces C-DSL to improve data mining outcomes by considering context.

problem Data collection ambiguities, data imbalance, hidden biases, lack of domain info, and data incompleteness.
method Developed Context-Driven Data Science Lifecycle (C-DSL) to address data quality issues.
result Tangible improvements to data mining outcomes were achieved through C-DSL.

A quantum circuit designed for efficient statistical model preparation and training.

problem Challenges in preparing and learning statistical models on quantum processors.
method Utilizes the maximum entropy principle to design a statistics-informed parameterized quantum circuit (SI-PQC).
result Improves trainability and interpretability for learning quantum states and classical model parameters.

Develops efficient algorithms for data science, tackling the curse of dimensionality.

problem Tackles the curse of dimensionality in large datasets.
method Focuses on feature extraction techniques and meta-heuristic algorithms, including evolutionary algorithms.
result Evolutionary algorithms are effective in solving optimization problems with a curse of dimensionality.

The authors seek financial datasets to benchmark feature engineering methods on US market data.

problem Improving predictive models for financial data science competitions.
method Feature engineering methods applied to multivariate time-series data from the US market.
result Predictive power of models tested against Numerai-Signals targets.

This paper analyzes social influence using causal data science.

problem Separating genuine causal processes from spurious correlations in social influence data.
method The approach involves partitioning data into groups with minimal contradiction, followed by constrained MLE for causal topology learning.
result The method can retrieve genuine causal arcs and improve influence spread prediction.

Python's tools drive machine learning advancements across industries.

problem Processing and analyzing large data sets for insights.
method Advancements in deep learning, classical ML, and GPU computing.
result Python's dominance in scientific computing boosts machine learning adoption.