This paper reflects on teaching Data Science to diverse students.
problem Adapting learning practices to diverse backgrounds in Data Science.
method Summarizes experiences from teaching a postgraduate Data Science module.
result Draws lessons relevant to teaching Data Science.
Data science teams collaborate extensively, using various tools and stakeholders.
problem Lack of understanding in how data science workers collaborate in practice.
method Conducted an online survey with 183 data science workers.
result Data science teams are highly collaborative and use multiple tools and stakeholders.
Baccalaureate institutions seek to integrate statistics into data science.
problem Graduate statisticians losing ground in data science.
method Reviewing historical contributions of statisticians and calling for integration.
result Baccalaureate institutions need to integrate statistics into data science education.
Foundation models alter medical data science workflow, challenging veridical data science principles.
problem Foundation models disrupt traditional data science practices in medicine.
method Critically examined the medical foundation model lifecycle and its deviation from veridical data science principles.
result Foundation models challenge veridical data science principles of predictability, computability, and stability.
Data science redefines causal inference from observational data, classifying tasks into description, prediction, and counterfactual prediction.
problem Widespread misunderstandings about data science's role in causal inference from observational data.
method Organizing data science tasks into three classes: Description, prediction, and counterfactual prediction (including causal inference).
result The necessity of subject-matter expert knowledge for causal analyses in data science.
Defines data science as a natural ecosystem with challenges and missions.
problem Challenges and missions in data science due to 5D complexities and data life cycle phases.
method Systemic and data-centric view of data science as a fusion of data universe and its challenges, formalizing a general-purpose architecture.
result Essential data science as a natural ecosystem integrating specific disciplines and high-impact applications.
Article calls for data scientists to join ethics debate.
problem Impact of data science on society and ethics.
method Systematic approach from CNIL, established ethical guidelines.
result Data science requires professional ethics to build trust.
This paper explores data science applications in economics using a taxonomy of models and hybrid models showing higher accuracy.
problem Investigating data science applications in economics.
method Systematic literature review using Prisma method.
result Hybrid models showed higher prediction accuracy than other algorithms.
Proposes a methodology to improve data science ROI by addressing key business questions.
problem Companies often fail to maximize data science value, focusing on basic analysis.
method Categorizes and answers 'The Big Three' questions using data science methods.
result Shows how to apply the methodology to real business use cases.
Data science reveals unexpected tax manipulation patterns in Italian regions.
problem Detecting tax income manipulation in Italian municipalities.
method Adopting Benford's first digit law to analyze tax data.
result Marked disparities found in tax distributions across regions, contrary to expectations.
Data science enhances knot theory by analyzing invariant relations.
problem Understanding the complex relations between knot invariants.
method Topological data analysis applied to knot theory.
result New insights into long-standing conjectures about knots.
This study analyzes data science vocabulary changes over 13 years.
problem Understanding evolution of data science terms over time.
method Exploratory Data Analysis, Latent Semantic Analysis, Latent Dirichlet Analysis, N-grams Analysis.
result Identified new vocabulary and its incorporation into scientific literature.
Kan extensions help in data science extrapolation and learning.
problem Generalizing functions over larger sets in data science.
method Kan extensions in category theory applied to data science problems.
result Kan extensions can be used to derive classification and clustering algorithms.
New geometric methods improve optimization and data science problems.
problem Improving optimization and data science problems.
method Geometric tools for high-dimensional optimization and statistical data science.
result New algorithms and statistical guarantees for optimization and data science.
Paper relaxes optimal transport using convex functions for data science.
problem Optimal transport problem on finite spaces.
method Relaxation via strictly convex functions (Kullback-Leibler divergence, Bregman divergences). Gradient descent iterative process.
result Mathematical foundations and iterative process for the relaxed optimal transport problem.
FinTech uses data science and AI to transform finance.
problem Transforming finance with data science and AI.
method DSAI techniques including complex system methods, quantitative methods, etc.
result DSAI enables smart FinTech for various financial sectors.
Data science models, although successful in a number of commercial domains, have had limited applicability in scientific problems involving complex physical phenomena. Theory-guided data science (TGDS) is an emerging paradigm that aims to leverage the wealth of scientific knowledge for improving the effectiveness of da…
Ridge regularization simplifies model complexity in data science.
problem Overfitting in statistical models.
method Adding a penalty on the magnitude of coefficients.
result Effective in reducing model complexity and improving generalization.
Proposes PCS framework for veridical data science results.
problem Ensuring reliable, reproducible, and transparent data science results.
method PCS workflow with predictability, computability, and stability principles.
result PCS inference procedures demonstrate favorable performance in high-dimensional simulations.
ADS automates data preparation for ML/AI, reducing human effort.
problem Manual and time-consuming data preparation for ML/AI.
method Data-driven approach using statistics and ML.
result ADS automates data exploration and processing steps.
The Prescriptive Canvas improves business outcomes by directly prescribing actions based on predictions.
problem Sub-optimal performance in business projects due to a two-step approach of prediction and decision-making.
method The Prescriptive Canvas methodology for framing and communicating actions directly based on predictions.
result Improves framing and communication across stakeholders for successful business impact.
Kaggle chronicles 15 years of competitions, innovation, and data science.
problem Exploring 15 years of data science competitions and innovations.
method Longitudinal trend analysis and exploratory data analysis of millions of kernels and discussion threads.
result Kaggle is a growing platform with diverse use cases and adaptable Kagglers.
Single parameter fits any dataset, simplifying complex data science.
problem Approximating any dataset with a single parameter.
method Adopting chaos theory concepts, adjusting a single real-valued parameter.
result Arbitrary precision fit to all data samples.
Data science reveals co-evolution of income inequality and savings across countries.
problem Understanding the co-evolution of income inequality and savings across countries.
method Time series data for Gini indices and Gross Domestic Savings (% of GDP) were used to construct correlation and similarity matrices, and a multi-dimensional scaling technique was applied. Linear regression was used to test the empirical linkage between income inequality and savings.
result The empirical model proposed by Chakraborti-Chakrabarti (2000) holds reasonably true for many economies of the world, showing a moderate relationship between income inequality and savings.
Data science predicts user interest for midwifery content.
problem Improving midwives' learning and preventing maternal and newborn deaths.
method Forecasting methods using user-generated logs from online learning apps.
result Determining future user interest in midwifery content types.
Causal inference is crucial for understanding data in Data Science.
problem Understanding causal effects in data science, even when data is non-causal.
method Review of causal roadmap, including scientific question, causal model, estimands, statistical estimators, and interpretation.
result Using the causal roadmap framework improves statistical analysis and interpretation in Data Science.
Project analyzes traffic videos to improve Jakarta's safety.
problem Improving traffic safety in Jakarta.
method Developed a pipeline to analyze traffic videos, turning them into usable databases.
result Better understanding of traffic challenges and safety risks.
Pipeline for comparing trading algorithms in finance and crypto.
problem Disconnected research and applications in algorithmic trading.
method General pipeline for designing, programming, and evaluating trading strategies.
result Systematic comparison of trading algorithms in finance and crypto.
Data science principles enhance AI interpretability for better user control.
problem Risks from opaque AI models without clear impacts.
method Synthesizes principles from interpretability literature, emphasizing audience goals.
result Illustrates basic techniques and criteria for evaluating interpretability.
Data science reveals patterns in elliptic curve ranks and coefficients.
problem Understanding rational points on elliptic curves via BSD conjecture.
method Data science, machine learning, topological data analysis.
result Patterns and distributions in rank versus Weierstrass coefficients.
Provides a compendium of data sources for various applications.
problem Lack of comprehensive data sources for data science, machine learning, and AI.
method Compilation of diverse data sources across multiple application areas.
result A comprehensive list of data sources for data scientists and machine learning experts.
Analyzes 6M Python notebooks and 2M enterprise DS pipelines to guide investments in data science.
problem Challenges in following the rapidly evolving landscape of data science technologies and applications.
method Downloaded and analyzed over 6M Python notebooks and 2M enterprise DS pipelines, performing statistical and comparative analyses.
result Identifies actionable conclusions for system builders and technology bets for practitioners based on current trends.
Lecture note for data science students on machine learning basics and advanced topics.
problem No specific problem stated; preparing students for advanced machine learning.
method Introduction of basic machine learning concepts and advanced topics.
result Students are prepared to study advanced machine learning topics.
Machine learning applied to particle physics, focusing on future research and development.
problem Improving particle physics analysis and event identification.
method Roadmap for machine learning applications, resource requirements, and collaborative initiatives.
result Identification of resource needs and areas for external collaboration.
This article guides data scientists on avoiding discrimination in machine learning.
problem Machine learning systems can create or exacerbate societal disparities.
method Provides a taxonomy of practices and measures to mitigate discrimination.
result Data scientists should be intentional about modeling and reducing discriminatory outcomes.
This paper introduces C-DSL to improve data mining outcomes by considering context.
problem Data collection ambiguities, data imbalance, hidden biases, lack of domain info, and data incompleteness.
method Developed Context-Driven Data Science Lifecycle (C-DSL) to address data quality issues.
result Tangible improvements to data mining outcomes were achieved through C-DSL.
A quantum circuit designed for efficient statistical model preparation and training.
problem Challenges in preparing and learning statistical models on quantum processors.
method Utilizes the maximum entropy principle to design a statistics-informed parameterized quantum circuit (SI-PQC).
result Improves trainability and interpretability for learning quantum states and classical model parameters.
Develops efficient algorithms for data science, tackling the curse of dimensionality.
problem Tackles the curse of dimensionality in large datasets.
method Focuses on feature extraction techniques and meta-heuristic algorithms, including evolutionary algorithms.
result Evolutionary algorithms are effective in solving optimization problems with a curse of dimensionality.
The paper uses data science to predict stock trends of Amazon, Apple, Google, and Microsoft.
problem Short-term market movement prediction for major tech stocks.
method Combination of technical analysis and machine/deep learning for trend classification.
result Generated labels for data set: +1 (buy), 0 (hold), -1 (sell).
Data science detects Ethereum honeypots using transaction behavior.
problem Detecting and identifying new Ethereum honeypots.
method Data science approach based on transaction behavior features.
result Data science approach detects new and previously unknown honeypots.
Randomized graph construction ensures giant component with fewer edges.
problem Efficiently constructing sparse graphs with good connectivity.
method Randomly connecting points to a subset of their nearest neighbors.
result A sparser graph with comparable connectivity properties.
The authors seek financial datasets to benchmark feature engineering methods on US market data.
problem Improving predictive models for financial data science competitions.
method Feature engineering methods applied to multivariate time-series data from the US market.
result Predictive power of models tested against Numerai-Signals targets.
OT theory optimizes moving sand piles to solve data science problems.
problem Optimizing the movement of data distributions.
method Comparing and transforming probability distributions to minimize cost.
result OT theory can be applied to various data science problems.
Paper evaluates synthetic retail data for fidelity, utility, and privacy.
problem Ensuring accurate synthetic data in retail.
method Differentiates between continuous and discrete data, measures fidelity and utility, and uses Differential Privacy for privacy.
result Validated framework for reliable and scalable synthetic data evaluation.
New estimator improves mutual information estimation.
problem Estimating mutual information in data science and machine learning.
method Proposes a new estimator that uses a preliminary estimate of the data distribution.
result A preliminary estimate helps in estimating mutual information more accurately.
This paper analyzes social influence using causal data science.
problem Separating genuine causal processes from spurious correlations in social influence data.
method The approach involves partitioning data into groups with minimal contradiction, followed by constrained MLE for causal topology learning.
result The method can retrieve genuine causal arcs and improve influence spread prediction.
Python's tools drive machine learning advancements across industries.
problem Processing and analyzing large data sets for insights.
method Advancements in deep learning, classical ML, and GPU computing.
result Python's dominance in scientific computing boosts machine learning adoption.
Teaches math, physics, and machine learning using Calabi-Yau spaces.
problem Understanding Calabi-Yau spaces in geometry, physics, and machine learning.
method Lecture series, colloquia, and seminars.
result Pedagogical introduction to computational geometry, physical implications, and data science of Calabi-Yau manifolds.