The reproducibility of scientific research has become a point of critical concern. We argue that openness and transparency are critical for reproducibility, and we outline an ecosystem for open and transparent science that has emerged within the human neuroimaging community. We discuss the range of open data sharing re…
Fine-tuned open-source LLMs match or exceed closed-source models in social science research.
problem Limited scalability and high costs of large LLMs in social science research.
method Fine-tuning open-source models for specific tasks, exploring training set size effects, proposing hybrid workflow.
result Small, fine-tuned open-source LLMs achieve equal or superior performance to commercial alternatives.
ML4Chem offers a user-friendly platform for developing and deploying machine learning models in chemistry.
problem Developing and deploying machine learning models in chemistry and materials science.
method User-experience design, six core building blocks: data, featurization, models, model optimization, inference, and visualization.
result Ease of use and functionality of the atomistic module for neural networks and kernel ridge regression.
Mathematical model predicts international trade and global economy dynamics.
problem Understanding complex international trade and economy interactions.
method Developed a mathematical model for non-equilibrium processes in open systems.
result Predicted model accurately reflects international trade and economy.
Summary of a talk given at the International Seminar "Analysis of spectral invariants and related operator theory", Tokyo University of Science, Unga Campus, 5-6 Oct. 2009
Review of automation's role in chemical discoveries.
problem Improving autonomous discovery in chemistry.
method Classification of discovery types, assessment of autonomy, case studies.
result Rapid advancements in automation and machine learning are transforming experimentation and modeling.
High-dimensional statistics advances in complex data domains.
problem Complex, rich datasets challenge traditional methods.
method Evolved to address sophisticated estimation and inference problems.
result Deepened connections with optimization, concentration, and information theory.
Pipeline for comparing trading algorithms in finance and crypto.
problem Disconnected research and applications in algorithmic trading.
method General pipeline for designing, programming, and evaluating trading strategies.
result Systematic comparison of trading algorithms in finance and crypto.
The AI2 Reasoning Challenge (ARC), a new benchmark dataset for question answering (QA) has been recently released. ARC only contains natural science questions authored for human exams, which are hard to answer and require advanced logic reasoning. On the ARC Challenge Set, existing state-of-the-art QA systems fail to s…
The scientific literature is a rich source of information for data mining with conceptual knowledge graphs; the open science movement has enriched this literature with complementary source code that implements scientific models. To exploit this new resource, we construct a knowledge graph using unsupervised learning me…
As machine learning systems become ubiquitous, there has been a surge of interest in interpretable machine learning: systems that provide explanation for their outputs. These explanations are often used to qualitatively assess other criteria such as safety or non-discrimination. However, despite the interest in interpr…
Optimizes control interventions in real-world networks using deep-learning and network science.
problem Optimizing control over socioeconomic networks subject to constraints.
method Integrates optimization tools from deep-learning with network science.
result Characterizes vulnerability of corporate networks to takeovers.
Survey on AI math foundations, focusing on neural networks.
problem Lack of rigorous mathematical foundation for AI.
method Survey and discussion of theoretical directions in AI.
result Discussion of open problems in AI math.
AutoML serves as the bridge between varying levels of expertise when designing machine learning systems and expedites the data science process. A wide range of techniques is taken to address this, however there does not exist an objective comparison of these techniques. We present a benchmark of current open source Aut…
Data science principles enhance AI interpretability for better user control.
problem Risks from opaque AI models without clear impacts.
method Synthesizes principles from interpretability literature, emphasizing audience goals.
result Illustrates basic techniques and criteria for evaluating interpretability.
We present three case studies of organizations using a data science competition to answer a pressing question. The first is in education where a nonprofit that creates smart school budgets wanted to automatically tag budget line items. The second is in public health, where a low-cost, nonprofit women's health care prov…
Federated learning collaborates clients to train models without sharing data.
problem Privacy and data sharing in machine learning.
method Central server orchestrates collaborative training of models on decentralized data.
result Recent advances and open problems in FL.
Open-FinLLMs tackle financial tasks with multimodal capabilities.
problem Financial LLMs lack multimodal capabilities and real-world applicability.
method Developed Open-FinLLMs, an open-source multimodal financial LLM suite.
result Open-FinLLMs outperform advanced financial and general LLMs in diverse tasks.
Deep learning excels in AI but struggles with causal physics.
problem Deep learning struggles with causal relationships in physical sciences.
method Combining Bayesian methods, physical constraints, and causal models.
result Deep learning can mislead in systems with unclear causal relationships.
Factor Engine simplifies financial factor computation and analysis in Python.
problem Efficient computation and analysis of financial factors.
method Modular, extensible Python library with decorators, integrates with data science ecosystem.
result Mispricing factors computed by Factor Engine and Stata implementation are highly similar.
OpenML is an online platform for open science collaboration in machine learning, used to share datasets and results of machine learning experiments. In this paper we introduce OpenML-Python, a client API for Python, opening up the OpenML platform for a wide range of Python-based tools. It provides easy access to all da…
U-aggregation combines multiple models without labels for better risk prediction.
problem Challenges in selecting best model for new populations due to limited data and lack of true labels.
method U-aggregation, an unsupervised model aggregation method that integrates pre-trained models without observed labels.
result U-aggregation improves genetic risk prediction of complex traits using publicly available models.
Julia accelerates machine learning in various fields with balance of efficiency and simplicity.
problem Efficiency and simplicity in machine learning algorithms.
method Developed and applied Julia language in machine learning.
result Julia balances efficiency and simplicity for machine learning.
StepMix estimates mixture models with covariates for social science applications.
problem Estimating latent classes with covariates in social science models.
method Pseudo-likelihood estimation using one-, two-, and three-step approaches.
result Unified framework for expectation-maximization subroutines.
Many systems of interest in science and engineering are made up of interacting subsystems. These subsystems, in turn, could be made up of collections of smaller interacting subsystems and so on. In a series of papers David Spivak with collaborators formalized these kinds of structures (systems of systems) as algebras o…
LLMs excel at summarizing and repairing complex models without needing full models.
problem Understanding and repairing complex models like GAMs.
method Hierarchical reasoning and extensive background knowledge.
result LLMs can detect anomalies, describe reasons, and suggest repairs.
GRETEL unifies GCE evaluation across various settings.
problem Lack of standardized evaluation for Graph Counterfactual Explanations.
method Unified framework for testing GCE methods in diverse settings.
result GRETEL promotes reproducible evaluations of GCE techniques.
AutoAIViz visualizes AutoAI's model generation process, improving user trust.
problem Limited transparency in AutoAI systems leads to lack of user understanding and trust.
method Developed and evaluated AutoAIViz, an experimental system that visualizes AutoAI's model generation process.
result AutoAIViz helps users complete data science tasks and increases their understanding, improving trust in AutoAI systems.
Faster, more accurate IRT model for large datasets.
problem Speed and accuracy challenges in fitting IRT models to large datasets.
method Variational Bayesian inference for IRT models.
result Higher log likelihoods and improved missing data imputation.
Python library for causal discovery from observational data.
problem Revealing causal relations from observational data.
method Comprehensive collection of causal discovery methods in Python.
result Ease of use for non-specialists and modular building blocks for developers.
Understanding the nature of dark energy, the mysterious force driving the accelerated expansion of the Universe, is a major challenge of modern cosmology. The next generation of cosmological surveys, specifically designed to address this issue, rely on accurate measurements of the apparent shapes of distant galaxies. H…
CGD improves diffusion models' out-of-distribution generalization.
problem Reliable sampling from high-value regions beyond training data.
method Context-guided diffusion (CGD) using unlabeled data and smoothness constraints.
result Substantial performance gains across various diffusion processes.
SurvSet offers a repository of 76 T2E datasets for ML benchmarking.
problem Lack of open-source T2E dataset repositories for ML benchmarking.
method Consistent formatting of datasets for various ML algorithms and statistical methods.
result SurvSet provides a comprehensive resource for T2E analysis.
OMLT combines ML and optimization for solving complex problems.
problem Solving complex decision-making problems in computer science and engineering.
method OMLT integrates neural networks and gradient-boosted trees into optimization problems using machine learning.
result OMLT seamlessly integrates with Pyomo and solves real-world problems.
Recent experimental advances in neuroscience have opened new vistas into the immense complexity of neuronal networks. This proliferation of data challenges us on two parallel fronts. First, how can we form adequate theoretical frameworks for understanding how dynamical network processes cooperate across widely disparat…
This work explores using deep NNs to learn quantum systems from probability distributions.
problem Learning quantum systems from limited probability distribution data.
method Using deep neural networks to reconstruct quantum Hamiltonian from probability distributions.
result Deep neural networks can learn quantum Hamiltonians from probability distributions.
GenSBI offers JAX-native SBI methods for natural sciences.
problem Lack of native SBI libraries in JAX for natural sciences.
method Flow and diffusion generative models in JAX.
result Near-ideal mean C2ST scores on SBIBM tasks.
FCM clustering adapts to persistence diagrams for topological data analysis.
problem Integrating topological data into machine learning workflows.
method Adapting Fuzzy c-Means to persistence diagrams.
result FCM clustering captures topological structure without additional processing.
Machine learning impacts computational math, offering new functions approximations.
problem Machine learning's black box nature hinders further progress in computational math.
method Analyzes machine learning's impact on computational math and vice versa.
result Integrating computational math with machine learning can enhance both fields.
Proposes LVGP for multi-source data fusion in science and engineering.
problem Differences in quality and comprehensiveness of data sources.
method Latent Variable Gaussian Process (LVGP) framework.
result Improved predictions for sparse-data problems.
With the large-scale penetration of the internet, for the first time, humanity has become linked by a single, open, communications platform. Harnessing this fact, we report insights arising from a unified internet activity and location dataset of an unparalleled scope and accuracy drawn from over a trillion (1.5$\times…
Develops c-GNF for personalized social science policy analysis.
problem Challenges in estimating causal effects and counterfactual inference in social sciences.
method causal-Graphical Normalizing Flow (c-GNF) method.
result c-GNF performs well in estimating causal effects and counterfactual inference.
PyODDS is an end-to end Python system for outlier detection with database support. PyODDS provides outlier detection algorithms which meet the demands for users in different fields, w/wo data science or machine learning background. PyODDS gives the ability to execute machine learning algorithms in-database without movi…
Data Science is currently a popular field of science attracting expertise from very diverse backgrounds. Current learning practices need to acknowledge this and adapt to it. This paper summarises some experiences relating to such learning approaches from teaching a postgraduate Data Science module, and draws some learn…
Data science teams collaborate extensively, using various tools and stakeholders.
problem Lack of understanding in how data science workers collaborate in practice.
method Conducted an online survey with 183 data science workers.
result Data science teams are highly collaborative and use multiple tools and stakeholders.
ML methods improve planetary science data analysis.
problem Insufficient use of ML in planetary science.
method Ten recommendations for integrating ML in planetary science.
result Expanding planetary science insights from large datasets.
Review of automation's role in chemical discovery, emphasizing future challenges.
problem Improving automation's contribution to chemical discovery.
method Analysis of exemplary studies and open research directions.
result Future autonomous systems need improvement in data handling, model building, and experiment automation.
We deliver a call to arms for probabilistic numerical methods: algorithms for numerical tasks, including linear algebra, integration, optimization and solving differential equations, that return uncertainties in their calculations. Such uncertainties, arising from the loss of precision induced by numerical calculation …