The reproducibility of scientific research has become a point of critical concern. We argue that openness and transparency are critical for reproducibility, and we outline an ecosystem for open and transparent science that has emerged within the human neuroimaging community. We discuss the range of open data sharing re…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Fine-tuned open-source LLMs match or exceed closed-source models in social science research.
ML4Chem offers a user-friendly platform for developing and deploying machine learning models in chemistry.
Mathematical model predicts international trade and global economy dynamics.
Summary of a talk given at the International Seminar "Analysis of spectral invariants and related operator theory", Tokyo University of Science, Unga Campus, 5-6 Oct. 2009
Review of automation's role in chemical discoveries.
High-dimensional statistics advances in complex data domains.
Pipeline for comparing trading algorithms in finance and crypto.
The AI2 Reasoning Challenge (ARC), a new benchmark dataset for question answering (QA) has been recently released. ARC only contains natural science questions authored for human exams, which are hard to answer and require advanced logic reasoning. On the ARC Challenge Set, existing state-of-the-art QA systems fail to s…
The scientific literature is a rich source of information for data mining with conceptual knowledge graphs; the open science movement has enriched this literature with complementary source code that implements scientific models. To exploit this new resource, we construct a knowledge graph using unsupervised learning me…
As machine learning systems become ubiquitous, there has been a surge of interest in interpretable machine learning: systems that provide explanation for their outputs. These explanations are often used to qualitatively assess other criteria such as safety or non-discrimination. However, despite the interest in interpr…
Optimizes control interventions in real-world networks using deep-learning and network science.
Federated learning (FL) is a machine learning setting where many clients (e.g. mobile devices or whole organizations) collaboratively train a model under the orchestration of a central server (e.g. service provider), while keeping the training data decentralized. FL embodies the principles of focused data collection an…
Survey on AI math foundations, focusing on neural networks.
AutoML serves as the bridge between varying levels of expertise when designing machine learning systems and expedites the data science process. A wide range of techniques is taken to address this, however there does not exist an objective comparison of these techniques. We present a benchmark of current open source Aut…
Data science principles enhance AI interpretability for better user control.
We present three case studies of organizations using a data science competition to answer a pressing question. The first is in education where a nonprofit that creates smart school budgets wanted to automatically tag budget line items. The second is in public health, where a low-cost, nonprofit women's health care prov…
Open-FinLLMs tackle financial tasks with multimodal capabilities.
Deep learning excels in AI but struggles with causal physics.
Factor Engine simplifies financial factor computation and analysis in Python.
OpenML is an online platform for open science collaboration in machine learning, used to share datasets and results of machine learning experiments. In this paper we introduce OpenML-Python, a client API for Python, opening up the OpenML platform for a wide range of Python-based tools. It provides easy access to all da…
U-aggregation combines multiple models without labels for better risk prediction.
Julia accelerates machine learning in various fields with balance of efficiency and simplicity.
StepMix estimates mixture models with covariates for social science applications.
Many systems of interest in science and engineering are made up of interacting subsystems. These subsystems, in turn, could be made up of collections of smaller interacting subsystems and so on. In a series of papers David Spivak with collaborators formalized these kinds of structures (systems of systems) as algebras o…
LLMs excel at summarizing and repairing complex models without needing full models.
GRETEL unifies GCE evaluation across various settings.
Faster, more accurate IRT model for large datasets.
Python library for causal discovery from observational data.
Understanding the nature of dark energy, the mysterious force driving the accelerated expansion of the Universe, is a major challenge of modern cosmology. The next generation of cosmological surveys, specifically designed to address this issue, rely on accurate measurements of the apparent shapes of distant galaxies. H…
CGD improves diffusion models' out-of-distribution generalization.
SurvSet offers a repository of 76 T2E datasets for ML benchmarking.
OMLT combines ML and optimization for solving complex problems.
Recent experimental advances in neuroscience have opened new vistas into the immense complexity of neuronal networks. This proliferation of data challenges us on two parallel fronts. First, how can we form adequate theoretical frameworks for understanding how dynamical network processes cooperate across widely disparat…
This work explores using deep NNs to learn quantum systems from probability distributions.
GenSBI offers JAX-native SBI methods for natural sciences.
FCM clustering adapts to persistence diagrams for topological data analysis.
Machine learning impacts computational math, offering new functions approximations.
Proposes LVGP for multi-source data fusion in science and engineering.
With the large-scale penetration of the internet, for the first time, humanity has become linked by a single, open, communications platform. Harnessing this fact, we report insights arising from a unified internet activity and location dataset of an unparalleled scope and accuracy drawn from over a trillion (1.5$\times…
Develops c-GNF for personalized social science policy analysis.
PyODDS is an end-to end Python system for outlier detection with database support. PyODDS provides outlier detection algorithms which meet the demands for users in different fields, w/wo data science or machine learning background. PyODDS gives the ability to execute machine learning algorithms in-database without movi…
Data Science is currently a popular field of science attracting expertise from very diverse backgrounds. Current learning practices need to acknowledge this and adapt to it. This paper summarises some experiences relating to such learning approaches from teaching a postgraduate Data Science module, and draws some learn…
ML methods improve planetary science data analysis.
Review of automation's role in chemical discovery, emphasizing future challenges.
We deliver a call to arms for probabilistic numerical methods: algorithms for numerical tasks, including linear algebra, integration, optimization and solving differential equations, that return uncertainties in their calculations. Such uncertainties, arising from the loss of precision induced by numerical calculation …
Machine learning improves wildfire science and management, but requires expert knowledge.
Causal inference from observational data is the goal of many data analyses in the health and social sciences. However, academic statistics has often frowned upon data analyses with a causal objective. The introduction of the term "data science" provides a historic opportunity to redefine data analysis in such a way tha…