Paper develops an attention mechanism for long-term scientific impact prediction.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Social media enhances or diminishes scientific status, depending on usage.
Machine learning impacts computational math, offering new functions approximations.
Before retiring, looking back to forty years of writing and publishing scientific papers, I decided to present to the scientific community a selection of my scientific works. I chose mostly articles published in prestigious journals or Proceedings that made a certain impact in the scientific world. I have selected thir…
Machine learning enhances cosmology through new tools and data analysis.
Prior work finds a diversity paradox: diversity breeds innovation, and yet, underrepresented groups that diversify organizations have less successful careers within them. Does the diversity paradox hold for scientists as well? We study this by utilizing a near-population of ~1.2 million US doctoral recipients from 1977…
Measuring the impact of scientific articles is important for evaluating the research output of individual scientists, academic institutions and journals. While citations are raw data for constructing impact measures, there exist biases and potential issues if factors affecting citation patterns are not properly account…
CausalBench aims to advance causal learning research with a transparent platform.
Targeted Learning uses robust statistics for reproducible research.
A research frontier has emerged in scientific computation, wherein numerical error is regarded as a source of epistemic uncertainty that can be modelled. This raises several statistical challenges, including the design of statistical methods that enable the coherent propagation of probabilities through a (possibly dete…
When society maintains a competitive system to promote an abstract goal, competition by necessity relies on imperfect proxy measures. For instance profit is used to measure value to consumers, patient volumes to measure hospital performance, or the Journal Impact Factor to measure scientific value. Here we note that \t…
The dependency structure of credit risk parameters is a key driver for capital consumption and receives regulatory and scientific attention. The impact of parameter imperfections on the quality of expected loss (EL) in the sense of a fair, unbiased estimate of risk expenses however is barely covered. So far there are n…
The scientific data about the state of our planet, presented at the 2012 (Rio+20) summit, documented that today's human family lives even less sustainably than it did in 1992. The data indicate furthermore that the environmental impacts from our current economic activities are so large, that we are approaching situatio…
New method improves spatial prediction validation accuracy.
Clarifies the various fairness definitions in ML.
Confirmation bias leads to biased estimates in noisy data analysis.
Interpretable classification models are built with the purpose of providing a comprehensible description of the decision logic to an external oversight agent. When considered in isolation, a decision tree, a set of classification rules, or a linear model, are widely recognized as human-interpretable. However, such mode…
MNIST-Nd offers synthetic datasets to benchmark clustering across dimensions.
xVal tokenizes numbers continuously for better scientific model training.
Galactica learns from scientific literature to help researchers.
Paper uses SciPhyRL for optimizing large institutional portfolios.
Data science models, although successful in a number of commercial domains, have had limited applicability in scientific problems involving complex physical phenomena. Theory-guided data science (TGDS) is an emerging paradigm that aims to leverage the wealth of scientific knowledge for improving the effectiveness of da…
A textbook on machine learning explaining patterns, predictions, and actions.
Introduces Quantum Data Center for quantum era benefits.
Advanced kernels improve Gaussian process accuracy by incorporating domain knowledge.
Study reveals model misspecification significantly impacts neural SBI algorithms.
Precise scientific analysis in collider-based particle physics is possible because of complex simulations that connect fundamental theories to observable quantities. The significant computational cost of these programs limits the scope, precision, and accuracy of Standard Model measurements and searches for new phenome…
Evidence shows that in a significant number of cases the current methods of research do not allow for reproducible and falsifiable procedures of scientific investigation. As a consequence, the majority of critical decisions at all levels, from personal investment choices to overreaching global policies, rely on some va…
There is significant interest in using modern neural networks for scientific applications due to their effectiveness in modeling highly complex, non-linear problems in a data-driven fashion. However, a common challenge is to verify the scientific plausibility or validity of outputs predicted by a neural network. This w…
Paper proposes a method to estimate scientific parameters in hybrid models without relying on model architecture.
In scientific inference problems, the underlying statistical modeling assumptions have a crucial impact on the end results. There exist, however, only a few automatic means for validating these fundamental modelling assumptions. The contribution in this paper is a general criterion to evaluate the consistency of a set …
AutoSciDACT detects scientific anomalies in noisy data.
MDNs offer a data-efficient alternative to diffusion and flow models for multimodal scientific learning.
Survey of deep learning models for scientific discovery.
New framework compresses and recovers scientific data efficiently.
We present a weakly-supervised data augmentation approach to improve Named Entity Recognition (NER) in a challenging domain: extracting biomedical entities (e.g., proteins) from the scientific literature. First, we train a neural NER (NNER) model over a small seed of fully-labeled examples. Second, we use a reference s…
Hurd's career overview and publications listed.
Paper extends causal inference to non-Euclidean data like images and distributions.
This thesis advances algorithms and software for QMC, GP, and sciML.
New AI approach improves quantum device calibration by leveraging prior scientific discoveries.
Why do nations produce scientific research? This is a fundamental problem in the field of social studies of science. The paper confronts this question here by showing vital determinants of science to explain the sources of social power and wealth creation by nations. Firstly, this study suggests a new general definitio…
Method analyzes hyperparameters using HSIC for better neural network performance.
AI+MPS workshop aims to strengthen AI's role in science.
New ML methods improve physical system understanding by quantifying uncertainty across diverse regimes.
Figures are an important channel for scientific communication, used to express complex ideas, models and data in ways that words cannot. However, this visual information is mostly ignored in analyses of the scientific literature. In this paper, we demonstrate the utility of using scientific figures as markers of knowle…
Paper presents a workflow for reliable unsupervised learning in science.
Develops a Bayesian framework for symbolic regression of scientific expressions.
Machine learning methods have been remarkably successful for a wide range of application areas in the extraction of essential information from data. An exciting and relatively recent development is the uptake of machine learning in the natural sciences, where the major goal is to obtain novel scientific insights and di…