Hurd's career overview and publications listed.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We describe the SemEval task of extracting keyphrases and relations between them from scientific documents, which is crucial for understanding which publications describe which processes, tasks and materials. Although this was a new task, we had a total of 26 submissions across 3 evaluation scenarios. We expect the tas…
SentiCite analyzes citations for sentiment and nature, improving on existing methods.
Automatic summarisation is a popular approach to reduce a document to its main arguments. Recent research in the area has focused on neural approaches to summarisation, which can be very data-hungry. However, few large datasets exist and none for the traditionally popular domain of scientific publications, which opens …
Author name disambiguation in bibliographic databases is the problem of grouping together scientific publications written by the same person, accounting for potential homonyms and/or synonyms. Among solutions to this problem, digital libraries are increasingly offering tools for authors to manually curate their publica…
Quantum computing promises new financial modeling.
Blockchain technology, and more specifically Bitcoin (one of its foremost applications), have been receiving increasing attention in the scientific community. The first publications with Bitcoin as a topic, can be traced back to 2012. In spite of this short time span, the production magnitude (1162 papers) makes it nec…
Researchers often summarize their work in the form of posters. Posters provide a coherent and efficient way to convey core ideas from scientific papers. Generating a good scientific poster, however, is a complex and time consuming cognitive task, since such posters need to be readable, informative, and visually aesthet…
New approach uses interpolation models and error bounds for verifiable scientific machine learning.
Paper develops an attention mechanism for long-term scientific impact prediction.
Improved object detection for scientific document images.
A new method validates generative models in high-dimensional data.
This study analyzes data science vocabulary changes over 13 years.
This paper reviews statistical and machine learning methods for anti-money laundering.
The machine learning community adopted the use of null hypothesis significance testing (NHST) in order to ensure the statistical validity of results. Many scientific fields however realized the shortcomings of frequentist reasoning and in the most radical cases even banned its use in publications. We should do the same…
NeuroQuery synthesizes brain mapping evidence across diverse concepts.
Breast cancer is the most common cancer and is the leading cause of cancer death among women worldwide. Detection of breast cancer, while it is still small and confined to the breast, provides the best chance of effective treatment. Computer Aided Detection (CAD) systems that detect cancer from mammograms will help in …
Autonomous driving is getting a lot of attention in the last decade and will be the hot topic at least until the first successful certification of a car with Level 5 autonomy. There are many public datasets in the academic community. However, they are far away from what a robust industrial production system needs. Ther…
Study shows Twitter sentiments predict stock price fluctuations.
Teaches reproducible research to medical students and postgrads.
New benchmarks show LLMs struggle with causal discovery.
Deployment-complete benchmarking assesses if evidence leads to consistent deployment actions.
New dataset for industrial machine sounds to aid maintenance.
Extracts roles of authors from biomedical papers.
Scientific discovery is limited by hypothesis redundancy, and hybrid methods can exploit non-local exploration.
Algorithms learned from data are increasingly used for deciding many aspects in our life: from movies we see, to prices we pay, or medicine we get. Yet there is growing evidence that decision making by inappropriately trained algorithms may unintentionally discriminate people. For example, in automated matching of cand…
Synthetic data can be used to ask more questions and accelerate discovery with provable validity guarantees.
Paper discusses ethical norms for machine learning to prevent misuse.
Peer review is the foundation of scientific publication, and the task of reviewing has long been seen as a cornerstone of professional service. However, the massive growth in the field of machine learning has put this community benefit under stress, threatening both the sustainability of an effective review process and…
Back cover text: Megaprojects and Risk provides the first detailed examination of the phenomenon of megaprojects. It is a fascinating account of how the promoters of multibillion-dollar megaprojects systematically and self-servingly misinform parliaments, the public and the media in order to get projects approved and b…
We summarize a book under publication with his title written by the three present authors, on the theory of Zipf's law, and more generally of power laws, driven by the mechanism of proportional growth. The preprint is available upon request from the authors. For clarity, consistence of language and conciseness, we disc…
This chapter introduces reproducibility in machine learning for medical imaging.
Accurate real time crime prediction is a fundamental issue for public safety, but remains a challenging problem for the scientific community. Crime occurrences depend on many complex factors. Compared to many predictable events, crime is sparse. At different spatio-temporal scales, crime distributions display dramatica…
Benchmarking recursive collapse claims with a new framework under false-positive control.
Scientific publications have evolved several features for mitigating vocabulary mismatch when indexing, retrieving, and computing similarity between articles. These mitigation strategies range from simply focusing on high-value article sections, such as titles and abstracts, to assigning keywords, often from controlled…
While all kinds of mixed data -from personal data, over panel and scientific data, to public and commercial data- are collected and stored, building probabilistic graphical models for these hybrid domains becomes more difficult. Users spend significant amounts of time in identifying the parametric form of the random va…
This work uses scientific constraints to validate neural network predictions in fusion physics.
In recent years, real estate industry has captured government and public attention around the world. The factors influencing the prices of real estate are diversified and complex. However, due to the limitations and one-sidedness of their respective views, they did not provide enough theoretical basis for the fluctuati…
Hypothesis testing is one of the most common types of data analysis and forms the backbone of scientific research in many disciplines. Analysis of variance (ANOVA) in particular is used to detect dependence between a categorical and a numerical variable. Here we show how one can carry out this hypothesis test under the…
Bitcoins and Blockchain technologies are attracting the attention of different scientific communities. In addition, their widespread industrial applications and the continuous introduction of cryptocurrencies are also stimulating the attention of the public opinion. The underlying structure of these technologies consti…
Extracts StarCraft II tournament data for AI and ML studies.
xVal tokenizes numbers continuously for better scientific model training.
Galactica learns from scientific literature to help researchers.
We present a Bayesian non-negative tensor factorization model for count-valued tensor data, and develop scalable inference algorithms (both batch and online) for dealing with massive tensors. Our generative model can handle overdispersed counts as well as infer the rank of the decomposition. Moreover, leveraging a repa…
Data science models, although successful in a number of commercial domains, have had limited applicability in scientific problems involving complex physical phenomena. Theory-guided data science (TGDS) is an emerging paradigm that aims to leverage the wealth of scientific knowledge for improving the effectiveness of da…
The paper explores learning with a mix of private and public data while maintaining privacy.
Social media enhances or diminishes scientific status, depending on usage.
We present a non-naive version of the Precautionary (PP) that allows us to avoid paranoia and paralysis by confining precaution to specific domains and problems. PP is intended to deal with uncertainty and risk in cases where the absence of evidence and the incompleteness of scientific knowledge carries profound implic…