EduQG generates better educational questions by pre-training on scientific text.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Scientific documents rely on both mathematics and text to communicate ideas. Inspired by the topical correspondence between mathematical equations and word contexts observed in scientific texts, we propose a novel topic model that jointly generates mathematical equations and their surrounding text (TopicEq). Using an e…
xVal tokenizes numbers continuously for better scientific model training.
Figures are an important channel for scientific communication, used to express complex ideas, models and data in ways that words cannot. However, this visual information is mostly ignored in analyses of the scientific literature. In this paper, we demonstrate the utility of using scientific figures as markers of knowle…
The scientific literature is a rich source of information for data mining with conceptual knowledge graphs; the open science movement has enriched this literature with complementary source code that implements scientific models. To exploit this new resource, we construct a knowledge graph using unsupervised learning me…
Since time immemorial, people have been looking for ways to organize scientific knowledge into some systems to facilitate search and discovery of new ideas. The problem was partially solved in the pre-Internet era using library classifications, but nowadays it is nearly impossible to classify all scientific and popular…
New method extracts biological concepts from cell microscopy images.
Successful multimodal search and retrieval requires the automatic understanding of semantic cross-modal relations, which, however, is still an open research problem. Previous work has suggested the metrics cross-modal mutual information and semantic correlation to model and predict cross-modal semantic relations of ima…
FreB protocol uses AI to infer hidden parameters with valid confidence regions.
Prior work finds a diversity paradox: diversity breeds innovation, and yet, underrepresented groups that diversify organizations have less successful careers within them. Does the diversity paradox hold for scientists as well? We study this by utilizing a near-population of ~1.2 million US doctoral recipients from 1977…
The study explores how machine learning can enhance scientific research.
In a variety of application domains the content to be recommended to users is associated with text. This includes research papers, movies with associated plot summaries, news articles, blog posts, etc. Recommendation approaches based on latent factor models can be extended naturally to leverage text by employing an exp…
Many research fields codify their findings in standard formats, often by reporting correlations between quantities of interest. But the space of all testable correlates is far larger than scientific resources can currently address, so the ability to accurately predict correlations would be useful to plan research and a…
Historically, games of all kinds have often been the subject of study in scientific works of Computer Science, including the field of machine learning. By using machine learning techniques and applying them to a game with defined rules or a structured dataset, it's possible to learn and improve on the already existing …
A novel multilayer network approach for text analysis.
Irony and sarcasm are two complex linguistic phenomena that are widely used in everyday language and especially over the social media, but they represent two serious issues for automated text understanding. Many labeled corpora have been extracted from several sources to accomplish this task, and it seems that sarcasm …
Due to recent technical and scientific advances, we have a wealth of information hidden in unstructured text data such as offline/online narratives, research articles, and clinical reports. To mine these data properly, attributable to their innate ambiguity, a Word Sense Disambiguation (WSD) algorithm can avoid numbers…
LIMEADE improves AI advice for opaque models, enhancing accuracy and user satisfaction.
NeuroQuery synthesizes brain mapping evidence across diverse concepts.
CNN accurately reconstructs lattice topology with strong thermal fluctuations.
One of the main computational and scientific challenges in the modern age is to extract useful information from unstructured texts. Topic models are one popular machine-learning approach which infers the latent topical structure of a collection of documents. Despite their success --- in particular of its most widely us…
A textbook on machine learning explaining patterns, predictions, and actions.
Natural language processing often involves computations with semantic or syntactic graphs to facilitate sophisticated reasoning based on structural relationships. While convolution kernels provide a powerful tool for comparing graph structure based on node (word) level relationships, they are difficult to customize and…
Many data sets contain rich information about objects, as well as pairwise relations between them. For instance, in networks of websites, scientific papers, and other documents, each node has content consisting of a collection of words, as well as hyperlinks or citations to other nodes. In order to perform inference on…
We describe a new method for visualizing topics, the distributions over terms that are automatically extracted from large text corpora using latent variable models. Our method finds significant -grams related to a topic, which are then used to help understand and interpret the underlying distribution. Compared with …
Hierarchical NMF organizes COVID-19 literature into a searchable tree.
Galactica learns from scientific literature to help researchers.
Data science models, although successful in a number of commercial domains, have had limited applicability in scientific problems involving complex physical phenomena. Theory-guided data science (TGDS) is an emerging paradigm that aims to leverage the wealth of scientific knowledge for improving the effectiveness of da…
Social media enhances or diminishes scientific status, depending on usage.
Leveraging new data sources is a key step in accelerating the pace of materials design and discovery. To complement the strides in synthesis planning driven by historical, experimental, and computed data, we present an automated method for connecting scientific literature to synthesis insights. Starting from natural la…
Back cover text: Megaprojects and Risk provides the first detailed examination of the phenomenon of megaprojects. It is a fascinating account of how the promoters of multibillion-dollar megaprojects systematically and self-servingly misinform parliaments, the public and the media in order to get projects approved and b…
Evidence shows that in a significant number of cases the current methods of research do not allow for reproducible and falsifiable procedures of scientific investigation. As a consequence, the majority of critical decisions at all levels, from personal investment choices to overreaching global policies, rely on some va…
There is significant interest in using modern neural networks for scientific applications due to their effectiveness in modeling highly complex, non-linear problems in a data-driven fashion. However, a common challenge is to verify the scientific plausibility or validity of outputs predicted by a neural network. This w…
Paper proposes a method to estimate scientific parameters in hybrid models without relying on model architecture.
AutoSciDACT detects scientific anomalies in noisy data.
MDNs offer a data-efficient alternative to diffusion and flow models for multimodal scientific learning.
Geometric method captures rare topics and temporal alignment in co-author networks.
Linguistic calibration improves long-form text confidence.
Survey of deep learning models for scientific discovery.
New framework compresses and recovers scientific data efficiently.
Hurd's career overview and publications listed.
This thesis advances algorithms and software for QMC, GP, and sciML.
New AI approach improves quantum device calibration by leveraging prior scientific discoveries.
Originally designed to model text, topic modeling has become a powerful tool for uncovering latent structure in domains including medicine, finance, and vision. The goals for the model vary depending on the application: in some cases, the discovered topics may be used for prediction or some other downstream task. In ot…
Before retiring, looking back to forty years of writing and publishing scientific papers, I decided to present to the scientific community a selection of my scientific works. I chose mostly articles published in prestigious journals or Proceedings that made a certain impact in the scientific world. I have selected thir…
Why do nations produce scientific research? This is a fundamental problem in the field of social studies of science. The paper confronts this question here by showing vital determinants of science to explain the sources of social power and wealth creation by nations. Firstly, this study suggests a new general definitio…
Paper presents a workflow for reliable unsupervised learning in science.
Develops a Bayesian framework for symbolic regression of scientific expressions.