HOoD detects near-out-of-distribution groups in correlated biomedical assays.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We introduce a new approach to variable selection, called Predictive Correlation Screening, for predictor design. Predictive Correlation Screening (PCS) implements false positive control on the selected variables, is well suited to small sample sizes, and is scalable to high dimensions. We establish asymptotic bounds f…
Paper uses transfer learning and Bayesian optimization to reduce DNA sequence design experiments.
Using different sources of information to support automated extracting of relations between biomedical concepts contributes to the development of our understanding of biological systems. The primary comprehensive source of these relations is biomedical literature. Several relation extraction approaches have been propos…
Robust machine learning models improve DNA regulatory sequence prediction under various shifts.
Model predicts future term connections in biomedical research.
Locally sparse neural networks improve interpretability for biomedical tabular data.
BIOMRC dataset improves MRC performance, especially for non-experts.
Enhancement attacks can falsely improve machine learning model performance in biomedical research.
Generates biomedical abstracts from titles, years, and keywords.
Review of modern computational optimal transport methods for biomedical applications.
Biophysical models explain deep learning in gene regulation.
Method integrates logical rules into neural multi-hop reasoning for drug repurposing.
Paper proposes a new method for imputing missing biomedical data.
Entity linking is the task of linking mentions of named entities in natural language text, to entities in a curated knowledge-base. This is of significant importance in the biomedical domain, where it could be used to semantically annotate a large volume of clinical records and biomedical literature, to standardized co…
A deep probabilistic model analyzes DNA-encoded library data for efficient screening.
There is a growing need for fast and accurate methods for testing developmental neurotoxicity across several chemical exposure sources. Current approaches, such as in vivo animal studies, and assays of animal and human primary cell cultures, suffer from challenges related to time, cost, and applicability to human physi…
The support vector machine (SVM) is a widely used machine learning tool for classification based on statistical learning theory. Given a set of training data, the SVM finds a hyperplane that separates two different classes of data points by the largest distance. While the standard form of SVM uses L2-norm regularizatio…
Computational chemists typically assay drug candidates by virtually screening compounds against crystal structures of a protein despite the fact that some targets, like the Opioid Receptor and other members of the GPCR family, traverse many non-crystallographic states. We discover new conformational states of …
Monitoring the biomedical literature for cases of Adverse Drug Reactions (ADRs) is a critically important and time consuming task in pharmacovigilance. The development of computer assisted approaches to aid this process in different forms has been the subject of many recent works. One particular area that has shown pro…
A multi-task learning model for slot tagging in biomedical domains.
Study improves LLMs for PPI analysis by addressing uncertainty.
Paper semantifies bioassay text using neural networks.
Method cleans noisy training labels for biomedical data.
Understanding the morphological changes of primary neuronal cells induced by chemical compounds is essential for drug discovery. Using the data from a single high-throughput imaging assay, a classification model for predicting the biological activity of candidate compounds was introduced. The image recognition model wh…
Deep learning has shown its great promise in various biomedical image segmentation tasks. Existing models are typically based on U-Net and rely on an encoder-decoder architecture with stacked local operators to aggregate long-range information gradually. However, only using the local operators limits the efficiency and…
New DL model handles missing data in biomedical datasets.
Biomedical data are widely accepted in developing prediction models for identifying a specific tumor, drug discovery and classification of human cancers. However, previous studies usually focused on different classifiers, and overlook the class imbalance problem in real-world biomedical datasets. There are a lack of st…
Measurement noise limits the advantage of nonlinear models over linear models in biomedical prediction
spex-LVM infers interpretable latent factors from biomedical data.
Biomedical text tagging systems are plagued by the dearth of labeled training data. There have been recent attempts at using pre-trained encoders to deal with this issue. Pre-trained encoder provides representation of the input text which is then fed to task-specific layers for classification. The entire network is fin…
Proposes a novel method to identify complex effects in multi-view datasets.
Motivation: State-of-the-art biomedical named entity recognition (BioNER) systems often require handcrafted features specific to each entity type, such as genes, chemicals and diseases. Although recent studies explored using neural network models for BioNER to free experts from manual feature engineering, the performan…
Deep learning methods have shown extraordinary potential for analyzing very diverse biomedical data, but their dissemination beyond developers is hindered by important computational hurdles. We introduce ImJoy (https://imjoy.io/), a flexible and open-source browser-based platform designed to facilitate widespread reuse…
VICatMix clusters categorical biomedical data efficiently and selects relevant variables.
Self-supervised method predicts clean signal and noise distribution from noisy images.
Machine learning speeds up FLIM analysis in biomedical research.
SiMLR reduces complex biomedical data into simpler, interpretable forms.
We present an analysis of the problem of identifying biological context and associating it with biochemical events in biomedical texts. This constitutes a non-trivial, inter-sentential relation extraction task. We focus on biological context as descriptions of the species, tissue type and cell type that are associated …
Unified Bayesian model for multi-modal, small sample size biomedical data classification.
MedCAT extracts valuable medical information from unstructured text.
New algorithm improves model generalization in structured biomedical domains.
New algorithm improves causal discovery in biomedical data.
M3E2 neural network estimates multiple treatment effects.
We propose a model for tagging unstructured texts with an arbitrary number of terms drawn from a tree-structured vocabulary (i.e., an ontology). We treat this as a special case of sequence-to-sequence learning in which the decoder begins at the root node of an ontological tree and recursively elects to expand child nod…
In medicine, visualizing chromosomes is important for medical diagnostics, drug development, and biomedical research. Unfortunately, chromosomes often overlap and it is necessary to identify and distinguish between the overlapping chromosomes. A segmentation solution that is fast and automated will enable scaling of co…
Proposes a new Lasso method with performance constraints.
It is common that a trained classification model is applied to the operating data that is deviated from the training data because of noise. This paper demonstrates that an ensemble classifier, Diversified Multiple Tree (DMT), is more robust in classifying noisy data than other widely used ensemble methods. DMT is teste…