Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

59119178237 · Jun 202019922001200920182026
48 results for molecular regression

Gaussian process regression loses locality in high dimensions, affecting molecular energy surface fitting.

problem Loss of locality in high-dimensional Gaussian process regression.
method Analysis of Matern family kernels and multi-zeta basis functions.
result The property of locality disappears in high dimensions, impacting regression quality.

Multitask Gaussian process regression reduces data generation costs for molecular property prediction.

problem Data bottleneck in training surrogate models for molecular properties.
method Multitask Gaussian process regression over heterogeneous data sources (CC and DFT).
result Predicts at CC-level accuracy with over an order of magnitude reduction in data generation cost.

We introduce multiscale invariant dictionaries to estimate quantum chemical energies of organic molecules, from training databases. Molecular energies are invariant to isometric atomic displacements, and are Lipschitz continuous to molecular deformations. Similarly to density functional theory (DFT), the molecule is re…

2016-05-16abs ↗pdf ↗

Framework learns physics-informed continuum models from molecular data.

problem Discovering accurate and robust data-driven continuum models from molecular simulation data.
method Operator regression framework using neural networks in modal space with physical inductive biases.
result Learned operators generalize to unseen system characteristics.

Paper proposes efficient training for normalizing flows in Boltzmann generators.

problem Training normalizing flows for Boltzmann generators is computationally challenging and unstable.
method Regression Training of Normalizing Flows (RegFlow) using 2\ell_2-regression.
result RegFlow enables efficient and stable training of normalizing flows for Boltzmann generators.

Solves the initial CV problem for molecular simulations using machine learning.

problem Selecting appropriate collective variables for enhancing sampling in molecular simulations.
method Data-driven approach inspired by supervised machine learning (SML).
result Various SML algorithms can be used as initial collective variables (SML_cv) for accelerated sampling.

A set of molecular descriptors whose length is independent of molecular size is developed for machine learning models that target thermodynamic and electronic properties of molecules. These features are evaluated by monitoring performance of kernel ridge regression models on well-studied data sets of small organic mole…

2017-01-23abs ↗pdf ↗

A new method uses IVA to fuse diverse molecular features for better machine learning predictions.

problem Challenges in selecting features for accurate molecular property prediction.
method Independent Vector Analysis (IVA) for fusing multiple molecular feature vectors into a single, compact set.
result Improved prediction performance of regression models for molecular properties.

Persistent homology provides a new, efficient molecular descriptor for protein dynamics.

problem Designing effective molecular descriptors for high-dimensional MD trajectories.
method Introduced masked Flood complex, a protein-tailored modification of simplicial complexes, for persistent homology.
result Persistent homology-based descriptors are competitive across protein dynamics tasks, including frame-level observable regression and MSM estimation.

All SMILES VAE learns molecule latent representations from SMILES strings.

problem Non-unique SMILES strings and high computational cost of graph convolutions hinder VAEs for molecular property optimization.
method Stacked recurrent neural networks encode multiple SMILES strings, pooling hidden representations, and attentional pooling builds a final latent representation.
result All SMILES VAE significantly surpasses state-of-the-art in molecular property optimization tasks.

Spherical CNNs tackle 3D data analysis, especially spherical images.

problem Learning problems involving spherical images, like omnidirectional vision and molecular regression.
method Defined spherical cross-correlation, developed a generalized FFT for efficient computation.
result Demonstrated spherical CNNs' effectiveness in 3D model recognition and atomization energy regression.

A new QSAR model selects relevant molecular descriptors for bioactivity prediction.

problem Redundant, noisy, and irrelevant descriptors in QSAR models.
method SPL-Logsum method using regularization and self-paced learning.
result SPL-Logsum method outperforms other methods in classification performance and model interpretability.

Study evaluates uncertainty quantification methods for molecular property prediction.

problem Uncertainty in neural models for molecular property prediction.
method Systematically evaluated several UQ methods on five benchmark datasets.
result No single method is unequivocally superior, and none provides reliable error ranking across datasets.

Simple linear models outperform complex BO methods in high dimensions.

problem Overcoming the curse of dimensionality in Bayesian optimization.
method Bayesian linear regression with linear kernels, applied to high-dimensional search spaces.
result Simple linear models match or outperform state-of-the-art BO methods in high-dimensional tasks.

Paper develops a method to predict cancer patient survival using molecular profiles.

problem Accurately predicting cancer patient survival with complex survival-molecular profile relationships.
method Kernel Cox partially linear regression with a novel regularized garrotized kernel machine (RegGKM) method.
result The proposed method outperforms other methods in predicting survival accuracy.

Unified theory linking atom-centered and message-passing models for molecular properties.

problem Combining atom-centered and message-passing models for accurate molecular property prediction.
method Generalizing ACDC framework to include multi-centered information, providing a complete linear basis for regression.
result Unified understanding of atom-centered and message-passing models, providing a coherent foundation.

LNK improves uncertainty estimation for molecular dynamics, reducing errors by up to 2.5 times.

problem Uncertainty estimation for molecular force fields to improve model reliability.
method LNK: Gaussian Process-based extension to GNNs addressing six desiderata.
result LNK reduces out-of-equilibrium detection errors by up to 2.5 times compared to existing methods.

MuML models predict molecular dipole moments using atomic partial charges and dipoles.

problem Predicting molecular dipole moments accurately and efficiently.
method Combining atomic partial charges and atomic dipoles within a physically inspired ML model.
result MuML models achieve excellent transferability and accuracy, approaching DFT results at a fraction of the computational cost.

A new method learns graph-level features for drug properties predication.

problem Predicting drug efficacy and toxicity from molecular graphs.
method Introducing a dummy super node connected to all nodes and modifying graph operations to learn graph-level features.
result The method improves molecular properties predication performance on MoleculeNet.

Deep generative models have been wildly successful at learning coherent latent representations for continuous data such as video and audio. However, generative modeling of discrete data such as arithmetic expressions and molecular structures still poses significant challenges. Crucially, state-of-the-art methods often …

2017-03-06abs ↗pdf ↗

A new model designs molecular latent vectors for drug discovery.

problem Designing effective molecular descriptors from molecular structures.
method Proposes a denoising diffusion probabilistic model (DDPM) for variational autoencoding molecular graphs.
result Demonstrates superior prediction performance and robustness compared to existing approaches.

DOCKSTRING simplifies docking simulations for better drug design benchmarks.

problem Lack of meaningful benchmarks for ligand design.
method Open-source Python package for docking scores, extensive dataset, and pharmaceutically-relevant tasks.
result Docking scores are more appropriate benchmarks than simple physicochemical properties.

MoFlow generates chemically valid molecular graphs from latent representations.

problem Generating chemically valid molecular graphs from latent representations is challenging.
method MoFlow uses a flow-based approach with Glow for bond generation and a novel graph conditional flow for atom generation, ensuring chemical validity and efficiency.
result MoFlow achieves state-of-the-art performance in molecular graph generation and optimization.

Study compares GNNs and classical molecular featurisations for molecular property and cliff prediction.

problem Comparing GNNs and classical featurisations for molecular property and cliff prediction.
method Systematic exploration and comparison of PDVs, ECFPs, and GNNs; introduction of substructure pooling.
result Sort & Slice outperforms hash-based folding in ECFP vectorization.

Gemini uses inexpensive measurements to correct biases in expensive property evaluations.

problem Accurate estimation of materials properties using expensive measurements is hindered in scientific discovery campaigns.
method Gemini is a data-driven model that corrects systematic biases between property evaluation methods using inexpensive measurements.
result Gemini reduces the number of expensive evaluations needed for Bayesian optimization in materials discovery.

Novel RL approach for molecular design using quantum mechanics.

problem Existing RL methods for molecular design are limited in scope and reward function.
method Formulation in Cartesian coordinates, direct use of quantum mechanics for reward function, translation and rotation invariant state-action space.
result Agent efficiently learns to solve molecular design tasks from scratch.

Automates molecule design with a novel variational autoencoder.

problem Designing molecules based on specific chemical properties.
method Junction tree variational autoencoder generating tree-structured scaffolds and combining them into molecules.
result Significantly outperforms previous models on molecular generation and optimization tasks.

Bayesian neural networks quantify uncertainties in molecular property predictions.

problem Poor predictions in molecular property predictions due to unreliable training data.
method Bayesian neural networks to estimate model-driven and data-driven uncertainties.
result Uncertainty quantification is necessary for reliable molecular applications.

Graph neural networks improve molecular property prediction.

problem Efficiently predicting molecular properties with high accuracy and scalability.
method Gated Graph Recursive Neural Networks (GGNN) with skip connections.
result GGNN achieves state-of-the-art performance on molecular property prediction benchmarks.

Machine learning improves molecular dynamics simulations by reducing costs and enhancing accuracy.

problem Inaccurate and costly molecular dynamics simulations hinder chemical system description.
method Adaptive sampling of reference data points and machine learning models for predicting molecular properties.
result Machine learning models can predict molecular dipole moments and infrared spectra accurately.

Generative model learns molecular geometry from graph representations.

problem Generating equilibrium states for molecular systems is computationally expensive.
method Probabilistic model based on Euclidean distance geometry.
result Generative model achieves state-of-the-art accuracy in molecular conformation generation.

Improved molecular property prediction using attention and gate mechanisms.

problem Predicting molecular properties from chemical data.
method Attention and gate mechanisms in graph convolutional networks.
result Improved prediction of molecular properties, including photovoltaic efficiency.

AniDS improves molecular force field modeling by learning anisotropic noise.

problem Molecular force field modeling suffers from oversimplified assumptions about atomic motions.
method AniDS introduces anisotropic noise generation for better modeling of directional and structural variability.
result AniDS outperforms existing methods on benchmarks, achieving significant improvements in force prediction accuracy.

Machine learning models simulate molecular spectra and reactions in solvents.

problem Accurate simulation of molecular spectra and reactions in solvent environments.
method Introduced FieldSchNet, a deep neural network for modeling molecular interactions with external fields.
result Demonstrated significant lowering of Claisen rearrangement reaction activation barrier using FieldSchNet.

LaPool improves molecular graph representation learning by capturing interaction importance.

problem Lack of efficient intermediate pooling steps in GNNs leads to poor molecular substructure representation.
method LaPool is a novel, data-driven, and interpretable hierarchical graph pooling method that considers node features and graph structure.
result LaPool outperforms recent GNNs on molecular graph prediction and understanding tasks.