Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,236 papers · 148 categories

Trend · papers per month

3.2%6.3%9.5%12.6% · Jul 202519922001200920182026
48 results for biomolecular simulation

Optimizes biomolecular simulations by ranking adaptive sampling policies.

problem Efficiently sampling biomolecular systems to capture complex dynamical behaviors.
method Metric-driven ranking of adaptive sampling policies to identify the optimal policy for each round.
result Different adaptive sampling policies lead to faster convergence and improved sampling performance.

Tabular in-context learners perform well on biomolecular tasks, but performance depends on the representation used.

problem Predicting biomolecular properties from limited labeled data.
method Evaluating tabular in-context learners on protein fitness regression and small-molecule classification tasks.
result Tabular in-context learners are competitive for protein fitness regression but not for small-molecule classification.

Review of mathematical representations for biomolecular data.

problem Complexity and high dimensionality of biomolecular datasets hinder ML applications.
method Developed low-dimensional and scalable mathematical representations using algebraic topology, differential geometry, and graph theory.
result Mathematical representations improve protein-ligand binding predictions and other biomolecular applications.

Framework integrates Markov and causal models for accurate counterfactual inference.

problem Lack of counterfactual inference in Markov models and identification in causal models.
method Defines structural causal models in terms of Markov process parameters and equilibrium dynamics, enabling consistent counterfactual inference.
result Proposed framework alleviates identifiability issues and improves accuracy of counterfactual inference.

Enhanced diffusion sampling improves rare event sampling in biomolecular simulations.

problem Efficiently sampling rare transition events in biomolecular systems.
method Quantitative steering protocols to generate biased ensembles and exact reweighting.
result Fast, accurate, and scalable estimation of equilibrium properties.

Enhanced diffusion sampling tackles rare event sampling in biomolecular simulations.

problem Efficiently sampling rare transition events in biomolecular simulations.
method Quantitative steering protocols to generate biased ensembles, followed by exact reweighting.
result Fast, accurate, and scalable estimation of equilibrium properties for folding free energies.

This review explores the use of machine learning in discovering collective variables for biomolecular dynamics.

problem Understanding the conformational dynamics and molecular recognition in biomolecules.
method Statistical analysis of high-dimensional spatiotemporal data generated from molecular dynamics simulations.
result Machine learning algorithms can be used to discover abstract collective variables that describe biomolecular dynamics.

Autoencoders discover and accelerate molecular dynamics simulations.

problem Efficient sampling of macromolecular folding landscapes with high free energy barriers.
method Employing auto-associative artificial neural networks to learn nonlinear collective variables (CVs) that are explicit and differentiable functions of atomic coordinates.
result Substantial speedups in exploration of configurational space and discovery of data-driven CVs.

Bayesian model clusters diverse 'omics data for disease subtyping.

problem Clustering diverse 'omics datasets conflates multiple structures.
method Multi-view Bayesian mixture model with semi-supervised learning.
result Identifies distinct clusters of patients for stratified medicine.

A deep learning model organizes RNA graphs to reveal folding patterns and properties.

problem Organizing and understanding the complex folding patterns of RNA secondary structures.
method Geometric scattering autoencoder (GSAE) network for learning graph embeddings.
result GSAE accurately reflects bistable RNA structures and can sample new folding trajectories.

Pipeline learns topological features for protein stability prediction.

problem Predicting protein stability using topological features.
method Data-driven method to learn topological features, comparing with expert features.
result Topological features achieve 92%-99% of SME-based models' performance.

Landmark Diffusion Maps reduce manifold learning complexity for high-volume data streams.

problem Complexity of out-of-sample extensions in manifold learning techniques.
method Landmark Diffusion Maps (L-dMaps) using pruned spanning trees or k-medoids to select landmark points.
result Up to 50-fold speedups in out-of-sample extension with less than 4% errors in manifold reconstruction.

Cryo-electron microscopy (cryo-EM) is an emerging experimental method to characterize the structure of large biomolecular assemblies. Single particle cryo-EM records 2D images (so-called micrographs) of projections of the three-dimensional particle, which need to be processed to obtain the three-dimensional reconstruct…

2013-11-29abs ↗pdf ↗

Generative model tailors anticancer drugs based on transcriptomic data.

problem Designing effective anticancer drugs considering genetic profiles.
method RL framework using pretrained VAEs to generate compounds conditioned on transcriptomic data.
result Generative model produces molecules with high predicted inhibitory effects.

This abstract reviews recent methods for predicting protein-ligand binding affinity.

problem Predicting protein-ligand binding affinity for various applications in life sciences.
method Traditional and deep learning models for binding affinity prediction.
result Improved predictive performance of AI-driven models.

New method for causal discovery in high dimensions with confounder blanket assumption.

problem Inferring causal relationships from observational data in high dimensions.
method Relaxes parametric restrictions and sparsity constraints, focusing on confounder blanket.
result Provable sound and complete structure learning algorithm with finite sample error control.

Unified mathematical theory for analyzing biomolecular geometry and flexibility.

problem Lack of a unified mathematical theory for analyzing biomolecular geometry and flexibility.
method Introducing de Rham-Hodge theory, Helmholtz-Hodge decomposition, and discrete exterior calculus.
result Unified framework for predicting macromolecular flexibility and natural modes.

K-Models clusters functional data with ordinal constraints for better interpretability.

problem Challenges in extracting meaningful insights from functional data due to lack of interpretability.
method Integrates ordinal constraints into clustering to improve interpretability and structure identification.
result Enhances interpretability of clustering results while maintaining performance.

BIDIFAC integrates multi-platform, multi-cohort data for shared and unique patterns.

problem Integration of multi-platform, multi-cohort data for shared and unique patterns.
method BIDIFAC integrates bidimensionally linked matrices into four components: globally shared, row-shared, column-shared, and single-matrix structural components.
result BIDIFAC reveals shared and unique patterns of variability in multi-platform, multi-cohort data.

We introduce a novel class of localized atomic environment representations, based upon the Coulomb matrix. By combining these functions with the Gaussian approximation potential approach, we present LC-GAP, a new system for generating atomic potentials through machine learning (ML). Tests on the QM7, QM7b and GDB9 biom…

2016-11-16abs ↗pdf ↗

Machine learning generates coarse-grained force fields for molecular dynamics.

problem Creating thermodynamically consistent coarse-grained models for larger systems.
method Hybrid architecture using graph neural networks to learn molecular features.
result Framework reproduces thermodynamics for small biomolecular systems.

Develops methods for spectral estimation and rare-event prediction in complex systems.

problem Challenges in understanding dynamics in complex systems with many degrees of freedom.
method Inexact iterative numerical linear algebra methods for spectral estimation and rare-event prediction.
result Demonstrates methods on low-dimensional and high-dimensional models, showing their effectiveness.

Develops an MS-inspired algorithm for regression mode finding and space partitioning.

problem Finding local modes of regression functions and partitioning input space.
method Mean-shift-inspired algorithm for iterative gradient ascent.
result Proves convergence and rates of convergence for estimated local modes.

Paper uses transfer learning and Bayesian optimization to reduce DNA sequence design experiments.

problem Designing many similar DNA sequences for specific applications is expensive and time-consuming.
method Combines transfer learning with Bayesian optimization to reduce experiment count.
result Total number of experiments can be significantly reduced by sharing information between tasks.

Study limits of circadian synchronization under different light signals.

problem Disruption of circadian rhythms due to misalignment with external light signals.
method Matrix-free approach for locating periodic steady states, numerical continuation, bifurcation diagrams, unsupervised learning.
result Limits of circadian synchronization to external light signals of different frequency and duty cycle.

Continuous-depth Evoformer reduces protein folding prediction time and resource usage.

problem Efficient protein structure prediction with reduced computational costs.
method Continuous-depth formulation of Evoformer using Neural Ordinary Differential Equations (Neural ODEs).
result The continuous-time Evoformer achieves constant memory cost and improved efficiency.

Unified framework for sampling and approximating high-dimensional energy landscapes.

problem Sampling and approximating complex energy landscapes in physical systems with constraints and energy barriers.
method Formulates a minimax optimization problem that jointly adapts surrogate approximation and adaptive sampling.
result Demonstrates effectiveness in biomolecular systems with up to 30 collective variables.

In this thesis we present the novel semi-supervised network-based algorithm P-Net, which is able to rank and classify patients with respect to a specific phenotype or clinical outcome under study. The peculiar and innovative characteristic of this method is that it builds a network of samples/patients, where the nodes …

2017-02-04abs ↗pdf ↗

The simulator is an R package that streamlines the process of performing simulations by creating a common infrastructure that can be easily used and reused across projects. Methodological statisticians routinely write simulations to compare their methods to preexisting ones. While developing ideas, there is a temptatio…

2016-06-30abs ↗pdf ↗

A new framework connects machine learning models with simulation models efficiently.

problem Interpreting complex machine learning models for real-world applications.
method Model-bridging framework using kernel mean embeddings.
result Simulations and machine learning models can be used together without high computational costs.

Smartfluidnet accelerates Eulerian fluid simulation with neural networks.

problem Current neural network methods for Eulerian fluid simulation lack flexibility and generalization.
method Smartfluidnet automates model generation and dynamic switching to meet user requirements.
result Smartfluidnet achieves 1.46x and 590x speedup compared to state-of-the-art models, with better simulation quality.

Proposes a new simulator for complex arrival processes.

problem Modeling and simulating complex arrival processes with non-stationary and multi-dimensional rates.
method Integrates Monte Carlo and GANs to model a broad class of arrival processes.
result Consistent and efficient estimation of the simulator using Wasserstein distance.