Deep learning predicts protein structures accurately.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Machine learning is often used in virtual screening to find compounds that are pharmacologically active on a target protein. The weave module is a type of graph convolutional deep neural network that uses not only features focusing on atoms alone (atom features) but also features focusing on atom pairs (pair features);…
DiAMoNDBack models protein backmapping from coarse-grained Cα traces.
AniDS improves molecular force field modeling by learning anisotropic noise.
AbDiffuser generates full-atom antibodies with sequence and structure fidelity.
Recent machine learning methods make it possible to model potential energy of atomic configurations with chemical-level accuracy (as calculated from ab-initio calculations) and at speeds suitable for molecular dynam- ics simulation. Best performance is achieved when the known physical constraints are encoded in the mac…
NucleusDiff models atomic nuclei interactions to prevent separation violations in drug design.
New approach uses contrastive learning for better wireless power control.
Improved covariance matrix estimation for portfolio optimization with guaranteed PSD and controlled conditioning.
The paper introduces a statistical test to assess and rank distance measures.
Graph neural networks fail to distinguish certain 3D atom configurations.
Two algorithms estimate Wasserstein distance matrices from few entries for manifold learning.
We find an upper bound for geodesic distances associated to monotone Riemannian metrics on positive definite matrices and density matrices.
In signal analysis and synthesis, linear approximation theory considers a linear decomposition of any given signal in a set of atoms, collected into a so-called dictionary. Relevant sparse representations are obtained by relaxing the orthogonality condition of the atoms, yielding overcomplete dictionaries with an exten…
New 3D protein analysis methods improve accuracy.
Generates valid Euclidean distance matrices for molecular structures.
The density matrices are positively semi-definite Hermitian matrices of unit trace that describe the state of a quantum system. The goal of the paper is to develop minimax lower bounds on error rates of estimation of low rank density matrices in trace regression models used in quantum state tomography (in particular, i…
Proposes a new Sliced-Wasserstein distance for covariance matrices in M/EEG signals.
Efficiently representing real world data in a succinct and parsimonious manner is of central importance in many fields. We present a generalized greedy pursuit framework, allowing us to efficiently solve structured matrix factorization problems, where the factors are allowed to be from arbitrary sets of structured vect…
Paper proposes new methods for improving interatomic potentials.
Study of strictly accretive matrices using Finsler geometry.
Discovery of atomistic systems with desirable properties is a major challenge in chemistry and material science. Here we introduce a novel, autoregressive, convolutional deep neural network architecture that generates molecular equilibrium structures by sequentially placing atoms in three-dimensional space. The model e…
This work incorporates topological features via persistence diagrams to classify point cloud data arising from materials science. Persistence diagrams are multisets summarizing the connectedness and holes of given data. A new distance on the space of persistence diagrams generates relevant input features for a classifi…
Optimal transport is #P-hard when components are independent, even with approximate solutions.
Geodesic distance matrices can reveal shape properties that are largely invariant to non-rigid deformations, and thus are often used to analyze and represent 3-D shapes. However, these matrices grow quadratically with the number of points. Thus for large point sets it is common to use a low-rank approximation to the di…
A set of molecular descriptors whose length is independent of molecular size is developed for machine learning models that target thermodynamic and electronic properties of molecules. These features are evaluated by monitoring performance of kernel ridge regression models on well-studied data sets of small organic mole…
Improves latent variable learning for complex data.
Model tracks structural changes in Brownian particle configurations on a sphere.
Study of metrics on positive-definite matrices from power potential, linking to power means.
We introduce an universum of the Polish (=complete separable metric) space - the convex cone of distance matrices and study its geometry. It happened that the generic Polish spaces in this sense of this universum is so called Urysohn spaces defined by P.S.Urysohn in 20-th, and generic metric triple (= metric space with…
Positive definite matrices abound in a dazzling variety of applications. This ubiquity can be in part attributed to their rich geometric structure: positive definite matrices form a self-dual convex cone whose strict interior is a Riemannian manifold. The manifold view is endowed with a "natural" distance function whil…
We analyze low rank tensor completion (TC) using noisy measurements of a subset of the tensor. Assuming a rank-, order-, tensor where , the best sampling complexity that was achieved is , which is obtained by solving a tensor nuclear-norm minimizatio…
A method for learning embeddings from multi-view data using Gromov-Wasserstein.
DimeNet uses directional message passing to improve molecular predictions.
This paper develops a low-nonnegative-rank approximation method to identify the state aggregation structure of a finite-state Markov chain under an assumption that the state space can be mapped into a handful of meta-states. The number of meta-states is characterized by the nonnegative rank of the Markov transition mat…
Paper introduces a new distance measure for Gaussian Mixture Models.
We detect the backbone of the weighted bipartite network of the Japanese credit market relationships. The backbone is detected by adapting a general method used in the investigation of weighted networks. With this approach we detect a backbone that is statistically validated against a null hypothesis of uniform diversi…
SPIRE enables efficient federated learning for diffusion models by separating client-specific embeddings from a shared backbone.
A new method compares unaligned datasets using log-Euclidean signatures of SPD matrices.
A new algorithm computes elastic shape distances between curves efficiently.
The problem of filtering information from large correlation matrices is of great importance in many applications. We have recently proposed the use of the Kullback-Leibler distance to measure the performance of filtering algorithms in recovering the underlying correlation matrix when the variables are described by a mu…
We propose a non-parametric regression methodology, Random Forests on Distance Matrices (RFDM), for detecting genetic variants associated to quantitative phenotypes representing the human brain's structure or function, and obtained using neuroimaging techniques. RFDM, which is an extension of decision forests, requires…
We study discrete time dynamical systems governed by the state equation . Here are weight matrices, is an activation function, and is the input data. This relation is the backbone of recurrent neural networks (e.g. LSTMs) which have broad applications in sequential learning tasks. …
Using Blanchfield pairings, we show that two Alexander polynomials cannot be realized by a pair of matrices with Gordian distance one if a corresponding quadratic equation does not have an integer solution. We also give an example of how our results help in calculating the Gordian distances, algebraic Gordian distances…
This paper studies convergence behavior of latent mixing measures that arise in finite and infinite mixture models, using transportation distances (i.e., Wasserstein metrics). The relationship between Wasserstein distances on the space of mixing measures and f-divergence functionals such as Hellinger and Kullback-Leibl…
Great computational effort is invested in generating equilibrium states for molecular systems using, for example, Markov chain Monte Carlo. We present a probabilistic model that generates statistically independent samples for molecules from their graph representations. Our model learns a low-dimensional manifold that p…
Transformer-M learns molecular data in 2D or 3D formats.
Let be the set of all density matrices (Hermitian positively semi-definite matrices of unit trace). Consider a problem of estimation of an unknown density matrix based on outcomes of measurements of observables ( bei…