This review explores the use of machine learning in discovering collective variables for biomolecular dynamics.
problem Understanding the conformational dynamics and molecular recognition in biomolecules.
method Statistical analysis of high-dimensional spatiotemporal data generated from molecular dynamics simulations.
result Machine learning algorithms can be used to discover abstract collective variables that describe biomolecular dynamics.
Modeling business cycles via collective risk fluctuations in economic agents' risk space.
problem Understanding and predicting business cycles through economic agents' risk dynamics.
method Continuous numerical risk grades for economic agents, modeling collective economic variables and flows as functions of risk coordinates, deriving equations for their evolution.
result Business and credit cycles are explained as fluctuations of collective economic variables and their mean risks in the risk space of economic agents.
Current economic theories miss most of economic dynamics.
problem Accuracy of economic theories and policies depend on economic variables and processes.
method Identify and analyze overlooked economic variables and processes.
result Many economic variables and processes not accounted for in current theories.
Generative model learns conditional distributions on collective variable levels.
problem Modeling conditional probability distributions on collective variable levels.
method General and efficient learning approach, data enrichment strategy.
result Effective generative models on different level-sets of collective variables.
ISOKANN learns collective variables and effective dynamics for metastable transitions.
problem Understanding metastable transitions in complex molecular systems.
method Integrates Koopman operators with neural networks to extract CVs and effective dynamics.
result Reconstructs coarse-grained kinetics and reproduces transition times across barriers.
New method learns collective variables using autoencoders for molecular simulations.
problem Learning low-dimensional slow degrees of freedom (collective variables) for molecular simulations.
method Iterative method involving CV learning with autoencoders and reweighting scheme.
result Achieves convergence of learned collective variables.
A neural network RG approach for efficient collective variable identification.
problem Identifying mutually independent collective variables in complex systems.
method A variational RG approach using normalizing flows and neural nets to map physical configurations to latent variables with reduced mutual information.
result Direct access to the renormalized energy function of latent variables for unbiased training and efficient sampling.
Autoencoders discover and accelerate molecular dynamics simulations.
problem Efficient sampling of macromolecular folding landscapes with high free energy barriers.
method Employing auto-associative artificial neural networks to learn nonlinear collective variables (CVs) that are explicit and differentiable functions of atomic coordinates.
result Substantial speedups in exploration of configurational space and discovery of data-driven CVs.
Scientific and business practices are increasingly resulting in large collections of randomized experiments. Analyzed together, these collections can tell us things that individual experiments in the collection cannot. We study how to learn causal relationships between variables from the kinds of collections faced by m…
Solves the initial CV problem for molecular simulations using machine learning.
problem Selecting appropriate collective variables for enhancing sampling in molecular simulations.
method Data-driven approach inspired by supervised machine learning (SML).
result Various SML algorithms can be used as initial collective variables (SML_cv) for accelerated sampling.
Proposes a new model for joint probability distributions in computer vision.
problem Limitation of existing models in meeting diverse downstream tasks.
method Uses parametric conditional probability distributions for each group of variables conditioned on the rest.
result Models can be used for any downstream task without task-specific design.
Proposes a method for selecting important variables in high-dimensional data.
problem High-dimensional classification problems with many noise variables.
method Probability-based nonparametric multiple-class classification method with variable selection.
result The method can have prediction power similar to Bayes rule and retains interpretability.
Time-lagged autoencoders improve molecular dynamics data analysis.
problem Analyzing slow collective variables in molecular kinetics.
method Modified autoencoder neural network for dimension reduction.
result Time-lagged autoencoders reliably capture slow dynamics.
Bayesian models discover CVs for complex systems, enhancing sampling methods.
problem Limitations in modeling complex systems in biochemistry and materials science.
method Formulated CV discovery as a Bayesian inference problem, using deep learning and variational inference.
result Discovered CVs improve predictive ability for alanine dipeptide and ALA-15 peptides.
Recourse explanations can become invalid if collective actions change statistical data.
problem Recourse explanations may become invalid due to collective behavior changing data statistics.
method Formal characterization of conditions under which recourse explanations remain valid under performativity.
result Recourse actions may become invalid if they are influenced by or intervene on non-causal variables.
This paper proposes a general adaptive procedure for budget-limited predictor design in high dimensions called two-stage Sampling, Prediction and Adaptive Regression via Correlation Screening (SPARCS). SPARCS can be applied to high dimensional prediction problems in experimental science, medicine, finance, and engineer…
New framework models uncertainty with uncertainty variables.
problem Modeling uncertainty in state estimation and inference.
method Developed uncertainty variables, sets, and graphical models.
result Preserves independence properties and builds useful concepts.
Develops a framework for analyzing multi-agent and many-body systems with feedback loops.
problem Optimal order of multi-agent and general many-body systems
method Derive macroscopic properties and optimal degree of order
result Optimal degree of order balances productivity, stability, and adaptability
New CVs preserve transition rates in molecular dynamics.
problem Designing CVs that accurately capture rare events in high-dimensional systems.
method Integrating manifold learning and group-invariant featurization to construct neural network-based CVs that satisfy orthogonality conditions.
result Achieved a CV for butane that reproduces the anti-gauche transition rate with less than ten percent relative error.
Optimizes feature selection for molecular kinetics models.
problem Selecting interpretable features for kinetic models from molecular dynamics simulations.
method Direct optimization of features without constructing a full kinetic model.
result Direct feature optimization leads to more efficient model selection.
A method for collecting human supervision that combines rules and instance labels.
problem Lack of labeled data and inefficient human supervision.
method Rule-exemplar method with training algorithm for joint denoising and model training.
result Our algorithm is more accurate than existing methods and effectively denoises rules.
Develops variable-lag Granger causality for more accurate time series analysis.
problem Fixed time delay assumption in Granger causality does not fit many real-world applications.
method Variable-lag Granger causality, inferring with arbitrary time delays.
result Performs better than existing methods in coordinated collective behavior studies.
Active learning method for binary classification with variable selection.
problem Efficiently label subjects in large datasets for binary classification.
method Model-based active learning with sequential variable selection.
result Proposed procedure reduces training cost/time and improves classification model.
Extracts intrinsic spatial coordinates for complex agent systems to learn PDEs.
problem Modeling collective dynamics of heterogeneous agents.
method Data-driven extraction of intrinsic spatial coordinates, learning PDEs in emergent space.
result Collective dynamics can be approximated through learned PDEs in emergent coordinates.
Enhances diffusion-based sampling for molecular systems.
problem Inefficiency and thermodynamic mode miss in diffusion-based samplers for molecular systems.
method Introduces a sequential bias along collective variables (CVs) to encourage exploration and increase temperature in the projected space.
result Improves efficiency, mode discovery, and free energy estimation; first to demonstrate reactive sampling.
If the twist numbers of a collection of oriented alternating link diagrams are bounded, then the Alexander polynomials of the corresponding links have bounded euclidean Mahler measure (see Definition 1.2). The converse assertion does not hold. Similarly, if a collection of oriented link diagrams, not necessarily altern…
Paper proposes a deep RL method for hedging variable annuities, outperforming misspecified models.
problem Model miscalibration in variable annuity contracts with GMMB and GMDB riders.
method Two-phase deep reinforcement learning approach: training phase in a controlled environment, online learning phase in real market.
result Trained reinforcement learning agent hedges equally well as correct Delta in training phase and outperforms misspecified Deltas.
Inference and learning of graphical models are both well-studied problems in statistics and machine learning that have found many applications in science and engineering. However, exact inference is intractable in general graphical models, which suggests the problem of seeking the best approximation to a collection of …
Study proposes a method for identifying important variables in multi-class classification problems.
problem Lack of studies on variable selection in nonparametric classification models, especially for multi-class problems.
method Sparse non-parametric density estimation approach for identifying high impacts variables.
result Proposed method identifies important variables for each class in multi-class classification problems.
Transfer neural networks for efficient protein dynamics sampling.
problem Efficiently sampling protein dynamics in related systems.
method Variational auto-encoder framework with latent embedding for collective variable.
result Transferable model trained on one protein can efficiently sample related mutants.
New pruning rules reduce search space for Bayesian network learning.
problem Learning Bayesian networks efficiently with BIC score.
method Entropy-based pruning rules to reduce candidate parent sets.
result Significant gains in structure learning with low computational cost.
SRV learns slow molecular modes from simulations.
problem Discovering slow collective motions in molecular dynamics.
method State-free reversible VAMPnets (SRV) for nonlinear CV approximation.
result SRVs capture slow dynamics in complex systems.
This work models variability in composite blades' vibrations.
problem Variability in structural dynamics due to manufacturing differences.
method Defining a general model for frequency response functions using mixtures of Gaussian processes.
result A general model for frequency response functions of composite blades.
We developed a method to learn Kolmogorov models for binary variables.
problem Interpreting complex relationships among binary random variables.
method Proposed a framework linking outcomes of binary variables and an algorithm for model computation.
result First-order optimality of the proposed algorithm despite combinatorial complexity.
Inference and learning of graphical models are both well-studied problems in statistics and machine learning that have found many applications in science and engineering. However, exact inference is intractable in general graphical models, which suggests the problem of seeking the best approximation to a collection of …
Adaptive PCR improves panel data analysis with uniform guarantees.
problem Adaptive data collection in panel data settings.
method Adapting PCR to online settings using martingale concentration.
result Time-uniform guarantees for adaptive PCR in panel data.
A new model predicts wind speed using multiple meteorological variables.
problem Precise wind speed forecasting for wind power producers and grid operators.
method Multi-variable Stacked Long Short-Term Memory (LSTM) network.
result The proposed MSLSTM model outperforms traditional methods in wind speed prediction.
Sparse models help in selecting fewer variables for efficient predictions.
problem Overfitting and high computational costs in learning models.
method Automated variable selection for sparse predictive models.
result Sparse models improve model efficiency and interpretability.
Develops variable-lag Granger causality and Transfer Entropy for time series analysis.
problem Fixed time delay assumption in Granger causality and Transfer Entropy does not hold in many applications.
method Variable-lag Granger causality and Transfer Entropy, using optimal warping path of Dynamic Time Warping (DTW).
result Proposed methods perform better than existing methods in both simulated and real-world datasets.
R package varrank ranks variables based on mutual information for multivariate data analysis.
problem Selecting and ranking variables for multivariate datasets.
method Minimum redundancy maximum relevance (mRMRe) model based on information theory.
result Flexible implementation for discrete and continuous data.
Optimizes resource allocation for distributed parameter estimation in sensor networks.
problem Maximizing accuracy in parameter estimation with limited resources.
method Formulates a data collection and collaboration policy design problem as a Fisher information maximization problem. Proposes multi-armed bandit algorithms for learning the optimal policy.
result Identifies optimal data collection and collaboration policies that balance resource use and estimation accuracy.
In this study, we propose an automatic learning method for variables selection based on Lasso in epidemiology context. One of the aim of this approach is to overcome the pretreatment of experts in medicine and epidemiology on collected data. These pretreatment consist in recoding some variables and to choose some inter…
A new method selects important variables for clustering from dependency networks.
problem Variable selection for clustering in high-cost data scenarios.
method Create dependency networks, rank variables by centrality, select top-n variables.
result Top-n variables improve clustering performance compared to existing methods.
Proposes an efficient method to select models under a budget constraint in cost-sensitive learning.
problem Cost-sensitive variable selection in classification problems.
method Ensemble of model schedules to find near optimal models under a budget constraint.
result Our approach outperforms existing methods in benchmark datasets.
Study on elastic curves with variable stiffness, derived from bending energy.
problem Modeling elastic wires with varying thickness.
method Derive Euler-Lagrange equations for curves with variable bending stiffness.
result Characterizations of elastic curves with variable stiffness.
Improves code2vec for Java classes by obfuscating variable names.
problem Code2vec's reliance on variable names makes it vulnerable to typos and attacks.
method Obfuscate variable names during code2vec training and aggregate method embeddings for class-level predictions.
result Obfuscated variable names improve model's robustness and accuracy.
Gradient boosting of regression trees is a competitive procedure for learning predictive models of continuous data that fits the data with an additive non-parametric model. The classic version of gradient boosting assumes that the data is independent and identically distributed. However, relational data with interdepen…
The paper introduces an inference algorithm for graded Bayesian networks.
problem Inference in graded Bayesian networks.
method Tropicalization of the marginal distribution of observed variables, rank-by-rank evaluation of hidden variables.
result Established an inference algorithm for graded Bayesian networks.