Optimizes molecular generation for chemist preferences.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Two regularization techniques improve GCNN explainability and preference from chemists.
Computational chemists typically assay drug candidates by virtually screening compounds against crystal structures of a protein despite the fact that some targets, like the Opioid Receptor and other members of the GPCR family, traverse many non-crystallographic states. We discover new conformational states of …
In this work, we present an application of Locally Interpretable Machine-Agnostic Explanations to 2-D chemical structures. Using this framework we are able to provide a structural interpretation for an existing black-box model for classifying biologically produced fuel compounds with regard to Research Octane Number. T…
Text-based representations of chemicals and proteins can be thought of as unstructured languages codified by humans to describe domain-specific knowledge. Advances in natural language processing (NLP) methodologies in the processing of spoken languages accelerated the application of NLP to elucidate hidden knowledge in…
Chemical structure elucidation is a serious bottleneck in analytical chemistry today. We address the problem of identifying an unknown chemical threat given its mass spectrum and its chemical formula, a task which might take well trained chemists several days to complete. Given a chemical formula, there could be over a…
We describe a fully data driven model that learns to perform a retrosynthetic reaction prediction task, which is treated as a sequence-to-sequence mapping problem. The end-to-end trained model has an encoder-decoder architecture that consists of two recurrent neural networks, which has previously shown great success in…
There is an intuitive analogy of an organic chemist's understanding of a compound and a language speaker's understanding of a word. Consequently, it is possible to introduce the basic concepts and analyze potential impacts of linguistic analysis to the world of organic chemistry. In this work, we cast the reaction pred…
Paper introduces a graph-based approach for retrosynthesis prediction.
In de novo drug design, computational strategies are used to generate novel molecules with good affinity to the desired biological target. In this work, we show that recurrent neural networks can be trained as generative models for molecular structures, similar to statistical language models in natural language process…
New method adapts to user preferences dynamically, improving recommendation models.
Many real-world engineering problems rely on human preferences to guide their design and optimization. We present PrefOpt, an open source package to simplify sequential optimization tasks that incorporate human preference feedback. Our approach extends an existing latent variable model for binary preferences to allow f…
Enhances preference learning by incorporating response times into binary choices.
Bayesian optimization learns DM preferences for multi-outcome experiments.
Deep generative models are able to suggest new organic molecules by generating strings, trees, and graphs representing their structure. While such models allow one to generate molecules with desirable properties, they give no guarantees that the molecules can actually be synthesized in practice. We propose a new molecu…
New study shows personalized content recommendations can lead to polarization of user preferences.
Computer-assisted synthesis planning aims to help chemists find better reaction pathways faster. Finding viable and short pathways from sugar molecules to value-added chemicals can be modeled as a retrosynthesis planning problem with a catalyst allowed. This is a crucial step in efficient biomass conversion. The tradit…
Bayesian optimization agent learns user preferences from pairwise comparisons.
New RLHF framework handles general preference oracles without reward functions.
This paper studies robust forward investment and consumption preferences within a zero-volatility context. Different from previous works, we consider an incomplete financial market model due to general investment portfolio constraints. We provide a new PDE characterization and a novel semi-explicit saddle-point constru…
Chemical reactions can be described as the stepwise redistribution of electrons in molecules. As such, reactions are often depicted using `arrow-pushing' diagrams which show this movement as a sequence of arrows. We propose an electron path prediction model (ELECTRO) to learn these sequences directly from raw reaction …
In preference-based reinforcement learning (RL), an agent interacts with the environment while receiving preferences instead of absolute feedback. While there is increasing research activity in preference-based RL, the design of formal frameworks that admit tractable theoretical analysis remains an open challenge. Buil…
Study on identifying most preferred policy in bandits with vector-valued rewards.
Stable and consistent model alignment for language models without assuming human preference models.
Dropping a tiny fraction of preferences can significantly alter the rankings of top LLMs.
Paper explores limits and possibilities of aligning LLMs with human preferences.
Paper improves parameter estimation of continuous distributions using preference feedback.
Paper investigates monotonicity issues in AI preference learning.
DOPL learns from preference feedback to solve RMAB problems.
Direct Density Ratio Optimization aligns LLMs with human preferences without assuming specific models.
New method uses Riemannian geometry to quantify molecular shapes.
Bayesian optimization with preference learning identifies preferred solutions in multi-objective problems.
Training models to prefer certain responses can unintentionally shift probability to harmful ones.
This work proves win rate is key to understanding preference learning.
IDT learns human preferences from uncertain decisions, even when humans are suboptimal.
The paper proposes a method to infer multi-objective rewards from preferences.
The paper explores game-theoretic alignment of LLMs with human preferences, finding limitations and conditions.
Fashion preference is a fuzzy concept that depends on customer taste, prevailing norms in fashion product/style, henceforth used interchangeably, and a customer's perception of utility or fashionability, yet fashion e-retail relies on algorithmically generated search and recommendation systems that process structured d…
Sparse molecular representations improve interpretability in graph neural networks.
Bal-PM reduces preference labeling costs for LLMs.
DPA aligns LLMs with multi-objective rewards for diverse user preferences.
Tutorials on preference learning with Gaussian Processes.
Robot motions in the presence of humans should not only be feasible and safe, but also conform to human preferences. This, however, requires user feedback on the robot's behavior. In this work, we propose a novel approach to leverage the user's brain signals as a feedback modality in order to decode the judgment of rob…
Given an incomplete ratings data over a set of users and items, the preference completion problem aims to estimate a personalized total preference order over a subset of the items. In practical settings, a ranked list of top- items from the estimated preference order is recommended to the end user in the decreasing …
The problem of retrosynthetic planning can be framed as one player game, in which the chemist (or a computer program) works backwards from a molecular target to simpler starting materials though a series of choices regarding which reactions to perform. This game is challenging as the combinatorial space of possible cho…
Improved DPO framework penalizes preference uncertainty to avoid overoptimization.
Response time improves alignment with diverse human preferences.
Derives a new objective to learn from human preferences without approximations.