LORL learns object-centric representations from vision and language.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proposes LLM-DCD for improved causal discovery from data.
Slot Attention extracts object-centric representations from images.
Transformer-based method discovers objects from images without labels.
The paper integrates statistical significance and discriminative power in pattern discovery.
Enhances drug discovery models by understanding human language.
The discovery of functional molecules is an expensive and time-consuming process, exemplified by the rising costs of small molecule therapeutic discovery. One class of techniques of growing interest for early-stage drug discovery is de novo molecular generation and optimization, catalyzed by the development of new deep…
New algorithm improves materials discovery using max K-Armed Bandit.
New method robustly discovers causal relationships from imperfect data.
We present the first evidence that adaptive learning techniques can boost the discovery of unusual objects within astronomical light curve data sets. Our method follows an active learning strategy where the learning algorithm chooses objects which can potentially improve the learner if additional information about them…
Frolicher and Nijenhuis recognized well in the middle of the previous century that the Lie bracket and its Jacobi identity could and should exist beyond Lie algebras. Nevertheless the conceptual meaning of their discovery has been obscured by the messy techniques they exploited. The principal objective in this paper is…
Concept Relation Discovery and Innovation Enabling Technology (CORDIET), is a toolbox for gaining new knowledge from unstructured text data. At the core of CORDIET is the C-K theory which captures the essential elements of innovation. The tool uses Formal Concept Analysis (FCA), Emergent Self Organizing Maps (ESOM) and…
The problem of secure friend discovery on a social network has long been proposed and studied. The requirement is that a pair of nodes can make befriending decisions with minimum information exposed to the other party. In this paper, we propose to use community detection to tackle the problem of secure friend discovery…
Study shows bottlenecks improve image segmentation quality.
SNAP efficiently identifies causal effects without needing full graph learning.
New methods identify concepts in trained embeddings reliably without human labels.
V-SysId identifies keypoints and 3D system from unlabeled videos.
Acquiring abilities in the absence of a task-oriented reward function is at the frontier of reinforcement learning research. This problem has been studied through the lens of empowerment, which draws a connection between option discovery and information theory. Information-theoretic skill discovery methods have garnere…
ERP improves drug discovery by balancing molecule generation quality and efficiency.
We study nonconvex optimization landscapes for learning overcomplete representations, including learning (i) sparsely used overcomplete dictionaries and (ii) convolutional dictionaries, where these unsupervised learning problems find many applications in high-dimensional data analysis. Despite the empirical success of …
This work improves molecular design by efficiently selecting diverse candidate molecules.
Learning from demonstration has been widely studied in machine learning but becomes challenging when the demonstrated trajectories are unstructured and follow different objectives. This short-paper proposes PODNet, Plannable Option Discovery Network, addressing how to segment an unstructured set of demonstrated traject…
In this paper we present our scientific discovery that good representation can be learned via continuous attention during the interaction between Unsupervised Learning(UL) and Reinforcement Learning(RL) modules driven by intrinsic motivation. Specifically, we designed intrinsic rewards generated from UL modules for dri…
Machine learning (ML) and artificial intelligence (AI) algorithms are now being used to automate the discovery of physics principles and governing equations from measurement data alone. However, positing a universal physical law from data is challenging without simultaneously proposing an accompanying discrepancy model…
Discovering novel materials can be greatly accelerated by iterative machine learning-informed proposal of candidates---active learning. However, standard \emph{global-scope error} metrics for model quality are not predictive of discovery performance, and can be misleading. We introduce the notion of \emph{Pareto shell-…
RAMBO optimizes multi-regime problems by discovering and modeling distinct energy basins.
A coreset is a subset of the training set, using which a machine learning algorithm obtains performances similar to what it would deliver if trained over the whole original data. Coreset discovery is an active and open line of research as it allows improving training speed for the algorithms and may help human understa…
New RL formulation for maximizing maximum reward in molecule generation.
STICC clusters geographic objects considering both spatial contiguity and attributes.
We introduce a framework for dynamic adversarial discovery of information (DADI), motivated by a scenario where information (a feature set) is used by third parties with unknown objectives. We train a reinforcement learning agent to sequentially acquire a subset of the information while balancing accuracy and fairness …
Bayesian network structure learning algorithms with limited data are being used in domains such as systems biology and neuroscience to gain insight into the underlying processes that produce observed data. Learning reliable networks from limited data is difficult, therefore transfer learning can improve the robustness …
We show that dropout training is best understood as performing MAP estimation concurrently for a family of conditional models whose objectives are themselves lower bounded by the original dropout objective. This discovery allows us to pick any model from this family after training, which leads to a substantial improvem…
We propose a novel unsupervised generative model that learns to disentangle object identity from other low-level aspects in class-imbalanced data. We first investigate the issues surrounding the assumptions about uniformity made by InfoGAN, and demonstrate its ineffectiveness to properly disentangle object identity in …
This paper establishes the existence of observable footprints that reveal the "causal dispositions" of the object categories appearing in collections of images. We achieve this goal in two steps. First, we take a learning approach to observational causal discovery, and build a classifier that achieves state-of-the-art …
The discovery of processes for the synthesis of new materials involves many decisions about process design, operation, and material properties. Experimentation is crucial but as complexity increases, exploration of variables can become impractical using traditional combinatorial approaches. We describe an iterative met…
Paper proposes conditional multidimensional scaling for better data reduction.
New algorithms improve causal graph discovery with adaptive interventions, even under worst-case interventional costs.
Accelerates optimization in asynchronous systems with sparse updates.
DDCD uses diffusion models to learn causal structures from noisy data.
The singly periodic genus-one helicoid was in the origin of the discovery of the first example of a complete minimal surface with finite topology but infinite total curvature, the celebrated Hoffman-Karcher-Wei's genus one helicoid. The objective of this paper is to give a uniqueness theorem for the singly periodic gen…
In this work we present Discrete Attend Infer Repeat (Discrete-AIR), a Recurrent Auto-Encoder with structured latent distributions containing discrete categorical distributions, continuous attribute distributions, and factorised spatial attention. While inspired by the original AIR model andretaining AIR model's capabi…
DOCKSTRING simplifies docking simulations for better drug design benchmarks.
Local discovery method uncovers direct unfairness in complex systems.
New method uses kernel deviance measures to discover causal relationships in heterogeneous data.
BEACON optimizes discovery by efficiently finding novel behaviors.
Proposes a method to select features for deep learning in noisy, high-dimensional data.
We discover subgroups for Cox model survival analysis, improving model accuracy.
Proposes MOGFNs for generating diverse Pareto optimal solutions in multi-objective optimization.