We investigate the problem of truth discovery based on opinions from multiple agents who may be unreliable or biased. We consider the case where agents' reliabilities or biases are correlated if they belong to the same community, which defines a group of agents with similar opinions regarding a particular event. An age…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The problem of secure friend discovery on a social network has long been proposed and studied. The requirement is that a pair of nodes can make befriending decisions with minimum information exposed to the other party. In this paper, we propose to use community detection to tackle the problem of secure friend discovery…
Neural node embeddings have recently emerged as a powerful representation for supervised learning tasks involving graph-structured data. We leverage this recent advance to develop a novel algorithm for unsupervised community discovery in graphs. Through extensive experimental studies on simulated and real-world data, w…
The false discovery rate (FDR)---the expected fraction of spurious discoveries among all the discoveries---provides a popular statistical assessment of the reproducibility of scientific studies in various disciplines. In this work, we introduce a new method for controlling the FDR in meta-analysis of many decentralized…
A communication-efficient method controls FDR in network settings.
Python library for causal discovery from observational data.
The last decade has seen great progress in both dynamic network modeling and topic modeling. This paper draws upon both areas to create a Bayesian method that allows topic discovery to inform the latent network model and the network structure to facilitate topic identification. We apply this method to the 467 top polit…
Paper presents a workflow for reliable unsupervised learning in science.
Decentralized detection avoids sharing data, controls false discoveries.
The analysis of temporal networks has a wide area of applications in a world of technological advances. An important aspect of temporal network analysis is the discovery of community structures. Real data networks are often very large and the communities are observed to have a hierarchical structure referred to as mult…
AI+MPS workshop aims to strengthen AI's role in science.
Proposes VILMAP for finding motifs and segmenting words in time series.
Aggregation distorts causal discovery results but recovery is possible with partial linearity or prior.
EDL discovers state-covering skills without relying on task rewards.
TimeGraph creates synthetic datasets for robust time-series causal discovery.
Machine learning enhances cosmology through new tools and data analysis.
Combines SBMs and graph neural nets for graph embeddings.
Develops a framework for causal structure learning using both interventional and observational data.
Improved error feedback method reduces communication complexity in distributed training.
Bayesian model detects communities in networks with covariates.
Study uses LLMs to automate data insights discovery.
New benchmarks show LLMs struggle with causal discovery.
Network detection is an important capability in many areas of applied research in which data can be represented as a graph of entities and relationships. Oftentimes the object of interest is a relatively small subgraph in an enormous, potentially uninteresting background. This aspect characterizes network detection as …
The waggle dance that honeybees perform is an astonishing way of communicating the location of food source. After over 60 years of its discovery, researchers still use manual labeling by watching hours of dance videos to detect different transitions between dance components thus extracting information regarding the dis…
Due to recent explosion of text data, researchers have been overwhelmed by ever-increasing volume of articles produced by different research communities. Various scholarly search websites, citation recommendation engines, and research databases have been created to simplify the text search tasks. However, it is still d…
Machine learning can help us in solving problems in the context big data analysis and classification, as well as in playing complex games such as Go. But can it also be used to find novel protocols and algorithms for applications such as large-scale quantum communication? Here we show that machine learning can be used …
Paper proposes a method to monitor research topic evolution.
Topic discovery has witnessed a significant growth as a field of data mining at large. In particular, time-evolving topic discovery, where the evolution of a topic is taken into account has been instrumental in understanding the historical context of an emerging topic in a dynamic corpus. Traditionally, time-evolving t…
Paper proposes a new dataset for group anomaly detection in physics.
AutoML struggles with climate change data, but offers potential improvements.
Causal autoregressive flows enable accurate causal inference and prediction.
Despite the overwhelming success of the existing Social Networking Services (SNS), their centralized ownership and control have led to serious concerns in user privacy, censorship vulnerability and operational robustness of these services. To overcome these limitations, Distributed Social Networks (DSN) have recently b…
We present a probabilistic framework for overlapping community discovery and link prediction for relational data, given as a graph. The proposed framework has: (1) a deep architecture which enables us to infer multiple layers of latent features/communities for each node, providing superior link prediction performance o…
XIMP improves molecular property prediction by integrating multiple graph representations.
Survey of deep learning models for scientific discovery.
Survey on robust data representation learning from a knowledge flow perspective.
With the recent popularity of graphical clustering methods, there has been an increased focus on the information between samples. We show how learning cluster structure using edge features naturally and simultaneously determines the most likely number of clusters and addresses data scale issues. These results are parti…
The paper evaluates methods for explaining deep learning in security.
Entropy regularization improves sparse model discovery in federated learning.
At what level should government or companies support research? This complex multi-faceted question encompasses such qualitative bonus as satisfying natural human curiosity, the quest for knowledge and the impact on education and culture, but one of its most scrutinized component reduces to the assessment of economic pe…
Introduces Q-structures for mechanics using advanced geometry.
The success of enhanced sampling molecular simulations that accelerate along collective variables (CVs) is predicated on the availability of variables coincident with the slow collective motions governing the long-time conformational dynamics of a system. It is challenging to intuit these slow CVs for all but the simpl…
EUREKA builds classifiers that use surprising features.
Differentiable causal discovery methods perform robustly under model violations.
New framework uses background knowledge to speed up causal discovery.
New method prevents invalid inference after causal discovery.
Paper proposes efficient sample collection strategy for RL.
Human communication takes many forms, including speech, text and instructional videos. It typically has an underlying structure, with a starting point, ending, and certain objective steps between them. In this paper, we consider instructional videos where there are tens of millions of them on the Internet. We propose a…