Proposes a method to cluster tasks for constructive cooperative multi-tasking.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper improves short text clustering by integrating semantic relationships into Optimal Transport.
HSACC improves multi-view clustering of incomplete data.
ClusTR improves clustering-based models' robustness without adversarial training.
Unsupervised segmentation learns features without labels, improving accuracy.
POTA improves short text clustering by generating reliable pseudo-labels.
Enhances cooperative multi-task SemCom for distributed users.
Proposes a new model for clustering passenger trajectories with graphs.
Clustering using neural networks has recently demonstrated promising performance in machine learning and computer vision applications. However, the performance of current approaches is limited either by unsupervised learning or their dependence on large set of labeled data samples. In this paper, we propose ClusterNet …
Techniques for data-mining, latent semantic analysis, contextual search of databases, etc. have long ago been developed by computer scientists working on information retrieval (IR). Experimental scientists, from all disciplines, having to analyse large collections of raw experimental data (astronomical, physical, biolo…
Improves clustering performance by mixing latent representations.
A new text clustering method using NMF and LSA improves stability and performance.
Learning image representations to capture fine-grained semantics has been a challenging and important task enabling many applications such as image search and clustering. In this paper, we present Graph-Regularized Image Semantic Embedding (Graph-RISE), a large-scale neural graph learning framework that allows us to tr…
Based on the Aristotelian concept of potentiality vs. actuality allowing for the study of energy and dynamics in language, we propose a field approach to lexical analysis. Falling back on the distributional hypothesis to statistically model word meaning, we used evolving fields as a metaphor to express time-dependent c…
Hybrid approach combines topic and graph embeddings for legal document clustering.
Unsupervised scheme ranks sentences in text documents based on semantic importance.
New method evaluates text-to-image synthesis for realism, variety, and semantic accuracy.
A new method uncovers intrinsic data structures for unsupervised domain adaptation.
New method identifies shared topics in LLM inputs and outputs for better detection of hallucinations.
Company2Vec creates embeddings from company websites for fine-grained business analytics.
New unsupervised learning framework for sound recognition.
Mixtures of Unigrams are one of the simplest and most efficient tools for clustering textual data, as they assume that documents related to the same topic have similar distributions of terms, naturally described by Multinomials. When the classification task is particularly challenging, such as when the document-term ma…
ContraSim learns financial headline similarities for market forecasting.
Representation of human actions as a sequence of human body movements or action attributes enables the development of models for human activity recognition and summarization. We present an extension of the low-rank representation (LRR) model, termed the clustering-aware structure-constrained low-rank representation (CS…
We propose Deep Feature Factorization (DFF), a method capable of localizing similar semantic concepts within an image or a set of images. We use DFF to gain insight into a deep convolutional neural network's learned features, where we detect hierarchical cluster structures in feature space. This is visualized as heat m…
Paper introduces SDM for detecting LLM hallucinations, improving on entropy tests.
The Adversarially Learned Mixture Model (AMM) is a generative model for unsupervised or semi-supervised data clustering. The AMM is the first adversarially optimized method to model the conditional dependence between inferred continuous and categorical latent variables. Experiments on the MNIST and SVHN datasets show t…
An increasing number of people are using online social networking services (SNSs), and a significant amount of information related to experiences in consumption is shared in this new media form. Text mining is an emerging technique for mining useful information from the web. We aim at discovering in particular tweets s…
Clustering using deep autoencoders has been thoroughly investigated in recent years. Current approaches rely on simultaneously learning embedded features and clustering the data points in the latent space. Although numerous deep clustering approaches outperform the shallow models in achieving favorable results on sever…
Early detection and precise characterization of emerging topics in text streams can be highly useful in applications such as timely and targeted public health interventions and discovering evolving regional business trends. Many methods have been proposed for detecting emerging events in text streams using topic modeli…
Our work improves VAE latent space clustering by enforcing invariant and equivariant learning.
In this thesis, we study the problem of feature learning on heterogeneous knowledge graphs. These features can be used to perform tasks such as link prediction, classification and clustering on graphs. Knowledge graphs provide rich semantics encoded in the edge and node types. Meta-paths consist of these types and abst…
SEMASIA provides a large dataset of latent representations for model comparison.
While it has become common to perform automated translations on natural language, performing translations between different representations of mathematical formulae has thus far not been possible. We implemented the first translator for mathematical formulae based on recursive neural networks. We chose recursive neural…
New MCMC method tackles label-switching problem for clustering.
Supervised object detection and semantic segmentation require object or even pixel level annotations. When there exist image level labels only, it is challenging for weakly supervised algorithms to achieve accurate predictions. The accuracy achieved by top weakly supervised algorithms is still significantly lower than …
HDGI learns node representations for heterogeneous graphs.
LEAK learns from mistakes to improve point cloud segmentation.
Deep generative models are tremendously successful in learning low-dimensional latent representations that well-describe the data. These representations, however, tend to much distort relationships between points, i.e. pairwise distances tend to not reflect semantic similarities well. This renders unsupervised tasks, s…
This paper tackles multi-modal label disentanglement in partition-based XMC.
Proposes a new model for clustering passenger trips considering hierarchical and multi-dimensional data.
There exist many approaches for description and recognition of unseen classes in datasets. Nevertheless, it becomes a challenging problem when we deal with multivariate time-series (MTS) (e.g., motion data), where we cannot apply the vectorial algorithms directly to the inputs. In this work, we propose a novel multiple…
Semi-supervised clustering seeks to augment traditional clustering methods by incorporating side information provided via human expertise in order to increase the semantic meaningfulness of the resulting clusters. However, most current methods are \emph{passive} in the sense that the side information is provided before…
Fractal Flow enhances normalizing flows with interpretable latent space and hierarchical modeling.
Corpus poisoning can manipulate word meanings in word embeddings, affecting natural language processing tasks.
Extends conformal prediction to contrastive learning for better coverage of positive samples.
This study improves sentence embeddings from BERT models.
With the recent success of embeddings in natural language processing, research has been conducted into applying similar methods to code analysis. Most works attempt to process the code directly or use a syntactic tree representation, treating it like sentences written in a natural language. However, none of the existin…