Unsupervised learning representations generalize better than supervised learning under distribution shifts.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper introduces metrics for robust unsupervised learning of vehicle interactions.
Unsupervised learning models can produce accurate but misleading predictions.
The paper tackles adversarial robustness by maximizing worst-case mutual information.
RAEUFS selects features from data without labels, improving robustness to outliers.
RKUM is an R package for robust kernel-based unsupervised methods.
Study on robustness of unsupervised representation learning in slightly misspecified settings.
A framework is presented for unsupervised learning of representations based on infomax principle for large-scale neural populations. We use an asymptotic approximation to the Shannon's mutual information for a large neural population to demonstrate that a good initial approximation to the global information-theoretic o…
A robust loss for anomaly mitigation and unsupervised contamination classification
Unified model explains AT's generative ability.
Improved training for VQ-VAE models with robust codebook learning.
Improves domain adaptation by combining multiple source domains and target domain data.
Study improves GMM learning performance through multi-task and transfer learning.
Improved disentanglement of data factors using recursive training.
Paper presents a workflow for reliable unsupervised learning in science.
In this paper, we reproduce the experiments of Artetxe et al. (2018b) regarding the robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings. We show that the reproduction of their method is indeed feasible with some minor assumptions. We further investigate the robustness of their m…
Paper proposes DRL for unsupervised IoT localization.
New model detects anomalies without needing clean data.
A novel unsupervised outlier detection method using Randomized PCA Forest.
Unsupervised learning techniques in computer vision often require learning latent representations, such as low-dimensional linear and non-linear subspaces. Noise and outliers in the data can frustrate these approaches by obscuring the latent spaces. Our main goal is deeper understanding and new development of robust ap…
Simple framework decouples word alignment and multilingual embedding mapping.
CoDAG combines domain adaptation and generalization for unsupervised continual domain shift learning.
Regression mixture models are widely studied in statistics, machine learning and data analysis. Fitting regression mixtures is challenging and is usually performed by maximum likelihood by using the expectation-maximization (EM) algorithm. However, it is well-known that the initialization is crucial for EM. If the init…
A new method combines OCSVM with representation learning for UAD.
We investigate the use of a non-parametric independence measure, the Hilbert-Schmidt Independence Criterion (HSIC), as a loss-function for learning robust regression and classification models. This loss-function encourages learning models where the distribution of the residuals between the label and the model predictio…
We present a sparse and invariant representation with low asymptotic complexity for robust unsupervised transient and onset zone detection in noisy environments. This unsupervised approach is based on wavelet transforms and leverages the scattering network from Mallat et al. by deriving frequency invariance. This frequ…
Disentangled representations have recently been shown to improve fairness, data efficiency and generalisation in simple supervised and reinforcement learning tasks. To extend the benefits of disentangled representations to more complex domains and practical applications, it is important to enable hyperparameter tuning …
Transformers solve Gaussian Mixture Models without supervision.
To the best of our knowledge, there are no general well-founded robust methods for statistical unsupervised learning. Most of the unsupervised methods explicitly or implicitly depend on the kernel covariance operator (kernel CO) or kernel cross-covariance operator (kernel CCO). They are sensitive to contaminated data, …
Proposes HOT method for robust multi-view learning.
Extends information bottleneck to multi-view unsupervised learning.
PhyloVAE learns tree topologies without supervision.
DART tackles adversarial robustness in domain adaptation without labeled target data.
We consider a general statistical learning problem where an unknown fraction of the training data is corrupted. We develop a robust learning method that only requires specifying an upper bound on the corrupted data fraction. The method minimizes a risk function defined by a non-parametric distribution with unknown prob…
UDA improves ABI robustness but fails under certain prior misspecifications.
Two new methods improve graph embedding without needing a complete graph structure.
SAMPLR optimizes for ground truth in aleatoric parameters to avoid curriculum-induced covariate shift.
MLS improves feature selection for imbalanced data.
Unsupervised domain adaptation seeks to learn an invariant and discriminative representation for an unlabeled target domain by leveraging the information of a labeled source dataset. We propose to improve the discriminative ability of the target domain representation by simultaneously learning tightly clustered target …
LOBSTUR-GNN adapts bootstrapping for unsupervised GNNs, improving node representation learning.
The uncertainty estimation is critical in real-world decision making applications, especially when distributional shift between the training and test data are prevalent. Many calibration methods in the literature have been proposed to improve the predictive uncertainty of DNNs which are generally not well-calibrated. H…
We introduce a novel co-learning paradigm for manifolds naturally equipped with a group action, motivated by recent developments on learning a manifold from attached fibre bundle structures. We utilize a representation theoretic mechanism that canonically associates multiple independent vector bundles over a common bas…
Many unsupervised kernel methods rely on the estimation of the kernel covariance operator (kernel CO) or kernel cross-covariance operator (kernel CCO). Both kernel CO and kernel CCO are sensitive to contaminated data, even when bounded positive definite kernels are used. To the best of our knowledge, there are few well…
Unsupervised machine translation---i.e., not assuming any cross-lingual supervision signal, whether a dictionary, translations, or comparable corpora---seems impossible, but nevertheless, Lample et al. (2018) recently proposed a fully unsupervised machine translation (MT) model. The model relies heavily on an adversari…
Smoothed analysis is a powerful paradigm in overcoming worst-case intractability in unsupervised learning and high-dimensional data analysis. While polynomial time smoothed analysis guarantees have been obtained for worst-case intractable problems like tensor decompositions and learning mixtures of Gaussians, such guar…
We consider learning of fundamental properties of communities in large noisy networks, in the prototypical situation where the nodes or users are split into two classes according to a binary property, e.g., according to their opinions or preferences on a topic. For learning these properties, we propose a nonparametric,…
Linear disentangled representations improve unsupervised action estimation.
Unsupervised learning is becoming more and more important recently. As one of its key components, the autoencoder (AE) aims to learn a latent feature representation of data which is more robust and discriminative. However, most AE based methods only focus on the reconstruction within the encoder-decoder phase, which ig…