Proposes an extended disentanglement framework with new metrics.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
ProSMIN improves representation quality through probabilistic self-supervised learning.
New method uses random convex polytopes to measure representation quality.
CDSSL improves representation quality by integrating linear and nonlinear dependencies.
ProductNet is a collection of high-quality product datasets for better product understanding. Motivated by ImageNet, ProductNet aims at supporting product representation learning by curating product datasets of high quality with properly chosen taxonomy. In this paper, the two goals of building high-quality product dat…
Learning knowledge representation is an increasingly important technology that supports a variety of machine learning related applications. However, the choice of hyperparameters is seldom justified and usually relies on exhaustive search. Understanding the effect of hyperparameter combinations on embedding quality is …
Semantic representations of words have been successfully extracted from unlabeled corpuses using neural network models like word2vec. These representations are generally high quality and are computationally inexpensive to train, making them popular. However, these approaches generally fail to approximate out of vocabul…
C-VAE improves VAE by resolving prior issues and generating better samples.
New autoencoder learns structured representations without regularization.
Generative model improves time series prediction quality.
Learning the right graph representation from noisy, multisource data has garnered significant interest in recent years. A central tenet of this problem is relational learning. Here the objective is to incorporate the partial information each data source gives us in a way that captures the true underlying relationships.…
Partial soft-matching distance improves neural representation comparison by allowing some neurons to remain unmatched.
Pre-training text representations have led to significant improvements in many areas of natural language processing. The quality of these models benefits greatly from the size of the pretraining corpora as long as its quality is preserved. In this paper, we describe an automatic pipeline to extract massive high-quality…
This work improves disentanglement in latent space models without sacrificing generation quality.
The recently introduced continuous Skip-gram model is an efficient method for learning high-quality distributed vector representations that capture a large number of precise syntactic and semantic word relationships. In this paper we present several extensions that improve both the quality of the vectors and the traini…
Contrastive learning struggles with class collapse and feature suppression, revealing bias towards simpler solutions.
Enhances DIM to match learned representations to a specific distribution.
Improved texture synthesis using wavelet-based statistics with rectifier non-linearity.
NeuCrowd creates high-quality samples from crowdsourced labels to improve representation learning.
New method learns high-quality Laplacian representations for reinforcement learning.
Bayesian SHMM models speech units from unannotated speech.
GGAN improves audio representation learning with fewer labels.
Bayesian algorithm improves word representations using semantic taxonomy.
Proposes GM Score to evaluate GANs considering diversity, disentanglement, and discriminability.
The paper examines the reliability of limit order book representations in the face of data perturbation.
This paper addresses two crucial problems of learning disentangled image representations, namely controlling the degree of disentanglement during image editing, and balancing the disentanglement strength and the reconstruction quality. To encourage disentanglement, we devise a distance covariance based decorrelation re…
In this work, we address the problem of musical timbre transfer, where the goal is to manipulate the timbre of a sound sample from one instrument to match another instrument while preserving other musical content, such as pitch, rhythm, and loudness. In principle, one could apply image-based style transfer techniques t…
Introduces TT-NF for more compact neural field representations.
Generates high-quality images using sparse DCT representations.
In many situations, we need to build and deploy separate models in related environments with different data qualities. For example, an environment with strong observation equipments (e.g., intensive care units) often provides high-quality multi-modal data, which are acquired from multiple sensory devices and have rich-…
Improves shared encoder representations for better multi-task learning performance.
We explore the question of whether the representations learned by classifiers can be used to enhance the quality of generative models. Our conjecture is that labels correspond to characteristics of natural data which are most salient to humans: identity in faces, objects in images, and utterances in speech. We propose …
In order to efficiently transmit and store speech signals, speech codecs create a minimally redundant representation of the input signal which is then decoded at the receiver with the best possible perceptual quality. In this work we demonstrate that a neural network architecture based on VQ-VAE with a WaveNet decoder …
Improves model classification accuracy in black-box settings.
Air quality forecasting has been regarded as the key problem of air pollution early warning and control management. In this paper, we propose a novel deep learning model for air quality (mainly PM2.5) forecasting, which learns the spatial-temporal correlation features and interdependence of multivariate air quality rel…
Conditional GANs are at the forefront of natural image synthesis. The main drawback of such models is the necessity for labeled data. In this work we exploit two popular unsupervised learning techniques, adversarial training and self-supervision, and take a step towards bridging the gap between conditional and uncondit…
Improves disentangled representation learning with multi-stage modeling.
SoftCLT improves time series representation learning by soft contrastive loss.
New methods improve feature extraction and representation quality in supervised and unsupervised DR.
DDMI generates high-quality INRs by adapting positional embeddings.
In this paper, we build an organization of high-dimensional datasets that cannot be cleanly embedded into a low-dimensional representation due to missing entries and a subset of the features being irrelevant to modeling functions of interest. Our algorithm begins by defining coarse neighborhoods of the points and defin…
Hierarchical graph clustering is a common technique to reveal the multi-scale structure of complex networks. We propose a novel metric for assessing the quality of a hierarchical clustering. This metric reflects the ability to reconstruct the graph from the dendrogram, which encodes the hierarchy. The optimal represent…
VCAE improves autoencoder quality on MNIST and CelebA.
Neural networks outperform kernels by learning features better.
Graph InfoClust learns node representations by capturing cluster-level information, improving graph mining tasks.
Representation learning becomes especially important for complex systems with multimodal data sources such as cameras or sensors. Recent advances in reinforcement learning and optimal control make it possible to design control algorithms on these latent representations, but the field still lacks a large-scale standard …
Survey on concept factorization methods for better feature learning.
Generative Adversarial Networks (GAN) have demonstrated impressive results in modeling the distribution of natural images, learning latent representations that capture semantic variations in an unsupervised basis. Beyond the generation of novel samples, it is of special interest to exploit the ability of the GAN genera…