This paper reviews data representation learning from traditional methods to deep learning.
problem Learning the intrinsic structure of data.
method Investigates traditional and deep learning methods.
result Deep learning models have achieved top results in various tasks.
Unified data representation learning improves non-parametric two-sample testing.
problem Improving non-parametric two-sample testing accuracy.
method Proposes RL-TST framework combining IRs and DRs for better test power.
result RL-TST outperforms existing methods by leveraging both IRs and DRs.
This article reviews statistical methods for learning data representations.
problem Learning meaningful representations of data.
method Statistical perspective on unsupervised and supervised representation learning.
result Recent advances in representation learning from a statistical viewpoint.
Proposes a new method for medical diagnosis using network-based representation learning.
problem Improving medical diagnosis accuracy through better data representation.
method Heterogeneous network-based model and modified metapath2vec algorithm for learning latent node representations.
result Significant performance boost in symptom/disease classification and disease prediction tasks.
Poincaré embeddings learn hierarchical symbolic data representations.
problem Learning hierarchical representations for complex symbolic data like text and graphs.
method Embedding into hyperbolic space (Poincaré ball) for efficient Riemannian optimization.
result Poincaré embeddings outperform Euclidean embeddings on data with latent hierarchies.
The paper examines the reliability of limit order book representations in the face of data perturbation.
problem The reliability of limit order book representations under data perturbation.
method Experimental analysis of existing representations and guidelines for future research.
result Existing representations of limit order book data are vulnerable to data perturbation.
Learn class-invariant and symmetry-equivariant representations for multi-class data.
problem Deep neural networks learn opaque representations; we aim to make them more transparent.
method Probabilistic modelling with two separate latent variables: invariant and equivariant.
result Qualitative and quantitative performance competitive with other methods, with little tuning.
RAMODO learns better representations for outlier detection in ultrahigh-dimensional data.
problem Suboptimal and unstable outlier detection in ultrahigh-dimensional data.
method Unified representation learning and outlier detection using a ranking model.
result RAMODO improves AUC performance and stability of random distance-based outlier detection.
A method to prevent image representation collapse through data-dependent augmentation.
problem Representation collapse due to image augmentations that damage information.
method Formalizing a stochastic encoding process with a tug-of-war between corruption and preserved information, using infoMax objective.
result Learning a data-dependent distribution of augmentations to avoid representation collapse.
UNTIE learns representations of coupled categorical data.
problem Challenges in learning from unlabeled categorical data with complex couplings.
method UNTIE approach for unsupervised representation learning of heterogeneous couplings.
result UNTIE significantly improves categorical data representations on 25 diverse datasets.
Discriminative clustering learns from both labeled and unlabeled data.
problem Clustering complex datasets with limited labeled data.
method Gradient-based stochastic training and optimal transport with entropic regularization.
result The method can learn feature representations even in fully unsupervised settings.
Unified framework for representation and causal structure learning using exchangeable data.
problem Identifying latent representations or causal structures in non-i.i.d. data.
method Identifiable Exchangeable Mechanisms (IEM) framework for representation and structure learning.
result New insights and identifiability results for causal structure and representation learning.
Paper presents a method to efficiently learn ordered representations of multi-agent data.
problem Challenges in learning consistent representations of multi-agent interactions.
method Dynamic alignment method to order multi-agent data for faster representation learning.
result Representation learning of multi-agent data is significantly accelerated.
A novel method for clustering multi-view data using dual representations.
problem Clustering multi-view data with consistent and unique information.
method One-step multi-view clustering method exploiting dual representations.
result The proposed method improves clustering performance on benchmark datasets.
S-VQ-VAE learns interpretable class-specific representations.
problem Learning interpretable representations of data.
method Supervised Vector Quantized Variational AutoEncoder (S-VQ-VAE).
result S-VQ-VAE learns interpretable class-specific representations.
Neural nets learn robust geometric data representations.
problem Ensuring neural networks are robust to adversarial attacks.
method Topological Data Analysis via persistence diagrams, Lipschitz stability.
result Certified ε-robustness on ORBIT5K dataset. Proposes a deep learning method for effective data representation.
problem Constructing effective data representations for prediction.
method A deep dimension reduction approach to learning representations with sufficiency, low dimensionality, and disentanglement.
result The proposed deep nonparametric representation is consistent and performs better than existing methods.
Better data representations can simplify learning tasks by aligning model distributions with true data distributions.
problem Learning complexity influenced by the alignment of model distributions with true data distributions.
method Analyzed the effect of data representations on learning complexity using a task complexity score and information coding length.
result Better representations can simplify learning tasks by aligning model distributions with true data distributions, improving learning outcomes.
This paper investigates how data augmentation improves linear separation of manifold data.
problem Understanding how data augmentation enhances linear separation of manifold data.
method Investigates the conditions under which self-supervised representations can linearly separate multi-manifold data.
result Self-supervised learning can linearly separate manifolds with a smaller distance than unsupervised learning.
Overview of structured data representation methods.
problem Structured data lacks vectorial form, complicating machine learning.
method Various approaches including kernel, distance, neural networks, and graph convolutional networks.
result New approaches like metric learning and recurrent decoder networks have emerged.
Transformative machine learning improves model accuracy and explainability with limited data.
problem Improving model accuracy and interpretability with limited data in scientific tasks.
method Transforming intrinsic data representations to extrinsic ones based on model predictions.
result Transformative machine learning significantly outperforms intrinsic representations in drug-design, gene expression prediction, and meta-learning.
IMSAT learns discrete representations by maximizing information and enforcing invariance.
problem Learning useful discrete representations from data.
method Information Maximizing Self-Augmented Training (IMSAT) with data augmentation and information-theoretic dependency maximization.
result IMSAT achieves state-of-the-art results for clustering and unsupervised hash learning.
Proposes a new method to learn data representations by modeling sample relations.
problem Lack of rich latent structural information in DAEs.
method Explicitly models and leverages sample relations as supervision for representation learning.
result Significantly improves clustering performance on benchmark datasets.
FairMixRep learns fair representations from mixed data types.
problem Representation learning in mixed numerical and categorical data with fairness constraints.
method Efficient encoder-decoder framework + fairness constraints.
result Excellent performance in preserving information and fairness in mixed data representations.
i-Mix improves contrastive learning across domains without domain-specific augmentations.
problem Improving contrastive representation learning for unlabeled data across diverse domains.
method i-Mix treats contrastive learning as a non-parametric classifier problem, mixing data in input and virtual label spaces.
result i-Mix consistently improves representation quality across image, speech, and tabular data domains.
Study presents a dataset and evaluation framework for representation learning in complex multimodal systems.
problem Lack of large-scale standard datasets for representation learning in complex multimodal systems.
method Implemented and compared several approaches to representation learning on a large-scale dataset for landing an airplane.
result Representations can be used for various applications including anomaly detection and optimal control.
Representation learning improves EHR data for healthcare tasks.
problem Transforming EHR data into useful representations for machine learning.
method Deep learning and disentangling underlying factors from EHR data.
result Better representations improve machine learning performance in healthcare.
Tile2Vec learns spatially meaningful representations without labels.
problem Lack of unsupervised methods for geospatial data.
method Unsupervised representation learning using the distributional hypothesis.
result Tile2Vec improves performance in spatial classification tasks.
VAE learns latent representations for bank customers' creditworthiness.
problem Improving marketing, CRM, and credit risk assessment in retail banks.
method Adopted Variational Autoencoder (VAE) to learn latent representations.
result VAE latent representations capture customers' creditworthiness and generalize to new data.
GGAN improves audio representation learning with fewer labels.
problem Learning representations for specific tasks from unlabelled data.
method Guided Generative Adversarial Neural Network (GGAN).
result GGAN learns better representations with fewer labelled data.
Integrates competitive learning into CNNs to enhance representation and speed up fine-tuning.
problem Efficient use of unlabeled data for CNNs' fine-tuning.
method Integrates unsupervised competitive learning into the convolutional layer of CNNs.
result Effective representation learning using unlabeled data, accelerated fine-tuning process.
Contrastive learning adapts to data intrinsic dimensions, learning low-dimensional representations.
problem Learning high-dimensional representations from multi-modal data.
method Multi-modal contrastive learning with temperature optimization.
result Contrastive learning adapts to intrinsic dimensions of data, not specified dimensions.
MPVAA learns holistic patient representations from mixed healthcare data.
problem Learning personalized patient representations from heterogeneous healthcare data.
method Mixed Pooling Multi-View Attention Autoencoder (MPVAA) that integrates non-linear relationships among multiple data modalities.
result MPVAA generates more effective patient representations than state-of-the-art methods.
A new unsupervised contrastive learning framework improves time series representation learning.
problem Lack of labeled data in time series data.
method Proposes an unsupervised contrastive learning framework using a novel contrastive loss and data augmentation.
result Framework outperforms other approaches on univariate and multivariate time series, and benefits transfer learning.
POLAR learns efficient data acquisition policies using pretrained belief representations.
problem Challenges in learning effective policies for adaptive data acquisition.
method POLAR decouples representation learning from policy learning by leveraging pretrained predictive foundation models as belief-state encoders.
result POLAR outperforms state-of-the-art methods across diverse tasks while requiring fewer training samples.
The paper formalizes criteria for non-spurious and disentangled representations using causal methods.
problem Formalizing criteria for non-spurious and disentangled representations in representation learning.
method Causal perspective, counterfactual quantities, observable consequences of causal assertions.
result Computable metrics for assessing representation learning based on observed data.
This paper formalizes representation learning and shows its benefits.
problem Understanding and formalizing the benefits of representation learning techniques.
method Introducing a formal framework to study representation learning and its utility.
result Representation learning can be performed provably and efficiently under plausible assumptions.
GANs learn good data representations without labels.
problem Learning good data representations without labeled data.
method Adversarial learning of latent space representations.
result GANs can learn good mappings from simple priors to target data distributions.
A new framework for semi-supervised learning using pseudo-representation labeling.
problem Improving deep learning models with limited labeled data.
method Pseudo-representation labeling framework integrating pseudo-labeling and self-supervised representation learning.
result Outperforms state-of-the-art semi-supervised learning methods in industrial classification problems.
Paper improves deep learning models for limit order book data.
problem Deep learning models' performance depends on robust input data representation.
method Identified and modified flaws in existing representations.
result Proposed modifications lead to state-of-the-art performance.
ExpCLR uses expert features to improve time-series representation learning.
problem Current representation learning approaches fail to ensure useful properties for time-series data.
method ExpCLR employs expert features to replace data transformations in contrastive learning, ensuring two useful properties for time-series representations.
result ExpCLR outperforms state-of-the-art methods on three real-world time-series datasets.
Perceptor Gradients learns symbolic representations from raw data.
problem Learning transferable symbolic representations from raw data.
method Decomposes policy into perceptor network and task encoding program.
result Efficiently learns symbolic representations for control tasks.
Generative models learn useful representations for complex sequential data.
problem Sequence prediction for high-dimensional input sequences.
method Three models based on Generative Stochastic Networks (GSN) for unsupervised sequence learning.
result GSNs provide evidence as a viable framework for complex sequential data.
Model learns disentangled static and dynamic data representations.
problem Learning disentangled representations from unordered data.
method Factorized graphical model exploiting sequential data regularities.
result Well-organized latent space for data dynamics.
The goal of unsupervised representation learning is to extract a new representation of data, such that solving many different tasks becomes easier. Existing methods typically focus on vectorized data and offer little support for relational data, which additionally describe relationships among instances. In this work we…
Framework for causal discovery using multi-modal data.
problem Failure of representation learning in causal tasks.
method Statistical and computational framework combining representation learning and causal inference.
result Effective use of observational and perturbational data for causal discovery.
IRM learns representations invariant to training distributions for better generalization.
problem Learning generalizable models across different training distributions.
method IRM learns a data representation that remains consistent across multiple training distributions, ensuring an optimal classifier matches across them.
result IRM enables out-of-distribution generalization by learning invariant correlations.
This paper introduces a new feature learning technique based on error representation.
problem Learning high-level features for classification from diverse and imbalanced data.
method Inverse feature learning using error representation approach.
result Significantly better performance compared to state-of-the-art techniques.