Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

18.6%37.2%55.8%74.3% · Jun 202019922001200920182026
48 results for data representation learning

Proposes a new method for medical diagnosis using network-based representation learning.

problem Improving medical diagnosis accuracy through better data representation.
method Heterogeneous network-based model and modified metapath2vec algorithm for learning latent node representations.
result Significant performance boost in symptom/disease classification and disease prediction tasks.

Poincaré embeddings learn hierarchical symbolic data representations.

problem Learning hierarchical representations for complex symbolic data like text and graphs.
method Embedding into hyperbolic space (Poincaré ball) for efficient Riemannian optimization.
result Poincaré embeddings outperform Euclidean embeddings on data with latent hierarchies.

The paper examines the reliability of limit order book representations in the face of data perturbation.

problem The reliability of limit order book representations under data perturbation.
method Experimental analysis of existing representations and guidelines for future research.
result Existing representations of limit order book data are vulnerable to data perturbation.

Learn class-invariant and symmetry-equivariant representations for multi-class data.

problem Deep neural networks learn opaque representations; we aim to make them more transparent.
method Probabilistic modelling with two separate latent variables: invariant and equivariant.
result Qualitative and quantitative performance competitive with other methods, with little tuning.

RAMODO learns better representations for outlier detection in ultrahigh-dimensional data.

problem Suboptimal and unstable outlier detection in ultrahigh-dimensional data.
method Unified representation learning and outlier detection using a ranking model.
result RAMODO improves AUC performance and stability of random distance-based outlier detection.

A method to prevent image representation collapse through data-dependent augmentation.

problem Representation collapse due to image augmentations that damage information.
method Formalizing a stochastic encoding process with a tug-of-war between corruption and preserved information, using infoMax objective.
result Learning a data-dependent distribution of augmentations to avoid representation collapse.

UNTIE learns representations of coupled categorical data.

problem Challenges in learning from unlabeled categorical data with complex couplings.
method UNTIE approach for unsupervised representation learning of heterogeneous couplings.
result UNTIE significantly improves categorical data representations on 25 diverse datasets.

Discriminative clustering learns from both labeled and unlabeled data.

problem Clustering complex datasets with limited labeled data.
method Gradient-based stochastic training and optimal transport with entropic regularization.
result The method can learn feature representations even in fully unsupervised settings.

Unified framework for representation and causal structure learning using exchangeable data.

problem Identifying latent representations or causal structures in non-i.i.d. data.
method Identifiable Exchangeable Mechanisms (IEM) framework for representation and structure learning.
result New insights and identifiability results for causal structure and representation learning.

Paper presents a method to efficiently learn ordered representations of multi-agent data.

problem Challenges in learning consistent representations of multi-agent interactions.
method Dynamic alignment method to order multi-agent data for faster representation learning.
result Representation learning of multi-agent data is significantly accelerated.

Proposes a deep learning method for effective data representation.

problem Constructing effective data representations for prediction.
method A deep dimension reduction approach to learning representations with sufficiency, low dimensionality, and disentanglement.
result The proposed deep nonparametric representation is consistent and performs better than existing methods.

Better data representations can simplify learning tasks by aligning model distributions with true data distributions.

problem Learning complexity influenced by the alignment of model distributions with true data distributions.
method Analyzed the effect of data representations on learning complexity using a task complexity score and information coding length.
result Better representations can simplify learning tasks by aligning model distributions with true data distributions, improving learning outcomes.

This paper investigates how data augmentation improves linear separation of manifold data.

problem Understanding how data augmentation enhances linear separation of manifold data.
method Investigates the conditions under which self-supervised representations can linearly separate multi-manifold data.
result Self-supervised learning can linearly separate manifolds with a smaller distance than unsupervised learning.

Transformative machine learning improves model accuracy and explainability with limited data.

problem Improving model accuracy and interpretability with limited data in scientific tasks.
method Transforming intrinsic data representations to extrinsic ones based on model predictions.
result Transformative machine learning significantly outperforms intrinsic representations in drug-design, gene expression prediction, and meta-learning.

IMSAT learns discrete representations by maximizing information and enforcing invariance.

problem Learning useful discrete representations from data.
method Information Maximizing Self-Augmented Training (IMSAT) with data augmentation and information-theoretic dependency maximization.
result IMSAT achieves state-of-the-art results for clustering and unsupervised hash learning.

FairMixRep learns fair representations from mixed data types.

problem Representation learning in mixed numerical and categorical data with fairness constraints.
method Efficient encoder-decoder framework + fairness constraints.
result Excellent performance in preserving information and fairness in mixed data representations.

i-Mix improves contrastive learning across domains without domain-specific augmentations.

problem Improving contrastive representation learning for unlabeled data across diverse domains.
method i-Mix treats contrastive learning as a non-parametric classifier problem, mixing data in input and virtual label spaces.
result i-Mix consistently improves representation quality across image, speech, and tabular data domains.

Study presents a dataset and evaluation framework for representation learning in complex multimodal systems.

problem Lack of large-scale standard datasets for representation learning in complex multimodal systems.
method Implemented and compared several approaches to representation learning on a large-scale dataset for landing an airplane.
result Representations can be used for various applications including anomaly detection and optimal control.

GGAN improves audio representation learning with fewer labels.

problem Learning representations for specific tasks from unlabelled data.
method Guided Generative Adversarial Neural Network (GGAN).
result GGAN learns better representations with fewer labelled data.

Integrates competitive learning into CNNs to enhance representation and speed up fine-tuning.

problem Efficient use of unlabeled data for CNNs' fine-tuning.
method Integrates unsupervised competitive learning into the convolutional layer of CNNs.
result Effective representation learning using unlabeled data, accelerated fine-tuning process.

Contrastive learning adapts to data intrinsic dimensions, learning low-dimensional representations.

problem Learning high-dimensional representations from multi-modal data.
method Multi-modal contrastive learning with temperature optimization.
result Contrastive learning adapts to intrinsic dimensions of data, not specified dimensions.

MPVAA learns holistic patient representations from mixed healthcare data.

problem Learning personalized patient representations from heterogeneous healthcare data.
method Mixed Pooling Multi-View Attention Autoencoder (MPVAA) that integrates non-linear relationships among multiple data modalities.
result MPVAA generates more effective patient representations than state-of-the-art methods.

A new unsupervised contrastive learning framework improves time series representation learning.

problem Lack of labeled data in time series data.
method Proposes an unsupervised contrastive learning framework using a novel contrastive loss and data augmentation.
result Framework outperforms other approaches on univariate and multivariate time series, and benefits transfer learning.

POLAR learns efficient data acquisition policies using pretrained belief representations.

problem Challenges in learning effective policies for adaptive data acquisition.
method POLAR decouples representation learning from policy learning by leveraging pretrained predictive foundation models as belief-state encoders.
result POLAR outperforms state-of-the-art methods across diverse tasks while requiring fewer training samples.

The paper formalizes criteria for non-spurious and disentangled representations using causal methods.

problem Formalizing criteria for non-spurious and disentangled representations in representation learning.
method Causal perspective, counterfactual quantities, observable consequences of causal assertions.
result Computable metrics for assessing representation learning based on observed data.

This paper formalizes representation learning and shows its benefits.

problem Understanding and formalizing the benefits of representation learning techniques.
method Introducing a formal framework to study representation learning and its utility.
result Representation learning can be performed provably and efficiently under plausible assumptions.

A new framework for semi-supervised learning using pseudo-representation labeling.

problem Improving deep learning models with limited labeled data.
method Pseudo-representation labeling framework integrating pseudo-labeling and self-supervised representation learning.
result Outperforms state-of-the-art semi-supervised learning methods in industrial classification problems.

ExpCLR uses expert features to improve time-series representation learning.

problem Current representation learning approaches fail to ensure useful properties for time-series data.
method ExpCLR employs expert features to replace data transformations in contrastive learning, ensuring two useful properties for time-series representations.
result ExpCLR outperforms state-of-the-art methods on three real-world time-series datasets.

IRM learns representations invariant to training distributions for better generalization.

problem Learning generalizable models across different training distributions.
method IRM learns a data representation that remains consistent across multiple training distributions, ensuring an optimal classifier matches across them.
result IRM enables out-of-distribution generalization by learning invariant correlations.

This paper introduces a new feature learning technique based on error representation.

problem Learning high-level features for classification from diverse and imbalanced data.
method Inverse feature learning using error representation approach.
result Significantly better performance compared to state-of-the-art techniques.