Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

185370555740 · Jun 202019922001200920172026
48 results for Representation Analysis

Hidden symmetry of a G'-space X is defined by an extension of the G'-action on X to that of a group G containing G' as a subgroup. In this setting, we study the relationship between the three objects: (A) global analysis on X by using representations of G (hidden symmetry); (B) global analysis on X by using representat…

2016-08-30abs ↗pdf ↗

MediEncoder learns nonlinear representations for causal mediation analysis.

problem High-dimensional noisy covariates and mediators in biomedical studies.
method Coupled encoder-decoder architecture with cross-factor network.
result Improves estimation accuracy in high-dimensional causal mediation analysis.

PRESTO maps latent representations across diverse ML models.

problem Understanding variability in latent representations across different ML models.
method Uses persistent homology to characterize latent spaces and measure their pairwise similarity.
result Preserves desirable properties and enables sensitivity analysis of latent representations.

The paper shows how to learn causal representations with few environments and finite samples.

problem Learning causal representations from limited data and environments.
method Explicit, finite-sample guarantees with a logarithmic number of interventions.
result Consistent recovery of latent causal graph, mixing matrix, and unknown intervention targets.

This work characterizes how data augmentation shapes neural representations.

problem Understanding the impact of data augmentation on neural network representations.
method Embedding neural network hidden representations into a metric space invariant to transformations, analyzing shape-space trajectories.
result Increasing data augmentation strength leads to well-behaved trajectories in the embedded space, and different augmentation types steer representations in distinct directions.

Novel tRSA combines geometry and topology for brain and model analysis.

problem Traditional RSA overlooks topological information in neural representations.
method Topological RSA (tRSA) using nonlinear monotonic transforms.
result Robust model comparisons and novel insights into neural computation.

OMBA learns product and user representations for better online market basket analysis.

problem Limited ability to uncover rarely occurring and temporal associations in MBA.
method Jointly learns product and user representations, captures temporal dynamics, scalable online method.
result OMBA outperforms state-of-the-art methods by 21% on real-world datasets.

We propose a metric, Layer Saturation, defined as the proportion of the number of eigenvalues needed to explain 99% of the variance of the latent representations, for analyzing the learned representations of neural network layers. Saturation is based on spectral analysis and can be computed efficiently, making live ana…

2019-07-19abs ↗pdf ↗

Correlated component analysis as proposed by Dmochowski et al. (2012) is a tool for investigating brain process similarity in the responses to multiple views of a given stimulus. Correlated components are identified under the assumption that the involved spatial networks are identical. Here we propose a hierarchical pr…

2018-02-07abs ↗pdf ↗

Graph Neural Networks outperform the Weisfeiler-Lehman algorithm in representation power.

problem Limited representation power of Graph Neural Networks compared to the Weisfeiler-Lehman algorithm.
method Algebraic analysis using eigenvalue decomposition of graph operators.
result Graph Neural Networks produce more discriminative representations than the Weisfeiler-Lehman algorithm.

The paper analyzes MAML's representation using RSA, revealing that feature reuse is not the primary reason for its success.

problem Understanding why model-agnostic meta-learning (MAML) works well in few-shot learning tasks.
method Representation similarity analysis (RSA) applied to MAML's few-shot learning instantiation.
result Feature reuse is not the primary reason for MAML's success; instead, it is the learning task itself that increases representation similarity.

A spherical topological manifold of dimension n-1 forms a prototile on its cover, the (n-1)-sphere. The tiling is generated by the fixpoint-free action of the group of deck transformations. By a general theorem, this group is isomorphic to the first homotopy group. Multiplicity and selection rules appear in the form of…

2008-10-19abs ↗pdf ↗

TV-SurvCaus improves causal inference for dynamic treatments in survival analysis.

problem Estimating causal effects of time-varying treatments on survival outcomes.
method Representation balancing techniques extended to time-varying treatment regimes with survival outcomes.
result TV-SurvCaus outperforms existing methods in estimating individualized treatment effects with time-varying covariates and treatments.

Proposes Fair Archetypal Analysis to reduce fairness concerns in data representation.

problem Inadvertent encoding of sensitive attributes in Archetypal Analysis.
method Integrates fairness regularization into Archetypal Analysis and its nonlinear extension.
result Reduces group separability without significantly compromising explained variance.

We present a novel method that can learn a graph representation from multivariate data. In our representation, each node represents a cluster of data points and each edge represents the subset-superset relationship between clusters, which can be mutually overlapped. The key to our method is to use formal concept analys…

2018-12-08abs ↗pdf ↗

Model learns code representations from comments for data analysis tasks.

problem Lack of descriptive labels for analyzing large code corpora.
method Weakly supervised transformer architecture for joint code and comment representation.
result Model achieves 38% accuracy increase over expert-supplied heuristics.

New method learns behavioral representations from mobility data.

problem Analyzing behavioral similarity of moving individuals from CDR trajectories.
method mob2vec framework combining segmentation, generalization, and unsupervised learning.
result Mob2vec generates low-dimensional vector representations preserving mobility behavior similarities.

Multimodal fusion is considered a key step in multimodal tasks such as sentiment analysis, emotion detection, question answering, and others. Most of the recent work on multimodal fusion does not guarantee the fidelity of the multimodal representation with respect to the unimodal representations. In this paper, we prop…

2019-08-13abs ↗pdf ↗

Proposes a neural network autoencoder for smoothing and representation learning of functional data.

problem Lack of sufficient nonlinear representations in existing methods for functional data analysis.
method Develops a neural network autoencoder architecture to process functional data directly, learning both smoothing and representation.
result Outperforms traditional methods in prediction, classification, and computational efficiency.

ProGraML uses graph-based machine learning to improve program optimization and analysis.

problem Improving program optimization and analysis with machine learning.
method Low-level, language agnostic graph representation and message passing neural networks.
result ProGraML achieves an average 94.0 F1 score on a benchmark dataset, significantly outperforming state-of-the-art approaches.

KL annealing helps VAEs avoid posterior collapse and overfitting.

problem Posterior collapse and overfitting in VAEs.
method Theoretical analysis of learning dynamics with KL annealing.
result Posterior collapse is inevitable when ββ exceeds a threshold.

Unified toolkit for comparing neural representations using SRTD and NTS.

problem Heuristic asymmetry and unbounded scores in existing divergences.
method Developed SRTD and NTS to address these issues.
result Unified, robust, and scale-invariant metric for comparing neural representations.

This paper reviews nonlinear ICA for disentangled representations in unsupervised learning.

problem Finding useful disentangled representations in unsupervised deep learning.
method Review of nonlinear ICA theory and algorithms for disentanglement.
result Nonlinear ICA can be shown to estimate useful disentangled representations.

DORA analyzes deep neural networks' internal representations to detect spurious correlations.

problem Detecting spurious correlations in deep neural networks' internal representations.
method DORA uses Extreme-Activation (EA) distance measure to assess representation similarities.
result Identifies internal representations capable of detecting spurious correlations.

The paper develops a new approach to conditional risk measures using modular convex analysis.

problem Developing a new method for conditional risk measures.
method Random modular approach to conditional certainty equivalents and niveloids in the conditional LL^{\infty}-space.
result Retrieves a conditional variational formula for optimized certainty equivalents and applies it to the conditional entropic risk measure.

Paper shows regularization improves robustness in domain generalization.

problem Improving robustness in domain generalization.
method Derives novel theoretical analysis to control representation smoothness and proposes a regularization method.
result Regularization improves robustness in domain generalization.

Word embeddings are representations of individual words of a text document in a vector space and they are often use- ful for performing natural language pro- cessing tasks. Current state of the art al- gorithms for learning word embeddings learn vector representations from large corpora of text documents in an unsu- pe…

2017-08-14abs ↗pdf ↗

Meta-learning for bandit tasks using shared representations.

problem Learning new bandit tasks efficiently using shared low-dimensional representations.
method Proposes a greedy policy to learn new bandit tasks leveraging a partially learned low-dimensional representation.
result Upper bound on regret of proposed policy, showing efficiency of learning new tasks.

Diffusion Maps framework is a kernel based method for manifold learning and data analysis that defines diffusion similarities by imposing a Markovian process on the given dataset. Analysis by this process uncovers the intrinsic geometric structures in the data. Recently, it was suggested to replace the standard kernel …

2015-11-19abs ↗pdf ↗

Geometric stability measures neural network robustness, distinguishing from similarity metrics.

problem Lack of robustness in neural network representations.
method Introduces geometric stability, quantified by Shesha metric measuring self-consistency.
result Stability and similarity are uncorrelated, revealing distinct properties of neural network robustness.

Temporal-difference and Q-learning learn feature representations that converge to optimal ones.

problem Understanding how feature representations evolve in temporal-difference and Q-learning with neural networks.
method Mean-field theory applied to overparameterized two-layer neural networks.
result The feature representation converges to the optimal one, generalizing previous results.

We present Deep Generalized Canonical Correlation Analysis (DGCCA) -- a method for learning nonlinear transformations of arbitrarily many views of data, such that the resulting transformations are maximally informative of each other. While methods for nonlinear two-view representation learning (Deep CCA, (Andrew et al.…

2017-02-08abs ↗pdf ↗

Learning representations of data is an important problem in statistics and machine learning. While the origin of learning representations can be traced back to factor analysis and multidimensional scaling in statistics, it has become a central theme in deep learning with important applications in computer vision and co…

2019-11-26abs ↗pdf ↗

New metric for disentangling multivariate representations, accounting for more complex entanglements.

problem Current disentanglement metrics fail to detect entanglements involving more than two variables.
method Partial Information Decomposition framework to analyze information sharing and propose a new disentanglement metric.
result The proposed metric correctly identifies entanglements in high-dimensional spaces.

In this paper we propose a function space approach to Representation Learning and the analysis of the representation layers in deep learning architectures. We show how to compute a weak-type Besov smoothness index that quantifies the geometry of the clustering in the feature space. This approach was already applied suc…

2017-10-09abs ↗pdf ↗