Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

50100150200 · May 202619922001200920182026
48 results for embeddings alignment

LPL optimizes embeddings to align local neighborhoods, improving cross-lingual word alignment.

problem Aligning embeddings across different datasets and languages.
method Locality Preserving Loss (LPL) optimizes model to project embeddings while maintaining local neighborhoods and aligning them.
result LPL-based alignment leads to better and consistent accuracy, especially in small training set settings.

Researchers mapped clinical jargon to consumer language using embeddings alignment.

problem Mapping and translating clinical jargon to consumer language for better communication.
method Trained embeddings on clinical and consumer language corpora, aligned them using Procrustes algorithm, and refined with adversarial training.
result Procrustes algorithm effectively aligned clinical and consumer language embeddings.

G-CREWE efficiently aligns large networks using node embeddings and compression.

problem Efficiently aligning large networks for various applications.
method Uses node embeddings and compression to align networks at fine and coarse resolutions.
result G-CREWE achieves efficient and accurate network alignment, twice as fast as existing methods.

Geometric approach for unsupervised word embedding alignment.

problem Learning alignment between word embeddings of source and target languages.
method Formulates alignment as domain adaptation on the manifold of doubly stochastic matrices, employing Riemannian conjugate gradient algorithm.
result Empirically outperforms state-of-the-art methods on bilingual lexicon induction tasks.

Paper proposes unsupervised knowledge graph alignment with adversarial learning.

problem Aligning knowledge graphs from different sources or languages without large amounts of aligned triplets.
method Adversarial learning framework to align entity and relation embeddings, with mutual information regularization.
result Framework effectively aligns knowledge graphs in unsupervised and weakly-supervised settings.

FedGTEA learns new tasks in federated learning with task embeddings and alignment.

problem Federated class-incremental learning with task-specific knowledge and model uncertainty.
method Cardinality-Agnostic Task Encoder (CATE) for Gaussian task embeddings, 2-Wasserstein distance for inter-task alignment.
result FedGTEA achieves superior classification performance and mitigates forgetting.

Proposes EOT eigenmaps for aligning and embedding multiple datasets.

problem Aligning and embedding multiple datasets with shared structures but individual distortions.
method Entropic Optimal Transport (EOT) eigenmaps, leveraging leading singular vectors of EOT plan matrix.
result Proves theoretical guarantees and favorable properties for aligning and embedding datasets.

This paper tackles UDA by learning domain-invariant embeddings using distribution alignment and pseudo-labels.

problem Unsupervised domain adaptation between two visual domains.
method Shared deep encoder, Sliced-Wasserstein Distance, deep classifier, pseudo-labels for class alignment.
result Effective solution for training deep classification networks on source domain to generalize to target domain.

The paper develops dynamic word embeddings to capture evolving language structures.

problem Capturing the evolving meanings and associations of words over time.
method Develops a dynamic statistical model to learn time-aware word vector representation, solving the alignment problem.
result The model reliably captures the evolution of language over time and outperforms state-of-the-art approaches.

Simple framework decouples word alignment and multilingual embedding mapping.

problem Learning multilingual embeddings without supervision.
method Two-stage approach: 1) unsupervised word alignment, 2) mapping embeddings to shared space.
result Robust performance across various multilingual tasks, including distant languages.

New graph kernel scales well with graph size and number, achieving state-of-the-art performance.

problem Graph kernels lose structure information when representing graphs.
method Proposes a positive-definite global alignment graph kernel using random features and random graph embeddings.
result Achieves quasi-linear scalability with respect to graph size and number.

DHGAK aligns substructures for better graph kernel performance.

problem Limited performance of traditional graph kernels due to missing substructure similarities.
method Hierarchically aligns relational substructures in deep embedding space, assigning same feature maps in RKHS.
result DHGAK outperforms state-of-the-art graph kernels on various benchmarks.

Proposes a new method for manifold alignment using geometry-regularized twin autoencoders.

problem Traditional MA methods lack out-of-sample extension and real-world applicability.
method Guided representation learning with geometry-regularized twin autoencoders.
result Improves cross-domain generalization and robustness while maintaining alignment fidelity.

Enhances thematic investing with stock embeddings from textual data.

problem Challenges in constructing thematic portfolios due to overlapping sector boundaries and evolving market dynamics.
method Introduces THEME, a framework that fine-tunes embeddings using hierarchical contrastive learning, aligning themes and stocks using their hierarchical relationship and incorporating stock returns.
result Theme-aligned portfolios demonstrate compelling performance, significantly outperforming large language models in thematic asset retrieval.

Proposes COALA method for learning audio representations aligned with tags.

problem Lack of annotated data for high-performance audio representation learning.
method Aligns latent representations of audio and tags using a contrastive loss.
result Audio embedding model captures both acoustic and semantic characteristics.

Proposes a method to align and differentiate feature clusters for unsupervised domain adaptation.

problem Difficulty in obtaining labeled data for domain adaptation.
method Label propagation and cycle consistency to align feature clusters.
result Successfully formed aligned and discriminative clusters for better domain adaptation.

A new method for unsupervised domain adaptation using manifold learning.

problem Leveraging rich source domain information to target domain without labeled data.
method Discriminative Manifold Embedding and Alignment framework.
result Consistent transferability and discriminability achieved through manifold metric alignment.

Proposes JK-EGW for multimodal alignment using shared latent space.

problem Aligning data from multiple modalities into a shared representation space.
method Joint kernel entropic Gromov--Wasserstein Optimal Transport (JK-EGW).
result Improved multimodal retrieval performance compared to baselines.

Proposes Gromov-Wasserstein methods for multi-view embedding.

problem Integrating multiple representations of the same samples in heterogeneous geometries.
method Gromov-Wasserstein optimal transport for multi-view embedding.
result Preserves intrinsic relational structure across views effectively.

Develops a topology-based test for AI model alignment and interpretability.

problem Testing and interpreting opaque AI models is difficult due to their complexity and lack of interpretability.
method Introduces a topology-based multi-modal alignment test to make AI models more interpretable.
result Demonstrates the effectiveness of the topology-based test in making AI model deployment and comparison more intuitive.

A method to simplify complex high-dimensional data visualization.

problem Difficult interpretation of linear projections in high-dimensional data.
method Decomposition of linear projections into axis-aligned projections using Dempster-Shafer theory.
result Linear projections can be effectively represented by a sparse set of axis-aligned projections, revealing more intuitive insights.

New post-processing methods improve word embedding performance.

problem Boosting the performance of word embeddings for similarity and analogy tasks.
method Optimizing a semi-Riemannian manifold with Centralised Kernel Alignment (CKA) to shrink the covariance matrix towards a scaled identity matrix.
result Improved performance on downstream tasks after smoothing the spectrum of word vectors.

Method embeds numeric tabular datasets into a shared vector space for similarity and retrieval.

problem Lack of meaningful representation for numeric tabular datasets in large language models.
method Structured exploratory data analysis descriptors, sentence transformer embedding, CCA for cross-dataset alignment.
result Total P@1 score of 0.9 across 15 datasets, robust nearest-neighbor retrieval and cluster structure.

Space-efficient feature maps improve string alignment kernel scalability.

problem String alignment kernels scale poorly with quadratic complexity, limiting large-scale applications.
method Presented SFMEDM, a space-efficient feature map for edit distance with moves using metric embedding and random Fourier features.
result Demonstrated superior performance of SFMEDM in prediction accuracy, scalability, and computation efficiency.

Enhancing spectral embedding for low-dimensional embeddings in rare disease cohorts

problem Representing clinical concepts and patients in electronic health records
method Spectral-based unsupervised learning with flexible knowledge transfer
result Outperforms competing approaches in challenging scenarios

Proposes a new neural network for text-dependent speaker verification.

problem Improves speaker verification by encoding phrase and speaker information.
method Uses differentiable alignment models to produce supervectors from utterances.
result Achieves competitive performance in text-dependent speaker verification tasks.

Transformer models align words through attention weights, closely approximating Optimal Transport.

problem Understanding the internal mechanism of transformer models in language processing.
method Empirical evidence and theoretical analysis of attention weights and their relation to Optimal Transport.
result Transformer models can simulate gradient descent on the dual of entropy-regularized OT problem, providing a theoretical foundation for token alignment.

A novel OT-based method for aligning hyperbolic representations.

problem Aligning different hyperbolic representations of hierarchical data.
method Optimal transport (OT) on the Poincaré model of hyperbolic spaces, using gyrobarycenter mapping.
result Both Euclidean and hyperbolic OT-based methods perform similarly in retrieval tasks.

Improved vision-language embeddings boost cross-task learning.

problem Creating general vision systems with better cross-task learning.
method Aligning image-word representations for better cross-task transfer.
result Improved inductive transfer from visual recognition to visual question answering.