LPL optimizes embeddings to align local neighborhoods, improving cross-lingual word alignment.
problem Aligning embeddings across different datasets and languages.
method Locality Preserving Loss (LPL) optimizes model to project embeddings while maintaining local neighborhoods and aligning them.
result LPL-based alignment leads to better and consistent accuracy, especially in small training set settings.
Develops techniques to align word embeddings from different sources.
problem Aligning word embeddings from diverse datasets or methods.
method Simple closed-form techniques for optimal rotation, translation, and scaling.
result Maximizes cosine similarity and minimizes root mean squared errors.
Study benchmarks embedding-based entity alignment methods for KGs.
problem Align entities across different KGs using embeddings.
method Survey and categorize 23 embedding-based methods, propose new KG sampling algorithm, develop open-source library.
result Understood strengths and limitations of embedding-based methods.
Researchers mapped clinical jargon to consumer language using embeddings alignment.
problem Mapping and translating clinical jargon to consumer language for better communication.
method Trained embeddings on clinical and consumer language corpora, aligned them using Procrustes algorithm, and refined with adversarial training.
result Procrustes algorithm effectively aligned clinical and consumer language embeddings.
G-CREWE efficiently aligns large networks using node embeddings and compression.
problem Efficiently aligning large networks for various applications.
method Uses node embeddings and compression to align networks at fine and coarse resolutions.
result G-CREWE achieves efficient and accurate network alignment, twice as fast as existing methods.
New method aligns word embeddings without supervision.
problem Aligning word embeddings trained on monolingual data.
method Wasserstein Procrustes approach with orthogonal and permutation matrices.
result Obtains state-of-the-art results in unsupervised word translation.
Proposes ACCA for better alignment of multiple data perspectives.
problem Unclear alignment between multiple data perspectives.
method Iteratively solves alignment and multi-view embedding.
result Improved alignment and embedding of multiple data perspectives.
Geometric approach for unsupervised word embedding alignment.
problem Learning alignment between word embeddings of source and target languages.
method Formulates alignment as domain adaptation on the manifold of doubly stochastic matrices, employing Riemannian conjugate gradient algorithm.
result Empirically outperforms state-of-the-art methods on bilingual lexicon induction tasks.
Paper proposes unsupervised knowledge graph alignment with adversarial learning.
problem Aligning knowledge graphs from different sources or languages without large amounts of aligned triplets.
method Adversarial learning framework to align entity and relation embeddings, with mutual information regularization.
result Framework effectively aligns knowledge graphs in unsupervised and weakly-supervised settings.
Paper proposes a method to align word embedding models in a joint latent space.
problem Challenges in aligning variations of word embedding models.
method Generative process using synthetic data points based on linguistic relationships.
result Substantial improvements in recovering embeddings of local neighborhoods.
Method trains shared embedding to align inputs and outputs for domain adaptation.
problem Unsupervised domain adaptation between different data domains.
method Regularized conditional alignment objective function and adversarial regularization.
result Improves classifier performance on unseen domain.
FedGTEA learns new tasks in federated learning with task embeddings and alignment.
problem Federated class-incremental learning with task-specific knowledge and model uncertainty.
method Cardinality-Agnostic Task Encoder (CATE) for Gaussian task embeddings, 2-Wasserstein distance for inter-task alignment.
result FedGTEA achieves superior classification performance and mitigates forgetting.
Proposes EOT eigenmaps for aligning and embedding multiple datasets.
problem Aligning and embedding multiple datasets with shared structures but individual distortions.
method Entropic Optimal Transport (EOT) eigenmaps, leveraging leading singular vectors of EOT plan matrix.
result Proves theoretical guarantees and favorable properties for aligning and embedding datasets.
This paper tackles UDA by learning domain-invariant embeddings using distribution alignment and pseudo-labels.
problem Unsupervised domain adaptation between two visual domains.
method Shared deep encoder, Sliced-Wasserstein Distance, deep classifier, pseudo-labels for class alignment.
result Effective solution for training deep classification networks on source domain to generalize to target domain.
The paper develops dynamic word embeddings to capture evolving language structures.
problem Capturing the evolving meanings and associations of words over time.
method Develops a dynamic statistical model to learn time-aware word vector representation, solving the alignment problem.
result The model reliably captures the evolution of language over time and outperforms state-of-the-art approaches.
Simple framework decouples word alignment and multilingual embedding mapping.
problem Learning multilingual embeddings without supervision.
method Two-stage approach: 1) unsupervised word alignment, 2) mapping embeddings to shared space.
result Robust performance across various multilingual tasks, including distant languages.
New graph kernel scales well with graph size and number, achieving state-of-the-art performance.
problem Graph kernels lose structure information when representing graphs.
method Proposes a positive-definite global alignment graph kernel using random features and random graph embeddings.
result Achieves quasi-linear scalability with respect to graph size and number.
DHGAK aligns substructures for better graph kernel performance.
problem Limited performance of traditional graph kernels due to missing substructure similarities.
method Hierarchically aligns relational substructures in deep embedding space, assigning same feature maps in RKHS.
result DHGAK outperforms state-of-the-art graph kernels on various benchmarks.
Proposes a new method for manifold alignment using geometry-regularized twin autoencoders.
problem Traditional MA methods lack out-of-sample extension and real-world applicability.
method Guided representation learning with geometry-regularized twin autoencoders.
result Improves cross-domain generalization and robustness while maintaining alignment fidelity.
Enhances thematic investing with stock embeddings from textual data.
problem Challenges in constructing thematic portfolios due to overlapping sector boundaries and evolving market dynamics.
method Introduces THEME, a framework that fine-tunes embeddings using hierarchical contrastive learning, aligning themes and stocks using their hierarchical relationship and incorporating stock returns.
result Theme-aligned portfolios demonstrate compelling performance, significantly outperforming large language models in thematic asset retrieval.
Proposes a method to align language and image data.
problem Aligning language and image data for better understanding.
method Uses triplet loss to learn consistent embeddings of language and images.
result Outperforms four baselines across multiple metrics.
DANE adapts network embeddings across multiple domains.
problem Learning embeddings for multiple networks without transferability.
method Graph Convolutional Network with adversarial learning.
result DANE achieves superior performance in cross-network domain adaptation.
BC-Aligner maintains backward compatibility of embeddings after frequent updates.
problem Updating embeddings without requiring consumer teams to retrain their models.
method Learning backward compatible embeddings through BC-Aligner.
result BC-Aligner maintains backward compatibility with existing unintended tasks after multiple model version updates.
Graph-Relational Domain Adaptation (GRDA) adapts domains based on their graph structure.
problem Uniform alignment of domains ignores topological structures.
method Uses a domain graph to encode adjacency and a novel graph discriminator.
result Empirically shows improved generalization and domain information incorporation.
Proposes COALA method for learning audio representations aligned with tags.
problem Lack of annotated data for high-performance audio representation learning.
method Aligns latent representations of audio and tags using a contrastive loss.
result Audio embedding model captures both acoustic and semantic characteristics.
RESTA defends LLMs against jailbreaking attacks by adding random noise to embeddings.
problem Vulnerability of LLMs to jailbreaking attacks that generate harmful outputs.
method Adds random noise to embedding vectors and aggregates during token generation.
result RESTA achieves superior robustness versus utility tradeoffs compared to baseline defenses.
Proposes a method to align and differentiate feature clusters for unsupervised domain adaptation.
problem Difficulty in obtaining labeled data for domain adaptation.
method Label propagation and cycle consistency to align feature clusters.
result Successfully formed aligned and discriminative clusters for better domain adaptation.
A new method for unsupervised domain adaptation using manifold learning.
problem Leveraging rich source domain information to target domain without labeled data.
method Discriminative Manifold Embedding and Alignment framework.
result Consistent transferability and discriminability achieved through manifold metric alignment.
DEAL model predicts links for new nodes with only attribute info.
problem Predicting links for new nodes with only attribute info.
method DEAL model with two encoders and alignment mechanism.
result DEAL significantly outperforms existing methods on inductive link prediction.
Model learns multilingual word representations robust to noise.
problem Learning multilingual word representations in noisy environments.
method Fit a generative latent variable model to a multilingual dictionary.
result Competitive multilingual embeddings across various tasks.
Proposes JK-EGW for multimodal alignment using shared latent space.
problem Aligning data from multiple modalities into a shared representation space.
method Joint kernel entropic Gromov--Wasserstein Optimal Transport (JK-EGW).
result Improved multimodal retrieval performance compared to baselines.
Proposes Gromov-Wasserstein methods for multi-view embedding.
problem Integrating multiple representations of the same samples in heterogeneous geometries.
method Gromov-Wasserstein optimal transport for multi-view embedding.
result Preserves intrinsic relational structure across views effectively.
Develops a topology-based test for AI model alignment and interpretability.
problem Testing and interpreting opaque AI models is difficult due to their complexity and lack of interpretability.
method Introduces a topology-based multi-modal alignment test to make AI models more interpretable.
result Demonstrates the effectiveness of the topology-based test in making AI model deployment and comparison more intuitive.
We solve a key problem in cross-lingual learning using a novel approach.
problem Aligning word embeddings across different languages.
method We devise a direct solution to the Wasserstein-Procrustes problem.
result Our method improves existing UCL approaches significantly.
A method to simplify complex high-dimensional data visualization.
problem Difficult interpretation of linear projections in high-dimensional data.
method Decomposition of linear projections into axis-aligned projections using Dempster-Shafer theory.
result Linear projections can be effectively represented by a sparse set of axis-aligned projections, revealing more intuitive insights.
New post-processing methods improve word embedding performance.
problem Boosting the performance of word embeddings for similarity and analogy tasks.
method Optimizing a semi-Riemannian manifold with Centralised Kernel Alignment (CKA) to shrink the covariance matrix towards a scaled identity matrix.
result Improved performance on downstream tasks after smoothing the spectrum of word vectors.
Unified deep architecture for domain-invariant network alignment.
problem Eliminate domain representation bias in network alignment.
method DANA (Domain Adversarial Network Alignment) using graph convolutional networks and semi-supervised learning.
result Achieves state-of-the-art alignment results on real-world social networks.
This paper improves cross-domain learning using random forests for manifold alignment.
problem Improving cross-domain learning and feature integration.
method Semi-supervised manifold alignment using random forest proximities.
result Random forest proximities enhance downstream classification accuracy.
Method embeds numeric tabular datasets into a shared vector space for similarity and retrieval.
problem Lack of meaningful representation for numeric tabular datasets in large language models.
method Structured exploratory data analysis descriptors, sentence transformer embedding, CCA for cross-dataset alignment.
result Total P@1 score of 0.9 across 15 datasets, robust nearest-neighbor retrieval and cluster structure.
Space-efficient feature maps improve string alignment kernel scalability.
problem String alignment kernels scale poorly with quadratic complexity, limiting large-scale applications.
method Presented SFMEDM, a space-efficient feature map for edit distance with moves using metric embedding and random Fourier features.
result Demonstrated superior performance of SFMEDM in prediction accuracy, scalability, and computation efficiency.
Enhancing spectral embedding for low-dimensional embeddings in rare disease cohorts
problem Representing clinical concepts and patients in electronic health records
method Spectral-based unsupervised learning with flexible knowledge transfer
result Outperforms competing approaches in challenging scenarios
A new method embeds distributions in a common space for optimal transport comparison.
problem Comparing distributions in different metric spaces.
method Sub-embedding robust Wasserstein (SERW) distance.
result SERW mimics GW distance properties and provides a cost relation.
Paper proposes KE-GCN for better graph embedding.
problem Efficiently leveraging complex graphs with heterogeneous entities and relations.
method Combines GCNs and knowledge embedding methods.
result Advantages in knowledge graph alignment and entity classification.
Proposes a new neural network for text-dependent speaker verification.
problem Improves speaker verification by encoding phrase and speaker information.
method Uses differentiable alignment models to produce supervectors from utterances.
result Achieves competitive performance in text-dependent speaker verification tasks.
Transformer models align words through attention weights, closely approximating Optimal Transport.
problem Understanding the internal mechanism of transformer models in language processing.
method Empirical evidence and theoretical analysis of attention weights and their relation to Optimal Transport.
result Transformer models can simulate gradient descent on the dual of entropy-regularized OT problem, providing a theoretical foundation for token alignment.
A novel OT-based method for aligning hyperbolic representations.
problem Aligning different hyperbolic representations of hierarchical data.
method Optimal transport (OT) on the Poincaré model of hyperbolic spaces, using gyrobarycenter mapping.
result Both Euclidean and hyperbolic OT-based methods perform similarly in retrieval tasks.
Improved vision-language embeddings boost cross-task learning.
problem Creating general vision systems with better cross-task learning.
method Aligning image-word representations for better cross-task transfer.
result Improved inductive transfer from visual recognition to visual question answering.
We introduce BilBOWA (Bilingual Bag-of-Words without Alignments), a simple and computationally-efficient model for learning bilingual distributed representations of words which can scale to large monolingual datasets and does not require word-aligned parallel training data. Instead it trains directly on monolingual dat…