This research shows how to learn shared representations from unpaired data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Enhancing spectral embedding for low-dimensional embeddings in rare disease cohorts
In deep neural nets, lower level embedding layers account for a large portion of the total number of parameters. Tikhonov regularization, graph-based regularization, and hard parameter sharing are approaches that introduce explicit biases into training in a hope to reduce statistical complexity. Alternatively, we propo…
The creation of social ties is largely determined by the entangled effects of people's similarities in terms of individual characters and friends. However, feature and structural characters of people usually appear to be correlated, making it difficult to determine which has greater responsibility in the formation of t…
Proposes EOT eigenmaps for aligning and embedding multiple datasets.
Recent works reveal that network embedding techniques enable many machine learning models to handle diverse downstream tasks on graph structured data. However, as previous methods usually focus on learning embeddings for a single network, they can not learn representations transferable on multiple networks. Hence, it i…
Kernel method embeds noisy datasets, capturing shared structures.
Word embeddings are a powerful approach for analyzing language, and exponential family embeddings (EFE) extend them to other types of data. Here we develop structured exponential family embeddings (S-EFE), a method for discovering embeddings that vary across related groups of data. We study how the word usage of U.S. C…
GDA-HIN adapts across heterogeneous networks by aligning shared and private node types.
Despite significant progress, deep reinforcement learning (RL) suffers from data-inefficiency and limited generalization. Recent efforts apply meta-learning to learn a meta-learner from a set of RL tasks such that a novel but related task could be solved quickly. Though specific in some ways, different tasks in meta-RL…
Bayesian interpretations of neural network have a long history, dating back to early work in the 1990's and have recently regained attention because of their desirable properties like uncertainty estimation, model robustness and regularisation. We want to discuss here the application of Bayesian models to knowledge sha…
ALF reduces network parameters and operations by 70% and 61%, respectively, on embedded hardware.
Simple framework decouples word alignment and multilingual embedding mapping.
DECAT framework evaluates multimodal models for shared biology, detecting confounders and false positives.
Acoustic word embeddings --- fixed-dimensional vector representations of arbitrary-length words --- have attracted increasing interest in query-by-example spoken term detection. Recently, on the fact that the orthography of text labels partly reflects the phonetic similarity between the words' pronunciation, a multi-vi…
Every day, hundreds of millions of new Tweets containing over 40 languages of ever-shifting vernacular flow through Twitter. Models that attempt to extract insight from this firehose of information must face the torrential covariate shift that is endemic to the Twitter platform. While regularly-retrained algorithms can…
Method embeds numeric tabular datasets into a shared vector space for similarity and retrieval.
Anchor PCA improves robustness in multi-domain PCA.
SPIRE enables efficient federated learning for diffusion models by separating client-specific embeddings from a shared backbone.
CMF is a technique for simultaneously learning low-rank representations based on a collection of matrices with shared entities. A typical example is the joint modeling of user-item, item-property, and user-feature matrices in a recommender system. The key idea in CMF is that the embeddings are shared across the matrice…
MCPCA analyzes shared factors across multiple data contexts.
Graph embedding leaks sensitive graph properties and subgraphs.
MPP trains a transformer to predict multiple physical systems, improving accuracy across various tasks.
Enhances generative model accuracy through knowledge transfer.
A new method combines multiple node embeddings using tensor decomposition.
In the recent years money laundering schemes have grown in complexity and speed of realization, affecting financial institutions and millions of customers globally. Strengthened privacy policies, along with in-country regulations, make it hard for banks to inner- and cross-share, and report suspicious activities for th…
We propose a novel framework for multi-task reinforcement learning (MTRL). Using a variational inference formulation, we learn policies that generalize across both changing dynamics and goals. The resulting policies are parametrized by shared parameters that allow for transfer between different dynamics and goal condit…
Method captures shared information across many views robustly.
Enhances graph classification with multiple graphs.
A method for learning embeddings from multi-view data using Gromov-Wasserstein.
This work explains how maximizing latent correlations across multiple data views helps in identifying shared and private components.
Proposes MANE for multi-view network embedding, improving node representations.
The paper constructs -hypersurfaces for and .
The hyperbolic manifold is a smooth manifold of negative constant curvature. While the hyperbolic manifold is well-studied in the literature, it has gained interest in the machine learning and natural language processing communities lately due to its usefulness in modeling continuous hierarchies. Tasks with hierarchica…
Study uses trajectory embedding to measure place function similarity at fine spatial granularity.
PIMA autoencoders discover shared features in multimodal scientific data.
We propose to formulate multi-label learning as a estimation of class distribution in a non-linear embedding space, where for each label, its positive data embeddings and negative data embeddings distribute compactly to form a positive component and negative component respectively, while the positive component and nega…
The heterogeneity-gap between different modalities brings a significant challenge to multimedia information retrieval. Some studies formalize the cross-modal retrieval tasks as a ranking problem and learn a shared multi-modal embedding space to measure the cross-modality similarity. However, previous methods often esta…
Representation learning has recently been successfully used to create vector representations of entities in language learning, recommender systems and in similarity learning. Graph embeddings exploit the locality structure of a graph and generate embeddings for nodes which could be words in a language, products of a re…
A new method is proposed to obtain the risk neutral probability of share prices without stochastic calculus and price modeling, via an embedding of the price return modeling problem in Le Cam's statistical experiments framework. Strategies-probabilities and are thus determined and used, respective…
For embedded 2-spheres in a 4-manifold sharing the same embedded transverse sphere homotopy implies isotopy, provided the ambient 4-manifold has no $\BZ_2$-torsion in the fundamental group. This gives a generalization of the classical light bulb trick to 4-dimensions, the uniqueness of spanning discs for a simple close…
Training-free source selection for LLM families with shared vocabularies
Traditionally, many text-mining tasks treat individual word-tokens as the finest meaningful semantic granularity. However, in many languages and specialized corpora, words are composed by concatenating semantically meaningful subword structures. Word-level analysis cannot leverage the semantic information present in su…
This paper studies how to find compact state embeddings from high-dimensional Markov state trajectories, where the transition kernel has a small intrinsic rank. In the spirit of diffusion map, we propose an efficient method for learning a low-dimensional state embedding and capturing the process's dynamics. This idea a…
Word embedding maps words into a low-dimensional continuous embedding space by exploiting the local word collocation patterns in a small context window. On the other hand, topic modeling maps documents onto a low-dimensional topic space, by utilizing the global word collocation patterns in the same document. These two …
The paper describes complex structures on Oeljeklaus-Toma manifolds.
IITK wins FinSim 2020 task on financial hypernym detection.
TransINT embeds KGs by preserving implication rules, outperforming existing methods.