Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

56113169225 · Jun 202019922001200920182026
48 results for shared embeddings

Enhancing spectral embedding for low-dimensional embeddings in rare disease cohorts

problem Representing clinical concepts and patients in electronic health records
method Spectral-based unsupervised learning with flexible knowledge transfer
result Outperforms competing approaches in challenging scenarios

Proposes SSE, a data-driven method to regularize embedding layers in neural nets.

problem Reduces overfitting in embedding layers of neural nets.
method Stochastically shares embeddings during SGD, integrating with existing algorithms.
result Improves generalization on various tasks, including recommender systems and natural language models.

Improved acoustic word embeddings using shared decoder in multi-view encoders.

problem Learning discriminative acoustic word embeddings from text labels.
method Combining Siamese multi-view encoders with a shared decoder network to maximize the relationship between acoustic and text embeddings.
result 11.1% relative improvement in average precision on acoustic word discrimination task with WSJ dataset.

AN2VEC disentangles feature and structural information in social networks.

problem Difficulty in separating feature and structural information in social networks.
method Graph Convolutional Networks (GCN) Variational Autoencoder for disentangled node embeddings.
result AN2VEC captures joint information of structure and features better than unshared information.

Proposes EOT eigenmaps for aligning and embedding multiple datasets.

problem Aligning and embedding multiple datasets with shared structures but individual distortions.
method Entropic Optimal Transport (EOT) eigenmaps, leveraging leading singular vectors of EOT plan matrix.
result Proves theoretical guarantees and favorable properties for aligning and embedding datasets.

Kernel method embeds noisy datasets, capturing shared structures.

problem Limited power in capturing nonlinear structures, noisiness, high-dimensionality, and interpretability issues.
method Kernel spectral joint embeddings using duo-landmark integral operators.
result Consistent recovery of low-dimensional noiseless signals and convergence to eigenfunctions of integral operators.

Word embeddings are a powerful approach for analyzing language, and exponential family embeddings (EFE) extend them to other types of data. Here we develop structured exponential family embeddings (S-EFE), a method for discovering embeddings that vary across related groups of data. We study how the word usage of U.S. C…

2017-09-28abs ↗pdf ↗

GDA-HIN adapts across heterogeneous networks by aligning shared and private node types.

problem Domain adaptation challenges in heterogeneous networks with shared and private node types.
method Generalized Domain Adaptive model across HINs (GDA-HIN) that aligns identical-type nodes and edges while utilizing different-type nodes and edges.
result GDA-HIN outperforms state-of-the-art methods in various domain adaptation tasks across heterogeneous networks.

ALF reduces network parameters and operations by 70% and 61%, respectively, on embedded hardware.

problem Efficient deployment of deep learning models on resource-constrained hardware.
method Autoencoder-based low-rank filter-sharing technique.
result ALF achieves significant compression with minimal accuracy loss.

Simple framework decouples word alignment and multilingual embedding mapping.

problem Learning multilingual embeddings without supervision.
method Two-stage approach: 1) unsupervised word alignment, 2) mapping embeddings to shared space.
result Robust performance across various multilingual tasks, including distant languages.

DECAT framework evaluates multimodal models for shared biology, detecting confounders and false positives.

problem Determining if multimodal models learn shared biology or just confounders.
method DECAT framework classifies multimodal representations into four diagnostic scenarios using null-referenced metrics.
result DECAT detects confounders and false positives in multimodal models, improving with larger cohorts and stronger representations.

Method embeds numeric tabular datasets into a shared vector space for similarity and retrieval.

problem Lack of meaningful representation for numeric tabular datasets in large language models.
method Structured exploratory data analysis descriptors, sentence transformer embedding, CCA for cross-dataset alignment.
result Total P@1 score of 0.9 across 15 datasets, robust nearest-neighbor retrieval and cluster structure.

Paper proposes a new topology for AML analysis using Poincaré embeddings.

problem Complex money laundering schemes and regulatory constraints hinder AML analysis and information sharing.
method Proposes a new topology for AML analysis using Poincaré embeddings.
result Demonstrates improved AML analysis and information sharing through Poincaré embeddings.

SPIRE enables efficient federated learning for diffusion models by separating client-specific embeddings from a shared backbone.

problem Large diffusion models are impractical for federated learning due to their size.
method SPIRE separates the network into a global backbone and client-specific embeddings, enabling efficient finetuning.
result SPIRE achieves parameter-efficient finetuning, updating only a small fraction of weights.

CMF is a technique for simultaneously learning low-rank representations based on a collection of matrices with shared entities. A typical example is the joint modeling of user-item, item-property, and user-feature matrices in a recommender system. The key idea in CMF is that the embeddings are shared across the matrice…

2013-12-20abs ↗pdf ↗

MPP trains a transformer to predict multiple physical systems, improving accuracy across various tasks.

problem Training models for specific physical systems is inefficient and requires fine-tuning.
method MPP trains a shared transformer on multiple heterogeneous physical systems, projecting fields into a shared embedding space.
result A single MPP-pretrained transformer outperforms task-specific models on all pretraining sub-tasks and downstream tasks.

VASE learns disentangled representations that generalize across domains.

problem Learning new knowledge from diverse data sources while preserving old knowledge.
method VASE uses shared embeddings and Minimum Description Length principle to disentangle representations.
result VASE achieves better cross-domain inference and disentangled representations.

Method captures shared information across many views robustly.

problem Modeling hundreds of views per event and learning robust embeddings without view knowledge.
method View bootstrapping using multi-view correlation and matrix concentration theory.
result View bootstrapping captures shared information across many views robustly.

New approach for sharing deep learning costs between devices and cloud.

problem Prohibitive deep learning computational requirements for embedded devices.
method Study of representation compressibility in MobileNetV2 for balancing computation, bandwidth, and accuracy.
result An optimal splitting layer for network can be found with a simple PCA-based compression scheme.

A method for learning embeddings from multi-view data using Gromov-Wasserstein.

problem Challenges in learning low-dimensional representations from multi-view relational data with differing geometries.
method Bary-GWMDS and Mean-GWMDS-C, Gromov-Wasserstein-based methods operating on distance matrices.
result Stable and geometrically meaningful embeddings learned from synthetic and real-world datasets.

This work explains how maximizing latent correlations across multiple data views helps in identifying shared and private components.

problem Understanding how to identify shared and private components in multiview data.
method An intuitive generative model of multiview data is adopted, and latent correlation maximization is shown to guarantee the extraction of shared components.
result Latent correlation maximization guarantees the extraction of shared components across views and disentangles private information.

Proposes MANE for multi-view network embedding, improving node representations.

problem Learning low-dimensional representations from multiple views of networks.
method MANE combines diversity and collaboration, including second-order collaboration, and attention-based extension MANE+.
result MANE+ outperforms state-of-the-art approaches on real-world multi-view networks.

The paper constructs λλ-hypersurfaces for λ>0λ>0 and λ<0λ<0.

problem Exploring λλ-hypersurfaces in different λλ-values and their properties.
method Constructing complete embedded and non-convex λλ-hypersurfaces diffeomorphic to a cylinder and doughnut-shaped.
result For λ>0λ>0, complete embedded and non-convex λλ-hypersurfaces are constructed, diffeomorphic to a cylinder.

This work learns low-rank hyperbolic embeddings for tasks with hierarchical structures.

problem Learning hyperbolic embeddings of tasks with hierarchical structures.
method Formulated as manifold optimization problems and proposed computationally efficient algorithms.
result Efficacy of the proposed approach demonstrated through empirical results.

PIMA autoencoders discover shared features in multimodal scientific data.

problem Discovering shared information in high-throughput scientific datasets.
method Physics-informed multimodal autoencoders (PIMA) with Gaussian mixture prior and product of experts formulation.
result Accurate cross-modal inference between images and mechanical stress-strain response in lattice metamaterials.

Study uses trajectory embedding to measure place function similarity at fine spatial granularity.

problem Measuring place function similarity at fine spatial granularity.
method Trajectory embedding to reduce dimensions and measure similarity of place functions.
result Embedding similarity can be a metric proxy for place functions at fine spatial granularity.

Chimera model combines link, content, and time for dynamic network analysis.

problem Community detection and prediction in evolving networks with dynamic changes.
method Shared factorization model that accounts for graph links, content, and temporal analysis.
result The approach simplifies temporal analysis and enables future community prediction.

The heterogeneity-gap between different modalities brings a significant challenge to multimedia information retrieval. Some studies formalize the cross-modal retrieval tasks as a ranking problem and learn a shared multi-modal embedding space to measure the cross-modality similarity. However, previous methods often esta…

2017-02-04abs ↗pdf ↗

DSE learns transferable skills across changing dynamics and goals.

problem Learning transferable skills across different reinforcement learning tasks.
method Variational inference for multi-task reinforcement learning with shared and task-specific latent spaces.
result Policies can generalize to unseen dynamics and goals conditions.

Paper improves geographic location embeddings using Flickr tags and structured data.

problem Lack of integration between Flickr metadata and structured scientific data.
method Learning vector space embeddings of geographic locations.
result Improved predictions of ecological features using the new method.

A new method is proposed to obtain the risk neutral probability of share prices without stochastic calculus and price modeling, via an embedding of the price return modeling problem in Le Cam's statistical experiments framework. Strategies-probabilities Pt0,nP_{t_0,n} and PT,nP_{T,n} are thus determined and used, respective…

2013-04-17abs ↗pdf ↗