Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

316394125 · May 202619922001200920182026
48 results for heterogeneous attributes

A new approach combines attributes and sequences for better item recommendations.

problem Difficulty in leveraging attribute information due to heterogeneity and sparseness.
method Heterogeneous Attribute Recurrent Neural Networks (HA-RNN) that incorporates heterogeneous attributes and captures sequential dependencies.
result Significant improvements over state-of-the-art models in item recommendation.

A new method improves semi-supervised learning by handling tasks with different attribute spaces.

problem Existing methods assume tasks share the same attribute space, limiting their applicability.
method Meta-learning approach that embeds labeled and unlabeled data in task-specific spaces using neural networks.
result Improves test performance on tasks with small labeled data using unlabeled and various task data.

Paper proposes TPathMine model for more accurate user attribute prediction.

problem Predicting user attributes from click data in heterogeneous networks.
method HetPathMine model with meta-path weights optimized for user emotional preferences.
result TPathMine model achieves higher accuracy in user attribute prediction.

Enhances machine learning for dynamic, interconnected entities.

problem Lack of systematic feature engineering for dynamic, interconnected entities.
method Augments current graph machine learning with comprehensive feature engineering in space and time.
result Improves supervised learning on heterogeneous, attributed entities interacting over time.

UNTIE learns representations of coupled categorical data.

problem Challenges in learning from unlabeled categorical data with complex couplings.
method UNTIE approach for unsupervised representation learning of heterogeneous couplings.
result UNTIE significantly improves categorical data representations on 25 diverse datasets.

Latent feature modeling allows capturing the latent structure responsible for generating the observed properties of a set of objects. It is often used to make predictions either for new values of interest or missing information in the original data, as well as to perform data exploratory analysis. However, although the…

2017-06-12abs ↗pdf ↗

A new method for name disambiguation in academic networks using multi-view attention and recurrent neural networks.

problem Disambiguating authors with the same name in large-scale academic networks.
method Multi-view Attention-based Pairwise Recurrent Neural Network (MA-PairRNN) that divides papers into blocks based on author attributes and merges blocks of the same author.
result MA-PairRNN significantly improves name disambiguation performance on real-world datasets.

Paper tackles embedding attributed sequences in unsupervised learning.

problem Mining tasks over attributed sequences with dependencies between sequences and attributes.
method Proposes a deep multimodal learning framework, NAS, for unsupervised learning of attributed sequences.
result NAS produces task-independent embeddings for various mining tasks on real-world datasets.

Paper proposes a method to detect fair communities in graphs considering demographic attributes.

problem Inconsistent community detection violates fairness constraints for nodes with demographic attributes.
method Develops an 1\ell_1-regularized pseudo-likelihood approach for fair graphical model selection.
result The method ensures demographic groups are fairly represented within detected communities.

A system for attributing ad effects using a neural network and Shapley values.

problem Attributing ad effects to individual ads in a complex, sequential environment.
method A two-step approach: response modeling with RNN and credit allocation with Shapley values.
result The system accurately allocates incremental ad effects to individual ads, handling sequence dependence.

SpecAE detects anomalies in attributed networks by projecting them into a tailored space.

problem Detecting anomalies in attributed networks with complex dependencies and nodal attributes.
method Spectral convolution and deconvolution framework, leveraging Laplacian sharpening and density estimation.
result SpecAE effectively detects global and community anomalies in attributed networks.

HL-VAE extends VAE for heterogeneous temporal and longitudinal data.

problem Handling heterogeneous data in temporal and longitudinal datasets.
method Proposes HL-VAE, an extension of existing VAEs for temporal and longitudinal data, incorporating likelihood models for various data types.
result HL-VAE achieves competitive performance in missing value imputation and predictive accuracy.

A method for trust evaluation of devices in human-device coexistence systems.

problem Efficient trust evaluation of devices in systems with diverse physical and social attributes.
method Canonical correlation analysis-enhanced hypergraph self-supervised learning (HSLCCA).
result The proposed HSLCCA method significantly outperforms baseline algorithms in identifying trusted devices.

HDGI learns node representations for heterogeneous graphs.

problem Challenges in learning node representations for heterogeneous graphs.
method HDGI uses meta-path structure, graph convolution, and semantic-level attention to maximize local-global mutual information.
result HDGI outperforms state-of-the-art methods on graph-related tasks.

VHGM-MAE generates synthetic humans from healthcare data.

problem Handling high-dimensional, sparse healthcare data with missing values.
method Masked autoencoder (MAE) tailored for healthcare data, addressing heterogeneity, missingness, and high-dimensionality.
result VHGM-MAE outperforms existing methods in missing value imputation and synthetic data generation.

Paper proposes a framework to predict therapeutic properties of compounds.

problem Predict therapeutic properties of compounds with heterogeneous data.
method Domain-adversarial multi-task framework using adversarial learning.
result Framework improves performance over competitive baselines.

In the presence of weak overall correlation, it may be useful to investigate if the correlation is significantly and substantially more pronounced over a subpopulation. Two different testing procedures are compared. Both are based on the rankings of the values of two variables from a data set with a large number n of o…

2015-04-21abs ↗pdf ↗

This paper models how funds choose between competing ESG rating methodologies based on investor preferences.

problem Competing ESG rating methodologies lead to different portfolio rewards and fund fees.
method Modeling funds with heterogeneous ESG priorities and analyzing portfolio changes and investor demand.
result Funds specialize more, but provider scores, investor participation, and equilibrium fees decrease in the benchmark equilibrium.

Two strategies for training network classifiers with feature heterogeneity.

problem Training network classifiers with agents having varying feature sizes and unreliable local decisions.
method Promotes global and local smoothing of classifier outputs.
result Output smoothing makes network classifier dynamics more complex, requiring regularization of parameters.

A new method corrects weight values to improve treatment effect estimation.

problem Estimating heterogeneous treatment effects in high-dimensional data with sample selection bias.
method Differentiable Pareto-Smoothed Weighting (DPSW) framework.
result Our method outperforms existing methods in treatment effect estimation.

TiFL divides clients into tiers to improve federated learning performance.

problem Heterogeneity in resource and data quality impacts FL performance.
method TiFL divides clients into tiers based on training performance and selects clients from the same tier in each training round.
result TiFL achieves faster training performance with comparable or better test accuracy.

ICAM creates interpretable feature attribution maps for brain images.

problem Challenges in predicting class relevance from brain images due to heterogeneity and background variation.
method A VAE-GAN framework for disentangling class relevance from background features.
result FA maps generated by ICAM outperform baseline methods and support phenotype variation exploration.

Study analyzes fintech terms in news and blogs, revealing specialized attributes of fintech companies.

problem Understanding specialized attributes of fintech companies through term analysis.
method Large scale analysis of fintech terms in news and blogs, using complex networks and statistically validated networks.
result Companies with fintech terms have over-expressions of specific attributes related to geography and economy.

Paper uses machine learning to model travel mode switching under a new transit system.

problem Modeling individual travel mode preferences and response to new mobility options.
method Interpretable machine learning approach to predict and interpret mode-switching behavior.
result Machine learning captures individual heterogeneity in travel mode choice.

Proposes LTD-RBM for robust and efficient latent truth discovery.

problem Discovering true values in noisy, conflicting or incomplete information.
method Restricted Boltzmann Machines (RBM) for a novel LTD algorithm.
result Superior to state-of-the-art LTD techniques in effectiveness, efficiency, and robustness.

Digital personas improve survey results for stable attributes but fail for subjective responses.

problem When can digital personas reliably approximate human survey findings?
method Using LISS panel, constructed personas from background variables and survey histories, tested against held-out post-cutoff answers.
result Digital personas improve alignment with human response distributions for stable attributes but fail for subjective responses.

ExCIR provides efficient, consistent, and scalable explainability for complex models.

problem Complex models lack transparency and require efficient, stable, and scalable explainability methods.
method ExCIR uses correlation-aware feature attribution with robust centering and groupwise aggregation.
result ExCIR delivers trustworthy agreement with global baselines and full model rankings, reduces runtime, and scales to large datasets.

SHAP Distance assesses semantic fidelity of synthetic tabular data.

problem Semantic fidelity of synthetic tabular data is not well evaluated.
method SHAP Distance, defined as cosine distance between global SHAP attribution vectors.
result SHAP Distance detects semantic discrepancies overlooked by standard measures.

K-Metamodes clusters security data without converting categorical attributes.

problem Clustering heterogeneous security data sets with categorical and numerical attributes.
method Frequency-based distance function for ensemble-based k-modes clustering, adapted feature discretisation.
result Higher effectiveness compared to previous methods on public security data sets.

FairMixRep learns fair representations from mixed data types.

problem Representation learning in mixed numerical and categorical data with fairness constraints.
method Efficient encoder-decoder framework + fairness constraints.
result Excellent performance in preserving information and fairness in mixed data representations.

Proposes a new model for joint probability distributions in computer vision.

problem Limitation of existing models in meeting diverse downstream tasks.
method Uses parametric conditional probability distributions for each group of variables conditioned on the rest.
result Models can be used for any downstream task without task-specific design.

Machine learning enhances wireless network authentication for diverse devices.

problem Complex dynamic wireless environments challenge conventional authentication methods.
method Intelligent authentication using machine learning for diverse physical layer attributes.
result Machine learning-based authentication provides cost-effective, reliable, and situation-aware security.

BIND removes background noise from binary matrices, improving detection accuracy and fairness.

problem Real data often violates the i.i.d assumption for binary matrix entries, leading to inaccurate detection.
method BIND optimizes detection by estimating row- and column-wise mixture distributions and eliminating background noise.
result BIND effectively removes background noise and increases detection accuracy and fairness.

HIRM models noisy, sparse, heterogeneous relational data using hierarchical clustering and Dirichlet processes.

problem Modeling noisy, sparse, and heterogeneous relational data.
method Hierarchical Chinese restaurant process and Dirichlet process mixture for clustering and modeling relation values.
result HIRM generalizes standard models and discovers relational structure in real-world datasets.

Paper proposes HIDAM model to improve MSE default risk assessment using heterogeneous information networks.

problem Default risk assessment for MSEs due to lack of credit information and diverse financial activities.
method HIDAM model incorporating heterogeneous information networks with multi-typed nodes and links, extracting interactive information through meta-paths, and using a hierarchical attention mechanism.
result HIDAM model outperforms state-of-the-art competitors on real-world banking data.

RUMBoost combines RUMs and deep learning for better choice modelling.

problem Creating interpretable and robust discrete choice models.
method Gradient Boosted Regression Trees for utility functions, with constraints for interpretability and monotonicity.
result RUMBoost outperforms ML and RUM benchmarks in predictive performance and interpretability.

Study copyright's impact on creative industries using AI-generated fonts.

problem Estimating supply and demand in creative industries with AI-generated content.
method Neural network embeddings, spatial regression, event-study analyses, structural model of supply and demand.
result Copyright can raise consumer welfare by encouraging product relocation.

The paper improves consumer preference modeling by considering multiple product categories.

problem Estimating consumer preferences across multiple product categories with varying attributes and price sensitivity.
method Extends matrix factorization techniques to account for time-varying product attributes and out-of-stock products, pooling information across categories to estimate heterogeneity in preferences.
result The model improves over traditional approaches, accurately estimating consumer preferences and price sensitivity.

CDOT optimizes transport between domains preserving both feature and geometric structure.

problem Optimizing transport between heterogeneous domains with preserved feature and geometric structure.
method CDOT uses operator-based regularization to align distance structures, proving pseudometric properties.
result CDOT improves robustness to local geometric variations and is provably convex.