A new method improves semi-supervised learning by handling tasks with different attribute spaces.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Attributes, such as metadata and profile, carry useful information which in principle can help improve accuracy in recommender systems. However, existing approaches have difficulty in fully leveraging attribute information due to practical challenges such as heterogeneity and sparseness. These approaches also fail to c…
Enhances machine learning for dynamic, interconnected entities.
CAP adapts optimization to class attributes for better fairness.
User-based attribute information, such as age and gender, is usually considered as user privacy information. It is difficult for enterprises to obtain user-based privacy attribute information. However, user-based privacy attribute information has a wide range of applications in personalized services, user behavior anal…
UNTIE learns representations of coupled categorical data.
Anomaly detection aims to distinguish observations that are rare and different from the majority. While most existing algorithms assume that instances are i.i.d., in many practical scenarios, links describing instance-to-instance dependencies and interactions are available. Such systems are called attributed networks. …
Flexible inference model for multilayer networks with heterogeneous data.
This paper introduces a general Bayesian non- parametric latent feature model suitable to per- form automatic exploratory analysis of heterogeneous datasets, where the attributes describing each object can be either discrete, continuous or mixed variables. The proposed model presents several important properties. First…
Latent feature modeling allows capturing the latent structure responsible for generating the observed properties of a set of objects. It is often used to make predictions either for new values of interest or missing information in the original data, as well as to perform data exploratory analysis. However, although the…
Mining tasks over sequential data, such as clickstreams and gene sequences, require a careful design of embeddings usable by learning algorithms. Recent research in feature learning has been extended to sequential data, where each instance consists of a sequence of heterogeneous items with a variable length. However, m…
A new method for name disambiguation in academic networks using multi-view attention and recurrent neural networks.
Paper proposes a method to detect fair communities in graphs considering demographic attributes.
ICAM creates interpretable feature attribution maps for brain images.
FairMixRep learns fair representations from mixed data types.
CoMGNN models heterogeneous graphs with evolving nodes and edges.
NFTs with diverse rare attributes sell at higher prices.
HL-VAE extends VAE for heterogeneous temporal and longitudinal data.
A method for trust evaluation of devices in human-device coexistence systems.
VHGM-MAE generates synthetic humans from healthcare data.
Graph representation learning is to learn universal node representations that preserve both node attributes and structural information. The derived node representations can be used to serve various downstream tasks, such as node classification and node clustering. When a graph is heterogeneous, the problem becomes more…
In the presence of weak overall correlation, it may be useful to investigate if the correlation is significantly and substantially more pronounced over a subpopulation. Two different testing procedures are compared. Both are based on the rankings of the values of two variables from a data set with a large number n of o…
Every design choice will have different effects on different units. However traditional A/B tests are often underpowered to identify these heterogeneous effects. This is especially true when the set of unit-level attributes is high-dimensional and our priors are weak about which particular covariates are important. How…
A new method for disentangled latent spaces in VAEs that can manipulate attributes.
CDOT optimizes transport between domains preserving both feature and geometric structure.
A new method corrects weight values to improve treatment effect estimation.
Nowadays processing of Big Security Data, such as log messages, is commonly used for intrusion detection purposed. Its heterogeneous nature, as well as combination of numerical and categorical attributes does not allow to apply the existing data mining methods directly on the data without feature preprocessing. Therefo…
Financial markets provide an ideal frame for studying decision making in crowded environments. Both the amount and accuracy of the data allows to apply tools and concepts coming from physics that studies collective and emergent phenomena or self-organised and highly heterogeneous systems. We analyse the activity of 29,…
Study copyright's impact on creative industries using AI-generated fonts.
With the rapid development of high-throughput technologies, parallel acquisition of large-scale drug-informatics data provides huge opportunities to improve pharmaceutical research and development. One significant application is the purpose prediction of small molecule compounds, aiming to specify therapeutic propertie…
Recent years have witnessed an increased focus on interpretability and the use of machine learning to inform policy analysis and decision making. This paper applies machine learning to examine travel behavior and, in particular, on modeling changes in travel modes when individuals are presented with a novel (on-demand)…
Study analyzes fintech terms in news and blogs, revealing specialized attributes of fintech companies.
RUMBoost combines RUMs and deep learning for better choice modelling.
Federated Learning (FL) enables learning a shared model across many clients without violating the privacy requirements. One of the key attributes in FL is the heterogeneity that exists in both resource and data due to the differences in computation and communication capacity, as well as the quantity and content of data…
SLOGAN improves GANs' conditional generation by balancing latent attribute distributions.
This paper presents a Semantic Attribute Modulation (SAM) for language modeling and style variation. The semantic attribute modulation includes various document attributes, such as titles, authors, and document categories. We consider two types of attributes, (title attributes and category attributes), and a flexible a…
Proposes VCLANC for attributed network clustering using node and attribute embeddings.
New insights into how to inspect and learn from multi-stage processes and AI reasoning.
Digital personas improve survey results for stable attributes but fail for subjective responses.
Cluster analysis of very high dimensional data can benefit from the properties of such high dimensionality. Informally expressed, in this work, our focus is on the analogous situation when the dimensionality is moderate to small, relative to a massively sized set of observations. Mathematically expressed, these are the…
ExCIR provides efficient, consistent, and scalable explainability for complex models.
The fifth generation (5G) and beyond wireless networks are critical to support diverse vertical applications by connecting heterogeneous devices and machines, which directly increase vulnerability for various spoofing attacks. Conventional cryptographic and physical layer authentication techniques are facing some chall…
WGNN learns graph representations from incomplete attribute data.
We address the problem of latent truth discovery, LTD for short, where the goal is to discover the underlying true values of entity attributes in the presence of noisy, conflicting or incomplete information. Despite a multitude of algorithms to address the LTD problem that can be found in literature, only little is kno…
Social network analysis is an important problem in data mining. A fundamental step for analyzing social networks is to encode network data into low-dimensional representations, i.e., network embeddings, so that the network topology structure and other attribute information can be effectively preserved. Network represen…
AI-enhanced product embeddings boost demand analysis accuracy.
SHAP Distance assesses semantic fidelity of synthetic tabular data.
Introduces FairCOCCO for fair learning with multitype, multivariate sensitive attributes.