Research in several fields now requires the analysis of data sets in which multiple high-dimensional types of data are available for a common set of objects. In particular, The Cancer Genome Atlas (TCGA) includes data from several diverse genomic technologies on the same cancerous tumor samples. In this paper we introd…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
ProJIVE integrates multiple data types to explain joint and individual variation.
This paper analyzes AJIVE for estimating shared subspace across multiple datasets, revealing its strengths and limitations.
sJIVE combines structure and prediction in multi-source data.
Paper uses JIVE to decompose word embeddings, improving sentiment analysis performance.
Integrative analysis of disparate data blocks measured on a common set of experimental subjects is a major challenge in modern data analysis. This data structure naturally motivates the simultaneous exploration of the joint and individual variation within each data block resulting in new insights. For instance, there i…
Unified model learns joint and individual features from brain imaging data.
Proposes HeteroJIVE for joint subspace estimation in multi-view data with statistical and structural heterogeneity.
Paper uses non-Euclidean analysis to classify brain structure variations.
This work develops a model to distinguish network and covariate information.
EB-VAE combines tumor growth and dropout data for personalized treatment response modeling.
We present a non-parametric prognostic framework for individualized event prediction based on joint modeling of both longitudinal and time-to-event data. Our approach exploits a multivariate Gaussian convolution process (MGCP) to model the evolution of longitudinal signals and a Cox model to map time-to-event data with…
Proposes a new method to explain complex machine learning models.
We study the cross-correlation matrix of inventory variations of the most active individual and institutional investors in an emerging market to understand the dynamics of inventory variations. We find that the distribution of cross-correlation coefficient has a power-law form in the bulk followed by …
In systems biomedicine, an experimenter encounters different potential sources of variation in data such as individual samples, multiple experimental conditions, and multi-variable network-level responses. In multiparametric cytometry, which is often used for analyzing patient samples, such issues are critical. While c…
Survey of factor analysis, PCA, variational inference, and VAE.
Jointly tuning ensemble models improves performance and uncertainty calibration.
Improves transparency and incorporates prior knowledge in Gaussian Process models.
MMVAE learns multi-modal data with shared and private latent spaces.
Measuring the impact of scientific articles is important for evaluating the research output of individual scientists, academic institutions and journals. While citations are raw data for constructing impact measures, there exist biases and potential issues if factors affecting citation patterns are not properly account…
MCPCA analyzes shared factors across multiple data contexts.
We introduce a factor analysis model that summarizes the dependencies between observed variable groups, instead of dependencies between individual variables as standard factor analysis does. A group may correspond to one view of the same set of objects, one of many data sets tied by co-occurrence, or a set of alternati…
We investigate the joint dynamics of spot and implied volatility from an empirical perspective. We focus on the equity market with the SPX Index our underlying of choice. Using only observable quantities, we extract the instantaneous variance curves implied by the market and study their daily variations jointly with sp…
Study joint invariants on symplectic spaces, extending group and space variations.
New method extracts joint and individual signals from multi-view data.
A method for identifying joint and individual subspaces from multi-view data.
Improved multimodal variational models capture more complex joint distributions.
We use deep neural networks to estimate an asset pricing model for individual stock returns that takes advantage of the vast amount of conditioning information, while keeping a fully flexible form and accounting for time-variation. The key innovations are to use the fundamental no-arbitrage condition as criterion funct…
We consider the problem of sufficient dimensionality reduction (SDR), where the high-dimensional observation is transformed to a low-dimensional sub-space in which the information of the observations regarding the label variable is preserved. We propose DVSDR, a deep variational approach for sufficient dimensionality r…
PACE explains ViTs by modeling patch-level concept distributions, surpassing existing methods.
Objective: Joint analysis of multi-subject brain imaging datasets has wide applications in biomedical engineering. In these datasets, some sources belong to all subjects (joint), a subset of subjects (partially-joint), or a single subject (individual). In this paper, this source model is referred to as joint/partially-…
VPP learns joint policies for multi-agent RL through interactions.
Newsroom in online ecosystem is difficult to untangle. With prevalence of social media, interactions between journalists and individuals become visible, but lack of understanding to inner processing of information feedback loop in public sphere leave most journalists baffled. Can we provide an organized view to charact…
Machine learning (ML) techniques such as (deep) artificial neural networks (DNN) are solving very successfully a plethora of tasks and provide new predictive models for complex physical, chemical, biological and social systems. However, in most cases this comes with the disadvantage of acting as a black box, rarely pro…
New method reduces inference variance for faster optimization.
This work gives an in-depth derivation of the trainable evidence lower bound obtained from the marginal joint log-Likelihood with the goal of training a Multi-Modal Variational Autoencoder (MVAE).
Local decision boundary approximation improves model explanations for complex models.
We investigate deep generative models that can exchange multiple modalities bi-directionally, e.g., generating images from corresponding texts and vice versa. A major approach to achieve this objective is to train a model that integrates all the information of different modalities into a joint representation and then t…
We study the problem of cooperative inference where a group of agents interact over a network and seek to estimate a joint parameter that best explains a set of observations. Agents do not know the network topology or the observations of other agents. We explore a variational interpretation of the Bayesian posterior de…
New model improves multimodal autoencoders by learning joint and conditional distributions.
Deep-learning CNN automates Cu alloy grain size evaluation.
Compared with shallow domain adaptation, recent progress in deep domain adaptation has shown that it can achieve higher predictive performance and stronger capacity to tackle structural data (e.g., image and sequential data). The underlying idea of deep domain adaptation is to bridge the gap between source and target d…
Efficiently estimates online variational learning using importance sampling.
This work improves multi-modal generative models by using permutation-invariant neural networks.
Despite a growing literature on explaining neural networks, no consensus has been reached on how to explain a neural network decision or how to evaluate an explanation. Our contributions in this paper are twofold. First, we investigate schemes to combine explanation methods and reduce model uncertainty to obtain a sing…
A new framework models multi-state events and biomarkers.
New framework for contesting algorithmic decisions, not just explaining them.
Customer Satisfaction is the most important factors in the industry irrespective of domain. Key Driver Analysis is a common practice in data science to help the business to evaluate the same. Understanding key features, which influence the outcome or dependent feature, is highly important in statistical model building.…