A new model generates samples with a succinct common representation using Wyner's common information.
problem Generating samples with a succinct common representation.
method Proposes a variational Wyner model trained to minimize symmetric Kullback-Leibler divergence with regularization terms.
result Demonstrates utility through joint and conditional generation experiments.
Efficiently estimates distributed mean with side information, near-optimal and universal.
problem Distributed mean estimation with side information in communication constrained settings.
method Wyner-Ziv estimators for communication and computation efficiency.
result Near-optimal and universal recovery guarantees for distributed optimization and compression.
We simplify information measure computation using learned features.
problem Computing information measures from raw data is computationally expensive.
method Developed a separable design for computing information measures from learned feature representations.
result A variety of information measures can be computed efficiently through learned feature representations.
Tutorial on information bottleneck problems with connections to coding and learning.
problem Information bottleneck problems and their connections to coding and learning.
method Information theoretic perspective, practical methods, connections to various problems.
result Optimal trade-offs between relevance and complexity in discrete and vector Gaussian frameworks.
Measuring the relationship between any pair of variables is a rich and active area of research that is central to scientific practice. In contrast, characterizing the common information among any group of variables is typically a theoretical exercise with few practical methods for high-dimensional data. A promising sol…
The paper identifies universal features for high-dimensional data inference.
problem Identifying universal low-dimensional features from high-dimensional data for inference tasks.
method Introduces natural notions of universality and shows a local equivalence among them, using information geometry.
result Reveals the complementary roles of various data analysis techniques.
This paper offers a characterization of fundamental limits on the classification and reconstruction of high-dimensional signals from low-dimensional features, in the presence of side information. We consider a scenario where a decoder has access both to linear features of the signal of interest and to linear features o…
Optimizes portfolios using neural network approximations of asset sensitivities to common drivers.
problem Optimizing portfolios with complex asset dynamics and common drivers.
method Model asset dynamics with PDEs, approximate sensitivities with neural networks, and use hierarchical clustering on sensitivity matrix for optimization.
result Achieves over-performance in portfolio optimization across various markets and datasets.
New method detects latent common causes from observational data.
problem Detecting latent common causes in observational data.
method Modified causal discovery algorithms to detect latent common causes.
result Successfully detects latent common causes in various noise regimes and real data.
SBPMT combines bagging and boosting for improved classification.
problem Improving classification performance in machine learning.
method Designing SBPMT, a hybrid of bagging and boosting, with PMT as base classifiers.
result SBPMT is consistent and can reduce generalization error with more subagging iterations.
We study the problem of discovering the simplest latent variable that can make two observed discrete variables conditionally independent. The minimum entropy required for such a latent is known as common entropy in information theory. We extend this notion to Renyi common entropy by minimizing the Renyi entropy of the …
MCCA extracts shared structure from multiple tensor datasets.
problem Extracting shared structure from multiple tensor datasets.
method Multilinear common component analysis (MCCA) using Kronecker products of mode-wise covariance matrices.
result MCCA constructs a common basis that retains information from multiple tensor datasets.
Study finds stocks with common firm fears earn lower returns.
problem Identifying and quantifying firm-level investor fears.
method Analysis of equity options to identify common firm-level fears and their impact on stock returns.
result Stocks with exposure to common bad fears earn lower returns and require higher compensation.
Hopformer combines common trends with series-specific details for better time series forecasting.
problem Forecasting multiple time-series with high-dimensional covariates while retaining series-specific information.
method Hopformer uses a two-stage framework: SPA for common trends and LoRA-fine-tuned Transformer for residual dependencies.
result Improves MASE by an average of 6.56% across synthetic and real-world benchmarks.
Proposes MV-Co-VH for multi-view clustering using visible and hidden views.
problem Lack of efficient algorithms for fully utilizing multi-view data.
method Projects multiple views to a common hidden space using NMF, then applies collaborative learning.
result Competitive clustering performance on UCI and real-world datasets.
NestedVAE isolates common factors from paired images without additional supervision.
problem Reduction of data-driven biases in machine learning models.
method Combines deep latent variable models with information bottleneck theory.
result NestedVAE significantly outperforms alternative methods in various tasks.
Solves a game between brokers and informed traders using stochastic differential equations.
problem Optimizing wealth in a game between brokers and informed traders with private signals.
method Closed-form solutions to a mean-field game using forward-backward SDEs.
result Optimal trading strategies for both brokers and informed traders are found.
MAGMA uses a common mean process to improve multi-step-ahead time series forecasting.
problem Improving multiple-step-ahead predictions for time series data.
method Proposes a novel multi-task Gaussian process framework with a common mean process for sharing information across tasks.
result Significantly improves predictive performances, even far from observations, and reduces computational complexity.
Study shows publicly available news impacts financial markets.
problem Impact of publicly available news on financial markets.
method Extracted news from Common Crawl, identified relevant companies, used sentiment analysis and information theory.
result Publicly available news has significant impact on financial markets.
Proposes adversarial normalization for multi-domain image segmentation.
problem Current image normalization is per-dataset, limiting multi-domain segmentation.
method Adversarial training to learn common normalizing functions across multiple datasets.
result Optimal normalizer improves segmentation accuracy and realism.
Spatial information is not always necessary for spatio-temporal models.
problem The necessity of including spatial information in spatio-temporal models.
method Comparison of spatial agnostic neural networks with state-of-the-art models on ten datasets.
result Spatial information is not always needed in most spatio-temporal models.
Improved ranking method for scarce data with feature info.
problem Ranking items with limited comparisons and feature data.
method Modified RankCentrality using diffusion methods for feature info.
result Meaningful rankings even with scarce comparisons.
In this work we propose a method for reducing the dimensionality of tensor objects in a binary classification framework. The proposed Common Mode Patterns method takes into consideration the labels' information, and ensures that tensor objects that belong to different classes do not share common features after the redu…
Motivation: Prediction of the interaction affinity between proteins and compounds is a major challenge in the drug discovery process. WideDTA is a deep-learning based prediction model that employs chemical and biological textual sequence information to predict binding affinity. Results: WideDTA uses four text-based inf…
This paper relaxes the common prior assumption in the public and private information game of Morris and Shin (2000, 2004). For the generalized game, where the agent's prior expectations are heterogenous, it derives a sharp condition for the emergence of unique/multiple equilibria. This condition indicates that unique e…
We consider the statistical problem of learning common source of variability in data which are synchronously captured by multiple sensors, and demonstrate that Siamese neural networks can be naturally applied to this problem. This approach is useful in particular in exploratory, data-driven applications, where neither …
Relationships between entities in datasets are often of multiple nature, like geographical distance, social relationships, or common interests among people in a social network, for example. This information can naturally be modeled by a set of weighted and undirected graphs that form a global multilayer graph, where th…
Study finds adding more information to robust option pricing does not improve bounds.
problem Exploring robust pricing of financial claims using minimal assumptions.
method Empirical study of variance options, incorporating intermediate market data.
result Incorporating more information does not improve robust pricing bounds.
HPCA improves PCA for portfolio management by interpreting sector-specific factors.
problem Difficult interpretation of PCA's higher eigenportfolios in practical portfolio management.
method Partitioning the market into sectors and applying Hierarchical PCA.
result HPCA leads to no loss of information and interpretable factors.
A new graph kernel uses LCS and Wasserstein distance for better graph comparisons.
problem Graph learning methods can be limited by information from distant vertices and path length constraints.
method Proposes a Graph Kernel based on LCS similarity and Wasserstein distance in a novel metric space.
result The new kernel emphasizes comparisons between similar paths and reduces information loss.
Multiple sets of measurements on the same objects obtained from different platforms may reflect partially complementary information of the studied system. The integrative analysis of such data sets not only provides us with the opportunity of a deeper understanding of the studied system, but also introduces some new st…
For common people, in contrast to brokers, bankers, and those who play on rising and falling prices of stocks, the stock market law is based on the simple fact that the depositors aim for financial profit at any given concrete stage. The common depositor cannot cause any significant variations in prices. This concept s…
Paper presents a technique using Spearman's Rank Correlation Coefficient for KE in TDs.
problem Extracting common characteristics and grouping similar TDs.
method Spearman's Rank Correlation Coefficient (SRCC) for KE.
result SRCC proves a comprehensive measure for high-quality KE.
Privacy-preserving distributed deep learning method for multiple classification.
problem Privacy issues in training deep learning models for various fields.
method Split learning into common extractor, cloud model, and local classifier.
result Average performance improvement of 2.63% over existing local training models.
This paper introduces a deep-learning based efficient classifier for common dermatological conditions, aimed at people without easy access to skin specialists. We report approximately 80% accuracy, in a situation where primary care doctors have attained 57% success rate, according to recent literature. The rationale of…
A new approach simplifies Sliced-Wasserstein distances to improve learning performance.
problem The concentration of measure phenomenon makes random projections uninformative in high dimensions.
method Propose rescaling the 1D Wasserstein distance to make all slices equally informative.
result The classical Sliced-Wasserstein, properly configured, can match or surpass complex variants.
Proposes an MTL method with clustering to improve regression accuracy.
problem Improving regression accuracy by sharing information among related tasks.
method Centroid parameter for clustering tasks, separating regression and clustering parameters.
result Improves estimation and prediction accuracy for regression coefficient vectors.
This paper presents an original approach for jointly fitting survival times and classifying samples into subgroups. The Coxlogit model is a generalized linear model with a common set of selected features for both tasks. Survival times and class labels are here assumed to be conditioned by a common risk score which depe…
Methods for analysis of principal components in discrete data have existed for some time under various names such as grade of membership modelling, probabilistic latent semantic analysis, and genotype inference with admixture. In this paper we explore a number of extensions to the common theory, and present some applic…
MALI aligns distinct domains using labeled data.
problem Aligning multi-domain data for machine learning.
method MALI learns manifold structure via diffusion and uses labeled data to guide alignment.
result MALI outperforms state-of-the-art methods across multiple datasets.
Finding relevant information from large document collections such as the World Wide Web is a common task in our daily lives. Estimation of a user's interest or search intention is necessary to recommend and retrieve relevant information from these collections. We introduce a brain-information interface used for recomme…
Eluder dimension and information gain are equivalent for reproducing kernel Hilbert spaces.
problem Complexity measures in bandit and reinforcement learning.
method Equivalence of eluder dimension and information gain for reproducing kernel Hilbert spaces.
result Eluder dimension and information gain are equivalent for reproducing kernel Hilbert spaces.
Ensembles of classification and regression trees remain popular machine learning methods because they define flexible non-parametric models that predict well and are computationally efficient both during training and testing. During induction of decision trees one aims to find predicates that are maximally informative …
Paper designs a penalty for model order selection using information criteria.
problem Selecting the correct model order from a set of candidate models.
method Designs a penalty for the generalized information criterion (GIC) to minimize underestimation.
result Optimal penalty minimizes underestimation while keeping overestimation below a specified level.
Boosted tree method improves MTL in heterogeneous domains.
problem Improving MTL in diverse, domain-specific tasks.
method Two-stage approach: common model for shared features, specific models for task-specific instances.
result Enhanced multi-task learning performance with interpretability.
The study challenges the notion that partial data annotation is inferior, suggesting it can sometimes outperform complete annotation.
problem The inefficiency and high cost of completely annotating structured data.
method Information theoretic formulation applied to three diverse structured learning tasks.
result Learning from partial structures can sometimes outperform learning from complete ones.
New research shows existing information-theoretic methods can't establish minimax rates for gradient descent in stochastic convex optimization.
problem Establishing minimax rates for gradient descent in stochastic convex optimization using information-theoretic methods.
method Examined several information-theoretic frameworks including input-output mutual information bounds, conditional mutual information bounds, PAC-Bayes bounds, and their variants.
result Proved that none of the examined information-theoretic frameworks can establish minimax rates for gradient descent in stochastic convex optimization.
Bayesian network modelling is a well adapted approach to study messy and highly correlated datasets which are very common in, e.g., systems epidemiology. A popular approach to learn a Bayesian network from an observational datasets is to identify the maximum a posteriori network in a search-and-score approach. Many sco…