Social media reduces individual investors' disposition effect through negative information.
problem The disposition effect in individual investors selling profitable assets too early and holding onto losing assets for too long.
method Analysis of post data and trading data from Xueqiu.com.
result Social media information significantly reduces the disposition effect.
We investigate the strength and the direction of information transfer in the U.S. stock market between the composite stock price index of stock market and prices of individual stocks using the transfer entropy. Through the directionality of the information transfer, we find that individual stocks are influenced by the …
The paper proposes methods to extract and analyze individual variable information from complex dependencies.
problem Analyzing and understanding complex dependencies between multiple variables.
method Reversible normalization and iterative dependency reduction to extract individual information, and use it for direct mutual information and multi-feature Granger causality analysis.
result Decoupling of variables to analyze their individual information and direct mutual information transfers.
Method estimates group structure in panel data using variance information.
problem Estimating group structure in panel data with unknown groups.
method Proposes a method to estimate unobserved groupings for panel data models using variance information.
result Superior performance compared to existing methods in simulations and empirical applications.
This work develops a model to distinguish network and covariate information.
problem Identifying unique network and covariate information.
method Low-rank model with two-step estimation: spectral method followed by refinement.
result The method accurately recovers joint and individual components.
Automated prediction of valence, one key feature of a person's emotional state, from individuals' personal narratives may provide crucial information for mental healthcare (e.g. early diagnosis of mental diseases, supervision of disease course, etc.). In the Interspeech 2018 ComParE Self-Assessed Affect challenge, the …
Proposes a new bound on generalization error using conditional mutual information.
problem Improving the generalization error bound in machine learning.
method Combines error decomposition and conditional mutual information techniques.
result New bound is order-wise better than previous ones in a simple Gaussian setting.
Optimizes personalized medicine prediction by individual tuning.
problem Predicting individual drug responses using genomic information.
method Introduces a new ridge estimator and tuning parameter calibration scheme.
result Optimal in terms of oracle inequalities, fast, and highly effective.
Adaptive truncation improves privacy in online Bayesian estimation.
problem Ensuring privacy in online Bayesian estimation of a static parameter.
method Sequential Monte Carlo, adaptive truncation, Thompson sampling.
result Adaptive truncation reduces privacy-preserving noise, enabling more accurate estimation.
We revisit the notion of individual fairness proposed by Dwork et al. A central challenge in operationalizing their approach is the difficulty in eliciting a human specification of a similarity metric. In this paper, we propose an operationalization of individual fairness that does not rely on a human specification of …
Paper introduces a method to learn physics between digital twins using imperfect models.
problem Learning physics from imperfect data and low-fidelity models.
method Bayesian Hierarchical modeling with physics-informed Gaussian processes.
result Models learning between digital twins are less uncertain than independent models but not over-confident.
Confidentiality of patient information is an essential part of Electronic Health Record System. Patient information, if exposed, can cause a serious damage to the privacy of individuals receiving healthcare. Hence it is important to remove such details from physician notes. A system is proposed which consists of a deep…
We study notions of fairness in decision-making systems when individuals have diverse preferences over the possible outcomes of the decisions. Our starting point is the seminal work of Dwork et al. which introduced a notion of individual fairness (IF): given a task-specific similarity metric, every pair of individuals …
Proposes a deep learning framework for estimating counterfactual outcomes.
problem Challenges in estimating individual outcomes under different treatments.
method Deep variational Bayesian framework integrating factual and similar subjects' outcomes.
result Rigorously integrates individual features and similar subjects' responses for counterfactual outcomes.
Meta-analysis improves personalized treatment rules across multiple sites.
problem Lack of generalizability in learning individualized treatment rules across different medical sites.
method Developed a method for individual-level meta-analysis of ITRs, borrowing sign-coherency information between sites.
result Jointly learned site-specific ITRs with improved generalizability.
Universal supervised learning is considered from an information theoretic point of view following the universal prediction approach, see Merhav and Feder (1998). We consider the standard supervised "batch" learning where prediction is done on a test sample once the entire training data is observed, and the individual s…
Algorithm samples fair rankings to ensure individual fairness while maintaining group fairness.
problem Fair ranking tasks with group fairness constraints and uncertainty in item utilities.
method Efficient algorithm that samples rankings from an individually-fair distribution ensuring group fairness.
result Expected utility of output ranking is at least α times optimal fair solution, where α depends on utilities and constraints.
Proposes a model to handle mobile health data with irregular measurements.
problem Handling heterogeneous, multi-resolution data in mobile health.
method Individualized dynamic latent factor model for irregular multi-resolution time series data.
result Superior performance compared to existing methods in simulation and smartwatch data applications.
In machine learning, classification models need to be trained in order to predict class labels. When the training data contains personal information about individuals, collecting training data becomes difficult due to privacy concerns. Local differential privacy is a definition to measure the individual privacy when th…
Before the massive spread of computer technology, information was far from complex. The development of technology shifted the paradigm: from individuals who faced scarce and costly information to individuals who face massive amounts of information accessible at low costs. Nowadays we are living in the era of big data a…
We propose a new problem formulation which is similar to, but more informative than, the binary multiple-instance learning problem. In this setting, we are given groups of instances (described by feature vectors) along with estimates of the fraction of positively-labeled instances per group. The task is to learn an ins…
The paper introduces CPICFs for better counterfactual explanations in high-dimensional spaces.
problem Creating useful counterfactual explanations for complex machine learning models.
method Modeling individual knowledge and using conformal prediction intervals to identify informative counterfactuals.
result CPICFs provide more informative counterfactuals by considering individual knowledge and prediction uncertainty.
New framework for forecasting psychological processes from ILD.
problem Forecasting psychological processes at the individual level from ILD.
method A novel modeling framework addressing challenges in ILD.
result Improved forecasting of psychological processes at the individual level.
New model quantifies how much machine learning models can reveal about individual data usage.
problem Measuring and reducing the leakage of membership information from machine learning models.
method Using information theory, conditional mutual information leakage, and Kullback-Leibler divergence to quantify and bound the leakage.
result The amount of membership information leakage is reduced by adding Gaussian (ε,δ)-differentially-private additive noises. Estimating individual level treatment effects (ITE) from observational data is a challenging and important area in causal machine learning and is commonly considered in diverse mission-critical applications. In this paper, we propose an information theoretic approach in order to find more reliable representations for e…
Information-theoretic bounded rationality describes utility-optimizing decision-makers whose limited information-processing capabilities are formalized by information constraints. One of the consequences of bounded rationality is that resource-limited decision-makers can join together to solve decision-making problems …
New method to quantify feature contributions to disparity without access to decision-making model.
problem Quantifying feature contributions to disparity when decision-making model is not accessible.
method Use information theory to measure redundant statistical dependency between protected attribute and feature.
result Quantify feature contributions to disparity using information theory.
New method refines prediction intervals for individual treatment effects using cross-world correlation.
problem Uncertainty in individual treatment effects for high-stakes decisions.
method Introduces cross-world correlation parameter ρ to refine prediction intervals for individual treatment effects.
result Achieves more stable and accurate coverage of prediction intervals for individual treatment effects.
When investors have heterogeneous attitudes towards risk, it is reasonable to assume that each investor has a pricing kernel, and that these individual pricing kernels are aggregated to form a market pricing kernel. The various investors are then buyers or sellers depending on how their individual pricing kernels compa…
GWIB improves counterfactual regression by balancing latent distributions and reducing selection bias.
problem Selection bias between control and treatment groups negatively impacts counterfactual regression performance.
method GWIB uses Gromov-Wasserstein information bottleneck to maximize mutual information between covariates and outcomes while penalizing kernelized mutual information between latent representations and covariates.
result GWIB consistently outperforms state-of-the-art CFR methods in ITE estimation tasks.
This work transfers causal knowledge between tasks for Individual Treatment Effect estimation.
problem Estimating Individual Treatment Effects (ITE) requires a large amount of data, making it challenging.
method The authors introduce a practical framework for efficient transfer of causal knowledge between tasks, using a Causal Inference Task Affinity (CITA) measure.
result ITE knowledge transfer can significantly reduce the amount of data needed for ITE estimation.
The definition of preferences assigned to individuals is a concept that concerns many disciplines, from economics, with the search of an acceptable outcome for an ensemble of individuals, to decision making an analysis of vote systems. We are concerned in the phenomena of good selection and economic fairness. In Arrow'…
Proposes a new model for estimating individual treatment effects.
problem Estimating individual treatment effects from observational data is challenging.
method Integrates diffusion modeling and conformal inference with propensity score and covariate approximation.
result Establishes rigorous theoretical guarantees and demonstrates competitive performance.
Specialists outperform generalists in ensemble classification.
problem Determining the accuracy of an ensemble of classifiers when individual classifier accuracies are known.
method Proved upper and lower bounds on ensemble accuracy, constructed specialist and generalist classifiers.
result Upper and lower bounds on ensemble accuracy, practical implications for classifier construction.
Improves representation learning for individual treatment effect estimation.
problem Estimating individual treatment effects with high accuracy.
method Introduces a structure keeper to maintain correlation between baseline covariates and representations, trains a discriminator to balance representation and information loss.
result Proposed SMRL algorithm minimizes treatment estimation error and outperforms state-of-the-art methods.
Estimating the largest community in a mixed population via sequential sampling.
problem Identifying the largest community in a mixed population with limited sampling.
method Sequential, random sampling of individuals across multiple boxes, optimizing sampling strategy and decision rule.
result Proposed algorithms achieve optimal error probability decay rates under fixed budget constraints.
Graphs are widely used as a natural framework that captures interactions between individual elements represented as nodes in a graph. In medical applications, specifically, nodes can represent individuals within a potentially large population (patients or healthy controls) accompanied by a set of features, while the gr…
In this work, we investigate the use of three information-theoretic quantities -- entropy, mutual information with the class variable, and a class selectivity measure based on Kullback-Leibler divergence -- to understand and study the behavior of already trained fully-connected feed-forward neural networks. We analyze …
This paper identifies and bounds ICE central moments using PO marginal central moments.
problem Identifying and characterizing treatment effect heterogeneity.
method Using only marginal central moments of potential outcomes, the paper identifies and bounds central moments of individual causal effects.
result Identification and bounding of central moments of ICE using marginal moments of POs.
Investor flows in Korean equity market transmit shared information, not private signals.
problem Whether investor flows transmit private information or only public signals.
method Transfer Entropy networks constructed from investor-type flows over
umNDates{} trading days.
result Investor flows transmit shared information, not private signals.
As algorithmic prediction systems have become widespread, fears that these systems may inadvertently discriminate against members of underrepresented populations have grown. With the goal of understanding fundamental principles that underpin the growing number of approaches to mitigating algorithmic discrimination, we …
Depression and anxiety are critical public health issues affecting millions of people around the world. To identify individuals who are vulnerable to depression and anxiety, predictive models have been built that typically utilize data from one source. Unlike these traditional models, in this study, we leverage a rich …
Differential privacy for simple linear regression protects small datasets from individual data leaks.
problem Protecting sensitive personal information in small datasets from individual data leaks.
method Differential privacy algorithms for simple linear regression tailored for small datasets (tens to hundreds of datapoints).
result Robust estimators like Theil-Sen perform well on small datasets, but standard algorithms improve as dataset size increases.
New method finds balanced clusters in graphs using auxiliary information.
problem Finding balanced clusters in graphs with population-level constraints.
method Proposes individual-level balancing constraint and develops spectral clustering algorithms.
result Establishes first statistical consistency result for constrained spectral clustering.
Study reveals how investor flows impact stock prices, especially during herding episodes.
problem Understanding how information transmits through prices and why it breaks down.
method Combining regularized deconvolution with Hawkes process analysis.
result Institutional price impact deteriorates sharply during herding episodes in small-cap stocks, while large-cap stocks maintain resilience.
Deep learning shows ETF imbalances are more informative than market imbalances.
problem Determining causality between ETF and market imbalances.
method Deep learning econometric methodology applied to stock and ETF transactions.
result ETF imbalance messages are more informative than market imbalance messages.
This paper evaluates heterogeneous information fusion using multi-task Gaussian processes in the context of geological resource modeling. Specifically, it empirically demonstrates that information integration across heterogeneous information sources leads to superior estimates of all the quantities being modeled, compa…
EB-VAE combines tumor growth and dropout data for personalized treatment response modeling.
problem Challenges in integrating longitudinal tumor measurements, dropout information, and genetic covariates.
method Extended EB-VAE framework to jointly model longitudinal and time-to-event data, incorporating dropout hazard and genetic covariates.
result Hybrid decoder formulation yields consistent treatment-effect parameters and prior predictive performance comparable to neural decoder.