New method combines regional HIV prevention trial data without sharing individual patient info.
problem Regional differences in HIV prevention efficacy, privacy concerns, and data sharing limitations.
method Federated learning approach that combines site-specific estimators via L1-regularization.
result Improved precision in estimating region-specific survival curves.
System predicts HIV patients at risk of dropping out of care.
problem High non-adherence and dropout rates among HIV patients.
method Predictive machine learning model based on resource constraints, stability, and fairness.
result Model performs 3x better than baseline for clinical use and 2.3x better for city-wide use.
HIV RNA viral load (VL) is an important outcome variable in studies of HIV infected persons. There exists only a handful of methods which classify patients by viral load patterns. Most methods place limits on the use of viral load measurements, are often specific to a particular study design, and do not account for com…
Model identifies key problems in HIV patients' records.
problem Complex and time-consuming task of identifying patient problems from electronic health records.
method Unsupervised phenotyping approach that jointly learns phenotypes from structured and unstructured data.
result Learned phenotypes and their relatedness are clinically valid and surpass existing methods.
The paper develops a causal machine learning framework to optimize aid allocation.
problem Optimizing aid allocation to reduce new HIV infections in poor countries.
method The framework uses a balancing autoencoder, counterfactual generator, and inference model to predict heterogeneous treatment effects.
result The framework predicts a reduction of up to 3.3% in new HIV infections, saving 50,000 lives.
Develops algorithm to find subgroups with different treatment effects in HIV patients.
problem Estimating treatment effects in EHR data with challenges like time-varying confounding.
method SDLD algorithm combining generalized interaction tree and longitudinal targeted maximum likelihood estimation.
result Identifies subgroups of HIV patients at higher risk of weight gain with dolutegravir-containing ARTs.
Paper derives policy rules from observational data for hepatitis C treatment.
problem Improving treatment guidelines for HIV/HCV co-infected patients.
method Weighted K-means algorithm for estimating CATEs, decision tree implementation.
result Identifies a subgroup with high spontaneous HCV clearance rate.
New framework quantifies variable importance across all good models and is stable across data distribution.
problem Conflicting variable importance conclusions from different models trained on the same data.
method Proposes a new variable importance framework that considers all good models and is stable across data distribution.
result Framework accurately estimates true variable importance and recovers rankings for complex setups.
New algorithms improve MPBART for HIV patient data.
problem Improving inference for multinomial outcomes in HIV patient data.
method Introduced two new algorithms for fitting MPBART.
result Better performance in MCMC convergence and predictive accuracy.
Motivation: HIV is difficult to treat because its virus mutates at a high rate and mutated viruses easily develop resistance to existing drugs. If the relationships between mutations and drug resistances can be determined from historical data, patients can be provided personalized treatment according to their own mutat…
Personalized models explain TB treatment outcomes considering patient context.
problem Heterogeneity in TB treatment outcomes due to co-morbidities.
method Multi-task learning approach encoding patient context into personalized models.
result Identifies anemia, age of onset, and HIV as influential for treatment efficacy.
Count data, for example the number of observed cases of a disease in a city, often arise in the fields of healthcare analytics and epidemiology. In this paper, we consider performing regression on multivariate data in which our outcome is a count. Specifically, we derive log-likelihood functions for finite mixtures of …
BSFP method reveals latent patterns in multi-omic data for predicting lung function in HIV-associated OLD.
problem Limited understanding of multi-omic molecular phenomena and clinical outcomes in obstructive lung disease.
method Bayesian Simultaneous Factorization and Prediction (BSFP) method for multi-omic data, accommodating imputation and full posterior inference.
result BSFP reveals distinct clusters of patients with OLD and multi-omic patterns related to lung function decline.
It is the main purpose of this paper to introduce a graph-valued stochastic process in order to model the spread of a communicable infectious disease. The major novelty of the SIR model we promote lies in the fact that the social network on which the epidemics is taking place is not specified in advance but evolves thr…
Affinity propagation is an exemplar-based clustering algorithm that finds a set of data-points that best exemplify the data, and associates each datapoint with one exemplar. We extend affinity propagation in a principled way to solve the hierarchical clustering problem, which arises in a variety of domains including bi…
Motivation: Proteins are known to undergo conformational changes in the course of their functions. The changes in conformation are often attributable to a small fraction of residues within the protein. Therefore identification of these variable regions is important for an understanding of protein function. Results: We …
A new graphical model for discrete data without parametric restrictions.
problem Discrete data modeling with restrictions.
method Additive conditional independence and penalized estimation of precision operator.
result Consistency of the estimator in ultrahigh-dimensional settings.
The aim of this paper is to construct a natural Riemann-Lagrange differential geometry on 1-jet spaces, in the sense of nonlinear connections, generalized Cartan connections, d-torsions, d-curvatures, jet electromagnetic fields and jet electromagnetic Yang-Mills energies, starting from some given nonlinear evolution OD…
Bayesian nonparametric method partitions shapes using curves.
problem Capturing complex shapes in multi-dimensional data.
method Proposes a novel spline partitioning approach using curves.
result Demonstrates improved shape modeling compared to existing methods.
A new method combines machine learning with mixed-effects models for better repeated measurement analysis.
problem Inference of linear coefficients in partially linear mixed-effects models with complex interactions and high-dimensional variables.
method Double machine learning approach to estimate nonparametrically nonlinear variables, then use standard linear mixed-effects techniques to estimate the linear coefficient.
result The estimated fixed effects coefficient converges at the parametric rate and is semiparametrically efficient.
SpINNEr uses matrix regression to analyze brain connectivity, improving accuracy over other methods.
problem Analyzing multi-dimensional data like brain imaging arrays using traditional scalar regression methods.
method SpINNEr applies matrix regression with nuclear norm and lasso norms to encourage low rank and sparse solutions.
result SpINNEr outperforms other methods in estimating brain connectivity, especially in well-connected regions.
We study the problem of off-policy policy evaluation (OPPE) in RL. In contrast to prior work, we consider how to estimate both the individual policy value and average policy value accurately. We draw inspiration from recent work in causal reasoning, and propose a new finite sample generalization error bound for value e…
The lack of interpretability remains a key barrier to the adoption of deep models in many applications. In this work, we explicitly regularize deep models so human users might step through the process behind their predictions in little time. Specifically, we train deep time-series models so their class-probability pred…
Exact selective inference with randomization for Gaussian regression models.
problem Exact selective inference in Gaussian regression models.
method Introduces a pivot for exact selective inference with randomization, reducing the problem to a bivariate truncated Gaussian distribution.
result Our pivot leads to exact inference and produces narrower confidence intervals than related methods.
Due to physiological variation, patients diagnosed with the same condition may exhibit divergent, but related, responses to the same treatments. Hidden Parameter Markov Decision Processes (HiP-MDPs) tackle this transfer-learning problem by embedding these tasks into a low-dimensional space. However, the original formul…
This paper proposes a new method for estimating sparse precision matrices in the high dimensional setting. It has been popular to study fast computation and adaptive procedures for this problem. We propose a novel approach, called Sparse Column-wise Inverse Operator, to address these two issues. We analyze an adaptive …
Framework evaluates AI proposals for drug discovery, finds no LLM advantage.
problem No principled framework exists for evaluating AI-guided scientific selection under budget constraints.
method Formally verified metric (BSDS/DQS) penalizes false discoveries and excessive abstention.
result LLMs provide no marginal value over existing classifiers in drug discovery.
SCQRNN prevents quantile crossing and improves computational efficiency.
problem Quantile crossing issue in regression models.
method Integrates ad hoc sorting in training to prevent quantile crossing and enhance computational efficiency.
result SCQRNN achieves faster convergence and non-intersecting quantiles.
Many immunization strategies have been proposed to prevent infectious viruses from spreading through a network. In this study, we propose efficient immunization strategies to prevent a default contagion that might occur in a financial network. An essential difference from the previous studies on immunization strategy i…
Paper presents a privacy-preserving algorithm for estimating peer effects using the Ising model.
problem Privacy concerns in estimating peer effects using network data.
method Developed a (ε,δ)-differentially private algorithm using Ising model. result Established regret bounds and validated performance on synthetic and real-world networks.
We consider the task of fitting a regression model involving interactions among a potentially large set of covariates, in which we wish to enforce strong heredity. We propose FAMILY, a very general framework for this task. Our proposal is a generalization of several existing methods, such as VANISH [Radchenko and James…
Modified model prevents volatility from approaching zero.
problem Volatility in the Gatheral model can approach zero, making it statistically indistinguishable.
method Proposed a modified model with Skorokhod reflection to prevent volatility from approaching zero.
result The modified model prevents volatility from approaching zero, preserving the model's flexibility.
This paper examines how skip connections prevent rank collapse in sequence models.
problem Rank collapse in sequence models, leading to reduced expressivity and training instabilities.
method Analytical and ablation studies of lambda-skip connections in SSMs.
result A sufficient condition to prevent rank collapse across various architectures.
The paper explores how sinks and diagonal patterns prevent attention oversmoothing.
problem Preventing attention oversmoothing in neural networks.
method Analyzing geometric conditions and conditions for dense vs. sparse attention, proving equivalence between sinks and hard attention switch, and comparing the costs of sinks vs. diagonal patterns.
result Sinks and diagonal patterns effectively prevent attention oversmoothing, and diagonal patterns provide a more flexible approach.
Super learner uses diverse screeners to improve prediction performance.
problem Performance issues with lasso screening in super learner.
method Used a diverse set of candidate screeners within the super learner ensemble.
result Diverse screeners protect against poor performance of any one screener.
The lack of interpretability remains a barrier to the adoption of deep neural networks. Recently, tree regularization has been proposed to encourage deep neural networks to resemble compact, axis-aligned decision trees without significant compromises in accuracy. However, it may be unreasonable to expect that a single …
Although recovering an Euclidean distance matrix from noisy observations is a common problem in practice, how well this could be done remains largely unknown. To fill in this void, we study a simple distance matrix estimate based upon the so-called regularized kernel estimate. We show that such an estimate can be chara…
Prevents sensitive data generation in diffusion models using labeled and unlabeled data.
problem Generating sensitive data in diffusion models using unlabeled data.
method Positive-Unlabeled Diffusion Models, approximating ELBO with labeled and unlabeled data.
result Prevents the generation of sensitive data without compromising image quality.
Deep Neural Networks are robust to minor perturbations of the learned network parameters and their minor modifications do not change the overall network response significantly. This allows space for model stealing, where a malevolent attacker can steal an already trained network, modify the weights and claim the new ne…
StratLearner learns strategies to prevent misinformation in social networks.
problem Learning strategies to protect against misinformation in social networks without knowing the diffusion model.
method Structured prediction framework using random features and large margin method.
result Our method produces near-optimal protectors without diffusion model information and outperforms other methods.
Proposes a method to avoid excessive exploration in reinforcement learning.
problem Avoiding excessive exploration in reinforcement learning to deploy it in practice.
method Designs a novel algorithm using UCB reinforcement learning policy with adaptive exploration constraints.
result Proves that the approach remains conservative while minimizing regret in tabular settings and validates on real-world tasks.
Tackling large approximate dynamic programming or reinforcement learning problems requires methods that can exploit regularities, or intrinsic structure, of the problem in hand. Most current methods are geared towards exploiting the regularities of either the value function or the policy. We introduce a general classif…
Mobile apps and machine learning improve malaria prevention and treatment.
problem High malaria cases and deaths in low-income countries.
method Adaptive interventions using mobile health apps and machine learning.
result Increased malaria testing, adherence, and provider skills.
New condition prevents hyperbolic spaces from matching curve complexes.
problem Identifying when hyperbolic spaces cannot match curve complexes.
method Analyzing specific hyperbolic complexes and identifying a condition.
result Identified a condition preventing quasi-isometry between hyperbolic spaces and curve complexes.
TAMD prevents degeneracy in finite mixtures, offering strong guarantees but modest practical improvements.
problem Degeneracy in maximum likelihood estimation of finite mixtures.
method Transcendental regularization with analytic barrier functions.
result Strong theoretical guarantees (identifiability, consistency, robustness) but modest practical improvements.
Large learning rates prevent memorization in denoising score matching.
problem Memorization of training data in diffusion-based generative models.
method Investigating the role of large learning rates in the small-noise regime, proving that they prevent convergence to the empirical optimal score.
result Large learning rates prevent memorization by making it impossible for the learned score to be arbitrarily close to the empirical optimal score.
New method predicts sets under unknown covariate shift with high confidence.
problem Adapting to unknown covariate shift in prediction sets.
method PredSet-1Step, a flexible distribution-free method.
result Achieves asymptotic probably approximately correct coverage.
Maker-taker fees can prevent algorithmic cooperation in market making, but not always.
problem Unexpected cooperation among independent algorithms in market making.
method Modeling market making as a repeated game, experimental analysis of transaction costs and rebates.
result Maker-taker fee models can destabilize cooperation, but not always with a specific relationship between costs and rebates.