New method uses unlabeled data to estimate intercept in case-control logistic regression.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A method for logistic regression inference using both internal and external data.
Estimates non-parametric logistic model using case-control data and external summary info.
For classification problems with significant class imbalance, subsampling can reduce computational costs at the price of inflated variance in estimating model parameters. We propose a method for subsampling efficiently for logistic regression by adjusting the class balance locally in feature space via an accept-reject …
New method uses geometric mean to avoid non-collapsibility in case-control studies.
Dimension reduction and variable selection are performed routinely in case-control studies, but the literature on the theoretical aspects of the resulting estimates is scarce. We bring our contribution to this literature by studying estimators obtained via L1 penalized likelihood optimization. We show that the optimize…
CIR method preserves relation for case-control studies.
Paper proposes a new method for estimating conditional densities using logistic regressions.
The causal assumptions, the study design and the data are the elements required for scientific inference in empirical research. The research is adequately communicated only if all of these elements and their relations are described precisely. Causal models with design describe the study design and the missing data mech…
A review of contrastive dimension reduction methods for treatment vs control studies.
Study causal inference under specific sampling methods with monotonicity assumptions.
Study improves treatment effect estimation using unlabeled covariates.
Bayesian hyperbolic MDS improves tree-like data representation.
Normative modeling has recently been proposed as an alternative for the case-control approach in modeling heterogeneity within clinical cohorts. Normative modeling is based on single-output Gaussian process regression that provides coherent estimates of uncertainty required by the method but does not consider spatial c…
Genome-wide association studies (GWASs) aim to detect genetic risk factors for complex human diseases by identifying disease-associated single-nucleotide polymorphisms (SNPs). The traditional SNP-wise approach along with multiple testing adjustment is over-conservative and lack of power in many GWASs. In this article, …
New methods for time-to-event prediction are proposed by extending the Cox proportional hazards model with neural networks. Building on methodology from nested case-control studies, we propose a loss function that scales well to large data sets, and enables fitting of both proportional and non-proportional extensions o…
regularized logistic regression has now become a workhorse of data mining and bioinformatics: it is widely used for many classification problems, particularly ones with many features. However, regularization typically selects too many features and that so-called false positives are unavoidable. In this pape…
A rigorous ML pipeline for binary classification in biomedical studies, focusing on pancreatic cancer.
Bayesian method models binary response and covariates for two groups, estimating causal relationships.
Discovering causal genetic variants from large genetic association studies poses many difficult challenges. Assessing which genetic markers are involved in determining trait status is a computationally demanding task, especially in the presence of gene-gene interactions. A non-parametric Bayesian approach in the form o…
We consider a problem of learning a binary classifier only from positive data and unlabeled data (PU learning) and estimating the class-prior in unlabeled data under the case-control scenario. Most of the recent methods of PU learning require an estimate of the class-prior probability in unlabeled data, and it is estim…
Novel approach uses quasi-conformal geometry for OSA classification from cephalometry.
Unified theory for semiparametric data fusion with individual-level data.
We develop a real-time anomaly detection algorithm for directed activity on large, sparse networks. We model the propensity for future activity using a dynamic logistic model with interaction terms for sender- and receiver-specific latent factors in addition to sender- and receiver-specific popularity scores; deviation…
Proposes spBART for risk prediction using epigenetic signatures and covariates.
Paper proposes a method for estimating tropical cyclone intensity distribution using deep learning.
The increasing access to brain signal data using electroencephalography creates new opportunities to study electrophysiological brain activity and perform ambulatory diagnoses of neuronal diseases. This work proposes a pairwise distance learning approach for Schizophrenia classification relying on the spectral properti…
Measuring the impact of scientific articles is important for evaluating the research output of individual scientists, academic institutions and journals. While citations are raw data for constructing impact measures, there exist biases and potential issues if factors affecting citation patterns are not properly account…
OpFlow predicts robust OD flows by learning choice potentials conditioned on spatial exposures.
Causal effect identification considers whether an interventional probability distribution can be uniquely determined without parametric assumptions from measured source distributions and structural knowledge on the generating system. While complete graphical criteria and procedures exist for many identification problem…
Algorithm reduces variance in causal effect estimation from multiple datasets.
gOMP algorithm selects features for various types of data.
Approaches for testing sets of variants, such as a set of rare or common variants within a gene or pathway, for association with complex traits are important. In particular, set tests allow for aggregation of weak signal within a set, can capture interplay among variants, and reduce the burden of multiple hypothesis te…
Gynaecologists and obstetricians visually interpret cardiotocography (CTG) traces using the International Federation of Gynaecology and Obstetrics (FIGO) guidelines to assess the wellbeing of the foetus during antenatal care. This approach has raised concerns among professionals with regards to inter- and intra-variabi…
New optimizer designs respect symmetry, improving deep learning models.
Enhanced Sampling Scheme improves masked generative modeling.
PRS improves rejection sampling by learning better proposals.
This paper reviews various sampling methods from statistics and machine learning.
RISA improves VFL by using imputed samples with low uncertainty.
Improved privacy-preserving methods for estimating multiple samples from distributions.
Paper introduces a new sampling method combining Consistency Models with importance sampling.
Neural network accuracy improves with denser training samples.
Wedge Sampling improves tensor completion with nearly-linear sample complexity.
Algorithm samples constrained stochastic differential equations.
In modern data analysis, random sampling is an efficient and widely-used strategy to overcome the computational difficulties brought by large sample size. In previous studies, researchers conducted random sampling which is according to the input data but independent on the response variable, however the response variab…
Generative models map simple samples to complex target samples.
Sampling is a fundamental problem in computer science and statistics. However, for a given task and stream, it is often not possible to choose good sampling probabilities in advance. We derive a general framework for adaptively changing the sampling probabilities via a collection of thresholds.In general, adaptive samp…
New sampling algorithm for non-smooth potentials.