Proposes BA method for unbiased time series anomaly detection evaluation.
problem Anomalies in time series data are rare, making F1-score unreliable.
method Introduces Balanced Point Adjustment (BA) to address F1-score bias.
result BA provides fairer evaluation of time series anomaly detectors.
Prognostic scores improve logistic regression analysis in RCTs with binary outcomes.
problem Non-collapsibility in logistic regression analysis of RCTs with binary endpoints.
method Prognostic score adjustment using AI predictions to address non-collapsibility.
result Prognostic score adjustment increases power or reduces sample size for estimating conditional odds ratios.
Paper introduces OCRR Score for quantifying DeFi wallet credit risk.
problem Inability to assess credit risk in decentralized finance.
method Probabilistic measure based on historical and predictive on-chain activity.
result Dynamic adjustment of LTV and LT based on wallet risk profile.
A new activation function improves credit scoring accuracy for imbalanced datasets.
problem Imbalanced datasets in credit scoring lead to underestimation of misclassification costs.
method Introduces ASIG, an asymmetric adjusted Sigmoid function.
result ASIG-embedded classifier outperforms traditional classifiers across various imbalance ratios.
New method detects and adjusts for temporal leakage in LLM backtests.
problem Standard backtest leakage detection is ineffective for modern models.
method Developed new methods to measure and adjust for temporal leakage.
result Demonstrated that models legitimately know more about times near their cutoffs, leading to structural leakage.
Proposes SD-KDE for density estimation using debiased kernel density with score-based adjustments.
problem Density estimation with bias in kernel density estimation.
method Adjusts data points by taking a step along the estimated score function, then applies standard KDE with modified bandwidth.
result Significantly reduces mean integrated squared error compared to standard Silverman KDE, especially with noisy score function estimates.
Bayesian method improves clinical trial efficiency.
problem Increase treatment effect estimates in clinical trials.
method Combines prognostic covariate adjustment with a Bayesian framework.
result Substantial increase in statistical power with controlled type I error.
New method improves sampling from score-based models by correcting bias.
problem Bias in sampling from score-based diffusion models.
method Metropolis-Hastings or Barker's accept-reject steps to correct bias, using the score function.
result Improves sample quality on synthetic and image datasets, yielding consistent gains in FID.
Bayesian method estimates QTEs from observational data.
problem Estimating nuanced characteristics of counterfactual distributions.
method Bayesian semiparametric conditional distribution regression model with double balancing score.
result Proposed method provides more accurate QTE estimates than other methods.
Improves trial efficiency by adjusting for historical prognostic scores.
problem Reducing statistical uncertainty in randomized trial estimates.
method Linear covariate adjustment using a prognostic model trained on historical data.
result Prognostic covariate adjustment achieves minimum variance and reduces mean-squared error.
Paper introduces Isotonic Mechanism for better item scoring.
problem Noisy reviewer scores; owner prefers not to disclose true scores.
method Uses owner's ranking of items and raw scores to adjust and improve accuracy.
result Adjusted scores are significantly more accurate than raw scores.
The paper presents an R package for learning Bayesian networks from epidemiological data.
problem Learning Bayesian networks from messy, highly correlated datasets in epidemiology.
method The paper introduces an R package abn that implements multiple frequentist scoring rules for learning Bayesian networks from observational data, addressing data separation and adjustment issues.
result The package abn is robust and efficient for learning Bayesian networks from epidemiological data.
CARD detects treatment responders with machine learning and adjustment.
problem Identifying responders in non-random treatment settings.
method Conformal prediction, machine learning, propensity score adjustment.
result High power responder detection in various scenarios.
New measure corrects news bias in NLP stock return forecasting.
problem Improving stock return and volatility forecasting accuracy.
method Hype-Adjusted Probability Measure, sentiment score equation.
result Significantly improved forecast accuracy for U.S. semiconductor tickers.
PCA method adjusted for local correlation in high-dimensional data.
problem PCA assumptions fail in datasets with local correlation.
method Generalized spiked population model, consistent estimation methods.
result Reduced bias and improved prediction accuracy.
Paper proposes embedding medical concepts from claims data for better risk adjustment models.
problem Lack of efficient representation of medical histories in risk adjustment models.
method Semantic embeddings of medical concepts from diagnostic, procedure, and prescription codes.
result Embedding-based models outperform commercial risk adjustment models in prospective risk score prediction.
Proposes a new method to handle data heterogeneity in causal inference.
problem Challenges of collaborating between different data centers due to heterogeneity.
method Collaborative inverse propensity score weighting estimator to adjust for distribution shift.
result Significant improvements over traditional meta-analysis methods when dealing with increased heterogeneity.
High-dimensional adjustment reduces bias in estimating peer effects from observational data.
problem Estimating peer effects from observational data is challenging due to confounding variables and high bias.
method Used high-dimensional adjustment with propensity score models to estimate peer effects.
result High-dimensional adjustment produces estimates of peer effects statistically indistinguishable from randomized experiments.
NROWAN-DQN improves stability and exploration in noisy networks.
problem Noisy networks struggle with stable exploration in complex tasks.
method Noise reduction and online weight adjustment for stable actions.
result NROWAN-DQN outperforms prior algorithms in stability and exploration.
Paper proposes a method to reduce hallucinations in diffusion models using Laplacian score sharpening.
problem Hallucinations in diffusion models create incoherent or unrealistic samples.
method Post-hoc adjustment to the score function during inference using Laplacian approximation.
result Significantly reduces the rate of hallucinated samples across various data types.
We revisit the problem of feature selection in linear discriminant analysis (LDA), that is, when features are correlated. First, we introduce a pooled centroids formulation of the multiclass LDA predictor function, in which the relative weights of Mahalanobis-transformed predictors are given by correlation-adjusted t…
The paper proposes a test to assess rater accuracy while accounting for rater covariates.
problem Assessing the accuracy of raters in medical imaging and forensic studies.
method Covariate-adjusted homogeneity test to determine differences in accuracy among multiple rater groups.
result The proposed test identifies statistically significant differences among five participant groups in a face recognition study.
MAFLA improves sampling from heavy-tailed distributions using MH-inspired corrections.
problem Sampling from heavy-tailed and multimodal distributions when neither target nor proposal densities can be evaluated.
method Metropolis-Adjusted Fractional Langevin Algorithm (MAFLA) with Score Balance Matching.
result MAFLA significantly improves finite-time sampling accuracy over unadjusted fractional Langevin dynamics.
Proposes measures for uncertainty quantification using proper scoring rules.
problem Uncertainty quantification for prediction tasks.
method Decomposes proper scoring rules into divergence and entropy components, tailoring uncertainty quantification to specific tasks.
result Flexibility in uncertainty quantification improves performance in selective prediction and active learning.
Cancer patients admitted to ICU had improved survival over 10 years.
problem To assess changes in survival of cancer patients admitted to ICU over 10 years.
method Retrospective analysis of MIMIC-III database, adjusted for confounders using logistic regression.
result Cancer patients had significantly lower 28-day and 1-year mortality rates over 10 years.
New method scores DAGs by identifying unobserved confounding.
problem Unobserved confounding complicates causal discovery.
method Score-based causal discovery algorithm that accounts for unobserved confounding.
result Sparse linear Gaussian DAGs can be recovered from observed data.
Proposes stabilized weights for causal inference using isotonic calibration.
problem Stability and bias issues in inverse propensity weighting.
method Post-hoc isotonic calibration of inverse propensity weights.
result Improves performance of doubly robust estimators of average treatment effect.
Proposes BSF to evaluate structure learning algorithms without bias.
problem Evaluation bias in structure learning algorithms.
method Balanced Scoring Function (BSF) to adjust reward based on edge difficulty.
result Eliminates bias in favour of underfitted graphs.
Proposes a method to balance fairness and utility in ranking models.
problem Systematic disparity across protected groups in ranking models.
method Model-agnostic post-processing framework using dynamic programming.
result Achieves a balance between fairness and utility across various metrics and datasets.
TQA improves prediction intervals for time series data by adjusting quantiles for both cross-sectional and longitudinal coverage.
problem Constructing reliable prediction intervals for cross-sectional time series data.
method Temporal Quantile Adjustment (TQA) method that adjusts the quantile in Conformal Prediction to account for both cross-sectional and longitudinal coverage.
result TQA improves longitudinal coverage while preserving cross-sectional coverage, as validated through extensive experimentation.
Weather2vec learns representations to adjust for non-local confounding in air pollution studies.
problem Non-local confounding in evaluating environmental policies and climate events on health outcomes.
method weather2vec framework using balancing scores to learn representations of non-local information.
result The framework effectively adjusts for confounding in air pollution studies.
Paper proposes a self-learning framework for reject inference in credit scoring.
problem Sample bias in credit scoring models due to training on accepted cases only.
method Develops a self-learning framework considering distinct training regimes for iterative labeling and model training, introduces a new evaluation measure.
result Demonstrates the superiority of the adjusted self-learning framework over regular self-learning and previous reject inference strategies.
Bayesian PROCOVA uses AI to adjust for covariates in RCTs.
problem Unbiased and precise treatment effect inferences from RCTs.
method Generative AI constructs digital twins for covariate adjustment, using an additive mixture prior.
result Efficiency gains in smaller RCTs compared to frequentist methods.
New framework assesses LLM security risks in BFSI.
problem Lack of domain-specific security evaluation for LLMs in BFSI.
method Risk-aware evaluation framework combining taxonomy, automated red-teaming, and ensemble judging.
result Higher decoding stochasticity and adaptive interaction lead to more severe disclosures.
Study introduces a new framework for policy learning without positivity assumption.
problem Learning optimal treatment assignment policies from observational data with constraints.
method Incremental propensity score policies and semiparametric efficiency theory.
result Validated framework's performance through numerical experiments.
Model analyzes cooccurrence data for recommender systems and item relevance.
problem High-dimensional cooccurrence data from online platforms.
method Shared parameter Alternating Tweedie (SA-Tweedie) model with Fisher scoring and learning rate adjustment.
result SA-Tweedie model outperforms other methods in optimizing parameters.
A neural-network model clusters subjects based on their lifetime distributions.
problem Clustering subjects into clusters based on their lifetime distributions.
method A neural-network based lifetime clustering model that maximizes divergence between empirical lifetime distributions of clusters.
result Significantly better lifetime clusters compared to competing approaches.
Proposes SDRG to adjust missingness in machine learning models.
problem Systemic missingness in observational data leads to biased parameter estimation.
method Introduces SDRG using two models: weight-corrected gradients and per-covariate control variates.
result Empirically demonstrates convergence in training image classifiers with missing data.
New metrics fail adversarial tests, with some more robust than others.
problem Evaluation metrics for time-series anomaly detection were improved but not fully robust.
method Adversarial stress-testing of 12 adopted metrics on real benchmarks.
result Some metrics are more robust than others, with ROC-based metrics being gamed more often.
New conformal prediction methods for long-tailed classification problems.
problem Rare classes are systematically omitted in existing conformal prediction methods.
method Introduced a new conformal score function and a new interpolation procedure.
result Smoothly trade off set size and class-conditional coverage.
New method improves counterfactual distribution learning for high-dimensional outcomes.
problem Counterfactual distribution learning for high-dimensional outcomes with concentrated structure.
method Geometry-adaptive diffusion-guided smoothing estimators combining causal nuisance adjustment and local outcome geometry.
result Geometry-adaptive methods show steeper error decay in semi-synthetic experiments.
Ablation studies show BCF model's propensity score is not essential for treatment effect estimation.
problem Understanding the necessity of propensity score in nonparametric treatment effect estimation.
method Partial ablation studies of Bayesian Causal Forest (BCF) model.
result Excluding estimated propensity score does not affect treatment effect estimation or uncertainty quantification.
The paper improves Lasso de-biasing methods to enhance confidence interval efficiency.
problem Improving confidence intervals for Lasso in high-dimensional linear models.
method Degrees-of-freedom adjustment to modify Lasso de-biasing schemes.
result The degrees-of-freedom adjustment ensures asymptotic efficiency for any direction a0 under certain conditions. Deconfounding scores improve causal effect estimation with weak overlap.
problem Poor overlap in treatment and control groups makes causal effect estimators brittle.
method Introduces feature representations that improve overlap without introducing bias.
result Deconfounding scores satisfy a zero-covariance condition that is identifiable in observed data.
CALVER verifies causal reasoning traces, improving over voting methods in complex queries.
problem Voting fails in causal reasoning due to repeated confounding errors and multiple valid answers.
method CALVER scores structured traces against causal criteria and selects the highest-scoring candidate.
result CALVER selects valid answers more accurately than voting methods, especially with larger sample sizes.
Develops a method to estimate quantiles in censored data using random forests.
problem Inability of random forests to handle censored data, leading to poor predictive performance.
method Censored Quantile Regression Forests (CQRF) based on local adaptive random forests.
result Consistent estimation of quantiles without parametric modeling assumptions.
Balance corrects biased survey data for more accurate insights.
problem Bias in survey data leads to inaccurate insights and underperforming models.
method Three steps: bias understanding, weight adjustment, and evaluation.
result Corrected data leads to more accurate ML model training and insights.
Propensity score matching improves fairness in machine learning models.
problem Bias in training data affects fairness metrics in machine learning models.
method Propensity score matching to evaluate and mitigate bias in test data.
result FairMatch significantly reduces bias in test data without sacrificing predictive performance.