Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

0.3%0.5%0.8%0.4% · Mar 201519922001200920172026
45 results for Meta-analysis

Develops exact and invariant study-based decompositions for network meta-analysis.

problem Lack of exact contribution decompositions in network meta-analysis.
method Contrast-space projection formulation of NMA, study-based definition of direct and indirect evidence.
result Exact covariance-aware decompositions of NMA estimator into direct and indirect contributions.

Meta-analysis improves personalized treatment rules across multiple sites.

problem Lack of generalizability in learning individualized treatment rules across different medical sites.
method Developed a method for individual-level meta-analysis of ITRs, borrowing sign-coherency information between sites.
result Jointly learned site-specific ITRs with improved generalizability.

Reaching a global view of brain organization requires assembling evidence on widely different mental processes and mechanisms. The variety of human neuroscience concepts and terminology poses a fundamental challenge to relating brain imaging results across the scientific literature. Existing meta-analysis methods perfo…

2020-02-21abs ↗pdf ↗

Proposes a method to correct for covariate shift in meta-analysis of randomized trials.

problem Invalidation of standard IPD meta-analysis due to covariate shift across studies.
method Placebo-anchored transport framework that treats source-trial outcomes as proxy signals and target-trial placebo outcomes as gold labels.
result Yields target-identified effect estimates in connected targets and a principled screen--then--transport procedure in disconnected targets.

Proposes a new method to handle data heterogeneity in causal inference.

problem Challenges of collaborating between different data centers due to heterogeneity.
method Collaborative inverse propensity score weighting estimator to adjust for distribution shift.
result Significant improvements over traditional meta-analysis methods when dealing with increased heterogeneity.

Universal approach combines OOD detection scores for robustness.

problem Combining diverse OOD detection scores for robustness.
method Quantile normalization to p-values, meta-analysis, probabilistic interpretation.
result Significantly improved robustness and performance across diverse OOD detection scenarios.

Large-scale automated meta-analysis of neuroimaging data has recently established itself as an important tool in advancing our understanding of human brain function. This research has been pioneered by NeuroSynth, a database collecting both brain activation coordinates and associated text across a large cohort of neuro…

2016-05-01abs ↗pdf ↗

The primary goal of this study is doing a meta-analysis research on two groups of published studies. First, the ones that focus on the evaluation of the United States Department of Agriculture (USDA) forecasts and second, the ones that evaluate the market reactions to the USDA forecasts. We investigate four questions. …

2018-01-19abs ↗pdf ↗

Mixed effects (ME) models inform a vast array of problems in the physical and social sciences, and are pervasive in meta-analysis. We consider ME models where the random effects component is linear. We then develop an efficient approach for a broad problem class that allows nonlinear measurements, priors, and constrain…

2019-09-24abs ↗pdf ↗

The study assesses machine learning generalization using various data set characteristics.

problem Estimating confidence in machine learning predictions and assessing generalization capabilities.
method Meta-analysis of 109 classification data sets, modeling generalization as a function of various characteristics.
result The convex hull of the training data is relevant for assessing machine learning generalization, challenging the common assumption about the curse of dimensionality.

This article provides a thorough meta-analysis of the anomaly detection problem. To accomplish this we first identify approaches to benchmarking anomaly detection algorithms across the literature and produce a large corpus of anomaly detection benchmarks that vary in their construction across several dimensions we deem…

2015-03-03abs ↗pdf ↗

Latent Noise Injection improves synthetic data generation for privacy and statistical alignment.

problem Slow convergence of generative models in high-dimensional settings.
method Latent Noise Injection using Masked Autoregressive Flows (MAF).
result Synthetic data closely reflects the underlying distribution, especially in high-dimensional settings.

The outcome of a functional genomics pipeline is usually a partial list of genomic features, ranked by their relevance in modelling biological phenotype in terms of a classification or regression model. Due to resampling protocols or just within a meta-analysis comparison, instead of one list it is often the case that …

2010-04-08abs ↗pdf ↗

We present a new data-driven benchmark system to evaluate the performance of new MCMC samplers. Taking inspiration from the COCO benchmark in optimization, we view this task as having critical importance to machine learning and statistics given the rate at which new samplers are proposed. The common hand-crafted exampl…

2017-12-16abs ↗pdf ↗

We present the Parallel, Forward-Backward with Pruning (PFBP) algorithm for feature selection (FS) in Big Data settings (high dimensionality and/or sample size). To tackle the challenges of Big Data FS PFBP partitions the data matrix both in terms of rows (samples, training examples) as well as columns (features). By e…

2017-08-23abs ↗pdf ↗

Paper proposes a novel method to assess treatment effect estimators using cross-validation.

problem Lack of ground truth to objectively assess treatment effect estimators in RCTs.
method Cross-validation-like methodology combining noisy difference-of-means estimate and aggregation across RCTs.
result Aggressive downweighting or truncation of large values reduces variance and improves treatment effect estimation.

We present the expected values from p-value hacking as a choice of the minimum p-value among mm independents tests, which can be considerably lower than the "true" p-value, even with a single trial, owing to the extreme skewness of the meta-distribution. We first present an exact probability distribution (meta-distrib…

2016-03-24abs ↗pdf ↗

Study finds little progress in medical machine learning benchmarks over 3 years.

problem Lack of meaningful progress in medical machine learning benchmarks for structured healthcare data.
method Comprehensive review and meta-analysis of benchmarks in medical machine learning for structured data.
result Deep recurrent models perform only better than logistic regression on certain clinical prediction tasks.

fedCI and fedCI-IOD enable federated causal discovery across diverse datasets with privacy and power enhancements.

problem Causal discovery across multiple datasets with privacy constraints and heterogeneity.
method federated conditional independence test (fedCI) and Integration of Overlapping Datasets (IOD) algorithm extension (fedCI-IOD).
result fedCI-IOD achieves comparable performance to fully pooled analyses, enhancing statistical power and privacy.

Paper introduces ZIPTF and C-ZIPTF for better tensor factorization of zero-inflated count data.

problem Inefficient tensor factorization for zero-inflated count data, especially in scRNA-seq.
method Zero Inflated Poisson Tensor Factorization (ZIPTF) and Consensus Zero Inflated Poisson Tensor Factorization (C-ZIPTF).
result ZIPTF and C-ZIPTF improve tensor factorization accuracy and consistency for zero-inflated count data.

Wisdom of the crowd, the collective intelligence derived from responses of multiple human or machine individuals to the same questions, can be more accurate than each individual, and improve social decision-making and prediction accuracy. This can also integrate multiple programs or datasets, each as an individual, for…

2017-11-28abs ↗pdf ↗

KM-GPT automates IPD reconstruction from KM plots with high accuracy and scalability.

problem Manual digitization of IPD from KM plots is error-prone and lacks scalability.
method KM-GPT integrates advanced image preprocessing, multi-modal reasoning, and iterative reconstruction algorithms.
result KM-GPT generates high-quality IPD without manual input or intervention, achieving superior accuracy.

This work classifies strategies to incorporate spatial dependence in Random Forest models.

problem Spatial and temporal dependence in environmental data not adequately modeled by standard Random Forest.
method Taxonomy and systematic review of strategies to adjust Random Forest for spatially dependent data.
result 32 scientific documents reviewed, providing a comprehensive classification of strategies.

This paper computes exact posterior distributions of mixture weights in hierarchical Bayesian models.

problem Uncertainty in class membership or data-generating processes in heterogeneous data.
method Exact marginalization of mixture weights using dynamic programming and FFT for two components, and joint dynamic program for K >= 3 components.
result Exact posterior distributions of mixture weights are finite mixtures of Beta distributions, providing credible intervals and per-observation local false-discovery rates.

New BGs use diffusion models to improve sampling from complex distributions.

problem Sampling from complex, multi-modal distributions is challenging.
method Combines diffusion models with annealed Monte Carlo for improved sampling.
result Second-order denoising kernels can improve performance in high-dimensional spaces.