Researchers analyze record statistics in correlated random walks and Lévy flights.
problem Understanding record statistics in correlated time series.
method Review of random walk models and Lévy flights, focusing on number of records and record ages.
result Effects of correlations on record statistics were observed and analyzed.
The study of record statistics of correlated series is gaining momentum. In this work, we study the records statistics of the time series of select stock market data and the geometric random walk, primarily through simulations. We show that the distribution of the age of records is a power law with the exponent α lyi…
We study the statistics of record-breaking events in daily stock prices of 366 stocks from the Standard and Poors 500 stock index. Both the record events in the daily stock prices themselves and the records in the daily returns are discussed. In both cases we try to describe the record statistics of the stock data with…
While records and order statistics of independent and identically distributed (i.i.d.) random variables X_1, ..., X_N are fully understood, much less is known for strongly correlated random variables, which is often the situation encountered in statistical physics. Recently, it was shown, in a series of works, that one…
We consider the occurrence of record-breaking events in random walks with asymmetric jump distributions. The statistics of records in symmetric random walks was previously analyzed by Majumdar and Ziff and is well understood. Unlike the case of symmetric jump distributions, in the asymmetric case the statistics of reco…
We study the statistics of records of a one-dimensional random walk of n steps, starting from the origin, and in presence of a constant bias c. At each time-step the walker makes a random jump of length ηdrawn from a continuous distribution f(η) which is symmetric around a constant drift c. We focus in particular on th…
Modeling maximum drawdown records in capital markets using PDMP.
problem Capturing the statistical properties of maximum drawdown records in financial markets.
method Piecewise Deterministic Markov Process (PDMP) for modeling, statistical analysis of mean and variance, simulation study, parameter estimation techniques.
result Derivation of statistical results including mean and variance of maximum drawdown records.
We investigate the statistics of records in a random sequence {xB(0)=0,xB(1),⋯,xB(n)=xB(0)=0} of n time steps. The sequence xB(k)'s represents the position at step k of a random walk `bridge' of n steps that starts and ends at the origin. At each step, the increment of the position is a random ju…
A new family of nonparametric statistics, the r-statistics, is introduced. It consists of counting the number of records of the cumulative sum of the sample. The single-sample r-statistic is almost as powerful as Student's t-statistic for Gaussian and uniformly distributed variables, and more powerful than the sign and…
The study sets performance limits for record linkage using KL divergence.
problem Efficiently merging records in large, noisy databases to remove duplicates.
method Assesses performance bounds using Kullback-Leibler divergence in a Bayesian record linkage framework.
result Provides upper and lower bounds on misclassification probability.
Improves Gaussian process factor models for multi-population recordings.
problem Cubic runtime scaling with trial length and group number limits application to large-scale recordings.
method Two approximate approaches: inducing variables and frequency domain.
result Achieved orders of magnitude speed-up with minimal statistical performance impact.
Proposes statistical inference for dependency knowledge graphs from EHR data.
problem Statistical uncertainty in linking entities in EHR data.
method Dynamic log-linear topic model with singular value decomposition.
result Established asymptotic normality for sparse graph edge recovery.
The extreme event statistics plays a very important role in the theory and practice of time series analysis. The reassembly of classical theoretical results is often undermined by non-stationarity and dependence between increments. Furthermore, the convergence to the limit distributions can be slow, requiring a huge am…
Big data from phone calls improves credit scoring models and profits.
problem Improving credit scoring models to enhance financial inclusion.
method Combining call-detail records and traditional data to build scorecards using social network analytics.
result Combining call-detail records with traditional data significantly increases model performance and profit.
Improved neural network detects heart sounds with 87.5% accuracy from noisy recordings.
problem Detecting cardiac abnormalities from noisy heart sound recordings.
method Segmental Convolutional Neural Network (CNN) architecture trained on noisy recordings.
result Best model achieved 87.5% accuracy on PhysioNet/CinC Challenge dataset.
Study predicts colorectal polyp recurrence using medical records and statistical models.
problem Identifying patient characteristics influencing colorectal polyp recurrence.
method Natural language processing for extracting polyp characteristics, Kaplan-Meier curves, Cox proportional hazards modeling, random survival forest models.
result Polyp size, number, location, and patient smoking status significantly influence recurrence risk.
A new method improves fitting neural data with spiking network models.
problem Fitting spiking network models to neural activity does not produce realistic data.
method Augment log-likelihood with dissimilarity terms measured by summary statistics and optimized via back-propagation.
result The new method generates more realistic neural activity statistics and improves network connectivity inference.
Paper cleans option price datasets by removing outliers.
problem Unusual option prices in datasets.
method Statistical techniques to identify and remove outliers.
result Removes option prices violating no arbitrage assumption.
Bayesian method improves record linkage accuracy.
problem Non-trivial record linkage without unique identifiers.
method Bayesian estimation of bipartite matchings.
result Bayesian approach outperforms traditional methods.
CorGAN generates synthetic healthcare records while preserving privacy.
problem Generating realistic synthetic healthcare records while maintaining privacy.
method Combining Convolutional Generative Adversarial Networks and Convolutional Autoencoders to capture correlations between medical features.
result CorGAN generates synthetic data with performance similar to real data in various ML settings.
Town hall discusses AI's impact on statistics, culture, and training.
problem Adapting statistics to AI advancements and infrastructure.
method Open panel discussion and audience Q&A.
result Candid perspectives on evolving statistical practices.
Enhances generative model for clinical data privacy and accuracy.
problem Data privacy in electronic patient records.
method Improves a time-series generative model with privacy safeguards.
result DP-TimeGAN achieves a mean authenticity of 0.778 on the CKD dataset.
Enhances detection of adverse drug events using diverse healthcare record data.
problem Detecting adverse drug events from mixed data types in electronic health records.
method Aggregate diagnosis codes, drug codes, and lab measurements; use recursive feature selection.
result Significant improvement in AUC using additional features, statistically significant.
We study the statistics of the number of records R_{n,N} for N identical and independent symmetric discrete-time random walks of n steps in one dimension, all starting at the origin at step 0. At each time step, each walker jumps by a random length drawn independently from a symmetric and continuous distribution. We co…
Meta-Dynamic models learn shared neural dynamics across tasks.
problem Learning latent dynamics from neural recordings across different tasks.
method Captures variabilities on a low-dimensional manifold to meta-learn dynamics.
result Meta-Dynamic models can rapidly learn latent dynamics from new recordings.
Learn embeddings from EHRs to predict ICD codes.
problem Predicting ICD codes from patient visits in EHRs.
method Deep neural network trained to predict ICD codes, capturing clinical information.
result Embeddings capture relevant clinical information and can be used in machine learning models.
Paper introduces 'plausible deniability' for privacy-preserving data synthesis.
problem Challenges in releasing full data records while preserving privacy.
method Introduces 'plausible deniability' criterion and mechanisms for generating synthetic datasets.
result Generative technique preserves utility of original data and is efficient for large datasets.
Method reconstructs glacier front trajectories from record moraine data.
problem Understanding past glacier dynamics from limited record data.
method Stochastic generator based on Brownian motion and NBI hyper parameter tuning.
result Reconstructed glacier front trajectories from moraine records.
EKG-based models show better stability across patient populations than EHR-based models.
problem Model generalization issues in EHR and EKG-based predictive models.
method Two tests to measure model generalization, comparing EHR and EKG data.
result EKG-based models are more stable across different patient populations.
Paper classifies heart sound recordings as normal or abnormal.
problem Classifying normal/abnormal heart sound recordings.
method Four steps: preprocessing, feature extraction, training, validation. Back propagation neural network used.
result Optimal threshold determined for distinguishing normal and abnormal.
This study analyzes inequality in Romanian income distribution using advanced statistical methods.
problem Characterizing inequality in Romania's income distribution.
method Advanced statistical techniques, specifically Theil index decomposition.
result Salient factors contributing to income inequality in Romania identified.
vLGP recovers neural dynamics from spike trains, improving prediction and capturing complex patterns.
problem Recovering latent neural trajectories from noisy spike trains is challenging.
method vLGP combines generative model, history-dependent point process observation, and smoothness prior.
result vLGP achieves higher performance in predicting omitted spike trains and capturing neural dynamics.
New method reconstructs images from fMRI data using unlabeled data.
problem Challenges in acquiring labeled data for fMRI-to-image reconstruction.
method Self-supervised training with Encoder-Decoder and Decoder-Encoder networks.
result Reconstruction network adapts to new unlabeled test data.
Study validates machine learning models for patient outcomes using various methods.
problem Validating machine learning models for patient outcomes in electronic health records.
method Used three state-of-the-art machine learning methods (random forest, gradient boosting, logistic regression) to predict patient outcomes and assess feature importance.
result Permutation tests applied to random forest and gradient boosting models showed the most agreement with clinical interpretation of feature importance.
This article reviews entity resolution methods and their applications.
problem Integrating information from multiple sources to clean and accurately link records.
method Clustering, semi- and fully supervised methods, canonicalization.
result Modern probabilistic record linkage has been foundational.
New method generates diverse EHR data types while maintaining privacy.
problem Privacy risks in sharing EHRs across large scales.
method Refined GAN model, feature constraints, and utility measures.
result New model preserves EHR statistics, correlations, and constraints.
Random projections simplify complex data for classification.
problem Handling high-dimensional data in classification problems.
method Two techniques: ensemble of random projections and hashing/sketching.
result Approximate statistical efficiency with reduced complexity.
The paper tackles uniform sampling from databases with duplicates.
problem Sampling uniformly from entities with duplicate records.
method Two-stage process: frequency estimation followed by rejection sampling.
result Efficient sampling algorithms under various data properties.
Deep learning predicts hospital readmission risk from notes.
problem Predicting hospital-wide readmission risk from doctors' notes.
method Deep learning and natural language processing on unstructured text.
result Model predicts readmission with c-statistic 0.70.
MLHO predicts COVID-19 adverse outcomes using past medical records.
problem Predicting adverse outcomes after COVID-19 infection.
method Iterative feature and algorithm selection, sequential representation mining.
result Mean AUC ROC of 0.91 for mortality prediction.
We study the return interval τ between price volatilities that are above a certain threshold q for 31 intraday datasets, including the Standard & Poor's 500 index and the 30 stocks that form the Dow Jones Industrial index. For different threshold q, the probability density function Pq(τ) scales with the mean i…
New method tackles MNAR missingness in domain adaptation.
problem Handling missingness in both source and target data.
method Reduces MNAR missingness to imputation problem, leveraging recent MNAR imputation methods.
result Developed a novel domain adaptation procedure for MNAR missingness shift.
New model captures patient-level EHR data efficiently.
problem Irregular EHR code timing and lack of temporal structure.
method Latent factor point process model with Fourier-Eigen embedding.
result Efficiently captures subgroup-specific temporal patterns.
Paper introduces max-plus statistical leverage scores for faster approximation of conventional scores.
problem Approximating statistical leverage scores of complex matrices efficiently.
method Max-plus algebraic analogue for statistical leverage scores.
result Max-plus statistical leverage scores can approximate conventional scores quickly and accurately.
Ancestry improves genealogy search results by ranking diverse record types.
problem Ranking diverse genealogy records equitably from various sources.
method Customized Coordinate Ascent, Stochastic Search, Normalized Cumulative Entropy.
result Demonstrated effectiveness of algorithms in improving relevance and diversity.
Study quantifies reproducibility of machine learning papers.
problem Lack of empirical reproducibility metrics in machine learning.
method Manual implementation of 255 papers from 1984-2017, analyzing features and results.
result Manual implementation revealed discrepancies between papers and their descriptions.
Proposes a method to derive knowledge graphs from EHR data.
problem Challenges in deriving generalizable knowledge from EHR data.
method Infer conditional dependency structure via a latent graphical block model (LGBM).
result Perfect recovery of block structure demonstrated.
Transformer learns hidden structure of EHR data for better prediction.
problem Lack of complete structure information in EHR data.
method Graph Convolutional Transformer using data statistics to learn structure.
result Consistently outperforms previous approaches on various prediction tasks.