Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

67134200267 · Jun 202019922001200920182026
48 results for record statistics

Researchers analyze record statistics in correlated random walks and Lévy flights.

problem Understanding record statistics in correlated time series.
method Review of random walk models and Lévy flights, focusing on number of records and record ages.
result Effects of correlations on record statistics were observed and analyzed.

The study of record statistics of correlated series is gaining momentum. In this work, we study the records statistics of the time series of select stock market data and the geometric random walk, primarily through simulations. We show that the distribution of the age of records is a power law with the exponent αα lyi…

2014-06-24abs ↗pdf ↗

We study the statistics of record-breaking events in daily stock prices of 366 stocks from the Standard and Poors 500 stock index. Both the record events in the daily stock prices themselves and the records in the daily returns are discussed. In both cases we try to describe the record statistics of the stock data with…

2013-07-08abs ↗pdf ↗

While records and order statistics of independent and identically distributed (i.i.d.) random variables X_1, ..., X_N are fully understood, much less is known for strongly correlated random variables, which is often the situation encountered in statistical physics. Recently, it was shown, in a series of works, that one…

2013-05-03abs ↗pdf ↗

We study the statistics of records of a one-dimensional random walk of n steps, starting from the origin, and in presence of a constant bias c. At each time-step the walker makes a random jump of length ηdrawn from a continuous distribution f(η) which is symmetric around a constant drift c. We focus in particular on th…

2012-06-29abs ↗pdf ↗

Modeling maximum drawdown records in capital markets using PDMP.

problem Capturing the statistical properties of maximum drawdown records in financial markets.
method Piecewise Deterministic Markov Process (PDMP) for modeling, statistical analysis of mean and variance, simulation study, parameter estimation techniques.
result Derivation of statistical results including mean and variance of maximum drawdown records.

We investigate the statistics of records in a random sequence {xB(0)=0,xB(1),,xB(n)=xB(0)=0}\{x_B(0)=0,x_B(1),\cdots, x_B(n)=x_B(0)=0\} of nn time steps. The sequence xB(k)x_B(k)'s represents the position at step kk of a random walk `bridge' of nn steps that starts and ends at the origin. At each step, the increment of the position is a random ju…

2015-05-22abs ↗pdf ↗

Improves Gaussian process factor models for multi-population recordings.

problem Cubic runtime scaling with trial length and group number limits application to large-scale recordings.
method Two approximate approaches: inducing variables and frequency domain.
result Achieved orders of magnitude speed-up with minimal statistical performance impact.

The extreme event statistics plays a very important role in the theory and practice of time series analysis. The reassembly of classical theoretical results is often undermined by non-stationarity and dependence between increments. Furthermore, the convergence to the limit distributions can be slow, requiring a huge am…

2011-05-31abs ↗pdf ↗

Big data from phone calls improves credit scoring models and profits.

problem Improving credit scoring models to enhance financial inclusion.
method Combining call-detail records and traditional data to build scorecards using social network analytics.
result Combining call-detail records with traditional data significantly increases model performance and profit.

Improved neural network detects heart sounds with 87.5% accuracy from noisy recordings.

problem Detecting cardiac abnormalities from noisy heart sound recordings.
method Segmental Convolutional Neural Network (CNN) architecture trained on noisy recordings.
result Best model achieved 87.5% accuracy on PhysioNet/CinC Challenge dataset.

Study predicts colorectal polyp recurrence using medical records and statistical models.

problem Identifying patient characteristics influencing colorectal polyp recurrence.
method Natural language processing for extracting polyp characteristics, Kaplan-Meier curves, Cox proportional hazards modeling, random survival forest models.
result Polyp size, number, location, and patient smoking status significantly influence recurrence risk.

A new method improves fitting neural data with spiking network models.

problem Fitting spiking network models to neural activity does not produce realistic data.
method Augment log-likelihood with dissimilarity terms measured by summary statistics and optimized via back-propagation.
result The new method generates more realistic neural activity statistics and improves network connectivity inference.

CorGAN generates synthetic healthcare records while preserving privacy.

problem Generating realistic synthetic healthcare records while maintaining privacy.
method Combining Convolutional Generative Adversarial Networks and Convolutional Autoencoders to capture correlations between medical features.
result CorGAN generates synthetic data with performance similar to real data in various ML settings.

Enhances detection of adverse drug events using diverse healthcare record data.

problem Detecting adverse drug events from mixed data types in electronic health records.
method Aggregate diagnosis codes, drug codes, and lab measurements; use recursive feature selection.
result Significant improvement in AUC using additional features, statistically significant.

We study the statistics of the number of records R_{n,N} for N identical and independent symmetric discrete-time random walks of n steps in one dimension, all starting at the origin at step 0. At each time step, each walker jumps by a random length drawn independently from a symmetric and continuous distribution. We co…

2012-04-23abs ↗pdf ↗

Paper introduces 'plausible deniability' for privacy-preserving data synthesis.

problem Challenges in releasing full data records while preserving privacy.
method Introduces 'plausible deniability' criterion and mechanisms for generating synthetic datasets.
result Generative technique preserves utility of original data and is efficient for large datasets.

Method reconstructs glacier front trajectories from record moraine data.

problem Understanding past glacier dynamics from limited record data.
method Stochastic generator based on Brownian motion and NBI hyper parameter tuning.
result Reconstructed glacier front trajectories from moraine records.

EKG-based models show better stability across patient populations than EHR-based models.

problem Model generalization issues in EHR and EKG-based predictive models.
method Two tests to measure model generalization, comparing EHR and EKG data.
result EKG-based models are more stable across different patient populations.

Paper classifies heart sound recordings as normal or abnormal.

problem Classifying normal/abnormal heart sound recordings.
method Four steps: preprocessing, feature extraction, training, validation. Back propagation neural network used.
result Optimal threshold determined for distinguishing normal and abnormal.

This study analyzes inequality in Romanian income distribution using advanced statistical methods.

problem Characterizing inequality in Romania's income distribution.
method Advanced statistical techniques, specifically Theil index decomposition.
result Salient factors contributing to income inequality in Romania identified.

vLGP recovers neural dynamics from spike trains, improving prediction and capturing complex patterns.

problem Recovering latent neural trajectories from noisy spike trains is challenging.
method vLGP combines generative model, history-dependent point process observation, and smoothness prior.
result vLGP achieves higher performance in predicting omitted spike trains and capturing neural dynamics.

Study validates machine learning models for patient outcomes using various methods.

problem Validating machine learning models for patient outcomes in electronic health records.
method Used three state-of-the-art machine learning methods (random forest, gradient boosting, logistic regression) to predict patient outcomes and assess feature importance.
result Permutation tests applied to random forest and gradient boosting models showed the most agreement with clinical interpretation of feature importance.

We study the return interval ττ between price volatilities that are above a certain threshold qq for 31 intraday datasets, including the Standard & Poor's 500 index and the 30 stocks that form the Dow Jones Industrial index. For different threshold qq, the probability density function Pq(τ)P_q(τ) scales with the mean i…

2005-11-11abs ↗pdf ↗

Proposes a method to derive knowledge graphs from EHR data.

problem Challenges in deriving generalizable knowledge from EHR data.
method Infer conditional dependency structure via a latent graphical block model (LGBM).
result Perfect recovery of block structure demonstrated.