We identify action representations from video data, proving their statistical benefits.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Statistical test evaluates if personalizing interventions is cost-effective.
Cost-benefit analysis often assumes accurate estimates, but this study finds significant inaccuracies.
Statistical arbitrage is a class of financial trading strategies using mean reversion models. The corresponding techniques rely on a number of assumptions which may not hold for general non-stationary stochastic processes. This paper presents an alternative technique for statistical arbitrage based on online learning w…
New algorithm balances user reward and statistical inference by mixing TS with UR based on difference size.
New framework improves cost-benefit analysis of policies.
Oil is perceived as a good diversification tool for stock markets. To fully understand this potential, we propose a new empirical methodology that combines generalized autoregressive score copula functions with high frequency data and allows us to capture and forecast the conditional time-varying joint distribution of …
Actuaries tackle loss of earning capacity in Denmark, balancing public benefits and private insurance.
We design and study a Contextual Memory Tree (CMT), a learning memory controller that inserts new memories into an experience store of unbounded size. It is designed to efficiently query for memories from that store, supporting logarithmic time insertion and retrieval operations. Hence CMT can be integrated into existi…
The paper studies the benefits of curriculum learning in linear regression tasks.
Score matching offers efficient estimation for certain distributions.
Data augmentation can achieve the same statistical benefits as full augmentation up to an approximation error.
Entropy-based model for hierarchical learning from multiscale data.
This paper explores how entropic regularization improves Wasserstein estimators' performance.
Early stopping improves logistic regression's calibration and consistency in high dimensions.
The paper introduces a method to measure the benefits of incidental supervision signals.
This paper derives -- considering a Gaussian setting -- closed form solutions of the statistics that Adrian and Brunnermeier and Acharya et al. have suggested as measures of systemic risk to be attached to individual banks. The statistics equal the product of statistic specific Beta-coefficients with the mean corrected…
National statistical systems are the enterprises tasked with collecting, validating and reporting societal attributes. These data serve many purposes - they allow governments to improve services, economic actors to traverse markets, and academics to assess social theories. National statistical systems vary in quality, …
Framework for discovering treatment benefits in user segments.
This dissertation shows that careful injection of noise into sample data can substantially speed up Expectation-Maximization algorithms. Expectation-Maximization algorithms are a class of iterative algorithms for extracting maximum likelihood estimates from corrupted or incomplete data. The convergence speed-up is an e…
Distributions over permutations arise in applications ranging from multi-object tracking to ranking of instances. The difficulty of dealing with these distributions is caused by the size of their domain, which is factorial in the number of considered entities (). It makes the direct definition of a multinomial dist…
This article introduces a framework to estimate the value of evidence-based decision making.
Data balancing reduces variance in machine learning models.
A wide variety of machine learning algorithms such as support vector machine (SVM), minimax probability machine (MPM), and Fisher discriminant analysis (FDA), exist for binary classification. The purpose of this paper is to provide a unified classification model that includes the above models through a robust optimizat…
One of the most surprising and exciting discoveries in supervised learning was the benefit of overparameterization (i.e. training a very large model) to improving the optimization landscape of a problem, with minimal effect on statistical performance (i.e. generalization). In contrast, unsupervised settings have been u…
The authors argue against the classification of forecasting methods as machine learning or statistical.
New methods stabilize EEG classification performance across subjects.
The Maximum Mean Discrepancy (MMD) has found numerous applications in statistics and machine learning, most recently as a penalty in the Wasserstein Auto-Encoder (WAE). In this paper we compute closed-form expressions for estimating the Gaussian kernel based MMD between a given distribution and the standard multivariat…
New methods solve tensor-on-tensor regression with unknown rank, revealing benefits of over-parameterization.
Enhanced ROOT-SGD optimizes stochastic optimization with diminishing stepsizes.
Research creates a machine learning model for predicting TAVI patient mortality.
Improved statistical inference for adaptive Thompson Sampling.
Extreme value theory enhances statistical learning extrapolation for rare events.
This paper presents a class of new algorithms for distributed statistical estimation that exploit divide-and-conquer approach. We show that one of the key benefits of the divide-and-conquer strategy is robustness, an important characteristic for large distributed systems. We establish connections between performance of…
Non-linear image reconstruction and signal analysis deal with complex inverse problems. To tackle such problems in a systematic way, I present information field theory (IFT) as a means of Bayesian, data based inference on spatially distributed signal fields. IFT is a statistical field theory, which permits the construc…
Paper proposes a statistical test for transfer learning in linear regression.
The paper establishes limits of transfer learning with neural networks.
Theoretical guarantees for neural estimators in parametric statistics are derived.
A number of results have recently demonstrated the benefits of incorporating various constraints when training deep architectures in vision and machine learning. The advantages range from guarantees for statistical generalization to better accuracy to compression. But support for general constraints within widely used …
This work presents a methodology to design trajectory tracking feedback control laws, which embed non-parametric statistical models, such as Gaussian Processes (GPs). The aim is to minimize unmodeled dynamics such as undesired slippages. The proposed approach has the benefit of avoiding complex terramechanics analysis …
The two key issues of modern Bayesian statistics are: (i) establishing principled approach for distilling statistical prior that is consistent with the given data from an initial believable scientific prior; and (ii) development of a Bayes-frequentist consolidated data analysis workflow that is more effective than eith…
Neural networks speed up statistical inference.
A neural network approach unifies Lasso for variable selection.
New algorithms improve Bayesian linear regression with spike-and-slab priors.
Causal ML predicts treatment outcomes, aiding personalized medicine.
Simple methods combine statistical tests for out-of-distribution detection.
PAS improves estimation of multiple means using ML predictions and shrinkage.
This article presents results from the first statistically significant study of cost escalation in transportation infrastructure projects. Based on a sample of 258 transportation infrastructure projects worth US$90 billion and representing different project types, geographical regions, and historical periods, it is fou…