We discuss the finite sample theoretical properties of online predictions in non-stationary time series under model misspecification. To analyze the theoretical predictive properties of statistical methods under this setting, we first define the Kullback-Leibler risk, in order to place the problem within a decision the…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper explores game-theoretic alignment of LLMs with human preferences, finding limitations and conditions.
LMs inevitably generate hallucinations, but can be made statistically negligible.
New framework for learning from imbalanced data with theoretical guarantees.
We integrate information-theoretic concepts into the design and analysis of optimistic algorithms and Thompson sampling. By making a connection between information-theoretic quantities and confidence bounds, we obtain results that relate the per-period performance of the agent with its information gain about the enviro…
Theoretical analysis of deep neural networks for time series data.
Gradient-based Monte Carlo sampling algorithms, like Langevin dynamics and Hamiltonian Monte Carlo, are important methods for Bayesian inference. In large-scale settings, full-gradients are not affordable and thus stochastic gradients evaluated on mini-batches are used as a replacement. In order to reduce the high vari…
Action chunking and data exploration improve behavior cloning in robotics.
Survey on multiplayer bandits, highlighting theoretical gaps and future directions.
Survey of deep learning methods for inverse problems, highlighting theoretical challenges.
Unified framework improves meta-learning generalization bounds.
Dimension reduction and variable selection are performed routinely in case-control studies, but the literature on the theoretical aspects of the resulting estimates is scarce. We bring our contribution to this literature by studying estimators obtained via L1 penalized likelihood optimization. We show that the optimize…
We study the risk performance of distributed learning for the regularization empirical risk minimization with fast convergence rate, substantially improving the error analysis of the existing divide-and-conquer based distributed learning. An interesting theoretical finding is that the larger the diversity of each local…
In this paper we study the learnability of deep random networks from both theoretical and practical points of view. On the theoretical front, we show that the learnability of random deep networks with sign activation drops exponentially with its depth. On the practical front, we find that the learnability drops sharply…
We study the task of online boosting--combining online weak learners into an online strong learner. While batch boosting has a sound theoretical foundation, online boosting deserves more study from the theoretical perspective. In this paper, we carefully compare the differences between online and batch boosting, and pr…
By establishing a connection between bi-directional Helmholtz machines and information theory, we propose a generalized Helmholtz machine. Theoretical and experimental results show that given \textit{shallow} architectures, the generalized model outperforms the previous ones substantially.
The VAE's reconstruction ability is studied using PAC-Bayes theory.
Theoretical study of random forests for nonlinear time series.
Graph matching---aligning a pair of graphs to minimize their edge disagreements---has received wide-spread attention from both theoretical and applied communities over the past several decades, including combinatorics, computer vision, and connectomics. Its attention can be partially attributed to its computational dif…
The study initiates a theoretical analysis of dynamic benchmarking models.
NCDEs improve predictions for irregular time series data.
M3PO improves model-based meta-RL with theoretical guarantees.
Proposes a new information-theoretic framework for analyzing deep neural networks.
Develops a new robustness criterion for VAEs and provides theoretical guarantees.
Theoretical analysis confirms non-conservative algorithms can converge to optimal policies.
New research shows existing information-theoretic methods can't establish minimax rates for gradient descent in stochastic convex optimization.
New model learns from random graph samples to estimate graph parameters.
Recovery of low-rank matrices from a small number of linear measurements is now well-known to be possible under various model assumptions on the measurements. Such results demonstrate robustness and are backed with provable theoretical guarantees. However, extensions to tensor recovery have only recently began to be st…
We provide a general theoretical analysis of expected out-of-sample utility, also referred to as decision-theoretic classification, for non-decomposable binary classification metrics such as F-measure and Jaccard coefficient. Our key result is that the expected out-of-sample utility for many performance metrics is prov…
Paper analyzes self-supervised image denoising with denatured data.
This work analyzes Fréchet regression using comparison geometry, providing theoretical and practical insights.
The paper explores the information-theoretic nature of excess risk in machine learning.
Paper explains why small-loss criterion works for learning from noisy labels.
In PU learning, a binary classifier is trained from positive (P) and unlabeled (U) data without negative (N) data. Although N data is missing, it sometimes outperforms PN learning (i.e., ordinary supervised learning). Hitherto, neither theoretical nor experimental analysis has been given to explain this phenomenon. In …
Understanding theoretical properties of deep and locally connected nonlinear network, such as deep convolutional neural network (DCNN), is still a hard problem despite its empirical success. In this paper, we propose a novel theoretical framework for such networks with ReLU nonlinearity. The framework explicitly formul…
This research provides theoretical guarantees for hyperparameter estimation in complex network dynamical systems.
We call a group FJ if it satisfies the - and -theoretic Farrell-Jones conjecture with coefficients in . We show that if is FJ, then the simple Borel conjecture (in dimensions ) holds for every group of the form . If in addition , which is true for …
Theoretical justification for asymmetric actor-critic algorithms in reinforcement learning.
Extended analysis of Q-learning's efficiency, matching optimal regret.
The paper explores stability and generalization of deep GCNs.
Paper proves stacking ensembling is effective and proposes a new family of stacked generalizations.
We show that model compression can improve the population risk of a pre-trained model, by studying the tradeoff between the decrease in the generalization error and the increase in the empirical risk with model compression. We first prove that model compression reduces an information-theoretic bound on the generalizati…
Unpublished results of S Straus and W Browder state that two notions of homotopy equivalence for manifolds with smooth group actions - isovariant and equivariant - often coincide under a condition called the Gap Hypothesis; the proofs use deep results in geometric topology. This paper analyzes the difference between th…
This paper investigates the theory of robustness against adversarial attacks. It focuses on the family of randomization techniques that consist in injecting noise in the network at inference time. These techniques have proven effective in many contexts, but lack theoretical arguments. We close this gap by presenting a …
Discrimination-aware classification is receiving an increasing attention in data science fields. The pre-process methods for constructing a discrimination-free classifier first remove discrimination from the training data, and then learn the classifier from the cleaned data. However, they lack a theoretical guarantee f…
Theoretical study shows AI models can recover from contaminated training data.
Bayesian MAML outperforms MAML in meta learning tasks with theoretical guarantees.
Study improves theoretical understanding of Bayesian deep learning for classification tasks.