Paper analyzes ECE bias and provides bounds for its estimation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We provide the first information theoretic tight analysis for inference of latent community structure given a sparse graph along with high dimensional node covariates, correlated with the same latent communities. Our work bridges recent theoretical breakthroughs in the detection of latent community structure without no…
New analysis improves generalization bounds for meta-learning.
We study the risk performance of distributed learning for the regularization empirical risk minimization with fast convergence rate, substantially improving the error analysis of the existing divide-and-conquer based distributed learning. An interesting theoretical finding is that the larger the diversity of each local…
Addresses theoretical and practical aspects of Gaussian differential privacy.
Signature kernel handles sequential data with theoretical and practical advantages.
This work analyzes Fréchet regression using comparison geometry, providing theoretical and practical insights.
Theoretical analysis improves understanding of Deep Q-Learning's behavior.
This paper provides theoretical guarantees for SPCA using the Elastic Net.
New method explains sensitivity of test data uncertainty in Bayesian inference.
We study the problem of supervised linear dimensionality reduction, taking an information-theoretic viewpoint. The linear projection matrix is designed by maximizing the mutual information between the projected signal and the class label (based on a Shannon entropy measure). By harnessing a recent theoretical result on…
Proposes a new CG interpretation of neural networks for better theoretical analysis.
Based on the new type of random walk process called the Potentials of Unbalanced Complex Kinetics (PUCK) model, we theoretically show that the price diffusion in large scales is amplified 2/(2 + b) times, where b is the coefficient of quadratic term of the potential. In short time scales the price diffusion depends on …
The information-theoretic analysis by Russo and Van Roy (2014) in combination with minimax duality has proved a powerful tool for the analysis of online learning algorithms in full and partial information settings. In most applications there is a tantalising similarity to the classical analysis based on mirror descent.…
Extended analysis of Q-learning's efficiency, matching optimal regret.
A game-theoretic framework identifies influential hyperparameters for neural networks.
We integrate information-theoretic concepts into the design and analysis of optimistic algorithms and Thompson sampling. By making a connection between information-theoretic quantities and confidence bounds, we obtain results that relate the per-period performance of the agent with its information gain about the enviro…
We present an operator-free, measure-theoretic approach to the conditional mean embedding (CME) as a random variable taking values in a reproducing kernel Hilbert space. While the kernel mean embedding of unconditional distributions has been defined rigorously, the existing operator-based approach of the conditional ve…
Theoretical analysis of cross-entropy loss functions and their robustness.
The paper analyzes set-to-set matching with neural networks, focusing on theoretical generalization.
This paper suggests a learning-theoretic perspective on how synaptic plasticity benefits global brain functioning. We introduce a model, the selectron, that (i) arises as the fast time constant limit of leaky integrate-and-fire neurons equipped with spiking timing dependent plasticity (STDP) and (ii) is amenable to the…
Stochastic Gradient Descent (SGD) has become popular for solving large scale supervised machine learning optimization problems such as SVM, due to their strong theoretical guarantees. While the closely related Dual Coordinate Ascent (DCA) method has been implemented in various software packages, it has so far lacked go…
High dimensional data analysis is known to be as a challenging problem. In this article, we give a theoretical analysis of high dimensional classification of Gaussian data which relies on a geometrical analysis of the error measure. It links a problem of classification with a problem of nonparametric regression. We giv…
The paper analyzes LIME for text data and provides theoretical guarantees.
Paper studies the theoretical equivalence between implicit and explicit neural networks in high dimensions.
Empirical analysis serves as an important complement to theoretical analysis for studying practical Bayesian optimization. Often empirical insights expose strengths and weaknesses inaccessible to theoretical analysis. We define two metrics for comparing the performance of Bayesian optimization methods and propose a ran…
Unified analysis of kernel-based methods under covariate shift.
RER improves sample complexity by updating in reverse order.
We analyze oversquashing in topological message-passing using relational structures.
This paper analyzes OGDA and EG methods for nonconvex minimax problems.
Max-Pooling operations are a core component of deep learning architectures. In particular, they are part of most convolutional architectures used in machine vision, since pooling is a natural approach to pattern detection problems. However, these architectures are not well understood from a theoretical perspective. For…
For multiple multivariate data sets, we derive conditions under which Generalized Canonical Correlation Analysis (GCCA) improves classification performance of the projected datasets, compared to standard Canonical Correlation Analysis (CCA) using only two data sets. We illustrate our theoretical results with simulation…
Properties of low-variability periods in the time series are analysed. The theoretical approach is used to show the relationship between the multi-scaling of low-variability periods and multi-affinity of the time series. It is shown that this technically simple method is capable of reveling more details about time-seri…
NCDEs improve predictions for irregular time series data.
Improved MTL-LSSVM for better multi-task learning performance.
Kernel methods have been among the most popular techniques in machine learning, where learning tasks are solved using the property of reproducing kernel Hilbert space (RKHS). In this paper, we propose a novel data analysis framework with reproducing kernel Hilbert -module (RKHM), which is another generalization of…
Proposes a new theoretical framework for PbRL that requires less human feedback.
New method shows spectral clustering is consistent with theoretical guarantees.
Theoretical analysis of deep neural networks for time series data.
Theoretical analysis shows MDMs can be efficient but not for all metrics.
Paper introduces a new generative learning model using Schrödinger bridge diffusion in latent space.
Paper introduces a new histogram estimator for nonparametric density estimation that improves performance.
DFM models are analyzed for generating distributions with provable convergence.
The paper analyzes how adversarial attacks affect sparse regression models.
While crowdsourcing has become an important means to label data, there is great interest in estimating the ground truth from unreliable labels produced by crowdworkers. The Dawid and Skene (DS) model is one of the most well-known models in the study of crowdsourcing. Despite its practical popularity, theoretical error …
This work analyzes SGGMs, offering convergence insights and practical design tips.
Machine learning is used more and more often for sensitive applications, sometimes replacing humans in critical decision-making processes. As such, interpretability of these algorithms is a pressing need. One popular algorithm to provide interpretability is LIME (Local Interpretable Model-Agnostic Explanation). In this…
The paper explores stability and generalization of deep GCNs.