We prove that square integrable holomorphic functions (with respect to a plurisubharmonic weight) can be extended in a square integrable manner from certain singular hypersurfaces (which include uniformly flat, normal crossing divisors) to entire functions in affine space. This provides evidence for a conjecture regard…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The CAP slope is Bayes' theorem in cumulative coordinates, unlocking the weight of evidence, Somers' D, and Gini coefficient.
We provide theoretical and empirical evidence that using tighter evidence lower bounds (ELBOs) can be detrimental to the process of learning an inference network by reducing the signal-to-noise ratio of the gradient estimator. Our results call into question common implicit assumptions that tighter ELBOs are better vari…
The standard interpretation of importance-weighted autoencoders is that they maximize a tighter lower bound on the marginal likelihood than the standard evidence lower bound. We give an alternate interpretation of this procedure: that it optimizes the standard variational lower bound, but using a more complex distribut…
Neural networks with learned biases can approximate any function.
Bayesian neural networks update beliefs with soft evidence, improving accuracy and calibration.
Unified framework for measuring concentration in weighted networks considering both weight distributions and network structure.
OPAA estimates probability densities using functional analysis.
Stochastic Bayesian Neural Network improves scalability and performance.
Improved classification model for high-cardinality categorical predictors.
Interpretability is an elusive but highly sought-after characteristic of modern machine learning methods. Recent work has focused on interpretability via , which justify individual model predictions. In this work, we take a step towards reconciling machine explanations with those that humans prod…
We introduce a Bayesian solution for the problem in forensic speaker recognition, where there may be very little background material for estimating score calibration parameters. We work within the Bayesian paradigm of evidence reporting and develop a principled probabilistic treatment of the problem, which results in a…
The Penrose theorem and Hawking's topology theorem are extended to weighted spacetimes.
Bayesian PINNs optimize loss weights for PDEs and data.
Efficiently recovers network community structure from clients' small subgraphs.
Bayesian neural networks show good correlation between out-of-sample performance and Bayesian evidence.
Expectation Maximization (EM) is among the most popular algorithms for maximum likelihood estimation, but it is generally only guaranteed to find its stationary points of the log-likelihood objective. The goal of this article is to present theoretical and empirical evidence that over-parameterization can help EM avoid …
Reweighting improves risk bounds in certain data regions.
New algorithm finds unbiased subnetworks in biased datasets.
Estimation of importance sampling weights for off-policy evaluation of contextual bandits often results in imbalance - a mismatch between the desired and the actual distribution of state-action pairs after weighting. In this work we present balanced off-policy evaluation (B-OPE), a generic method for estimating weights…
The paper analyzes constrained optimal portfolios in high dimensions using novel statistical learning techniques.
Weight Decay induces low-rank weight matrices in neural networks, improving generalization.
Develops exact and invariant study-based decompositions for network meta-analysis.
Martingale Doppelgänger-Eval benchmarks VLMs on candlestick evidence vs. trend extrapolation
A new method trains deep networks by separating weight locations from values.
This paper evaluates heuristics and hyperparameters in weight-sharing NAS methods.
Optimizes sliding window approach for tracking Gaussian densities.
Principal component analysis (PCA) is a useful tool when trying to construct factor models from historical asset returns. For the implied volatilities of U.S. equities there is a PCA-based model with a principal eigenportfolio whose return time series lies close to that of an overarching market factor. The authors show…
Much of the focus in machine learning research is placed in creating new architectures and optimization methods, but the overall loss function is seldom questioned. This paper interprets machine learning from a multi-objective optimization perspective, showing the limitations of the default linear combination of loss f…
Neural networks' weights don't converge to stationary points but training loss stabilizes.
Adversarial weighting improves regression task adaptation.
We investigate robustness of deep feed-forward neural networks when input data are subject to random uncertainties. More specifically, we consider regularization of the network by its Lipschitz constant and emphasize its role. We highlight the fact that this regularization is not only a way to control the magnitude of …
Boltzmann machines are powerful distributions that have been shown to be an effective prior over binary latent variables in variational autoencoders (VAEs). However, previous methods for training discrete VAEs have used the evidence lower bound and not the tighter importance-weighted bound. We propose two approaches fo…
Empirical evidence suggests that neural networks with ReLU activations generalize better with over-parameterization. However, there is currently no theoretical analysis that explains this observation. In this work, we provide theoretical and empirical evidence that, in certain cases, overparameterized convolutional net…
We examine how recently documented, fundamental phenomena in deep learning models subject to pruning are affected by changes in the pruning procedure. Specifically, we analyze differences in the connectivity structure and learning dynamics of pruned models found through a set of common iterative pruning techniques, to …
New loss function restores importance weighting in overparameterized models.
As a model problem for clustering, we consider the densest k-disjoint-clique problem of partitioning a weighted complete graph into k disjoint subgraphs such that the sum of the densities of these subgraphs is maximized. We establish that such subgraphs can be recovered from the solution of a particular semidefinite re…
In this paper we explore the role of sample mean in building a neural network for classification. This role is surprisingly extensive and includes: direct computation of weights without training, performance monitoring for samples without known classification, and self-training for unlabeled data. Experimental computat…
Recently theoretical guarantees have been obtained for matrix completion in the non-uniform sampling regime. In particular, if the sampling distribution aligns with the underlying matrix's leverage scores, then with high probability nuclear norm minimization will exactly recover the low rank matrix. In this article, we…
Algorithm reduces audit costs by identifying best service configurations from biased textual evidence.
Dual-edge spatial Jacobian image graph for interpretable diabetic retinopathy grading
Diffusion models optimize objectives similar to ELBO with Gaussian noise augmentation.
Membership in the Russell 1000 and 2000 Indices is based on a ranking of market capitalization in May. Each index is separately value weighted such that firms just inside the Russell 2000 are comparable in size to firms just outside (i.e. at the bottom of the Russell 1000) but have much higher index weights. These feat…
This paper proposes a continuous timing strategy for growth vs. defensive style allocation.
Our work presents extensive empirical evidence that layer rotation, i.e. the evolution across training of the cosine distance between each layer's weight vector and its initialization, constitutes an impressively consistent indicator of generalization performance. In particular, larger cosine distances between final an…
This paper compares gradient estimators in importance-weighted VI and justifies the superiority of DREP over REP.
Bayesian neural networks are compressed using feature and weight pruning based on posterior inclusion probabilities.
Variational Bayesian neural networks (BNNs) perform variational inference over weights, but it is difficult to specify meaningful priors and approximate posteriors in a high-dimensional weight space. We introduce functional variational Bayesian neural networks (fBNNs), which maximize an Evidence Lower BOund (ELBO) defi…