Optimal feature transfer identified through bias-variance analysis.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study tight offline learning bounds for linear MDPs using variance information.
We focus on mean-variance hedging problem for models whose asset price follows an exponential additive process. Some representations of mean-variance hedging strategies for jump type models have already been suggested, but none is suited to develop numerical methods of the values of strategies for any given time up to …
Language models allocate information storage, not collapsing into uniform representations.
Societal bias towards certain communities is a big problem that affects a lot of machine learning systems. This work aims at addressing the racial bias present in many modern gender recognition systems. We learn race invariant representations of human faces with an adversarially trained autoencoder model. We show that …
In this paper we propose an efficient variance reduction approach for additive functionals of Markov chains relying on a novel discrete time martingale representation. Our approach is fully non-asymptotic and does not require the knowledge of the stationary distribution (and even any type of ergodicity) or specific str…
The paper extends a variance gamma model to quadratic functions, reducing arbitrage and computational costs.
In a financial market model, we consider the variance-optimal semi-static hedging of a given contingent claim, a generalization of the classic variance-optimal hedging. To obtain a tractable formula for the expected squared hedging error and the optimal hedging strategy, we use a Fourier approach in a general multidime…
Neural--LinearUCB improves regret in neural contextual bandits.
Optimal estimator derived for partially observable LTI systems.
NP-PROV separates mean and variance spaces to improve function uncertainty.
In the paper, a mean-square minimization problem under terminal wealth constraint with partial observations is studied. The problem is naturally connected to the mean-variance hedging problem under incomplete information. A new approach to solving this problem is proposed. The paper provides a solution when the underly…
We analyze the variance of Fisher information estimators in deep learning models.
Unified method for MMD variance estimation improves accuracy and computational efficiency.
Statistical characteristics of deep network representations, such as sparsity and correlation, are known to be relevant to the performance and interpretability of deep learning. When a statistical characteristic is desired, often an adequate regularizer can be designed and applied during the training phase. Typically, …
New method estimates shape distance in neural representations with limited data.
Develops a new self-supervised learning method combining contrastive and non-contrastive approaches.
The paper tackles model collapse in GPLVMs by improving kernel flexibility and projection variance.
We give a new proof of the representation of implied volatility as a time-average of weighted expectations of local or stochastic volatility. With this proof we clarify the question of existence of 'forward implied variance' in the original derivation of Gatheral, who introduced this representation in his book 'The Vol…
Federated learning is a method of training models on private data distributed over multiple devices. To keep device data private, the global model is trained by only communicating parameters and updates which poses scalability challenges for large models. To this end, we propose a new federated learning algorithm that …
Hierarchical probabilistic models are able to use a large number of parameters to create a model with a high representation power. However, it is well known that increasing the number of parameters also increases the complexity of the model which leads to a bias-variance trade-off. Although it is a classical problem, t…
Bayesian deep neural networks converge to processes with α-stable marginals under infinite variance weights.
We propose a metric, Layer Saturation, defined as the proportion of the number of eigenvalues needed to explain 99% of the variance of the latent representations, for analyzing the learned representations of neural network layers. Saturation is based on spectral analysis and can be computed efficiently, making live ana…
In this paper we study mean-variance hedging under the G-expectation framework. Our analysis is carried out by exploiting the G-martingale representation theorem and the related probabilistic tools, in a contin- uous financial market with two assets, where the discounted risky one is modeled as a symmetric G-martingale…
This work explains the structural origins of attention sinks in LLMs.
In this paper we analyze a dynamic recursive extension of the (static) notion of a deviation measure and its properties. We study distribution invariant deviation measures and show that the only dynamic deviation measure which is law invariant and recursive is the variance. We also solve the problem of optimal risk-sha…
WMPG reduces policy gradient variance using world models.
New method proves non-contrastive self-supervised learning learns useful features.
VCAE improves autoencoder quality on MNIST and CelebA.
We consider hedging of a contingent claim by a 'semi-static' strategy composed of a dynamic position in one asset and static (buy-and-hold) positions in other assets. We give general representations of the optimal strategy and the hedging error under the criterion of variance-optimality and provide tractable formulas u…
We consider a Poisson process on a measurable space $(\BY,\mathcal{Y})$ equipped with a partial ordering, assumed to be strict almost everwhwere with respect to the intensity measure of . We give a Clark-Ocone type formula providing an explicit representation of square integrable martingales (defined with re…
To study how mental object representations are related to behavior, we estimated sparse, non-negative representations of objects using human behavioral judgments on images representative of 1,854 object categories. These representations predicted a latent similarity structure between objects, which captured most of the…
Deep Learning has revolutionized vision via convolutional neural networks (CNNs) and natural language processing via recurrent neural networks (RNNs). However, success stories of Deep Learning with standard feed-forward neural networks (FNNs) are rare. FNNs that perform well are typically shallow and, therefore cannot …
Paper explains contrastive learning using cosine similarity and proposes mitigations for batch size effects.
Flexible framework for modeling predictive distributions of time series
In this paper, we argue that, once the costs of maintaining the hedging portfolio are properly taken into account, semi-static portfolios should more properly be thought of as separate classes of derivatives, with non-trivial, model-dependent payoff structures. We derive new integral representations for payoffs of exot…
Optimizes option portfolios for skewed-t returns using VaR and variance measures.
This study introduces axioms to assess regression uncertainty measures.
Estimating and optimizing Mutual Information (MI) is core to many problems in machine learning; however, bounding MI in high dimensions is challenging. To establish tractable and scalable objectives, recent work has turned to variational bounds parameterized by neural networks, but the relationships and tradeoffs betwe…
Probabilistic principal component analysis (PPCA) seeks a low dimensional representation of a data set in the presence of independent spherical Gaussian noise, Sigma = (sigma^2)*I. The maximum likelihood solution for the model is an eigenvalue problem on the sample covariance matrix. In this paper we consider the situa…
We consider a market impact game for risk-averse agents that are competing in a market model with linear transient price impact and additional transaction costs. For both finite and infinite time horizons, the agents aim to minimize a mean-variance functional of their costs or to maximize the expected exponential u…
Prediction-powered causal inference achieves smaller asymptotic variance than traditional methods.
In this paper, we study the mean-variance portfolio selection problem under partial information with drift uncertainty. First we show that the market model is complete even in this case while the information is not complete and the drift is uncertain. Then, the optimal strategy based on partial information is derived, …
We determine the variance-optimal hedge when the logarithm of the underlying price follows a process with stationary independent increments in discrete or continuous time. Although the general solution to this problem is known as backward recursion or backward stochastic differential equation, we show that for this cla…
Proposes NUV priors for half-space and box constraints.
SupSiam and SupBYOL improve supervised representation learning with ANCL.
This paper proposes a multi-layer neural network structure for few-shot image recognition of novel categories. The proposed multi-layer neural network architecture encodes transferable knowledge extracted from a large annotated dataset of base categories. This architecture is then applied to novel categories containing…
Kernel VICReg improves SSL in RKHS, capturing nonlinear structures.