The paper presents a multi-power law for predicting loss curves across different learning rate schedules.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
LLMs learn peaked distributions slowly due to power-law losses.
Power of network tests degrades when vertices are misaligned.
We develop a novel method, called PoWER-BERT, for improving the inference time of the popular BERT model, while maintaining the accuracy. It works by: a) exploiting redundancy pertaining to word-vectors (intermediate encoder outputs) and eliminating the redundant vectors. b) determining which word-vectors to eliminate …
New findings show neural network training loss follows a power law over time.
Novel power transform unifies various mathematical functions.
Electronic power inverters are capable of quickly delivering reactive power to maintain customer voltages within operating tolerances and to reduce system losses in distribution grids. This paper proposes a systematic and data-driven approach to determine reactive power inverter output as a function of local measuremen…
Explains deep learning models and their geometric properties.
The paper predicts loss scaling across different datasets and compute scales.
Boosts change-point detection power with optimal sub-sampling.
Federated learning calibrates insurance indices from renewable energy producers' data.
The paper analyzes a private likelihood-ratio test for frequency tables under differential privacy constraints.
Correntropy is a second order statistical measure in kernel space, which has been successfully applied in robust learning and signal processing. In this paper, we define a nonsecond order statistical measure in kernel space, called the kernel mean-p power error (KMPE), including the correntropic loss (CLoss) as a speci…
This paper reformulates systemic risk measures and finds new properties and estimators.
We discover scaling laws for kernel regression loss under various learning rate schedules.
A new AMM design reduces impermanent loss and retains more liquidity.
Ensemble techniques are powerful approaches that combine several weak learners to build a stronger one. As a meta-learning framework, ensemble techniques can easily be applied to many machine learning methods. Inspired by ensemble techniques, in this paper we propose an ensemble loss functions applied to a simple regre…
Deep neural networks (DNNs) have great expressive power, which can even memorize samples with wrong labels. It is vitally important to reiterate robustness and generalization in DNNs against label corruption. To this end, this paper studies the 0-1 loss, which has a monotonic relationship with an empirical adversary (r…
Estimation of the operational risk capital under the Loss Distribution Approach requires evaluation of aggregate (compound) loss distributions which is one of the classic problems in risk theory. Closed-form solutions are not available for the distributions typically used in operational risk. However with modern comput…
We consider the problem of option hedging in a market with proportional transaction costs. Since super-replication is very costly in such markets, we replace perfect hedging with an expected loss constraint. Asymptotic analysis for small transactions is used to obtain a tractable model. A general expansion theory is de…
Paper introduces new loss functions for Siamese networks using FDA.
The one-bit quantization is implemented by one single comparator that operates at low power and a high rate. Hence one-bit compressive sensing (1bit-CS) becomes attractive in signal processing. When measurements are corrupted by noise during signal acquisition and transmission, 1bit-CS is usually modeled as minimizing …
A meticulous assessment of the risk of impacts associated with extreme wind events is of great necessity for populations, civil authorities as well as the insurance industry. Using the concept of spatial risk measure and related set of axioms introduced by Koch (2017, 2019), we quantify the risk of losses due to extrem…
The paper sets lower bounds for adversarial robustness in multiclass classification.
Domain adaptation provides a powerful set of model training techniques given domain-specific training data and supplemental data with unknown relevance. The techniques are useful when users need to develop models with data from varying sources, of varying quality, or from different time ranges. We build CrossTrainer, a…
This paper introduces Stochastic Gradient Langevin Boosting (SGLB) - a powerful and efficient machine learning framework that may deal with a wide range of loss functions and has provable generalization guarantees. The method is based on a special form of the Langevin diffusion equation specifically designed for gradie…
Value functions struggle to represent transition dynamics, impacting statistical efficiency.
Operational risk is the risk relative to monetary losses caused by failures of bank internal processes due to heterogeneous causes. A dynamical model including both spontaneous generation of losses and generation via interactions between different processes is presented; the efforts made by the bank to avoid the occurr…
We stabilize MI-based losses by adding a regularization term, improving their performance and stability.
Neural network predicts daily power consumption with high accuracy.
The classical asymptotic theory for parametric -estimators guarantees that, in the limit of infinite sample size, the excess risk has a chi-square type distribution, even in the misspecified case. We demonstrate how self-concordance of the loss allows to characterize the critical sample size sufficient to guarantee …
Anonymization reduces economic signal extraction from financial texts.
We propose a dynamical model for the estimation of Operational Risk in banking institutions. Operational Risk is the risk that a financial loss occurs as the result of failed processes. Examples of operational losses are the ones generated by internal frauds, human errors or failed transactions. In order to encompass t…
This work improves diffusion models by estimating the optimal loss value for better training diagnostics.
Paper explores challenges in training PINNs and loss landscape effects.
Loss-calibrated EP improves Bayesian decision-making by focusing on utility-sensitive posterior approximations.
We study the problem of nonparametric dependence detection. Many existing methods may suffer severe power loss due to non-uniform consistency, which we illustrate with a paradox. To avoid such power loss, we approach the nonparametric test of independence through the new framework of binary expansion statistics (BEStat…
Optimal exit strategies of CPT gamblers in unfair gambles
Method detects and locates eavesdropping in optical links.
Generative Adversarial Networks (GANs) were intuitively and attractively explained under the perspective of game theory, wherein two involving parties are a discriminator and a generator. In this game, the task of the discriminator is to discriminate the real and generated (i.e., fake) data, whilst the task of the gene…
SGD's escape rate depends on log loss barrier, not linear loss barrier.
We present a powerful new loss function and training scheme for learning binary hash codes with any differentiable model and similarity function. Our loss function improves over prior methods by using log likelihood loss on top of an accurate approximation for the probability that two inputs fall within a Hamming dista…
Study compares deep learning stock trading strategies in adverse market conditions.
Ensemble techniques are powerful approaches that combine several weak learners to build a stronger one. As a meta learning framework, ensemble techniques can easily be applied to many machine learning techniques. In this paper we propose a neural network extended with an ensemble loss function for text classification. …
This paper develops a novel graph convolutional network (GCN) framework for fault location in power distribution networks. The proposed approach integrates multiple measurements at different buses while taking system topology into account. The effectiveness of the GCN model is corroborated by the IEEE 123 bus benchmark…
Locational Marginal Pricing aims to free UK power markets.
We propose a reinforcement learning (RL) based closed loop power control algorithm for the downlink of the voice over LTE (VoLTE) radio bearer for an indoor environment served by small cells. The main contributions of our paper are to 1) use RL to solve performance tuning problems in an indoor cellular network for voic…
OptCS optimizes model selection after conformal inference, controlling FDR and power loss.