The paper tackles online learning with two types of losses and shows it's impossible without certain assumptions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Fuses ITRs for primary and secondary outcomes to minimize harm.
Aux-NAS uses auxiliary labels to improve primary task performance without extra inference cost.
DPC uses physics and neural nets to solve SDEs.
Fine-tuning LLMs improves capability but harms safety, study finds.
Learning with auxiliary tasks can improve the ability of a primary task to generalise. However, this comes at the cost of manually labelling auxiliary data. We propose a new method which automatically learns appropriate labels for an auxiliary task, such that any supervised learning task can be improved without requiri…
Hill-ADAM optimizes loss landscapes by exploring state space deterministically.
This research analyzes the consistency of convex and nonconvex surrogate losses for adversarially robust classification.
We develop a model for contagion in reinsurance networks by which primary insurers' losses are spread through the network. Our model handles general reinsurance contracts, such as typical excess of loss contracts. We show that simpler models existing in the literature--namely proportional reinsurance--greatly underesti…
Bayesian neural networks improved with scalable approximate inference.
Paper examines risk measure expansions under FGM dependence, improving accuracy at extreme levels.
Informer model with GMADL loss outperforms benchmarks in high frequency Bitcoin trading.
Statistical inference is considered for variables of interest, called primary variables, when auxiliary variables are observed along with the primary variables. We consider the setting of incomplete data analysis, where some primary variables are not observed. Utilizing a parametric model of joint distribution of prima…
A significant advance in accelerating neural network training has been the development of normalization methods, permitting the training of deep models both faster and with better accuracy. These advances come with practical challenges: for instance, batch normalization ties the prediction of individual examples with o…
Improved similarity search in embeddings using InfoNCE loss.
Proposes a probabilistic framework for smart contract risk quantification.
Over the past few years many research efforts have been devoted to the field of affect analysis. Various approaches have been proposed for: i) discrete emotion recognition in terms of the primary facial expressions; ii) emotion analysis in terms of facial Action Units (AUs), assuming a fixed expression intensity; iii) …
Study blockchain's impact on primary financial market challenges.
SGD updates align with a low-rank subspace but do not lead to further loss reduction.
A new geometric model for V1 hypercolumns combines symplectic and spherical models.
We address primary decomposition conjectures for knot concordance groups, which predict direct sum decompositions into primary parts. We show that the smooth concordance group of topologically slice knots has a large subgroup for which the conjectures are true and there are infinitely many primary parts each of which h…
Augmenting a neural network with memory that can grow without growing the number of trained parameters is a recent powerful concept with many exciting applications. We propose a design of memory augmented neural networks (MANNs) called Labeled Memory Networks (LMNs) suited for tasks requiring online adaptation in class…
Classifies Real primary Hopf surfaces and their associated groups.
We introduce new forecast encompassing tests for the risk measure Expected Shortfall (ES). The ES currently receives much attention through its introduction into the Basel III Accords, which stipulate its use as the primary market risk measure for the international banking regulation. We utilize joint loss functions fo…
In most machine learning applications, classification accuracy is not the primary metric of interest. Binary classifiers which face class imbalance are often evaluated by the score, area under the precision-recall curve, Precision at K, and more. The maximization of many of these metrics can be expressed as a con…
Improved online learning with time-varying constraints for complex domains.
In machine learning, statistics, econometrics and statistical physics, cross-validation (CV) is used asa standard approach in quantifying the generalisation performance of a statistical model. A directapplication of CV in time-series leads to the loss of serial correlations, a requirement of preserving anynon-stationar…
Flaky performance found in GNN SSL on RDBs, leading to worse linear evaluation.
Super Learner combines dynamic predictions from various models to improve survival estimates.
A new method detects concept drift without true labels.
Bayesian framework for policy learning in decision problems.
New loss function reduces outage probability in ML-assisted resource allocation.
Translation-based embedding models have gained significant attention in link prediction tasks for knowledge graphs. TransE is the primary model among translation-based embeddings and is well-known for its low complexity and high efficiency. Therefore, most of the earlier works have modified the score function of the Tr…
Understanding the morphological changes of primary neuronal cells induced by chemical compounds is essential for drug discovery. Using the data from a single high-throughput imaging assay, a classification model for predicting the biological activity of candidate compounds was introduced. The image recognition model wh…
Learning with a primary objective, such as softmax cross entropy for classification and sequence generation, has been the norm for training deep neural networks for years. Although being a widely-adopted approach, using cross entropy as the primary objective exploits mostly the information from the ground-truth class f…
We address feature interpretation and reproducibility issues in dense nets, proposing a modified loss function.
The paper introduces PD learning to improve deep learning theory.
We study holomorphic locally homogeneous geometric structures modelled on line bundles over the projective line. We classify these structures on primary Hopf surfaces. We write out the developing map and holonomy morphism of each of these structures explicitly on each primary Hopf surface.
Paper proposes a method to speed up discrete diffusion models by distilling many steps into few.
Large data collections required for the training of neural networks often contain sensitive information such as the medical histories of patients, and the privacy of the training data must be preserved. In this paper, we introduce a dropout technique that provides an elegant Bayesian interpretation to dropout, and show…
In this study, we generalize double tangent bundles to double jet bundles. We present a secondary vector bundle structure on a 1-jet of a vector bundle. We show that 1-jet of a vector bundle carries two vector bundle structures, namely primary and secondary structures. We also show that the manifold charts induced by p…
Improved PAC-Bayesian bounds by considering example difficulty.
Paper proposes a deep hedging method for Bermudan swaptions to manage residual profit and loss.
Automorphisms of Kodaira surfaces are shown to be affine transformations.
For all n > 0 there is a homomorphism from the smooth concordance group of knots in dimension 2n + 1 to an algebraically defined group called the rational algebraic concordance group. This algebraic concordance group splits as a direct sum of groups indexed by polynomials. For n > 1 the homomorphism is injective. This …
In several natural language tasks, labeled sequences are available in separate domains (say, languages), but the goal is to label sequences with mixed domain (such as code-switched text). Or, we may have available models for labeling whole passages (say, with sentiments), which we would like to exploit toward better po…
Randomly chosen primary hidden units and derived secondary units reduce neural network complexity.
The Softmax function is used in the final layer of nearly all existing sequence-to-sequence models for language generation. However, it is usually the slowest layer to compute which limits the vocabulary size to a subset of most frequent types; and it has a large memory footprint. We propose a general technique for rep…