This paper tackles noise in raw datasets to improve representation learning efficiency.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Cosine similarity can force points to grow in magnitude, causing convergence issues.
Two things seem to be indisputable in the contemporary deep learning discourse: 1. The categorical cross-entropy loss after softmax activation is the method of choice for classification. 2. Training a CNN classifier from scratch on small datasets does not work well. In contrast to this, we show that the cosine loss fun…
Person recognition aims at recognizing the same identity across time and space with complicated scenes and similar appearance. In this paper, we propose a novel method to address this task by training a network to obtain robust and representative features. The intuition is that we directly compare and optimize the cosi…
The main purpose of incremental learning is to learn new knowledge while not forgetting the knowledge which have been learned before. At present, the main challenge in this area is the catastrophe forgetting, namely the network will lose their performance in the old tasks after training for new tasks. In this paper, we…
Paper explains contrastive learning using cosine similarity and proposes mitigations for batch size effects.
SPEQ improves quantized neural networks by stochastic precision sharing and cosine similarity loss.
The paper presents a multi-power law for predicting loss curves across different learning rate schedules.
Modified cosine distance improves similarity performance in data with variance and correlation.
Improved MoE performance through perturbing cosine router.
Beta-SOD detects and corrects noisy object re-identification using cosine similarity and Beta mixtures.
Traditionally, multi-layer neural networks use dot product between the output vector of previous layer and the incoming weight vector as the input to activation function. The result of dot product is unbounded, thus increases the risk of large variance. Large variance of neuron makes the model sensitive to the change o…
Researchers establish bounds and continuity of decomposed Möbius energies using cosine formula.
Feature normalization prevents collapse in non-contrastive learning dynamics.
New insights show embedding lengths correlate with semantic properties.
Cosine schedule is optimal for discrete diffusion models.
One approach to deal with the statistical inefficiency of neural networks is to rely on auxiliary losses that help to build useful representations. However, it is not always trivial to know if an auxiliary task will be helpful for the main task and when it could start hurting. We propose to use the cosine similarity be…
Derives hyperbolic laws of cosines and sines with fermionic corrections.
This work analyzes when contrastive models are close to PCA or kernel methods.
The study introduces anytime learning schedules for large language models without fixed horizons.
We study the rigidity of polyhedral surfaces using variational principle. The action functionals are derived from the cosine laws. The main focus of this paper is on the cosine law for a non-triangular region bounded by three possibly disjoint geodesics. Several of these cosine laws were first discovered and used by Fe…
New method for European option pricing faster and more robust.
Study evaluates relevance metrics for similarity-based model explanations.
Efficiently applies NTK to large-scale datasets using random features.
New method extracts brain age from MRI sequences over time.
Feedback alignment methods need to be evaluated for accuracy and gradient cosine similarity.
Mode connectivity is a recently introduced frame- work that empirically establishes the connected- ness of minima by finding a high accuracy curve between two independently trained models. To investigate the limits of this setup, we examine the efficacy of this technique in extreme cases where the input models are trai…
Two binary Sine Cosine Algorithms improve feature selection in medical datasets.
CWGD measures gradient diversity weighted by curvature, improving SGD convergence.
T-PSDA improves speaker recognition accuracy on toroidal submanifolds.
We do further investigation in a certain cosine function defined for smooth Minkowski spaces. We prove that such function is symmetric if and only if the referred space is Euclidean, and also that it can be given in terms of the Gateaux derivative of the norm. As an application we use it to study the ratio between the …
The spherical Radon transform on the unit sphere can be regarded as a member of the analytic family of suitably normalized generalized cosine transforms. We derive new formulas for these transforms and apply them to study classes of intersections bodies in convex geometry.
Study compares metric learning loss functions for speaker verification.
Ensembling word embeddings to improve distributed word representations has shown good success for natural language processing tasks in recent years. These approaches either carry out straightforward mathematical operations over a set of vectors or use unsupervised learning to find a lower-dimensional representation. Th…
The convergence rate and final performance of common deep learning models have significantly benefited from heuristics such as learning rate schedules, knowledge distillation, skip connections, and normalization layers. In the absence of theoretical underpinnings, controlled experiments aimed at explaining these strate…
WSD schedule improves model training efficiency by adapting learning rates dynamically.
We extend the Fourier cosine method to discrete probability distributions, achieving faster convergence rates.
In this short article, we extend the cosine formula for the Möbius energy to generalized O'Hara energies. The newly derived formula gives us a condition for which the right circle minimizes the energy under the length-constraint. Furthermore, it shows us how far the energy is from the Möbius invariant property.
We present a method for training multi-label, massively multi-class image classification models, that is faster and more accurate than supervision via a sigmoid cross-entropy loss (logistic regression). Our method consists in embedding high-dimensional sparse labels onto a lower-dimensional dense sphere of unit-normed …
The standard loss function used to train neural network classifiers, categorical cross-entropy (CCE), seeks to maximize accuracy on the training data; building useful representations is not a necessary byproduct of this objective. In this work, we propose clustering-oriented representation learning (COREL) as an altern…
The study compares Euclidean and cosine distances in medical drug prescription prediction.
A large body of research into semantic textual similarity has focused on constructing state-of-the-art embeddings using sophisticated modelling, careful choice of learning signals and many clever tricks. By contrast, little attention has been devoted to similarity measures between these embeddings, with cosine similari…
Sensors which use electromagnetic induction (EMI) to excite a response in conducting bodies have long been investigated for subsurface explosive hazard detection. In particular, EMI sensors have been used to discriminate between different types of objects, and to detect objects with low metal content. One successful, p…
Proves properties of periodic billiard orbits in ellipses.
Improved text classification performance through conformal transformations of kernels.
The COS method proposed in Fang and Oosterlee (2008), although highly efficient, may lack robustness for a number of cases. In this paper, we present a Stable pricing of call options based on Fourier cosine series expansion. The Stability of the pricing methods is demonstrated by error analysis, as well as by a series …
Improved barrier option pricing in Heston model using COS-BEM method.
Paper introduces a new method for efficient portfolio risk quantification.