Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

105209314418 · Jun 202019922001200920182026
48 results for vectorization loss

Paper explores vector embeddings, distributional hypothesis, and PIP loss for natural language processing.

problem Understanding the effect of dimensionality on vector embeddings and their functionality.
method Formulates a theoretical framework, proposes PIP loss, and reveals bias-variance trade-off.
result Discovers robustness and forward stability of vector embeddings, answers dimensionality selection problem.

The paper explores MAE as a loss function for DNN vector-to-vector regression, proving its advantages over MSE.

problem Improving loss function for deep neural network based vector-to-vector regression.
method Presenting performance bounds and new properties of MAE, deriving generalized upper bounds, and interpreting MAE as a Laplacian distribution.
result MAE is a more suitable loss function than MSE for DNN based vector-to-vector regression, especially when errors follow a Laplacian distribution.

The paper tackles multi-armed bandits with vector losses, focusing on minimizing the \ell^\infty-norm of relative losses.

problem Minimizing the \ell^\infty-norm of relative losses in multi-armed bandits with multiple losses.
method Defines relative loss vector, derives lower bounds, and provides matching algorithms for both fixed-confidence best-arm identification and regret minimization.
result Derives problem-dependent sample complexity lower bound and matching algorithms for fixed-confidence best-arm identification.

Classification and regression tasks in overparameterized models show different generalization properties.

problem Comparing classification and regression in overparameterized models.
method Comparison of least-squares minimum-norm interpolation and hard-margin SVM using different loss functions.
result Interpolating solutions generalize well with 0-1 loss but not with square loss.

A new SVM classifier using L0/1L_{0/1} soft-margin loss for improved performance.

problem Improving SVM performance in binary classification tasks.
method Introducing L0/1L_{0/1} soft-margin loss and using the alternating direction method of multipliers.
result The new L0/1L_{0/1}-SVM model generates better performance with shorter computational time and fewer support vectors.

PoWER-BERT speeds up BERT inference by eliminating redundant word-vectors.

problem Improving BERT inference speed without sacrificing accuracy.
method Eliminating redundant word-vectors using a self-attention-based significance measure and learning the number of vectors to eliminate.
result Up to 4.5x reduction in inference time with <1% loss in accuracy on GLUE benchmark.

SGD converges to zero loss for separable data with fixed learning rate.

problem Optimizing homogeneous linear classifiers with SGD on linearly separable data.
method Proved convergence of SGD with fixed learning rate for separable data.
result SGD converges to zero loss for separable data with fixed learning rate.

New approach estimates personalized treatment effects using surrogate losses.

problem Estimating personalized treatment effects with binary outcomes and limited data.
method Proposes surrogate loss functions that incorporate both treatment and control data.
result Minimax support vector machine formulation yields tighter bounds.

New empirical Bayes estimator outperforms soft-thresholding for high-dimensional sparse vectors.

problem Estimating high-dimensional sparse vectors from noisy observations.
method Empirical Bayes shrinkage estimator using a Bernoulli-Gaussian prior.
result Hybrid estimator outperforms soft-thresholding in compressed sensing applications.

Study compares metric learning loss functions for speaker verification.

problem Comparing metric learning loss functions for end-to-end speaker verification.
method Cross entropy loss, cosine loss, angular margin loss, center loss, contrastive loss, triplet loss.
result Additive angular margin loss outperforms other loss functions.

The paper interprets VQ-VAE loss as a form of information bottleneck.

problem Understanding the VQ-VAE loss function.
method Interpreted VQ-VAE loss as variational deterministic information bottleneck (VDIB) and variational information bottleneck (VIB).
result VQ-VAE loss can be derived from VDIB and approximated by VIB.

New SVM model balances sparsity and robustness in noisy data.

problem Noise sensitivity and lack of sparsity in traditional SVM models.
method Combines elastic net loss with robust loss framework, integrates with SVM, uses half-quadratic algorithm.
result Proves sparsity and robustness, outperforms traditional SVMs in noisy environments.

A novel linear classification method that possesses the merits of both the Support Vector Machine (SVM) and the Distance-weighted Discrimination (DWD) is proposed in this article. The proposed Distance-weighted Support Vector Machine method can be viewed as a hybrid of SVM and DWD that finds the classification directio…

2013-10-11abs ↗pdf ↗

This paper extends neural collapse to imbalanced data under cross-entropy loss.

problem Analyzing neural collapse in deep networks with imbalanced data.
method Using the unconstrained feature model and cross-entropy loss, the paper studies neural collapse in imbalanced datasets.
result Feature vectors within the same class collapse to a single mean vector, but angles between them depend on sample size.

Sampling a fraction of pairs can match full evaluation in machine learning losses.

problem High computational cost of full pairwise loss evaluation.
method Survey sampling techniques targeting informative pairs.
result Performance close to full pairwise evaluation achieved with frugal sampling.

This paper consider penalized empirical loss minimization of convex loss functions with unknown non-linear target functions. Using the elastic net penalty we establish a finite sample oracle inequality which bounds the loss of our estimator from above with high probability. If the unknown target is linear this inequali…

2013-12-12abs ↗pdf ↗

RFSVM with random features achieves faster learning rates.

problem Improving the learning rate of SVM with random features.
method Support Vector Machine with NmN\ll m random features, optimized feature map, and reweighted feature selection.
result RFSVM achieves faster learning rates than O(1/m)O(1/\sqrt{m}) under low noise assumptions.

The paper studies consistency of surrogate loss procedures under constrained classifiers.

problem Consistency of surrogate loss approaches under constrained classifiers without correct specification.
method The paper develops theoretical results and hinge loss based procedures for a constrained classification problem.
result Hinge losses are the only surrogate losses that preserve consistency in second-best scenarios.

The stochastic linear bandit problem proceeds in rounds where at each round the algorithm selects a vector from a decision set after which it receives a noisy linear loss parameterized by an unknown vector. The goal in such a problem is to minimize the (pseudo) regret which is the difference between the total expected …

2016-06-17abs ↗pdf ↗

We consider the problem of designing a sparse Gaussian process classifier (SGPC) that generalizes well. Viewing SGPC design as constructing an additive model like in boosting, we present an efficient and effective SGPC design method to perform a stage-wise optimization of a predictive loss function. We introduce new me…

2012-06-26abs ↗pdf ↗

Study uses SVM to predict weather-induced home insurance claims and losses.

problem Assessing future weather-induced home insurance claims and losses for disaster preparedness.
method Support Vector Machine (SVM) regression for forecasting future claim dynamics.
result Illustrates SVM approach in forecasting weather-induced home insurance claims in a Canadian city.

A new method for stochastic optimal control improves accuracy over existing techniques.

problem Improving the accuracy of stochastic optimal control for noisy systems.
method Stochastic Optimal Control Matching (SOCM) using Iterative Diffusion Optimization (IDO) with path-wise reparameterization trick.
result SOCM achieves lower error than existing techniques for three out of four control problems, sometimes by an order of magnitude.

This paper aims at refined error analysis for binary classification using support vector machine (SVM) with Gaussian kernel and convex loss. Our first result shows that for some loss functions such as the truncated quadratic loss and quadratic loss, SVM with Gaussian kernel can reach the almost optimal learning rate, p…

2017-02-28abs ↗pdf ↗

Paper analyzes SVM behavior in high dimensions with exact formulas.

problem Characterizing SVM behavior in high-dimensional data with fixed ratio of features to samples.
method Exact asymptotic formulas derived through heuristic leave-one-out calculations.
result Exact formulas for variability of optimal coefficients, support vectors, objective function value, and misclassification error.

We present a framework to derive risk bounds for vector-valued learning with a broad class of feature maps and loss functions. Multi-task learning and one-vs-all multi-category learning are treated as examples. We discuss in detail vector-valued functions with one hidden layer, and demonstrate that the conditions under…

2016-06-05abs ↗pdf ↗

The paper introduces a new FOR framework using Huber and ε-insensitive losses.

problem Handling outliers and sparsity in functional output regression.
method Proposes a flexible FOR framework with infimal convolution losses and computable algorithms.
result Demonstrates efficiency and effectiveness on synthetic and real-world data.

Recently, fully-connected and convolutional neural networks have been trained to achieve state-of-the-art performance on a wide variety of tasks such as speech recognition, image classification, natural language processing, and bioinformatics. For classification tasks, most of these "deep learning" models employ the so…

2013-06-02abs ↗pdf ↗