Enhances ENet's prediction accuracy while maintaining uncertainty estimation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We address feature interpretation and reproducibility issues in dense nets, proposing a modified loss function.
New loss function improves convergence rate for neural networks.
SGD with large learning rates can achieve better test accuracy than expected.
Introduces Fitzpatrick losses, tighter than Fenchel-Young losses.
The overarching goal of this paper is to derive excess risk bounds for learning from exp-concave loss functions in passive and sequential learning settings. Exp-concave loss functions encompass several fundamental problems in machine learning such as squared loss in linear regression, logistic loss in classification, a…
In many applications of classifier learning, training data suffers from label noise. Deep networks are learned using huge training data where the problem of noisy labels is particularly relevant. The current techniques proposed for learning deep networks under label noise focus on modifying the network architecture and…
The predict-then-optimize framework is fundamental in many practical settings: predict the unknown parameters of an optimization problem, and then solve the problem using the predicted values of the parameters. A natural loss function in this environment is to consider the cost of the decisions induced by the predicted…
Characterizes learnability of forgiving 0-1 loss functions in multiclass settings.
Real-world large-scale datasets usually contain noisy labels and are imbalanced. Therefore, we propose derivative manipulation (DM), a novel and general example weighting approach for training robust deep models under these adverse conditions. DM has two main merits. First, loss function and example weighting are commo…
Label smoothing improves model robustness against misspecification.
Paper introduces a new topological loss for better convergence.
Fine-grained analysis of gradient descent with momentum provides modified loss equations.
We study the rates of convergence from empirical surrogate risk minimizers to the Bayes optimal classifier. Specifically, we introduce the notion of \emph{consistency intensity} to characterize a surrogate loss function and exploit this notion to obtain the rate of convergence from an empirical surrogate risk minimizer…
Develops a new density ratio estimator for causal inference.
It is known that Boosting can be interpreted as a gradient descent technique to minimize an underlying loss function. Specifically, the underlying loss being minimized by the traditional AdaBoost is the exponential loss, which is proved to be very sensitive to random noise/outliers. Therefore, several Boosting algorith…
Modified AUC improves CNN training by considering model confidence.
Image generating neural networks are mostly viewed as black boxes, where any change in the input can have a number of globally effective changes on the output. In this work, we propose a method for learning disentangled representations to allow for localized image manipulations. We use face images as our example of cho…
Gradient descent converges with arbitrary stepsize for separable data under Fenchel-Young losses.
Data compression speeds up machine learning loss calculations.
Improves forecast calibration for extreme events using modified loss functions.
Modified Jones-Faddy skew t-distribution captures asymmetry in stock returns.
This paper tackles noise in raw datasets to improve representation learning efficiency.
New insights into training machine learning models with momentum.
New method wraps black-box classifiers to reduce bias.
This paper presents a novel CNN-based approach for synthesizing high-resolution LiDAR point cloud data. Our approach generates semantically and perceptually realistic results with guidance from specialized loss-functions. First, we utilize a modified per-point loss that addresses missing LiDAR point measurements. Secon…
Translation-based embedding models have gained significant attention in link prediction tasks for knowledge graphs. TransE is the primary model among translation-based embeddings and is well-known for its low complexity and high efficiency. Therefore, most of the earlier works have modified the score function of the Tr…
In this study, we consider classification problems based on neural networks in data-imbalanced environment. Learning from an imbalanced data set is one of the most important and practical problems in the field of machine learning. A weighted loss function based on cost-sensitive approach is a well-known effective metho…
Modified ReLU networks improve regression estimation rates.
Understanding and evaluating the robustness of neural networks under adversarial settings is a subject of growing interest. Attacks proposed in the literature usually work with models trained to minimize cross-entropy loss and output softmax probabilities. In this work, we present interesting experimental results that …
New method estimates model risk without knowing function class.
We give a characterization of relative Ding stable toric Fano manifolds in terms of the behavior of the modified Ding functional. We call the corresponding behavior of the modified Ding functional the pseudo-boundedness from below. We also discuss the pseudo-boundedness of the Ding / Mabuchi functional of general Fano …
Previous research has shown that for stock indices, the most likely time until a return of a particular size has been observed is longer for gains than for losses. We establish that this so-called gain/loss asymmetry is present also for individual stocks and show that the phenomenon is closely linked to the well-known …
Gradient Descent (GD) approximators often fail in the solution space with multiple scales of convexities, i.e., in subspace learning and neural network scenarios. To handle that, one solution is to run GD multiple times from different randomized initial states and select the best solution over all experiments. However,…
While momentum-based accelerated variants of stochastic gradient descent (SGD) are widely used when training machine learning models, there is little theoretical understanding on the generalization error of such methods. In this work, we first show that there exists a convex loss function for which the stability gap fo…
Motivated by Kyprianou and Zhou (2009), Wang and Hu (2012), Avram et al. (2017), Li et al. (2017) and Wang and Zhou (2018), we consider in this paper the problem of maximizing the expected accumulated discounted tax payments of an insurance company, whose reserve process (before taxes are deducted) evolves as a spectra…
Enhances graph embeddings by preserving graph topology.
We demonstrate that the primal-dual witness proof method may be used to establish variable selection consistency and -bounds for sparse regression problems, even when the loss function and/or regularizer are nonconvex. Using this method, we derive two theorems concerning support recovery and -…
BCD algorithm finds global minima in neural networks.
Paper analyzes trade-offs in top-k classification accuracies and proposes a new loss function.
Existing approaches to online convex optimization (OCO) make sequential one-slot-ahead decisions, which lead to (possibly adversarial) losses that drive subsequent decision iterates. Their performance is evaluated by the so-called regret that measures the difference of losses between the online solution and the best ye…
Ridge regularized linear models (RRLMs), such as ridge regression and the SVM, are a popular group of methods that are used in conjunction with coefficient hypothesis testing to discover explanatory variables with a significant multivariate association to a response. However, many investigators are reluctant to draw ca…
A method for faster neural architecture search using low-fidelity training.
Data discretization is an important step in the process of machine learning, since it is easier for classifiers to deal with discrete attributes rather than continuous attributes. Over the years, several methods of performing discretization such as Boolean Reasoning, Equal Frequency Binning, Entropy have been proposed,…
NDI aims to forecast future natural disasters risk for insurers.
Proposes a new regularization technique for neural networks using elliptic operators.
We show a quite simple second variation formula for Perelman's -functional along the modified Kähler-Ricci flow over Fano manifolds.
The machine learning literature contains several constructions for prediction intervals that are intuitively reasonable but ultimately ad-hoc in that they do not come with provable performance guarantees. We present methods from the statistics literature that can be used efficiently with neural networks under minimal a…