Poisson learning doesn't solve graph semi-supervised learning issues.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new method assesses regression models' global optimality.
Adaptive rerouting reshapes impacts of maritime chokepoint disruptions
As financial instruments grow in complexity more and more information is neglected by risk optimization practices. This brings down a curtain of opacity on the origination of risk, that has been one of the main culprits in the 2007-2008 global financial crisis. We discuss how the loss of transparency may be quantified …
Distributed learning of probabilistic models from multiple data repositories with minimum communication is increasingly important. We study a simple communication-efficient learning framework that first calculates the local maximum likelihood estimates (MLE) based on the data subsets, and then combines the local MLEs t…
We propose SEARNN, a novel training algorithm for recurrent neural networks (RNNs) inspired by the "learning to search" (L2S) approach to structured prediction. RNNs have been widely successful in structured prediction applications such as machine translation or parsing, and are commonly trained using maximum likelihoo…
Global implicit function theorem for Fréchet spaces, solving derivative loss problems.
This paper introduces a scalable benchmark for evaluating local posterior sampling in neural networks.
Partial Label Learning (PLL) aims to learn from the data where each training instance is associated with a set of candidate labels, among which only one is correct. Most existing methods deal with such problem by either treating each candidate label equally or identifying the ground-truth label iteratively. In this pap…
A new learning scheme improves model efficiency and performance.
In this paper we revisit the weighted likelihood bootstrap, a method that generates samples from an approximate Bayesian posterior of a parametric model. We show that the same method can be derived, without approximation, under a Bayesian nonparametric model with the parameter of interest defined as minimising an expec…
Recent works have shown that on sufficiently over-parametrized neural nets, gradient descent with relatively large initialization optimizes a prediction function in the RKHS of the Neural Tangent Kernel (NTK). This analysis leads to global convergence results but does not work when there is a standard regulari…
A hybrid loss framework improves time series forecasting by balancing global and component errors.
SGD converges globally to logistic loss minima for two-layer nets.
This paper examines how different loss functions affect neural network features and performance.
We study the error landscape of deep linear and nonlinear neural networks with the squared error loss. Minimizing the loss of a deep linear neural network is a nonconvex problem, and despite recent progress, our understanding of this loss surface is still incomplete. For deep linear networks, we present necessary and s…
New neural network solves Nirenberg problem for curvature on sphere.
Improved time series forecasting with expert loss integration.
BCD algorithm finds global minima in neural networks.
This work justifies neural collapse under MSE loss and analyzes the optimization landscape.
Proposes a new contrastive loss for semi-supervised medical image segmentation.
Hill-ADAM optimizes loss landscapes by exploring state space deterministically.
The past decade has witnessed a successful application of deep learning to solving many challenging problems in machine learning and artificial intelligence. However, the loss functions of deep neural networks (especially nonlinear networks) are still far from being well understood from a theoretical aspect. In this pa…
Autoregressive feedback is considered a necessity for successful unconditional text generation using stochastic sequence models. However, such feedback is known to introduce systematic biases into the training process and it obscures a principle of generation: committing to global information and forgetting local nuanc…
This paper shows that every sublevel set of the loss function of a class of deep over-parameterized neural nets with piecewise linear activation functions is connected and unbounded. This implies that the loss has no bad local valleys and all of its global minima are connected within a unique and potentially very large…
Most existing word embedding methods can be categorized into Neural Embedding Models and Matrix Factorization (MF)-based methods. However some models are opaque to probabilistic interpretation, and MF-based methods, typically solved using Singular Value Decomposition (SVD), may incur loss of corpus information. In addi…
Due to the success of deep learning to solving a variety of challenging machine learning tasks, there is a rising interest in understanding loss functions for training neural networks from a theoretical aspect. Particularly, the properties of critical points and the landscape around them are of importance to determine …
Study shows how information loss and operation loss are related in feature representations.
New methods connect low-loss points on neural network surfaces.
We consider deep linear networks with arbitrary convex differentiable loss. We provide a short and elementary proof of the fact that all local minima are global minima if the hidden layers are either 1) at least as wide as the input layer, or 2) at least as wide as the output layer. This result is the strongest possibl…
In this study, a novel feature coding method that exploits invariance for transformations represented by a finite group of orthogonal matrices is proposed. We prove that the group-invariant feature vector contains sufficient discriminative information when learning a linear classifier using convex loss minimization. Ba…
L2G2G improves graph autoencoder accuracy without sacrificing scalability.
Deep ReLU networks with extra parameters have mostly good loss landscapes.
PhysicsFormer improves TSF models for GSWF with WEATHER-5K dataset.
GCBS improves contrastive learning performance efficiently.
Gradient-informed BNNs improve Bayesian optimization performance.
Graph kernels are widely used for measuring the similarity between graphs. Many existing graph kernels, which focus on local patterns within graphs rather than their global properties, suffer from significant structure information loss when representing graphs. Some recent global graph kernels, which utilizes the align…
While optimizing convex objective (loss) functions has been a powerhouse for machine learning for at least two decades, non-convex loss functions have attracted fast growing interests recently, due to many desirable properties such as superior robustness and classification accuracy, compared with their convex counterpa…
Informer model with GMADL loss outperforms benchmarks in high frequency Bitcoin trading.
GSP improves global average pooling for deep metric learning by learning weights and selecting semantic entities.
The paper decomposes probabilistic scores into reliability, uncertainty, and information loss.
Paper develops methods for non-quadratic loss low-rank matrix recovery.
This paper presents novel Gaussian process decentralized data fusion algorithms exploiting the notion of agent-centric support sets for distributed cooperative perception of large-scale environmental phenomena. To overcome the limitations of scale in existing works, our proposed algorithms allow every mobile sensing ag…
This paper investigates, from information theoretic grounds, a learning problem based on the principle that any regularity in a given dataset can be exploited to extract compact features from data, i.e., using fewer bits than needed to fully describe the data itself, in order to build meaningful representations of a re…
Study reveals properties of local minima in ReLU networks.
AdaLoss optimizes adaptive learning rates for efficient convergence in various models.
Global graph structure improves GNN performance.
Topology-enhanced loss improves 3D object reconstruction from 2D images.