Unintended effects from scaling neural network outputs with adaptive learning rates.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Two-layer CNNs can overfit well if initialized correctly.
Multi-output Gaussian processes (MOGPs) leverage the flexibility and interpretability of GPs while capturing structure across outputs, which is desirable, for example, in spatio-temporal modelling. The key problem with MOGPs is their computational scaling , which is cubic in the number of both inputs (e…
The paper presents a novel approach to multi-output regression using probabilistic circuits.
Sketching accelerates structured prediction methods for large datasets.
Typically, Softmax is used in the final layer of a neural network to get a probability distribution for output classes. But the main problem with Softmax is that it is computationally expensive for large scale data sets with large number of possible outputs. To approximate class probability efficiently on such large sc…
Proposes GPLFR for predicting high-dimensional outputs with few data.
SFM resolves small-scale physics challenges in weather data.
This work bridges two views of feature learning in neural networks.
Principal components analysis (PCA) is a standard tool for identifying good low-dimensional approximations to data in high dimension. Many data sets of interest contain private or sensitive information about individuals. Algorithms which operate on such data should be sensitive to the privacy risks in publishing their …
Normalization effects on deep neural networks impact output variance and test accuracy.
Applications such as weather forecasting and personalized medicine demand models that output calibrated probability estimates---those representative of the true likelihood of a prediction. Most models are not calibrated out of the box but are recalibrated by post-processing model outputs. We find in this work that popu…
Optimal scaling found to depend on operator norm across large models and datasets.
Sig-PCA integrates model outputs and observations to correct model biases.
Scaling ResNets requires careful consideration of the layer depth and output scaling factors.
A new UNet variant reduces spectral artifacts in image transformations.
In this paper, we propose hybrid building/floor classification and floor-level two-dimensional location coordinates regression using a single-input and multi-output (SIMO) deep neural network (DNN) for large-scale indoor localization based on Wi-Fi fingerprinting. The proposed scheme exploits the different nature of th…
Unified framework for scale-invariant representation learning using MAPCA.
A method for constructing tight prediction intervals for multiple numerical outputs.
While a typical supervised learning framework assumes that the inputs and the outputs are measured at the same levels of granularity, many applications, including global mapping of disease, only have access to outputs at a much coarser level than that of the inputs. Aggregation of outputs makes generalization to new in…
This study reveals the critical role of scale vectors in large language models, improving optimization and expressivity.
Derives a family of hyperparameter scaling strategies for neural networks.
Paper improves generalization bounds for structured output prediction problems.
Extreme classification problems are multiclass and multilabel classification problems where the number of outputs is so large that straightforward strategies are neither statistically nor computationally viable. One strategy for dealing with the computational burden is via a tree decomposition of the output space. Whil…
The study addresses negative transfer in multi-output Gaussian processes by proposing latent structures.
Study on neural networks' performance under different normalizations as N grows.
Recurrent Neural Networks (RNNs) are among the most popular models in sequential data analysis. Yet, in the foundational PAC learning language, what concept class can it learn? Moreover, how can the same recurrent unit simultaneously learn functions from different input tokens to different output tokens, without affect…
Wide neural networks with asymmetrical node scaling converge globally and learn features.
Multi-output Gaussian processes (MOGPs) are an extension of Gaussian Processes (GPs) for predicting multiple output variables (also called channels, tasks) simultaneously. In this paper we use the convolution theorem to design a new kernel for MOGPs, by modeling cross channel dependencies through cross convolution of t…
Adjoint SA speeds up bioprocess parameter learning.
A main goal of regression is to derive statistical conclusions on the conditional distribution of the output variable Y given the input values x. Two of the most important characteristics of a single distribution are location and scale. Support vector machines (SVMs) are well established to estimate location functions …
Recently, self-normalizing neural networks (SNNs) have been proposed with the intention to avoid batch or weight normalization. The key step in SNNs is to properly scale the exponential linear unit (referred to as SELU) to inherently incorporate normalization based on central limit theory. SELU is a monotonically incre…
Recently, we proposed to transform the outputs of each hidden neuron in a multi-layer perceptron network to have zero output and zero slope on average, and use separate shortcut connections to model the linear dependencies instead. We continue the work by firstly introducing a third transformation to normalize the scal…
Proposes a deep tree-ensemble model for multi-output prediction.
A scalable method for Bayesian inference in large linear models.
Efficiently approximates uncertainty in classification models using Dirichlet distributions.
The multi-label classification framework, where each observation can be associated with a set of labels, has generated a tremendous amount of attention over recent years. The modern multi-label problems are typically large-scale in terms of number of observations, features and labels, and the amount of labels can even …
EPFGNN models graph connections for better node classification.
The study examines spectral dynamics in deep neural networks, predicting how outliers evolve during training.
Duel-Evolve uses LLM self-preferences for test-time optimization of discrete outputs.
Economic systems, traditionally analyzed as almost independent national systems, are increasingly connected on a global scale. Only recently becoming available, the World Input-Output Database (WIOD) is one of the first efforts to construct the multi-regional input-output (MRIO) tables at the global level. By viewing t…
Productions functions map the inputs of a firm or a productive system onto its outputs. This article expounds generalizations of the production function that include state variables, organizational structures and increasing returns to scale. These extensions are needed in order to explain the regularities of the empiri…
MixDiff detects OOD samples in constrained access environments by comparing perturbed samples.
We introduce a new regression framework, Gaussian process regression networks (GPRN), which combines the structural properties of Bayesian neural networks with the non-parametric flexibility of Gaussian processes. This model accommodates input dependent signal and noise correlations between multiple response variables,…
The Fisher information matrix (FIM) plays an essential role in statistics and machine learning as a Riemannian metric tensor or a component of the Hessian matrix of loss functions. Focusing on the FIM and its variants in deep neural networks (DNNs), we reveal their characteristic scale dependence on the network width, …
Novel framework for unbiased confidence estimates in object detection.
Remasking improves the quality of discrete diffusion models for natural language and image generation.
Improving predictive understanding of Earth system variability and change requires data-model integration. Efficient data-model integration for complex models requires surrogate modeling to reduce model evaluation time. However, building a surrogate of a large-scale Earth system model (ESM) with many output variables i…