A method to reduce memory usage in deep learning models by adding inducing weights.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Weight Decay induces low-rank weight matrices in neural networks, improving generalization.
Global inducing points improve Bayesian neural network performance.
Any Sasakian structure can be closely mimicked by embeddings into weighted spheres.
Gradient descent with random weights in linear regression analyzed for various noise types.
Replacing MSE with f-divergence in diffusion models improves robustness under data contamination.
In recent years, deep neural networks (DNNs) have been applied to various machine leaning tasks, including image recognition, speech recognition, and machine translation. However, large DNN models are needed to achieve state-of-the-art performance, exceeding the capabilities of edge devices. Model reduction is thus nee…
We investigate the relation between weighted quasi-metric Spaces and Finsler Spaces. We show that the induced metric of a Randers space with reversible geodesics is a weighted quasi-metric space.
Deep Bayesian neural nets can use simpler weight approximations without sacrificing performance.
Enhances trading signals using image analysis and weighted moving averages.
We extend the notion of intersection graphs for knots in the theory of finite type invariants to string links. We use our definition to develop weight systems for string links via the adjacency matrix of the intersection graph, and show that these weight systems are related to the weight systems induced by the Conway a…
Characterizes inductive bias in multi-channel linear CNNs with bounded weight norm.
Learning representation on graph plays a crucial role in numerous tasks of pattern recognition. Different from grid-shaped images/videos, on which local convolution kernels can be lattices, however, graphs are fully coordinate-free on vertices and edges. In this work, we propose a Gaussian-induced convolution (GIC) fra…
Corrects distribution shift in target shift scenarios using importance weighting.
Optimal Euclidean structure minimizes energy in weighted toroidal graphs.
This work aims at solving the problems with intractable sparsity-inducing norms that are often encountered in various machine learning tasks, such as multi-task learning, subspace clustering, feature selection, robust principal component analysis, and so on. Specifically, an Iteratively Re-Weighted method (IRW) with so…
HALO learns to prune neural networks by adaptively shrinking weights.
In our previous paper, we discussed the hyperbolization of the configuration space of n(> 4) marked points with weights in the projective line up to projective transformations. A variation of the weights induces a deformation. It was shown that this correspondence of the set of the weights to the Teichmüller space when…
SURF steers scalarization weights to uniformly traverse the Pareto front.
Study on feature learning dynamics in infinite-depth neural networks, focusing on ResNets.
Let be a compact connected oriented dimensional manifold without boundary. In this work, shape space is the orbifold of unparametrized immersions from to . The results of \cite{Michor118}, where mean curvature weighted metrics were studied, suggest incorporating Gauß curvature weights in the …
Regularization can induce grokking in neural networks, improving generalization.
We investigate deep Bayesian neural networks with Gaussian weight priors and a class of ReLU-like nonlinearities. Bayesian neural networks with Gaussian priors are well known to induce an L2, "weight decay", regularization. Our results characterize a more intricate regularization effect at the level of the unit activat…
In the present paper, we study deformations of polar weighted homogeneous polynomials which are also polar weighted homogeneous polynomials. We describe a round handle decomposition of the Milnor fibration of a deformation of a polar weighted homogeneous polynomial concretely and give the number of round handles by the…
Importance weighted variational inference (Burda et al., 2015) uses multiple i.i.d. samples to have a tighter variational lower bound. We believe a joint proposal has the potential of reducing the number of redundant samples, and introduce a hierarchical structure to induce correlation. The hope is that the proposals w…
Cautious Weight Decay modifies weight decay for better optimization.
B-cos transforms improve neural network interpretability by aligning weights.
Paper investigates rigidity phenomena for weighted Ricci curvature bounds with Laplacian comparison theorem.
Algorithmic approaches endow deep learning systems with implicit bias that helps them generalize even in over-parametrized settings. In this paper, we focus on understanding such a bias induced in learning through dropout, a popular technique to avoid overfitting in deep learning. For single hidden-layer linear neural …
The paper analyzes prediction error in nonstationary settings using weighted risk minimization.
SGD and weight decay encourage neural networks to learn low-rank weight matrices.
Prior-weighted logistic regression has become a standard tool for calibration in speaker recognition. Logistic regression is the optimization of the expected value of the logarithmic scoring rule. We generalize this via a parametric family of proper scoring rules. Our theoretical analysis shows how different members of…
Bayesian framework improves minority class performance in class-imbalanced data.
A new method corrects weight values to improve treatment effect estimation.
With the success of deep neural networks, Neural Architecture Search (NAS) as a way of automatic model design has attracted wide attention. As training every child model from scratch is very time-consuming, recent works leverage weight-sharing to speed up the model evaluation procedure. These approaches greatly reduce …
Model for material elasticity and plasticity using networks.
Not all instances in a data set are equally beneficial for inducing a model of the data. Some instances (such as outliers or noise) can be detrimental. However, at least initially, the instances in a data set are generally considered equally in machine learning algorithms. Many current approaches for handling noisy and…
In importance sampling (IS)-based reinforcement learning algorithms such as Proximal Policy Optimization (PPO), IS weights are typically clipped to avoid large variance in learning. However, policy update from clipped statistics induces large bias in tasks with high action dimensions, and bias from clipping makes it di…
Koopman mode analysis applied to neural networks for training optimization.
Optimal discrete harmonic maps between hyperbolic surfaces are found via minimizing energy.
We prove several results about chordal graphs and weighted chordal graphs by focusing on exposed edges. These are edges that are properly contained in a single maximal complete subgraph. This leads to a characterization of chordal graphs via deletions of a sequence of exposed edges from a complete graph. Most interesti…
TILT improves target domain performance by penalizing an auxiliary component on unlabeled target inputs.
Optimizes pruning masks for neural networks using probabilistic fine-tuning and PAC-Bayes bounds.
Inverse classification uses an induced classifier as a queryable oracle to guide test instances towards a preferred posterior class label. The result produced from the process is a set of instance-specific feature perturbations, or recommendations, that optimally improve the probability of the class label. In this work…
STR reparameterizes DNN weights with soft thresholds for better sparsity and accuracy.
Bayesian neural networks can be simplified by parameterizing weights as rank- matrices, reducing parameter count and improving performance.
We consider a setting where an agent's uncertainty is represented by a set of probability measures, rather than a single measure. Measure-bymeasure updating of such a set of measures upon acquiring new information is well-known to suffer from problems; agents are not always able to learn appropriately. To deal with the…
Kähler information manifolds for signal filters in weighted Hardy spaces are explored.