Heuristic weighting improves denoising score matching without requiring noise distribution assumptions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper evaluates heuristics and hyperparameters in weight-sharing NAS methods.
A heuristic minimizes tardy jobs' total weight on single-machine scheduling.
Optimizes sample weights for representative data averages.
The use of automatic methods, often referred to as Neural Architecture Search (NAS), in designing neural network architectures has recently drawn considerable attention. In this work, we present an efficient NAS approach, named HM- NAS, that generalizes existing weight sharing based NAS approaches. Existing weight shar…
Optimizes sliding window approach for tracking Gaussian densities.
Learning sentence vectors from an unlabeled corpus has attracted attention because such vectors can represent sentences in a lower dimensional and continuous space. Simple heuristics using pre-trained word vectors are widely applied to machine learning tasks. However, they are not well understood from a theoretical per…
SLERP interpolation optimizes dynamic weight rebalancing in AMMs.
Develops a method for learning proposals in nested importance samplers.
Method aggregates models with different hyper-parameters to adapt to target domain.
Hybrid quantum-classical method optimizes financial index tracking.
Develops Heuristic Portfolio Optimization (HPO) as an information-restricted projection of Markowitz/tangency solution
Optimizes weights for better model performance in shifting data.
In this article we prove a family of local (in time) weighted Strichartz estimates with derivative losses for the Klein-Gordon equation on asymptotically de Sitter spaces and provide a heuristic argument for the non-existence of a global dispersive estimate on these spaces. The weights in the estimates depend on the ma…
STR reparameterizes DNN weights with soft thresholds for better sparsity and accuracy.
Proposes neuron alignment to optimize mode connectivity in neural networks.
New priors improve robustness and interpretability in penalized regression.
We show implicit filter level sparsity manifests in convolutional neural networks (CNNs) which employ Batch Normalization and ReLU activation, and are trained with adaptive gradient descent techniques and L2 regularization or weight decay. Through an extensive empirical study (Mehta et al., 2019) we hypothesize the mec…
Unweighted matrix factorization can match or outperform weighted methods in recommender systems.
In lexicon-based classification, documents are assigned labels by comparing the number of words that appear from two opposed lexicons, such as positive and negative sentiment. Creating such words lists is often easier than labeling instances, and they can be debugged by non-experts if classification performance is unsa…
The paper tackles efficient exploration in MDPs to learn accurate models.
We introduce NoisyNet, a deep reinforcement learning agent with parametric noise added to its weights, and show that the induced stochasticity of the agent's policy can be used to aid efficient exploration. The parameters of the noise are learned with gradient descent along with the remaining network weights. NoisyNet …
SoftAdapt dynamically adjusts loss weights for multi-part functions.
Combining information from different sources is a common way to improve classification accuracy in Brain-Computer Interfacing (BCI). For instance, in small sample settings it is useful to integrate data from other subjects or sessions in order to improve the estimation quality of the spatial filters or the classifier. …
BetaDataWeighter learns weights for unlabelled data to improve self-supervised learning accuracy.
New algorithm for weighted low rank approximation with provable guarantees.
Proposes a new method to initialize neural networks by estimating global curvature of weights.
Adapting neural networks to guide program optimization for better classifiers.
Aggregating multiple learners through an ensemble of models aim to make better predictions by capturing the underlying distribution of the data more accurately. Different ensembling methods, such as bagging, boosting, and stacking/blending, have been studied and adopted extensively in research and practice. While baggi…
Weight decay stabilizes training dynamics by slowing progressive sharpening.
Thompson Sampling, one of the oldest heuristics for solving multi-armed bandits, has recently been shown to demonstrate state-of-the-art performance. The empirical success has led to great interests in theoretical understanding of this heuristic. In this paper, we approach this problem in a way very different from exis…
We examine a class of deep learning models with a tractable method to compute information-theoretic quantities. Our contributions are three-fold: (i) We show how entropies and mutual informations can be derived from heuristic statistical physics methods, under the assumption that weight matrices are independent and ort…
During the last few years, there has been an interest in comparing simple or heuristic procedures for portfolio selection, such as the naive, equal weights, portfolio choice, against more "sophisticated" portfolio choices, and in explaining why, in some cases, the heuristic choice seems to outperform the sophisticated …
TDS provides exact samples for conditional distributions in diffusion models.
A method to combine saliency metrics for better CNN pruning decisions.
Recent pruning methods at initialization fall short of random pruning's accuracy.
In this paper, we develop a novel weighted Laplacian method, which is partially inspired by the theory of graph Laplacian, to study recent popular graph problems, such as multilevel graph partitioning and balanced minimum cut problem, in a more convenient manner. Since the weighted Laplacian strategy inherits the virtu…
Introduce a variance-weighted batch distribution for diverse sampling in diffusion models.
We consider applications of the theory of balanced weight filtrations and iterated logarithms, initiated in arXiv:1706.01073, to PDEs. The main result is a complete description of the asymptotics of the Yang--Mills flow on the space of metrics on a holomorphic bundle over a Riemann surface. A key ingredient in the argu…
New collapsing mechanism for G2-manifolds discovered.
We describe -MLE, a fast and efficient local search algorithm for learning finite statistical mixtures of exponential families such as Gaussian mixture models. Mixture models are traditionally learned using the expectation-maximization (EM) soft clustering technique that monotonically increases the incomplete (expec…
Analyzes layer-wise quantization effects in neural networks.
Federated Learning is a new subfield of machine learning that allows fitting models without collecting the training data itself. Instead of sharing data, users collaboratively train a model by only sending weight updates to a server. To improve the ranking of suggestions in the Firefox URL bar, we make use of Federated…
New loss function improves classification for imbalanced and sensitive groups.
A new test validates ensemble models against the null hypothesis.
Corrupting the input and hidden layers of deep neural networks (DNNs) with multiplicative noise, often drawn from the Bernoulli distribution (or 'dropout'), provides regularization that has significantly contributed to deep learning's success. However, understanding how multiplicative corruptions prevent overfitting ha…
In this paper we use a diffeo-geometric framework based on manifolds that are locally modeled on "convenient" vector spaces to study the geometry of some infinite dimensional spaces. Given a finite dimensional symplectic manifold , we construct a weak symplectic structure on each leaf of a foli…
A new design methodology for neural networks that is guided by traditional algorithm design is presented. To prove our point, we present two heuristics and demonstrate an algorithmic technique for incorporating additional weights in their signal-flow graphs. We show that with training the performance of these networks …