Improved UCB method for stochastic bandits using distance tuning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper introduces a Hessian-based method to improve generalization in fine-tuned deep neural networks.
New metric captures individual neuron tuning across neural networks.
BDC uses Distance Correlation for efficient Bayesian optimization of expensive functions.
A new method constrains deep networks during fine-tuning to improve generalization.
Paper defines a new distance metric for comparing learning tasks.
Improved fine-tuning with regularization and robustness for noisy labels.
New study reveals surprising adaptive rates in model selection for transfer learning.
DoWG optimizer automatically adapts to convex and nonsmooth problems without tuning.
New classifiers for HDLSS data classify without tuning, robustly.
It has been reported repeatedly that discriminative learning of distance metric boosts the pattern recognition performance. A weak point of ITML-based methods is that the distance threshold for similarity/dissimilarity constraints must be determined manually and it is sensitive to generalization performance, although t…
TAWT improves cross-task learning efficiency and guarantees.
This paper relates parameter distance to gradient breakdown for a broad class of nonlinear compositional functions. The analysis leads to a new distance function called deep relative trust and a descent lemma for neural networks. Since the resulting learning rule seems to require little to no learning rate tuning, it m…
New method tunes SMC samplers efficiently without high costs.
Our work presents extensive empirical evidence that layer rotation, i.e. the evolution across training of the cosine distance between each layer's weight vector and its initialization, constitutes an impressively consistent indicator of generalization performance. In particular, larger cosine distances between final an…
Distances are fundamental primitives whose choice significantly impacts the performances of algorithms in machine learning and signal processing. However selecting the most appropriate distance for a given task is an endeavor. Instead of testing one by one the entries of an ever-expanding dictionary of {\em ad hoc} dis…
New distances measure mixtures of Gaussians, useful in machine learning.
Region-based classification of PolSAR data can be effectively performed by seeking for the assignment that minimizes a distance between prototypes and segments. Silva et al (2013) used stochastic distances between complex multivariate Wishart models which, differently from other measures, are computationally tractable.…
Prodigy estimates learning rate without tuning, improving convergence.
Generative adversarial networks (GANs) have received a tremendous amount of attention in the past few years, and have inspired applications addressing a wide range of problems. Despite its great potential, GANs are difficult to train. Recently, a series of papers (Arjovsky & Bottou, 2017a; Arjovsky et al. 2017b; and Gu…
Tuning machine learning models, particularly deep learning architectures, is a complex process. Automated hyperparameter tuning algorithms often depend on specific optimization metrics. However, in many situations, a developer trades one metric against another: accuracy versus overfitting, precision versus recall, smal…
PWIL learns agent behavior from expert using Wasserstein distance.
GANs mode collapse solved with Bures distance.
We present an accelerated algorithm for hierarchical density based clustering. Our new algorithm improves upon HDBSCAN*, which itself provided a significant qualitative improvement over the popular DBSCAN algorithm. The accelerated HDBSCAN* algorithm provides comparable performance to DBSCAN, while supporting variable …
Fingerprinting techniques, which are a common method for indoor localization, have been recently applied with success into outdoor settings. Particularly, the communication signals of Low Power Wide Area Networks (LPWAN) such as Sigfox, have been used for localization. In this rather recent field of study, not many pub…
Reshef & Reshef recently published a paper in which they present a method called the Maximal Information Coefficient (MIC) that can detect all forms of statistical dependence between pairs of variables as sample size goes to infinity. While this method has been praised by some, it has also been criticized for its lack …
Paper proposes an algorithm to learn DAGs with indirect dependencies.
We introduce an asymmetric distance in the space of learning tasks, and a framework to compute their complexity. These concepts are foundational for the practice of transfer learning, whereby a parametric model is pre-trained for a task, and then fine-tuned for another. The framework we develop is non-asymptotic, captu…
Bayesian Optimisation (BO) refers to a class of methods for global optimisation of a function which is only accessible via point evaluations. It is typically used in settings where is expensive to evaluate. A common use case for BO in machine learning is model selection, where it is not possible to analytically…
As a highlighting research topic in the multimedia area, cross-media retrieval aims to capture the complex correlations among multiple media types. Learning better shared representation and distance metric for multimedia data is important to boost the cross-media retrieval. Motivated by the strong ability of deep neura…
Proposes a new RL method to fine-tune flow-based models with arbitrary rewards.
SCoreBO improves Bayesian optimization by learning hyperparameters and self-correcting.
Implicit Generative Models (IGMs) such as GANs have emerged as effective data-driven models for generating samples, particularly images. In this paper, we formulate the problem of learning an IGM as minimizing the expected distance between characteristic functions. Specifically, we minimize the distance between charact…
The paper proves convergence of graph Laplacian with kNN self-tuned kernels.
New robust method for optimal transportation improves statistical inference.
Global optimization problems whose objective function is expensive to evaluate can be solved effectively by recursively fitting a surrogate function to function samples and minimizing an acquisition function to generate new samples. The acquisition step trades off between seeking for a new optimization vector where the…
New estimators for intrinsic dimension and Wasserstein distance improve OT accuracy.
A new algorithm is proposed which accelerates the mini-batch k-means algorithm of Sculley (2010) by using the distance bounding approach of Elkan (2003). We argue that, when incorporating distance bounds into a mini-batch algorithm, already used data should preferentially be reused. To this end we propose using nested …
A permutation-based SW test achieves minimax-optimal power for two-sample testing.
Researchers establish bounds for SGMs' KL and Wasserstein divergences under various noise schedules.
New algorithm estimates task affinities without repeated training, improving model performance and efficiency.
This work investigates how neural collapse improves transfer learning for large-scale models.
RLMH improves adaptive MCMC by optimizing contrastive divergence reward.
POTNet uses penalized optimal transport to generate data without mode collapse.
Bayesian hierarchical clustering (BHC) is an agglomerative clustering method, where a probabilistic model is defined and its marginal likelihoods are evaluated to decide which clusters to merge. While BHC provides a few advantages over traditional distance-based agglomerative clustering algorithms, successive evaluatio…
pGMM kernel outperforms ordinary ridge regression and RBF kernel ridge regression without tuning.
A new method optimizes spatial sampling for level set estimation in one dimension.
We propose a fast method with statistical guarantees for learning an exponential family density model where the natural parameter is in a reproducing kernel Hilbert space, and may be infinite-dimensional. The model is learned by fitting the derivative of the log density, the score, thus avoiding the need to compute a n…