PNN-smoothing improves -means clustering by merging subsets' clusterings.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We provide initial seedings to the Quick Shift clustering algorithm, which approximate the locally high-density regions of the data. Such seedings act as more stable and expressive cluster-cores than the singleton modes found by Quick Shift. We establish statistical consistency guarantees for this modification. We then…
Discrete knot theory models use lattice-filtered graphs to detect merging knot components.
New method expands seed genes to functionally related clusters.
Improved K-Means++ and K-Means with faster run-time.
A fast regime-split Black-Scholes implied volatility solver
We study a well known noisy model of the graph isomorphism problem. In this model, the goal is to perfectly recover the vertex correspondence between two edge-correlated Erdős-Rényi random graphs, with an initial seed set of correctly matched vertex pairs revealed as side information. For seeded problems, our result pr…
A new method averages neural network parameters to rank features robustly.
Regularization improves stability and consistency of sparse autoencoders.
PPM improves graph matching for correlated Gaussian Wigner models with high probability.
New method speeds up k-means clustering for large k by improving nearest-neighbor search.
Confidence-based filtering reveals latent structure in diffusion models.
New method uses Rashomon sets to improve Bayesian inference in factorial designs.
We consider the problem of \emph{influence maximization}, the problem of maximizing the number of people that become aware of a product by finding the `best' set of `seed' users to expose the product to. Most prior work on this topic assumes that we know the probability of each user influencing each other user, or we h…
Paper tackles graph matching with partially correct seeds, improving performance guarantees.
In this paper, we present a novel approach for initializing deep neural networks, i.e., by turning PCA into neural layers. Usually, the initialization of the weights of a deep neural network is done in one of the three following ways: 1) with random values, 2) layer-wise, usually as Deep Belief Network or as auto-encod…
The paper shows how shared random seeds can reduce variance in machine learning evaluations.
Residual networks (ResNet) and weight normalization play an important role in various deep learning applications. However, parameter initialization strategies have not been studied previously for weight normalized networks and, in practice, initialization methods designed for un-normalized networks are used as a proxy.…
New method stabilizes machine learning predictions across random seeds.
Fairness audits fail under missing protected labels, especially at zero access.
Improved graph matching using covariates for network data integration.
Study finds many Lagrangian fillings for certain Legendrian links.
New method fills cluster seeds with exact Lagrangian structures.
Many methods for automated software test generation, including some that explicitly use machine learning (and some that use ML more broadly conceived) derive new tests from existing tests (often referred to as seeds). Often, the seed tests from which new tests are derived are manually constructed, or at least simpler t…
IIC decouples causal identification into two phases, significantly reducing the HTC gap in linear SEMs.
Mixed datasets consist of both numeric and categorical attributes. Various k-means-based clustering algorithms have been developed for these datasets. Generally, these algorithms use random partition as a starting point, which tends to produce different clustering results for different runs. In this paper, we propose, …
We reproduced the results of CheXNet with fixed hyperparameters and 50 different random seeds to identify 14 finding in chest radiographs (x-rays). Because CheXNet fine-tunes a pre-trained DenseNet, the random seed affects the ordering of the batches of training data but not the initialized model weights. We found subs…
Bayesian optimization is a powerful tool for expensive stochastic black-box optimization problems such as simulation-based optimization or machine learning hyperparameter tuning. Many stochastic objective functions implicitly require a random number seed as input. By explicitly reusing a seed a user can exploit common …
Proposes a method to accelerate safe sequential learning using offline data.
In this paper, we focus on quantifying model stability as a function of random seed by investigating the effects of the induced randomness on model performance and the robustness of the model in general. We specifically perform a controlled study on the effect of random seeds on the behaviour of attention, gradient-bas…
New method finds 198,846 toric-colorable seeds of Picard number 5.
Optimized biopharmaceutical seed train design reduces variability and saves time.
Study on how intraclass variability affects Temporal Ensembling accuracy.
Given two graphs, the graph matching problem is to align the two vertex sets so as to minimize the number of adjacency disagreements between the two graphs. The seeded graph matching problem is the graph matching problem when we are first given a partial alignment that we are tasked with completing. In this paper, we m…
Data-aware methods for dimensionality reduction and matrix decomposition aim to find low-dimensional structure in a collection of data. Classical approaches discover such structure by learning a basis that can efficiently express the collection. Recently, "self expression", the idea of using a small subset of data vect…
Community detection is, at its core, an attempt to attach an interpretable function to an otherwise indecipherable form. The importance of labeling communities has obvious implications for identifying clusters in social networks, but it has a number of equally relevant applications in product recommendations, biologica…
We present a novel approximate graph matching algorithm that incorporates seeded data into the graph matching paradigm. Our Joint Optimization of Fidelity and Commensurability (JOFC) algorithm embeds two graphs into a common Euclidean space where the matching inference task can be performed. Through real and simulated …
Efficiently selects seed nodes to maximize content influence in unknown social networks.
New method designs joint initial noises for diffusion models to improve diversity and alignment.
Eradicating hunger and malnutrition is a key development goal of the 21st century. We address the problem of optimally identifying seed varieties to reliably increase crop yield within a risk-sensitive decision-making framework. Specifically, we introduce a novel hierarchical machine learning mechanism for predicting c…
We prove that for a generic -dimensional integrable rolling distribution of contact elements (excluding developable seed and isotropic developable leaves) isometric correspondence of leaves of a general nature (independent of the shape of the seed) requires the Bäcklund transformation.
Paper proves uniqueness of a complex construction.
Clustering is an extensive research area in data science. The aim of clustering is to discover groups and to identify interesting patterns in datasets. Crisp (hard) clustering considers that each data point belongs to one and only one cluster. However, it is inadequate as some data points may belong to several clusters…
We present a parallelized bijective graph matching algorithm that leverages seeds and is designed to match very large graphs. Our algorithm combines spectral graph embedding with existing state-of-the-art seeded graph matching procedures. We justify our approach by proving that modestly correlated, large stochastic blo…
New method uses cluster shapes to improve track finding in particle collisions.
Randomly initialized transformers show extreme token preferences.
The paper finds Koopman invariant subspaces using personalized PageRank.
TOO optimizes stochastic epidemiological models by finding both parameter settings and random seeds.