Proposes provenance and pseudo-provenance for automated test generation.
problem Invalidation of provenance in generated tests.
method Annotation of generated tests with provenance trails and pseudo-provenance.
result Validates the reliability of generated tests and their relation to seeds.
The paper explains how many random seeds are needed for statistical significance in deep reinforcement learning experiments.
problem Ensuring statistical significance in deep reinforcement learning experiments.
method Theoretical guidelines for determining the number of random seeds for t-tests and bootstrap confidence intervals.
result Deviations from statistical test assumptions can lead to inaccurate evaluations of statistical errors.
OmniMatch algorithm perfectly matches graphs without edge correlation.
problem Graph matching in the absence of edge correlation.
method OmniMatch algorithm for seeded multiple graph matching.
result OmniMatch aligns O ( s α ) O(s^α) O ( s α ) unseeded vertices across multiple networks efficiently and perfectly. The paper proposes a method to stabilize predictions by identifying causal variables using a seed variable.
problem Stable prediction across unknown test data with potential spurious correlations.
method Conditional independence test based algorithm using a seed variable to separate causal from non-causal variables.
result The algorithm precisely separates causal and non-causal variables for stable prediction across test data.
A new protocol evaluates small machine learning improvements conservatively.
problem Uncertainty in small gains reported in machine learning papers.
method Paired bootstrap protocol with BCa confidence intervals and sign-flip permutation tests.
result Conservative evaluation reduces over-claiming of small improvements.
Fairness audits fail under missing protected labels, especially at zero access.
problem Understanding the reliability of fairness audits with incomplete protected-label data.
method Introduced a seed-calibrated stress test to separate missingness effects from seed-to-seed movement.
result Missing protected labels do not significantly alter fairness mitigation methods, but they can lead to harmful intersectional outcomes.
PNN-smoothing improves k k k -means clustering by merging subsets' clusterings.
problem Improving k k k -means clustering initialization efficiency and effectiveness. method Split dataset into subsets, cluster each subset, merge with PNN method.
result PNN-smoothing enhances k k k -means++ seeding, reducing costs. Seed selection affects Solvency II ratio, impacting insurance company stability.
problem Impact of RNG seed selection on Solvency II ratio stability.
method Detailed explanation of Solvency II ratio, theoretical background, and quality criteria for RNGs.
result True randomness is essential for stable Solvency II ratio results.
FIESTA optimizes model selection by efficiently evaluating multiple splits and seeds.
problem Inefficient model selection leading to unreliable performance comparisons.
method Adaptive bandit algorithms to determine optimal number of data splits and random seeds.
result Significantly reduces model evaluations while ensuring correct optimal model selection.
Non-determinism from GPUs dominates ResNet training accuracy variability.
problem Variability in ResNet training accuracy due to GPU non-determinism.
method Analysis of TensorFlow ResNet training on GPUs, comparing fixed seeds vs. different seeds.
result 74% of ResNet model variability is due to GPU non-determinism.
Empirical study shows SGD's random seed impacts model weights more than training examples, suggesting intrinsic privacy.
problem Understanding and leveraging the intrinsic randomness of SGD for improved privacy and utility.
method Large-scale empirical study on 120,000 models across four datasets, focusing on convex and non-convex objectives.
result Intrinsic randomness of SGD can reduce the need for additional noise to achieve privacy guarantees, with estimated ε i ( D ) ε_i(\mathcal{D}) ε i ( D ) as low as 6.3. New method measures model variability from stochastic optimization.
problem Measuring model quality obscured by stochastic optimization.
method Robust hypothesis testing and novel summary statistic.
result Shows α α α -trimming level is more expressive than performance metrics. Paper tackles graph matching with partially correct seeds, improving performance guarantees.
problem Graph matching with partially correct seeds.
method Proposes algorithms for matching vertices based on 1-hop and 2-hop neighborhoods, analyzing their performance guarantees.
result New 2-hop algorithm requires fewer correct seeds than the 1-hop algorithm, especially for sparse graphs.
The paper shows how shared random seeds can reduce variance in machine learning evaluations.
problem The statistical structure of comparative evaluation under shared random seeds is not well understood.
method An extended learning-based multi-agent economic simulator was used to demonstrate the effects of shared random seeds on variance reduction.
result Pairing seeds can reduce variance in machine learning evaluations, especially when outcomes are positively correlated at the seed level.
New method stabilizes machine learning predictions across random seeds.
problem Machine learning predictions vary across random seeds, causing instability.
method Introduces adaptive cross-bagging to eliminate seed dependence.
result Adaptive cross-bagging achieves targeted stability in debiased machine learning.
Enhances k-NN accuracy through randomized hyperstructure.
problem Improves k-NN accuracy by optimizing neighbor selection.
method Constructs a random n-dimensional hyperstructure around test instances to refine neighbor selection.
result 85.71% accuracy on Haberman's Cancer Survival dataset, compared to 80.95% for conventional k-NN.
Study finds many Lagrangian fillings for certain Legendrian links.
problem Understanding Lagrangian fillings for Legendrian links of finite type.
method Use of N-graphs and combinatorics of seed patterns.
result Proves existence of at least seeds many exact embedded Lagrangian fillings for Legendrian links of type ADE.
The paper examines how random seed affects model stability and proposes ASWA and NASWA techniques to improve model robustness.
problem The impact of random seed on model performance and stability.
method A controlled study on attention, gradient-based, and surrogate model interpretations. Proposed techniques: Aggressive Stochastic Weight Averaging (ASWA) and Norm-filtered Aggressive Stochastic Weight Averaging (NASWA).
result Improvement in model robustness with ASWA and NASWA techniques, reducing standard deviation of model performance by 72%.
New method fills cluster seeds with exact Lagrangian structures.
problem Surjectivity of map from Lagrangian fillings to cluster seeds.
method Construction of quiver with potential and Lagrangian disk surgeries.
result Trivial deformation space of quiver allows manipulation of fillings.
Improved community detection using effective resistance-based seed sets.
problem Low quality of seed nodes hinders community detection.
method Two-step process: effective resistance-based germination followed by PageRank.
result Improves precision and recall in community detection.
Study shows variability in CNN predictions for medical images, suggesting ensemble averaging to reduce it.
problem Variability in CNN predictions for medical images due to data ordering during training.
method Reproduced CheXNet results with random seeds to identify variability in predictions.
result Averaging predictions from multiple models reduces variability by nearly 70%.
Study finds optimal learning rate schedules for sub-100M quantization-aware training across bit-widths.
problem Optimal learning rate schedules for quantization-aware training depend on bit-width.
method Factorial grid testing over bit-width, warmdown fraction, LR magnitude, model size, and seed.
result INT6 QAT requires a different schedule than higher-precision training, falsifying the primary hypothesis.
Bayesian optimization improves performance with common random numbers.
problem Optimizing expensive stochastic functions with common random numbers.
method Proposes a novel Gaussian process model and Knowledge Gradient for Common Random Numbers.
result Significant performance improvements with moderate computational cost.
Optimized biopharmaceutical seed train design reduces variability and saves time.
problem Designing robust biotechnological processes with cell cultures.
method Coupling uncertainty-based upstream simulation and Bayes optimization using Gaussian processes.
result Optimized seed train design results in lower cell density variability and reduced process duration.
New method finds 198,846 toric-colorable seeds of Picard number 5.
problem Enumerating toric-colorable seeds of Picard number 5.
method Binary matroid approach and dynamic programming algorithm.
result 198,846 mod 2 toric-colorable seeds of dimension four and Picard number five.
Study on how intraclass variability affects Temporal Ensembling accuracy.
problem Effect of intraclass variability on Temporal Ensembling accuracy.
method Investigated through experiments with varying seed sizes and types on different datasets.
result Significant drop in accuracy with high intraclass variability datasets, more seed images improve accuracy, and seed type impacts overall efficiency.
Data-aware methods for dimensionality reduction and matrix decomposition aim to find low-dimensional structure in a collection of data. Classical approaches discover such structure by learning a basis that can efficiently express the collection. Recently, "self expression", the idea of using a small subset of data vect…
Given two graphs, the graph matching problem is to align the two vertex sets so as to minimize the number of adjacency disagreements between the two graphs. The seeded graph matching problem is the graph matching problem when we are first given a partial alignment that we are tasked with completing. In this paper, we m…
Seeds help match noisy graphs faster and more efficiently.
problem Matching noisy graphs with side information.
method Large neighborhood statistics for seeded graph matching.
result Achieves information-theoretic limit in polynomial time.
Model predicts crop yields to help farmers choose varieties that balance risk and reward.
problem Optimally selecting seed varieties to increase crop yield while managing risk.
method Hierarchical machine learning for yield prediction, integrated with weather forecasting, and decision-making models.
result Achieved a median absolute error of 3.74 bushels per acre in yield predictions.
Quickshift++ improves clustering stability and performance.
problem Improving initial seedings for clustering algorithms.
method Provably good initial seedings for Quick Shift clustering.
result Statistical consistency and strong clustering performance.
We present a novel approximate graph matching algorithm that incorporates seeded data into the graph matching paradigm. Our Joint Optimization of Fidelity and Commensurability (JOFC) algorithm embeds two graphs into a common Euclidean space where the matching inference task can be performed. Through real and simulated …
New algorithm for fuzzy clustering improves clustering quality.
problem Seeding iterative fuzzy algorithms for high-quality clustering.
method MaxMin Linear initialization algorithm for fuzzy C-Means.
result Validation through experiments on various datasets.
Efficiently selects seed nodes to maximize content influence in unknown social networks.
problem Maximizing content spread in social networks with unknown network model.
method Formulated as an infinite-horizon discounted MDP, uses model-based reinforcement learning to select seed users adaptively.
result Established a regret bound of O ~ ( T ) \widetilde O(\sqrt{T}) O ( T ) for the algorithm. Paper proves uniqueness of a complex construction.
problem Proving uniqueness of a complex construction.
method Defined a total complex and proved uniqueness of terms.
result Total complex C T o t ( L ) CTot(L) C T o t ( L ) is unique. PPM improves graph matching for correlated Gaussian Wigner models with high probability.
problem Graph matching in the Correlated Gaussian Wigner model with edge correlations.
method Seeded projected power method (PPM) for iterative improvement of initial partial matches.
result PPM recovers ground-truth matching with high probability in O(log n) iterations if seed is close enough.
ECL improves sequence modeling by learning evolutionary structure.
problem Ignoring evolutionary structure in sequence modeling leads to suboptimal performance.
method ECL progressively exposes models to sequences of increasing evolutionary distance.
result ECL improves performance across multiple biological domains and tasks.
We present a parallelized bijective graph matching algorithm that leverages seeds and is designed to match very large graphs. Our algorithm combines spectral graph embedding with existing state-of-the-art seeded graph matching procedures. We justify our approach by proving that modestly correlated, large stochastic blo…
Discrete knot theory models use lattice-filtered graphs to detect merging knot components.
problem Detecting merging knot components in discrete models.
method Lattice-filtered move graphs to model knot types, identifying connected components and merge scales.
result Merge scale defined by connected components of lattice-filtered move graphs, with specific examples for the figure-eight knot.
New method uses cluster shapes to improve track finding in particle collisions.
problem Combining timing and additional detector information for efficient track finding.
method Neural networks to analyze cluster shapes for track seeding.
result Cluster shapes reduce fake combinatorial backgrounds while maintaining high track efficiency.
Improved aspect detection from few seed keywords.
problem Fine-grained aspect detection from user reviews is labor-intensive.
method Weakly supervised co-training with student-teacher approach.
result Significant improvement in F1 scores over previous methods.
Paper proposes GSSNMF for legal document classification and topic modeling.
problem Lack of methods that can both classify and model topics with guidance.
method Guided Semi-Supervised Non-negative Matrix Factorization (GSSNMF).
result Improves both classification accuracy and topic coherence.
TOO optimizes stochastic epidemiological models by finding both parameter settings and random seeds.
problem Calibrating stochastic epidemiological models to match empirical observations.
method Gaussian process surrogates and Thompson sampling for optimization.
result Produces actual trajectories consistent with ground truth.
A cluster variety of Fock and Goncharov is a scheme constructed from the data related to the cluster algebras of Fomin and Zelevinsky. A seed is a combinatorial data which can be encoded as an n × n n\times n n × n matrix with integer entries, or as a quiver in special cases, together with n n n formal variables. A mutation is a c…
Proves isometric correspondence of leaves for generic rolling distributions.
problem Isometric correspondence of leaves in generic rolling distributions.
method Uses Bäcklund transformation for proof.
result Requires Bäcklund transformation for isometric correspondence of leaves.
RG-TTA adapts neural forecasters to streaming time series shifts by modulating adaptation intensity.
problem Adapting neural forecasters to distribution shifts in streaming time series data.
method RG-TTA uses a meta-controller that continuously modulates adaptation intensity based on distributional similarity.
result RG-TTA achieves the lowest MSE in 156 of 224 seed-averaged experiments, reducing MSE by 5.7% vs TTA.
EML-CD discovers causal mechanisms from neural networks in a structured way.
problem Extracting causal mechanisms from neural network weights is ill-posed.
method Integrates EML operator into causal structure learning, representing each edge mechanism as a gated EML binary tree.
result Achieves SHD=11.2 +/- 0.4 on real data, matching or outperforming existing methods.
SEED RL accelerates deep RL training on modern accelerators.
problem Training deep RL agents on large datasets at high speeds and low cost.
method Centralized inference, IMPALA/V-trace, R2D2, modern accelerators.
result Significant cost reduction and state-of-the-art performance on various games.