A new approach for anytime prediction using thin sub-networks and sparsity.
problem Efficient anytime prediction for deep neural networks.
method Training thin sub-networks and forcing sparsity on multi-branch network parameters.
result Thin sub-networks significantly outperform state-of-the-art dense architectures for anytime prediction.
We find faster-converging sub-networks that significantly reduce adversarial training time.
problem Finding optimal sub-networks for adversarial training is costly and time-consuming.
method We identify a subset of sub-networks that converge faster during training.
result Sub-networks can reduce adversarial training time by up to 49%.
The lottery ticket hypothesis finds multiple winning sub-networks in neural networks.
problem Finding a single winning sub-network in neural networks.
method Analyzing neural networks trained in isolation and on different tasks.
result Neural networks contain multiple sub-networks that match the accuracy of the original network, not just one.
Sparse sub-networks win transfer learning tasks.
problem Transfer learning with deep networks.
method Unstructured magnitude pruning to find winning tickets.
result Sparse sub-networks achieve similar or better accuracy than original networks.
Study finds differences in LTs across tasks and architectures, proposing a consensus-based method for generating refined lottery tickets.
problem Understanding the variability and uniqueness of Lottery Tickets across different image classification tasks and architectures.
method 28 combinations of image classification tasks and architectures, iterative pruning techniques, consensus-based method for generating refined lottery tickets.
result Disproves the uniqueness of Lottery Tickets and connects emergent mask structure to the choice of pruning.
New methods improve Laplace approximations for deep neural networks by selecting key parameters.
problem Improving uncertainty quantification in deep neural networks using computationally feasible approximations.
method Gradient-Laplace and Greedy-Laplace methods for selecting parameters in sub-network Laplace approximations.
result Gradient-Laplace method outperforms existing heuristic approaches and provides formal optimality guarantees.
Simplifies neural regression by combining two sub-networks for predictions and uncertainties.
problem Neural networks underestimate uncertainty, leading to overly confident predictions.
method Extends IRLS to a two-sub-network approach with shared representations and complementary loss functions.
result Proposed network is simpler to implement and more robust to uncertainty variations.
AlphaNet improves supernets training with alpha-divergence.
problem Improving the uncertainty distillation in weight-sharing NAS.
method Proposes alpha-divergence for better uncertainty distillation in supernets.
result Significant improvements in model performance across various FLOPs regimes.
This paper considers a Bayesian view for estimating a sub-network in a Markov random field. The sub-network corresponds to the Markov blanket of a set of query variables, where the set of potential neighbours here is big. We factorize the posterior such that the Markov blanket is conditionally independent of the networ…
DC-NAS improves neural architecture search by clustering and evaluating sub-networks.
problem Inaccurate evaluation of neural architectures in large search spaces.
method Divide-and-Conquer approach: feature representation, clustering, and evaluation of clusters.
result Achieved 75.1% top-1 accuracy on ImageNet, surpassing state-of-the-art methods.
The paper analyzes the role of ReLU gates in deep learning networks.
problem Understanding the role of gates in deep learning networks.
method Developed neural path features (NPF) and neural path values (NPV) to characterize the active sub-networks during training.
result The neural path kernel associated with NPFs is a fundamental quantity that characterizes the information stored in the gates of a DNN.
The paper examines rigidity of thin domains under specific boundary conditions.
problem Linear geometric rigidity of shallow thin domains with zero Dirichlet boundary conditions.
method Analyzes two scaling regimes for ε in (h, √h] and (√h, 1), proving rigidity formulas.
result Rigidity does not depend on curvature in the small parameter regime ε ∈ (h, √h].
If a tangle, K, in the 3-ball has no planar, meridional, essential surfaces in its exterior then thin position for K has no thin levels.
We define a new notion of thin position for a graph in a 3-manifold which combines the ideas of thin position for manifolds first originated by Scharlemann and Thompson with the idea of thin position for knots first originated by Gabai. This thin position has the property that connect summing annuli and pairs-of-pants …
A distributed SGD method for heterogeneous networks with hubs and workers.
problem Learning in heterogeneous multi-level networks with worker heterogeneity and varying communication.
method Multi-Level Local SGD: distributed SGD with hub-and-spoke paradigm and hub averaging.
result The method converges with error dependent on worker heterogeneity, hub network topology, and iterations.
New proof shows most thin knots satisfy Cabling Conjecture.
problem Cabling Conjecture for thin knots using Heegaard Floer homology.
method Heegaard Floer homology and immersed curves techniques.
result Almost all thin knots satisfy the Cabling Conjecture.
We produce embeddings of knots in thin position that admit compressible thin levels. We also find the bridge number of tangle sums where each tangle is high distance.
The paper finds many thin subgroups isomorphic to Gromov-Piatetski-Shapiro lattices.
problem Understanding thin subgroups in special linear groups.
method Constructing and embedding non-arithmetic hyperbolic manifolds into SL(n+1)(R).
result Non-arithmetic lattices in SO(n,1) can be embedded into SL(n+1)(R) as thin subgroups.
We describe a natural decomposition of a normal complex surface singularity (X,0) into its "thick" and "thin" parts. The former is essentially metrically conical, while the latter shrinks rapidly in thickness as it approaches the origin. The thin part is empty if and only if the singularity is metrically conical; the…
Let k be a knot in S3. In [8], H.N. Howards and J. Schultens introduced a method to construct a manifold decomposition of double branched cover of (S3, k) from a thin position of k. In this article, we will prove that if a thin position of k induces a thin decomposition of double branched cover of (S3,k) by Howards and…
Abby Thompson proved that if a link K is in thin position but not in bridge position then the knot complement contains an essential meridional planar surface, and she asked whether some thin level surface must be essential. This note is to give a positive answer to this question, showing that the if a link is in thin…
New method characterizes thin links via Conway spheres and tangle decompositions.
problem Characterize thin links without relying on specific knot invariants.
method Developed a relative version of thinness for tangles and used it to characterize thinness via tangle decompositions along Conway spheres.
result Characterized thin links via Conway spheres and tangle decompositions.
Regularized Stein thinning improves MCMC output approximations.
problem Pathologies in Stein thinning leading to poor approximations.
method Theoretical analysis and regularization to improve KSD.
result Regularized Stein thinning alleviates pathologies and improves efficiency.
Pruning FCNs reveals sub-networks that match CNNs' performance.
problem Understanding the inductive bias of pruning in neural networks.
method Iterative magnitude pruning of a simple FCN followed by analysis of the resulting architecture.
result Pruned FCNs exhibit key features of CNNs, suggesting new architectural biases.
The paper studies eigenvalues in negatively curved spaces, showing their relationship to the thick-thin decomposition.
problem Investigating the relationship between eigenvalues and the geometric structure of negatively curved spaces.
method Analyzing the eigenvalues of the Laplacian on the thick part of a manifold's thick-thin decomposition.
result Eigenvalues in the thick part of a negatively curved manifold are significantly smaller than those in the entire manifold.
Thin groups found in specific lattices.
problem Embedding right-angled Coxeter groups in arithmetic lattices.
method Using Agol's unpublished argument, embedding in indefinite orthogonal groups.
result Irreducible right-angled Coxeter groups embed as thin subgroups.
KSD Thinning uses KSD to thin MCMC samples efficiently.
problem Efficiently representing posterior distributions in Bayesian inference.
method KSD Thinning: retains only samples exceeding a KSD threshold.
result Established convergence and complexity tradeoffs for KSD Thinning.
Adapts CNN for robust medical image segmentation across different scanners and protocols.
problem Performance degradation of CNNs in medical image segmentation due to mismatch between training and test images.
method Designs a segmentation CNN as a concatenation of a shallow normalization CNN and a deep CNN. At test time, adapts the normalization sub-network for each test image using a denoising autoencoder.
result Consistently improves performance on multi-center MRI datasets of brain, heart, and prostate.
Data thinning splits observations into independent parts for convolution-closed distributions.
problem Validation of unsupervised learning results in settings with limited data.
method Data thinning, splitting observations into independent parts following the same distribution.
result Data thinning provides an attractive alternative to cross-validation in settings with limited sample splitting.
Study reveals trade dynamics in dry bulk shipping networks, highlighting their randomness and periodic changes.
problem Understanding the randomness and periodic changes in dry bulk shipping networks.
method Analysis of micro-level trade flow data from 2015 to 2023, focusing on grain, coal, and iron ore networks.
result Dry bulk shipping networks exhibit small-world phenomena and periodic life cycles, influenced by importing ports and global events.
Kernel thinning compresses distributions more effectively than i.i.d. sampling or standard thinning.
problem Efficiently compressing distributions for better sampling and integration accuracy.
method Introduces kernel thinning, a procedure that compresses an n-point approximation of a distribution into a sqrt(n)-point approximation with comparable integration error.
result Kernel thinning achieves a maximum discrepancy in integration error of O_d(n^(-1/2) sqrt(log n)) in probability for compactly supported distributions and O_d(n^(-1/2) (log n)^(d+1/2) sqrt(log log n)) for sub-exponential distributions.
Study shows (p,q)-cables of non-trivial knots are not thin.
problem Proving (p,q)-cables of non-trivial knots are not thin. method Bordered Floer theory of Lipshitz-Ozsváth-Thurston and Zemke's theorem.
result Proves (p,q)-cables of non-trivial knots are not thin. Generalizes data thinning for various distributions.
problem Thinning random variables without losing information.
method Relaxes summation requirement to sufficiency.
result Generalizes thinning to more distributions.
New method reduces summary points for datasets while maintaining quality.
problem Thinning datasets to reduce summary points while maintaining quality.
method Low-rank analysis of sub-Gaussian thinning.
result Guarantees high-quality compression for any distribution and kernel.
Proposes efficient training method for deep thin networks.
problem Deploying deep learning models with accuracy and compactness.
method Three-stage method: widen, warm up, fine tune.
result Deep thin networks trained with method outperform standard deep networks.
Wu has shown that if a link or a knot L in S3 in thin position has thin spheres, then the thin sphere of lowest width is an essential surface in the link complement. In this paper we show that if we further assume that L⊂S3 is prime, then the thin sphere of lowest width also does not have any vertical c…
This paper uses graph convolutional networks to improve the accuracy of neural architecture search.
problem Improving the precision of sampled sub-networks in weight-sharing NAS.
method Training a graph convolutional network to fit the performance of sampled sub-networks.
result Achieved higher rank correlation coefficient and better final architecture performance.
Arithmetic spaces' thin parts are negligible, impacting Betti numbers.
problem Understanding the structure of arithmetic locally symmetric spaces.
method Analyzing thin parts and deducing asymptotic results on Betti numbers.
result Arithmetic spaces' thin parts are negligible, impacting Betti numbers.
We give a method for searching for thin positions of a given link.
New thin subgroups found in special linear groups via bending techniques.
problem Finding thin subgroups of lattices in special linear groups.
method Techniques from convex projective geometry.
result Infinitely many non-commensurable lattices with thin subgroups.
We show that every thin position for a connected sum of small knots is obtained in an obvious way: place each summand in thin position so that no two summands intersect the same level surface, then connect the lowest minimum of each summand to the highest maximum of the adjacent summand below.
Early neural network training reveals important sub-networks and weight distributions.
problem Understanding the early phases of neural network training.
method Extensive measurements and quantitative probing of weight distribution and dataset reliance.
result Deep networks are not robust to reinitializing with random weights while maintaining signs, and weight distributions are highly non-independent.
Let L be a link in the 3-sphere that is in thin position but not in bridge position and let P be a thin level sphere. We generalize a result of Wu by giving a bound on the number of disjoint irreducible compressing disks that P can have, including identifying thin spheres with unique compressing disks. We also give con…
Compactifies group representations into thin triangle spaces.
problem Compactify group representations into geometric spaces.
method Geometric reformulation and extension of Culler-Morgan-Shalen theory.
result Ideal points are actions on real trees.
Survey of mathematical developments in gauge theory using thin homotopy.
problem Formalizing gauge fields as group homomorphisms on smooth connections.
method Introducing group structures on spaces of loops via thin homotopy equivalence.
result Clarified difference between thin and retrace equivalence for loops.
In this paper, we give an algorithm to build all compact orientable atoroidal Haken 3-manifolds with tori boundary or closed orientable Haken 3-manifolds, so that in both cases, there are embedded closed orientable separating incompressible surfaces which are not tori. Next, such incompressible surfaces are related to …
Study thin hyperbolic reflection groups and their properties.
problem Characterize and enumerate thin hyperbolic reflection groups.
method Analyze Zariski dense subgroups of hyperbolic isometries, apply Vinberg algorithm.
result All thin hyperbolic reflection groups are enumerable.
We introduce a method for creating a special type of tree, called a tree position, from a weighted graph. Leaves of the tree correspond to vertices of the original graph, and the tree edges contain information which can be used to partition these vertices. By repeatedly applying reducing operations to the tree position…