SGD updates align with a low-rank subspace but do not lead to further loss reduction.
problem Understanding the training dynamics of deep neural networks, particularly the role of the dominant subspace.
method Exploring whether neural networks can be trained within the dominant subspace of the loss Hessian.
result SGD updates, when projected onto the dominant subspace, do not decrease the training loss further, suggesting spurious alignment.
New method accelerates neural network training by focusing on flat directions.
problem Improving neural network training speed and stability.
method Bulk-SGD, interpolated gradient methods.
result Updates along the Dominant subspace can accelerate convergence but compromise stability.
Active sampling selects few points for accurate model reduction of high-fidelity systems.
problem Efficiently identify dominant subspaces for model reduction of large training sets.
method Proposes an active sampling strategy to select a few points from the training set to estimate dominant subspaces accurately.
result Active sampling can provide 17x speed-up without sacrificing accuracy.
A geometric analysis of the time series of returns has been performed in the past and it implied that the most of the systematic information of the market is contained in a space of small dimension. Here we have explored subspaces of this space to find out the relative performance of portfolios formed from the companie…
The paper explores anti-hyperbolicity for hyperkähler varieties.
problem Anti-hyperbolicity of hyperkähler varieties.
method Exploring various examples and criteria for meromorphic and holomorphic dominability by C^m.
result Generalizing known results about K3 surfaces to hyperkähler manifolds.
The paper extends hypothesis testing to non-diagonalizable matrices, improving network statistics inference.
problem Testing on non-diagonalizable matrices for network statistics.
method Generalizes Wald and t-tests to non-symmetric matrices, controlling convergence rates.
result Improved inference on network statistics from directed networks.
This paper considers the problem of estimating a high-dimensional vector of parameters θ∈Rn from a noisy observation. The noise vector is i.i.d. Gaussian with known variance. For a squared-error loss function, the James-Stein (JS) estimator is known to dominate the simple maximum-likelihood (…
We extend the theoretical analysis of a recently proposed single subspace learning algorithm, called Dual Principal Component Pursuit (DPCP), to the case where the data are drawn from of a union of hyperplanes. To gain insight into the properties of the ℓ1 non-convex problem associated with DPCP, we develop a geo…
ISOKANN learns collective variables and effective dynamics for metastable transitions.
problem Understanding metastable transitions in complex molecular systems.
method Integrates Koopman operators with neural networks to extract CVs and effective dynamics.
result Reconstructs coarse-grained kinetics and reproduces transition times across barriers.
LASER compresses recursive model activations by exploiting their low-dimensional structure.
problem Understanding and optimizing the geometric structure of recursive reasoning trajectories.
method Dynamic low-rank basis tracking via matrix-free subspace tracking with a fidelity-triggered reset mechanism.
result Recursive activations occupy a linear, low-dimensional subspace that can be compressed efficiently.
Memory-efficient optimizers fail to track a subspace, leading to unpredictable model performance.
problem Memory-efficient optimizers fail to track a subspace, leading to unpredictable model performance.
method Analyzing the behavior of memory-efficient optimizers like GaLore, which project gradients onto a rank-r subspace recomputed every T steps.
result Memory-efficient optimizers fail to track a subspace, leading to unpredictable model performance.
The paper develops methods to reduce deployment risk under dynamic covariate shifts.
problem Reduction of deployment risk under dynamic covariate shifts.
method Time-domain Poincare inequality and Jacobian-velocity theorem to identify and control directional tangent energy.
result Drift-aligned tangent regularization (DTR) reduces risk volatility and directional gain in low-rank drift regimes.
Introduces new limit spaces for degenerating Calabi-Yau families.
problem Understanding degenerating Calabi-Yau families and their limit structures.
method Introduces galaxy spaces as dense subspace of infinite open Calabi-Yau varieties.
result Galaxy spaces are projective limits of toroidal compactifications.
This paper proposes a new method to adapt ROMs for new parameter settings.
problem ROMs lack robustness when applied to new parameter settings.
method Regression trees on Grassmann Manifold to learn the mapping between parameters and POD bases.
result The proposed method is capable of establishing the mapping between parameters and POD bases, thus adapting ROMs for new parameters.
New method speeds up kernel-based machine learning for force field reconstruction.
problem Scalability issues in kernel-based machine learning for force field reconstruction.
method Nyström-type methods to construct preconditioners based on low-rank approximations of the kernel matrix.
result Effective preconditioners lead to super-linear convergence in kernel-based machine learning.
We study conformal symmetry breaking differential operators which map differential forms on Rn to differential forms on a codimension one subspace Rn−1. These operators are equivariant with respect to the conformal Lie algebra of the subspace Rn−1. They correspond to homomorphism…
PLUMAGE improves large model training efficiency and stability.
problem Accelerator memory and networking constraints during large model training.
method Probabilistic Low rank Unbiased Minimum Variance Gradient Estimator (PLUMAGE) that resolves bias and variance issues.
result PLUMAGE reduces training loss by 28% on average across the GLUE benchmark.
Gradient descent biases linear models in next-token prediction towards data entropy.
problem Optimization bias in next-token prediction models.
method Analysis of gradient descent on linear models with sparse conditional distributions.
result Gradient descent selects parameters that equate token logits differences to log-odds in the data subspace.
Massive MIMO is a variant of multiuser MIMO where the number of base-station antennas M is very large (typically 100), and generally much larger than the number of spatially multiplexed data streams (typically 10). Unfortunately, the front-end A/D conversion necessary to drive hundreds of antennas, with a signal band…
Connected domination numbers found for plane triangulations up to 13 vertices.
problem Finding connected domination numbers for plane triangulations.
method Analyzing triangulations of up to 13 vertices and proving the difference between connected and regular domination numbers can be arbitrarily large.
result Connected domination numbers for triangulations up to 13 vertices and upper bound for larger triangulations.
W2S FT often outperforms weak teachers due to low intrinsic dimensionality.
problem Understanding why weak-to-strong finetuning outperforms weak models.
method Analyzing W2S in ridgeless regression setting, focusing on variance reduction.
result Weak teacher's variance is inherited by strong student in shared feature subspace, reduced in discrepancy subspace.
Manifolds can be dominated by hypersurfaces in a sphere.
problem Dominating manifolds with hypersurfaces.
method Proving any smooth, closed, oriented manifold can be dominated by a codimension 1 submanifold of the sphere.
result Any smooth, closed, oriented manifold can be dominated by a codimension 1 submanifold of the sphere.
A new EnKF method for elliptic PDEs reduces dimensionality for accurate state estimation.
problem Elliptic PDEs in fluid flows make traditional EnKF regularization ineffective.
method Low-rank factorization of the Kalman gain based on the Jacobian spectrum.
result Inference can be performed in a low-dimensional subspace of the state space.
New method ranks multivariate distributions in SMOOP using q-dominance.
problem Lack of reliable methods to rank multivariate distributions in SMOOP.
method Introduces center-outward q-dominance and develops empirical test procedures.
result Proves q-dominance implies FSD and establishes a sample size threshold.
We show that non-domination results for targets that are not dominated by products are stable under Cartesian products.
New framework for ranking distributions using variable fractional parameters.
problem Ordering distributions with varying steepness and local non-concavities.
method Introducing a function γ:Ro[0,1] to replace the fixed parameter in fractional SD. result Enables ranking of a broader range of distributions and incorporates dynamic greediness.
Research examines when 4-manifolds are dominated by geometric ones.
problem When is an orientable closed 4-manifold dominated by another?
method Focuses on geometric or fibred cases.
result Characterizes conditions for domination.
3-manifolds can virtually dominate others with positive simplicial volume.
problem Domination of 3-manifolds with positive simplicial volume.
method Proving existence of finite covers with degree-1 maps.
result Virtual domination of 3-manifolds with positive simplicial volume.
Paper bounds subspace estimator error from noisy projections.
problem Estimating subspaces from noisy data.
method Derives perturbation bound on optimal subspace estimator.
result Fundamental result with implications in matrix completion and clustering.
PCA is one of the most widely used dimension reduction techniques. A related easier problem is "subspace learning" or "subspace estimation". Given relatively clean data, both are easily solved via singular value decomposition (SVD). The problem of subspace learning or PCA in the presence of outliers is called robust su…
Union of Subspaces (UoS) is a popular model to describe the underlying low-dimensional structure of data. The fine details of UoS structure can be described in terms of canonical angles (also known as principal angles) between subspaces, which is a well-known characterization for relative subspace positions. In this pa…
Study Sp(n)-orbits in complex and Σ-complex subspaces of Hermitian quaternionic vector spaces.
problem Characterize Sp(n)-orbits in Grassmannians of complex and Σ-complex subspaces. method Decompose subspaces into 4-dimensional complex addends and 2-dimensional totally complex subspace. Use properties of isoclinic subspaces and principal angles.
result Determine full set of invariants for Sp(n)-orbits in GrR(2k,4n). In subspace clustering, a group of data points belonging to a union of subspaces are assigned membership to their respective subspaces. This paper presents a new approach dubbed Innovation Pursuit (iPursuit) to the problem of subspace clustering using a new geometrical idea whereby subspaces are identified based on the…
Study domination between non-Fuchsian surface group representations and anti-de Sitter geometry.
problem Domination problem between non-Fuchsian representations of closed surface groups.
method Analysis of branched harmonic immersions and construction of anti-de Sitter 3-manifolds.
result Found that representations admitting branched harmonic immersions dominate other representations, and constructed large families of branched anti-de Sitter 3-manifolds.
The paper improves alignment methods for deep neural networks using geometric and spectral analysis.
problem Improving alignment methods for deep neural networks.
method Geometric and spectral analysis of residual Jacobian chains.
result Deterministic and margin-verified results on the transport of dominant singular subspaces across layers.
We consider the problem of detecting whether a tensor signal having many missing entities lies within a given low dimensional Kronecker-Structured (KS) subspace. This is a matched subspace detection problem. Tensor matched subspace detection problem is more challenging because of the intertwined signal dimensions. We s…
Paper shows affine constraint is unnecessary for high-dimensional data.
problem The necessity of an affine constraint in affine subspace clustering.
method Theoretical and empirical analysis of conditions for correctness of affine subspace clustering methods.
result Affine constraint has negligible effect on clustering performance for high-dimensional data.
A new method learns outcome-aware spectral features for causal effect estimation.
problem Estimation of causal effects in the presence of hidden confounders.
method Augmented Spectral Feature Learning framework that minimizes a contrastive loss derived from an augmented operator incorporating outcome information.
result Our method remains effective even under spectral misalignment.
A low-rank transformation learning framework for subspace clustering and classification is here proposed. Many high-dimensional data, such as face images and motion sequences, approximately lie in a union of low-dimensional subspaces. The corresponding subspace clustering problem has been extensively studied in the lit…
This paper investigates the generalization of Principal Component Analysis (PCA) to Riemannian manifolds. We first propose a new and general type of family of subspaces in manifolds that we call barycentric subspaces. They are implicitly defined as the locus of points which are weighted means of k+1 reference points.…
Dominant knots have isomorphic Seifert and Tait graphs.
problem Understanding knot dominance through graph isomorphism.
method Examined alternating knots and their Seifert and Tait graphs.
result Isomorphic Seifert and Tait graphs indicate dominant knots.
Flow Matching models help generative models stay within the subspace of real data.
problem How do generative models stay within the subspace of real data?
method Flow Matching models using a learned velocity field to transform a simple prior into a complex target distribution.
result Generated samples memorize real data points and represent the sample data subspace exactly.
New homology theory connects graph domination to subtle algebraic structures.
problem Understanding graph domination through algebraic homology.
method Interpreting überhomology as poset homology and showing its functorial properties.
result The Euler characteristic of bold homology equals the evaluation of the connected domination polynomial.
3-manifolds can be virtually dominated by maps of degree 8.
problem Understanding virtual domination of 3-manifolds.
method Proving existence of finite covers with specific map properties.
result Virtual 8-dominance of 3-manifolds.
In this letter, we consider two sets of observations defined as subspace signals embedded in noise and we wish to analyze the distance between these two subspaces. The latter entails evaluating the angles between the subspaces, an issue reminiscent of the well-known Procrustes problem. A Bayesian approach is investigat…
Give deterministic necessary and sufficient conditions to guarantee that if a subspace fits certain partially observed data from a union of subspaces, it is because such data really lies in a subspace. Furthermore, Give deterministic necessary and sufficient conditions to guarantee that if a subspace fits certain parti…
In this paper, we exhibit the tradeoffs between the (training) sample, computation and storage complexity for the problem of supervised classification using signal subspace estimation. Our main tool is the use of tensor subspaces, i.e. subspaces with a Kronecker structure, for embedding the data into lower dimensions. …
Odd-dimensional manifolds have contact maps of non-zero degree.
problem Contact domination in odd-dimensional manifolds.
method Proving existence of maps from tight contact manifolds.
result Existence of non-zero degree maps from Liouville-fillable but not Weinstein-fillable contact manifolds.