Memory-efficient optimizers fail to track a subspace, leading to unpredictable model performance.
problem Memory-efficient optimizers fail to track a subspace, leading to unpredictable model performance.
method Analyzing the behavior of memory-efficient optimizers like GaLore, which project gradients onto a rank-r subspace recomputed every T steps.
result Memory-efficient optimizers fail to track a subspace, leading to unpredictable model performance.
Study optimizes step size for Metropolis algorithm in non-identifiable cases.
problem Optimizing step size for Metropolis algorithm in non-identifiable models.
method Analytical derivation of average acceptance rate for non-identifiable cases.
result Developed optimization principle for step size based on average acceptance rate.
Latent feature models (LFM)s are widely employed for extracting latent structures of data. While offering high, parameter estimation is difficult with LFMs because of the combinational nature of latent features, and non-identifiability is a particularly difficult problem when parameter estimation is not unique and ther…
New research shows LLMs can't be explained by statistical generalization alone.
problem Understanding why large language models (LLMs) perform well despite statistical generalization limitations.
method Examined the non-identifiability of AR probabilistic models and their implications for LLMs.
result Non-identifiability of LLMs leads to different behaviors and requires a separate theoretical explanation.
This paper tackles non-identifiability in financial market simulations using multivariate time series data.
problem Non-identifiability issue in social simulation models, leading to indistinguishable simulated time series data.
method Proposes a maximization-based aggregation function to form a new calibration objective function using multiple time series features.
result Significant improvements in alleviating non-identifiability and achieving higher simulation fidelity.
Hypothesis testing in singular models is fundamentally about identifiable vs. non-identifiable parameters.
problem Testing in singular models is inherently problematic due to non-identifiability and degeneracy of Fisher information.
method Formalized the overlap obstruction and showed that hypotheses over non-identifiable parameters are untestable, while those over identifiable parameters reduce to classical testing.
result Hypotheses over non-identifiable parameters are untestable, while those over identifiable parameters reduce to classical testing.
New method mitigates bias in BNN+LVs due to non-identifiability.
problem Non-identifiability in BNN+LVs causes biased posterior mode.
method Developed novel inference procedure to mitigate bias.
result Inference method yields high-quality predictions and uncertainty estimates.
Variational autoencoders often collapse, showing latent variables are non-identifiable.
problem Posterior collapse in variational autoencoders due to non-identifiable latent variables.
method Proves latent variable non-identifiability causes posterior collapse. Proposes latent-identifiable models using Brenier maps and input convex neural networks.
result Latent-identifiable models resolve posterior collapse and provide meaningful representations.
New method warns of counterfactual non-identifiability in DSCMs.
problem Counterfactual inference from observational data is non-identifiable even without unobserved confounding.
method Prove counterfactual identifiability for monotonic generation mechanisms, provide impossibility result for general mechanisms, propose method for estimating worst-case errors.
result Non-identifiability of counterfactual inference from observational data, even in absence of unobserved confounding.
A new method resolves non-identifiability in reward modeling using anchor labels.
problem Non-identifiability in reward modeling from pairwise preferences alone.
method Anchor-guided Variance-aware Reward Modeling (AVRM) framework.
result AVRM resolves non-identifiability and improves reward modeling performance.
IMA addresses non-identifiability in nonlinear ICA by assuming orthogonal Jacobian columns.
problem Non-identifiability in nonlinear ICA.
method IMA assumes orthogonal Jacobian columns and extends to manifold settings.
result IMA circumvents non-identifiability issues and can be beneficial for higher-dimensional observations.
Overparametrized neural networks retain significant epistemic uncertainty even with sufficient data.
problem Epistemic uncertainty in overparametrized neural networks persists despite model identifiability.
method Analysis of non-identifiability and characterization of residual uncertainty in one-hidden-layer ReLU networks.
result Substantial parameter uncertainty remains even when the underlying function is fully identified.
New method uses logical relations to derive bounds and inequality constraints from causal models.
problem Recovering bounds and inequality constraints from unobserved confounding.
method Using rules of probability and restrictions on counterfactuals implied by causal graphical models.
result Powerful method to recover known and novel bounds and constraints.
Analysis of DPPs and k-DPPs via spectral decomposition reveals identifiable parameters and non-identifiability gaps.
problem Identifying parameters of DPPs and k-DPPs through spectral decomposition.
method Spectral decomposition of the covariance matrix, analysis of invariances, and counting arguments.
result Identifiability of parameters changes fundamentally for k-DPPs, with specific invariances and non-identifiability gaps.
Solves parameter non-identifiability in Bayesian LTI system identification.
problem Parameter non-identifiability in standard Bayesian approaches for LTI system identification.
method Embedding canonical forms of LTI systems within the Bayesian framework.
result Unlocking the use of meaningful priors and robust uncertainty estimates.
This work explores how overparametrization and priors affect Bayesian neural network posteriors.
problem Symmetries, non-identifiabilities, and weight-space priors fragment and inflate BNN posteriors.
method We study the interplay between overparametrization and priors in BNN posteriors, deriving key phenomena and validating through experiments.
result Overparametrization induces structured, prior-aligned weight posterior distributions.
Neural networks can learn relationships that traditional models cannot.
problem Identifying factors that differentiate neural networks from traditional models.
method Proving non-identifiability of neural networks compared to smooth parametric models.
result Neural networks can learn nontrivial relationships that traditional models cannot.
New framework estimates treatment effects based on preferences.
problem Estimating treatment effects with flexible outcomes.
method Preference-based Conditional Treatment Effect (CPTE) framework.
result CPTE provides interpretable targets and new identifiability conditions.
A new method for binary ICA using non-stationary sources.
problem Independent component analysis of binary data.
method Linear mixing model in latent space, followed by binary observation model with non-stationary sources.
result Proves non-identifiability with few observed variables but identifies with more variables.
Proposes efficient bounds for causal effect estimation under weak confounding.
problem Estimating causal effects with weakly confounded variables.
method Develops an efficient linear program to derive upper and lower bounds on causal effect under small entropy of unobserved confounders.
result Bounds are consistent and tighter for weakly confounded variables.
Transformers without skip connections collapse token representations to a single direction.
problem Rapid convergence of token representations to a single direction in self-attention-only Transformers.
method Analysis of layer normalization, residual connections, and multi-head attention mechanisms.
result Residual connections prevent rank collapse in real Transformers, while MLPs generate new feature directions.
The Rashomon effect shows many models can perform similarly, explored in this paper.
problem Why do many models perform similarly in machine learning?
method Categorized causes into statistical, structural, and procedural sources.
result Structural multiplicity persists and cannot be resolved without additional assumptions.
Proposes clustering and pruning to simplify causal data fusion models.
problem Combining observational and experimental data to identify causal effects.
method Generalizes pruning and clustering operations for multiple data sources.
result Derives conditions for inferring causal effects from simplified models.
POSCMs extend SCMs for causal modeling with latent contexts.
problem Causal modeling with latent contexts and endogenous mechanisms.
method Kolmogorov-Arnold-Sprecher edge-functional decomposition for explicit parametrization.
result Identifiability of structure and mechanisms under latent context.
Nonnegative matrix factorization (NMF) is a popular dimension reduction technique that produces interpretable decomposition of the data into parts. However, this decompostion is not generally identifiable (even up to permutation and scaling). While other studies have provide criteria under which NMF is identifiable, we…
We investigate the non-identifiability issues associated with bidirectional adversarial training for joint distribution matching. Within a framework of conditional entropy, we propose both adversarial and non-adversarial approaches to learn desirable matched joint distributions for unsupervised and supervised tasks. We…
Unified framework for singular statistical models using observable charts.
problem Non-identifiability and breakdown of classical asymptotic theory in singular models.
method Invariant framework based on observable charts to define local coordinate systems in model space.
result Observable order provides a lower bound on KL divergence vanishing rate in singular models.
The paper proposes a new method to calibrate multiple computer models simultaneously.
problem Calibrating multiple computer models one at a time is inefficient.
method Developed a probabilistic framework using customized neural networks.
result Simultaneous calibration improves predictive accuracy but can be non-identifiable in high dimensions.
Paper tackles non-identifiability of mixture models in partial order datasets.
problem Non-identifiability of mixture models in datasets with partial orders.
method Proved non-identifiability conditions and proposed GMM algorithms.
result GMM algorithms for learning mixtures of two Plackett-Luce models are consistent.
We introduce thermodynamic response functions for singular Bayesian models.
problem Singular Bayesian models violate regular asymptotics due to non-identifiability and degenerate Fisher geometry.
method Posterior tempering induces thermodynamic response functions, linking WAIC, WBIC, and singular fluctuation.
result WAIC, WBIC, and singular fluctuation are unified within a thermodynamic response framework.
Single sample estimation for hard-constrained models like SAT and coloring problems.
problem Estimating parameters of Markov Random Fields with hard constraints using a single sample.
method Pseudo-likelihood estimator with coupling techniques.
result Single-sample estimation is not always possible for hard constraints, and existence of an estimator is related to satisfiability.
Paper bounds subspace estimator error from noisy projections.
problem Estimating subspaces from noisy data.
method Derives perturbation bound on optimal subspace estimator.
result Fundamental result with implications in matrix completion and clustering.
PCA is one of the most widely used dimension reduction techniques. A related easier problem is "subspace learning" or "subspace estimation". Given relatively clean data, both are easily solved via singular value decomposition (SVD). The problem of subspace learning or PCA in the presence of outliers is called robust su…
Union of Subspaces (UoS) is a popular model to describe the underlying low-dimensional structure of data. The fine details of UoS structure can be described in terms of canonical angles (also known as principal angles) between subspaces, which is a well-known characterization for relative subspace positions. In this pa…
Unified framework for SGMoE resolves estimation and selection issues.
problem Non-identifiability, coupled differential relations, and tight coupling in softmax-Gated models.
method Unified statistical framework with Voronoi-type loss functions and dendrograms of mixing measures.
result Consistent selection of the number of experts without model sweeps, optimal parameter rates under overfitting.
Study Sp(n)-orbits in complex and Σ-complex subspaces of Hermitian quaternionic vector spaces.
problem Characterize Sp(n)-orbits in Grassmannians of complex and Σ-complex subspaces. method Decompose subspaces into 4-dimensional complex addends and 2-dimensional totally complex subspace. Use properties of isoclinic subspaces and principal angles.
result Determine full set of invariants for Sp(n)-orbits in GrR(2k,4n). In subspace clustering, a group of data points belonging to a union of subspaces are assigned membership to their respective subspaces. This paper presents a new approach dubbed Innovation Pursuit (iPursuit) to the problem of subspace clustering using a new geometrical idea whereby subspaces are identified based on the…
We consider the problem of detecting whether a tensor signal having many missing entities lies within a given low dimensional Kronecker-Structured (KS) subspace. This is a matched subspace detection problem. Tensor matched subspace detection problem is more challenging because of the intertwined signal dimensions. We s…
Paper shows affine constraint is unnecessary for high-dimensional data.
problem The necessity of an affine constraint in affine subspace clustering.
method Theoretical and empirical analysis of conditions for correctness of affine subspace clustering methods.
result Affine constraint has negligible effect on clustering performance for high-dimensional data.
A low-rank transformation learning framework for subspace clustering and classification is here proposed. Many high-dimensional data, such as face images and motion sequences, approximately lie in a union of low-dimensional subspaces. The corresponding subspace clustering problem has been extensively studied in the lit…
This paper investigates the generalization of Principal Component Analysis (PCA) to Riemannian manifolds. We first propose a new and general type of family of subspaces in manifolds that we call barycentric subspaces. They are implicitly defined as the locus of points which are weighted means of k+1 reference points.…
Flow Matching models help generative models stay within the subspace of real data.
problem How do generative models stay within the subspace of real data?
method Flow Matching models using a learned velocity field to transform a simple prior into a complex target distribution.
result Generated samples memorize real data points and represent the sample data subspace exactly.
In this letter, we consider two sets of observations defined as subspace signals embedded in noise and we wish to analyze the distance between these two subspaces. The latter entails evaluating the angles between the subspaces, an issue reminiscent of the well-known Procrustes problem. A Bayesian approach is investigat…
Give deterministic necessary and sufficient conditions to guarantee that if a subspace fits certain partially observed data from a union of subspaces, it is because such data really lies in a subspace. Furthermore, Give deterministic necessary and sufficient conditions to guarantee that if a subspace fits certain parti…
In this paper, we exhibit the tradeoffs between the (training) sample, computation and storage complexity for the problem of supervised classification using signal subspace estimation. Our main tool is the use of tensor subspaces, i.e. subspaces with a Kronecker structure, for embedding the data into lower dimensions. …
A new method for faster optimization in high dimensions.
problem Slow convergence in high-dimensional optimization problems.
method Subspace cubic regularized Newton method within Krylov subspace.
result Achieves a dimension-independent convergence rate of O(1/mk + 1/k^2).
We assume data sampled from a mixture of d-dimensional linear subspaces with spherically symmetric distributions within each subspace and an additional outlier component with spherically symmetric distribution within the ambient space (for simplicity we may assume that all distributions are uniform on their correspondi…
An elliptic theory is constructed for operators acting in subspaces defined via odd pseudodifferential projections. Subspaces of this type arise as Calderon subspaces for first order elliptic differential operators on manifolds with boundary, or as spectral subspaces for self-adjoint elliptic differential operators of …