DONUT improves treatment effect estimation by enforcing orthogonality constraints.
problem Estimating treatment effects from observational data is challenging due to unobserved outcomes.
method DONUT uses a regularization framework that formalizes unconfoundedness as orthogonality, leading to deep orthogonal networks.
result DONUT outperforms state-of-the-art methods in estimating average treatment effects.
Orthogonal random features approximate a Bessel kernel, offering sharper bounds than random Fourier features.
problem Approximating Gaussian kernel efficiently for large datasets.
method Use of Haar orthogonal matrices to construct orthogonal random features and analyze their bias and variance.
result Orthogonal random features approximate a Bessel kernel, not the Gaussian kernel, with sharper bounds.
A notion of orthogonality in multisymplectic geometry has been developed by Cantrijn, Ibort and de León and used by many authors. In this paper, we review this concept and propose a new type of orthogonality in multisymplectic geometry; we prove a number of results regarding this orthogonality and its associated subspa…
A recent theoretical analysis shows the equivalence between non-negative matrix factorization (NMF) and spectral clustering based approach to subspace clustering. As NMF and many of its variants are essentially linear, we introduce a nonlinear NMF with explicit orthogonality and derive general kernel-based orthogonal m…
Optimizing over the set of orthogonal matrices is a central component in problems like sparse-PCA or tensor decomposition. Unfortunately, such optimization is hard since simple operations on orthogonal matrices easily break orthogonality, and correcting orthogonality usually costs a large amount of computation. Here we…
AuON is a linear-time optimizer that improves upon Muon's performance without approximate orthogonal matrices.
problem High memory and computational costs of orthogonal momentum updates.
method AuON uses normalized nonlinear scaling and a 'emergency brake' to handle exploding attention logits.
result AuON achieves strong performance without approximate orthogonal matrices, preserving structural alignment and reconditioning.
Based on the orthogonal Labastida-Mari{ñ}o-Ooguri-Vafa conjecture made by L. Chen & Q. Chen [5], we derive an infinite product formula for Chern-Simons partition functions, which generalizes the Liu-Peng's [19] recent results to the orthogonal case. Symmetry property of this new infinite product structure is also discu…
New method evaluates feature interactions using orthogonal variance decomposition.
problem Feature selection fails to account for interactions between features.
method Orthogonal variance decomposition to evaluate feature subsets considering interactions.
result Our method accurately identifies relevant features and improves model accuracy.
L1-orthogonal regularization improves decision tree explainability of deep neural networks.
problem Lack of explainability in deep neural networks.
method L1-orthogonal regularization during training of decision trees.
result Decision trees closely approximate trained deep neural networks with improved accuracy and fidelity.
New ONMF model with NCP improves clustering efficiency.
problem Improving clustering performance using ONMF model.
method Transformed ONMF into norm-based non-convex constraints and applied NCP approach.
result Proposed NCP methods efficiently solve the ONMF clustering problem.
Proposes FOAGP for efficient orthogonal effect decomposition of black-box computer experiments.
problem Challenges in sensitivity analysis of black-box computer experiments with complex, nonlinear functional outputs.
method Functional-output orthogonal additive Gaussian process (FOAGP) with conditional orthogonality constraint.
result Demonstrates effectiveness in orthogonal effect decomposition and variance decomposition through simulations and real-world application.
Semi-parametric framework for nonlinear system identification
problem Nonlinear system identification
method Orthogonal Gaussian process regression
result Interpretable models from incomplete physics
VRSGT algorithm reduces orthogonality constraints in decentralized optimization.
problem Decentralized optimization with orthogonality constraints.
method VRSGT algorithm with variance reduction and orthogonal techniques.
result VRSGT achieves convergence rate of O(1 / k) for orthogonality constraints.
A new optimizer preserves orthogonality constraints on matrices efficiently.
problem Optimization on Stiefel manifold with orthogonality constraints.
method Interplay between continuous and discrete dynamics leading to a gradient-based optimizer with momentum.
result The method optimizes matrices on Stiefel manifold efficiently and accurately.
MuonEq improves training of matrix-valued parameters by rebalancing momentum before orthogonalization.
problem Training matrix-valued parameters with orthogonalized-update optimizers like Muon.
method MuonEq introduces three lightweight pre-orthogonalization equilibration schemes: two-sided row/column normalization (RC), row normalization (R), and column normalization (C).
result Row/column normalization acts as a zeroth-order surrogate for whitening and improves the geometry seen by orthogonalization.
A new method for optimizing neural networks with orthogonal constraints.
problem Optimizing neural networks with orthogonal constraints.
method Parametrization using the exponential map to transform constrained optimization into unconstrained.
result Faster, more accurate, and stable convergence in RNNs with orthogonal recurrent weights.
LOFT separates subspace rotation and transformation for orthogonal fine-tuning.
problem Conflating subspace rotation and transformation in orthogonal fine-tuning.
method LOFT explicitly separates subspace rotation and transformation, using task-aware support selection.
result LOFT recovers principal-subspace orthogonal adaptation and improves efficiency-performance trade-off.
Derives new orthogonal coordinates for evolving surfaces and curves.
problem Accounting for geometric effects in boundary layer asymptotics.
method Elementary derivation of orthogonal signed-distance coordinates.
result Provides vector calculus identities for these coordinates.
SOFARI improves inference on multi-task learning latent factors.
problem Challenges in precise inference on multi-task learning latent factor matrices.
method High-dimensional manifold-based Neyman near-orthogonality inference on Stiefel manifold structure.
result Easy-to-use bias-corrected estimators for latent factor vectors and singular values with asymptotic normal distributions.
Study uses orthogonal polynomials to solve option pricing equations.
problem Solving complex option pricing equations for various models.
method Galerkin-based method with Hermite and Laguerre polynomials.
result Compared solutions to existing semi-closed formulas.
New algorithms learn sparse set functions in non-orthogonal Fourier bases.
problem Learning sparse set functions in non-orthogonal Fourier bases.
method Novel algorithms using non-orthogonal Fourier transforms.
result At most nk−klog2k+k queries for k non-zero Fourier coefficients. Paper improves feature selection accuracy using transfer learning.
problem Improving feature selection accuracy in information criteria-based methods.
method Proposes TLCp, a transfer learning procedure based on Mallows' Cp.
result TLCp outperforms conventional Cp in accuracy and stability.
New differential geometry perspective on orthogonal RNNs.
problem Mitigating exploding and vanishing gradients in RNNs.
method Using tools from differential geometry, parameterizing vector fields via directional derivatives of scalar functions.
result Our approach achieves comparable or better results on benchmark tasks.
A l1-norm penalized orthogonal forward regression (l1-POFR) algorithm is proposed based on the concept of leaveone- out mean square error (LOOMSE). Firstly, a new l1-norm penalized cost function is defined in the constructed orthogonal space, and each orthogonal basis is associated with an individually tunable regulari…
New method for sampling orthogonal matrices using Hamiltonian Monte-Carlo.
problem Sampling from posterior distributions of orthogonal matrices in Bayesian models.
method Proposes a new sampling scheme based on Hamiltonian Monte-Carlo and Riemannian optimization.
result New method is comparable or faster in time per iteration and more sample-efficient than conventional methods.
EigenVI uses orthogonal function expansions for efficient variational inference.
problem Efficiently approximate complex distributions in variational inference.
method EigenVI constructs variational approximations using orthogonal function expansions, minimizing Fisher divergence.
result EigenVI provides more accurate approximations than existing methods for Gaussian BBVI.
New method learns complete orthogonal dictionary from samples with theoretical guarantees and efficiency.
problem Learning a complete orthogonal dictionary from sparsely generated signals.
method Maximizes the \(\ell^4\)-norm over the orthogonal group, using a novel algorithm based on matching, stretching, and projection (MSP).
result The MSP algorithm provably converges locally at a superlinear (cubic) rate and is significantly more efficient than existing methods.
Paper optimizes tensor deflation for non-orthogonal signals.
problem Recovering low-rank signals from noisy tensors with correlated components.
method Developed an asymptotic analysis and optimized deflation procedure using random tensor theory.
result Proposed an efficient tensor deflation algorithm that optimizes a parameter introduced in the deflation mechanism.
SpecNet2 improves spectral embedding without orthogonalization, achieving better performance and efficiency.
problem Improving spectral embedding methods for better performance and efficiency.
method Optimizes an equivalent objective of the eigen-problem without orthogonalization, allowing separate row and column sampling.
result Local and global convergence of the new objective using batch-based gradient descent is proven, and improved performance and efficiency are demonstrated on simulated and image datasets.
A method for interpreting SVMs using polynomial kernels, revealing model complexity.
problem Interpreting SVMs built with truncated orthogonal polynomial kernels.
method Orthogonal Representation Contribution Analysis (ORCA) with normalized Orthogonal Kernel Contribution (OKC) indices.
result The method reveals structural aspects of model complexity not captured by predictive accuracy.
A machine learning method selects optimal orthonormal bases for functional data analysis.
problem Lack of formal criteria for choosing initial orthonormal bases in functional data methods.
method Proposes a machine learning algorithm to learn and place knots for efficient orthogonal spline bases (splinets).
result Demonstrates efficiency, especially for sparse functional data and complex physical systems.
Principal component analysis (PCA) is an unsupervised method for learning low-dimensional features with orthogonal projections. Multilinear PCA methods extend PCA to deal with multidimensional data (tensors) directly via tensor-to-tensor projection or tensor-to-vector projection (TVP). However, under the TVP setting, i…
NS-RGS improves orthogonal group synchronization with faster convergence.
problem Orthogonal group synchronization from pairwise measurements.
method Newton-Schulz iteration for Riemannian gradient optimization.
result NS-RGS achieves linear convergence and near-optimal accuracy.
Efficient algorithm for orthogonal canonical correlation analysis (OCCA).
problem Solving the OCCA problem with orthogonality constraints.
method Sub-maximization problem with self-consistent-field (SCF) iteration for trace-fractional structure and orthogonal linear projections.
result Proposed algorithm converges globally to a KKT point and is more efficient.
Method estimates heterogeneous causal effects on networks using orthogonal learning.
problem Challenges in estimating causal effects on networks due to treatment effects on both treated and neighbors, and network homophily.
method Two-stage orthogonal learning framework: first stage uses graph neural networks for nuisance components, second stage residualizes and interpretable attention-based model for causal effects.
result Improves heterogeneous effect estimation and supports interpretable analyses.
Pion optimizes LLMs by preserving weight matrix singular values.
problem Training large language models (LLMs) with standard optimizers leads to unstable weight matrices.
method Pion uses orthogonal transformations to update weight matrices, preserving their singular values.
result Pion offers a stable alternative to standard optimizers for LLM pretraining and finetuning.
We solve the equivalence problem for the orthogonally separable webs on the three-sphere under the action of the isometry group. This continues a classical project initiated by Olevsky in which he solved the corresponding canonical forms problem. The solution to the equivalence problem together with the results by Olev…
Orthogonal Wasserstein GANs improve image quality without gradient norm regularization.
problem Wasserstein-GANs' gradient norm regularization limits the distribution's fidelity.
method Substituted gradient norm regularization with orthogonality constraints on weight matrices.
result Orthogonal Wasserstein GANs achieve better image quality and generalization.
Method preserves correlations in synthetic data.
problem Preserving dependence structure of original data.
method Orthogonal Procrustes problem for restoring Pearson correlation.
result Restores Pearson correlation structure while preserving feature distributions and downstream tasks performance.
KOOW method provides optimal covariate balance for continuous treatments.
problem Estimating effects of continuous treatments with robustness to model misspecification and extreme weights.
method Kernel Optimal Orthogonality Weighting (KOOW) using convex optimization.
result KOOW provides optimal covariate balance and controls for extreme weights.
A new quasi-Newton method tackles NMF with transform learning on orthogonal manifolds.
problem Efficiently learning transforms for NMF in non-convex optimization on orthogonal manifolds.
method Derives a quasi-Newton method on the orthogonal matrix manifold using sparse approximations of the Hessian.
result Outperforms state-of-the-art methods by orders of magnitude in experiments on synthetic and real audio data.
Improves model predictability by mixing forecasts and orthogonalizing models.
problem Redundant models contaminate model space and degrade predictive performance.
method Principal Component Analysis for model orthogonalization in Bayesian forecast mixing.
result Better prediction accuracy and excellent uncertainty quantification.
Orthogonal deep models defend against black-box attacks by ensuring internal representations are nearly orthogonal.
problem Vulnerability of deep learning models to black-box adversarial attacks.
method Introduce a gradient regularization scheme to encourage deep models' internal representations to be orthogonal to another model's.
result Orthogonal deep models significantly boost robustness against transferable black-box adversarial attacks.
The study extends Jacobi-orthogonality to indefinite scalar product spaces.
problem Generalizing Jacobi-orthogonality to indefinite scalar product spaces.
method Comparing principles, investigating tensor relations, proving properties.
result Every quasi-Clifford tensor is Jacobi-orthogonal; certain tensors are Jacobi-dual or Osserman.
Proposes a new method for analyzing multimodal neuroimaging data.
problem Combining interpretability and flexibility in multimodal data analysis.
method Orthogonalized kernel debiased machine learning approach.
result Established consistency and asymptotic normality of the estimated primary parameter.
New characterization of Osserman tensors using Jacobi-orthogonality.
problem Characterizing Osserman tensors.
method Introducing Jacobi-orthogonality as a new potential characterization.
result Jacobi-orthogonal tensors are Osserman, and all known Osserman tensors are Jacobi-orthogonal.
Proposes a new RNN structure to improve expressivity without sacrificing stability.
problem Exploding and vanishing gradient problems in RNNs and reduced expressivity.
method Introduces a non-normal RNN structure using Schur decomposition and splitting.
result Enhances expressivity while maintaining stability and training speed.
DFSOS improves sparse discriminant analysis for high-dimensional data.
problem Sparse discriminant analysis in high-dimensional settings with feature selection.
method Deflation-Free Sparse Optimal Scoring (DFSOS) using Bregman iteration and orthogonality-constrained optimization.
result DFSOS achieves comparable or better classification accuracy than deflation-based methods.