The paper argues for using Neyman orthogonal score for balancing in debiased machine learning.
problem Debiased machine learning requires a proper approach to balance covariates.
method The paper advocates for using Riesz regression with basis functions of X for balancing.
result Covariate balancing is only valid when the score-relevant regression error is a function of covariates alone.
DFSOS improves sparse discriminant analysis for high-dimensional data.
problem Sparse discriminant analysis in high-dimensional settings with feature selection.
method Deflation-Free Sparse Optimal Scoring (DFSOS) using Bregman iteration and orthogonality-constrained optimization.
result DFSOS achieves comparable or better classification accuracy than deflation-based methods.
Proposes a novel method for detecting novelty in multi-modal data.
problem Challenges in detecting novelty in high-dimensional, multi-modal data.
method Orthogonalized latent space for disentangling features and defining novelty score.
result Proposed method outperforms state-of-the-art algorithms in novelty detection.
ScoreMatchingRiesz improves debiased machine learning and policy effects estimation.
problem Improving debiased machine learning and policy effects estimation.
method Score matching and Riesz representer estimation.
result Estimates policy path for continuous treatments, improving interpretability.
A common method of generalizing binary to multi-class classification is the error correcting code (ECC). ECCs may be optimized in a number of ways, for instance by making them orthogonal. Here we test two types of orthogonal ECCs on seven different datasets using three types of binary classifier and compare them with t…
EigenVI uses orthogonal function expansions for efficient variational inference.
problem Efficiently approximate complex distributions in variational inference.
method EigenVI constructs variational approximations using orthogonal function expansions, minimizing Fisher divergence.
result EigenVI provides more accurate approximations than existing methods for Gaussian BBVI.
P-OCS detects OOD samples in a low-dimensional subspace, outperforming existing methods.
problem Efficient OOD detection for deep learning models in open-world environments.
method P-OCS operates in the orthogonal complement of the principal subspace, applying a single projected perturbation.
result P-OCS achieves state-of-the-art OOD detection with negligible computational cost and without requiring model retraining.
Paper develops robust econometric methods for staggered adoption studies.
problem Estimation challenges in event studies with staggered adoption.
method Design-first framework with exact probability limits, diagnostics, and orthogonal score constructions.
result Uniformly valid inference under restricted violations of parallel trends.
Statistical leverage scores emerged as a fundamental tool for matrix sketching and column sampling with applications to low rank approximation, regression, random feature learning and quadrature. Yet, the very nature of this quantity is barely understood. Borrowing ideas from the orthogonal polynomial literature, we in…
This work defines observation-specific explanations for black-box models.
problem Assigning importance to data points in black-box model predictions.
method Surrogate model construction using scattered data approximation and orthogonal matching pursuit.
result Validated approach on simulated and real-world datasets.
Many scientific questions require estimating the effects of continuous treatments. Outcome modeling and weighted regression based on the generalized propensity score are the most commonly used methods to evaluate continuous effects. However, these techniques may be sensitive to model misspecification, extreme weights o…
The paper develops methods for causal function estimation and inference with multiway clustered data.
problem Estimation and inference for causal functions under multiway clustering.
method Two-step procedure using machine learning for nuisance parameters and projection onto basis functions.
result Rejects the null hypothesis of uniformly zero effects and reveals heterogeneous treatment effects.
New method offsets DML's error-compounding issue and provides more stable causal parameter estimates.
problem Estimating ATE from observational data with robustness and stability.
method Robust Causal Learning (RCL) method to offset DML's deficiencies.
result RCL estimators are more stable and perform better than DML and traditional estimators.
FedSPDnet improves federated learning for SPD matrices, outperforming existing methods.
problem Federated learning for SPD matrices with orthogonality constraints.
method Two efficient aggregation strategies: ProjAvg and RLAvg, preserving geometric structure.
result FedSPDnet outperforms federated EEGnet in F1 score and robustness to federation and partial participation.
The paper uses double machine learning to estimate dynamic treatment effects robustly.
problem Estimating causal effects of dynamic treatments with time-varying covariates.
method Double machine learning with Neyman-orthogonal score functions for robustness.
result Asymptotic normality and n \sqrt{n} n -consistency of the estimators under specific conditions. Sparse Principal Component Analysis (sPCA) is a popular matrix factorization approach based on Principal Component Analysis (PCA) that combines variance maximization and sparsity with the ultimate goal of improving data interpretation. When moving from PCA to sPCA, there are a number of implications that the practition…
Unified framework for debiased machine learning using Riesz representer and Bregman divergence.
problem Estimating causal and structural parameters in machine learning.
method Generalized Riesz regression for fitting Riesz representer via Bregman divergence minimization.
result Automatic covariate balancing and Neyman orthogonality properties for debiased estimation.
This paper proposes a Lasso-type estimator for a high-dimensional sparse parameter identified by a single index conditional moment restriction (CMR). In addition to this parameter, the moment function can also depend on a nuisance function, such as the propensity score or the conditional choice probability, which we es…
The paper interprets diffusion models as gradient descent and proposes a new sampler.
problem Improving the efficiency and quality of diffusion models.
method Interprets diffusion models as gradient descent and proposes a new sampler.
result The new sampler achieves state-of-the-art FID scores and generates high quality samples.
A new stable similarity measure for time series using persistent homology.
problem Constructing a robust measure of time series similarity.
method Persistent homology for stability, bi-conditional periodicity score for similarity.
result Stability of the bi-conditional periodicity score under perturbations and dimension reduction.
A method for finding most influential sets reduces a complex problem to a sequence of simpler top- k k k problems.
problem Identifying most influential subsets in complex models.
method Reduces the problem to a sequence of top- k k k problems using Dinkelbach's method. result The method returns a globally optimal set for the univariate ratio objective, including partial linear models.
Chernozhukov, Chetverikov, Demirer, Duflo, Hansen, and Newey (2016) provide a generic double/de-biased machine learning (DML) approach for obtaining valid inferential statements about focal parameters, using Neyman-orthogonal scores and cross-fitting, in settings where nuisance parameters are estimated using a new gene…
Proposes a Bayesian framework for causal inference without explicit likelihood modeling.
problem Challenges in principled Bayesian inference for causal effects.
method Generalized Bayesian framework that places priors directly on causal estimands and updates using identification-driven loss functions.
result Yields generalized posteriors for causal effects with uncertainty quantification.
Framework disentangles deep feature uncertainty for efficient inference.
problem Inference-time uncertainty estimation for reliable decision-making.
method Uncertainty-Guided Inference-Time Selection framework.
result Significantly tighter prediction intervals and 60% compute reduction.
The study extends Jacobi-orthogonality to indefinite scalar product spaces.
problem Generalizing Jacobi-orthogonality to indefinite scalar product spaces.
method Comparing principles, investigating tensor relations, proving properties.
result Every quasi-Clifford tensor is Jacobi-orthogonal; certain tensors are Jacobi-dual or Osserman.
New characterization of Osserman tensors using Jacobi-orthogonality.
problem Characterizing Osserman tensors.
method Introducing Jacobi-orthogonality as a new potential characterization.
result Jacobi-orthogonal tensors are Osserman, and all known Osserman tensors are Jacobi-orthogonal.
A new method for learning manifolds efficiently using canonical basis functions.
problem Learning manifolds in high-dimensional data with efficient and distinct latent dimensions.
method Proposes a novel optimization objective to enforce a transformation matrix with a few prominent and non-degenerate basis functions.
result Demonstrates that minimizing the off-diagonal manifold metric elements ℓ 1 \ell_1 ℓ 1 -norm results in a more efficient latent space representation. Study isotropy groups for complex orthogonal and skew-symmetric matrices.
problem Understanding isotropy subgroups of orthogonal similarity transformations.
method Analysis of group structure of nonsingular block matrices.
result Group structure of isotropy subgroups related to block Toeplitz matrices.
New findings on Kähler manifolds restrict orthogonal coordinates existence.
problem Existence of orthogonal coordinates on Kähler manifolds.
method Algebraic and geometric techniques applied to Kähler manifolds.
result No nontrivial self-dual Kähler 4-manifolds or Ricci-flat Kähler 4-manifolds support orthogonal coordinates.
Constructs orthogonal coordinates in curved spaces.
problem Separating variables in curved spaces.
method Explicit construction of orthogonal coordinates and transformations.
result Explicit formulas for Killing tensors and Stäckel matrices.
OPT framework improves neural network generalization by learning an orthogonal transformation.
problem Improving neural network generalization.
method Orthogonal over-parameterized training (OPT) framework that minimizes hyperspherical energy.
result OPT framework provably minimizes hyperspherical energy and improves empirical generalization.
The paper studies surfaces in a bounded domain with orthogonal boundaries and proves curvature estimates.
problem Estimating the area of surfaces with orthogonal boundaries in a bounded domain.
method Weak formulation of orthogonality for curvature varifolds, classification of vanishing curvature varifolds.
result Existence of an orthogonal 2-varifold that minimizes L 2 L^2 L 2 curvature in the integer rectifiable class. Orthogonal random features approximate a Bessel kernel, offering sharper bounds than random Fourier features.
problem Approximating Gaussian kernel efficiently for large datasets.
method Use of Haar orthogonal matrices to construct orthogonal random features and analyze their bias and variance.
result Orthogonal random features approximate a Bessel kernel, not the Gaussian kernel, with sharper bounds.
Proposes FOAGP for efficient orthogonal effect decomposition of black-box computer experiments.
problem Challenges in sensitivity analysis of black-box computer experiments with complex, nonlinear functional outputs.
method Functional-output orthogonal additive Gaussian process (FOAGP) with conditional orthogonality constraint.
result Demonstrates effectiveness in orthogonal effect decomposition and variance decomposition through simulations and real-world application.
Novel prior for orthogonal functions improves functional component estimation.
problem Improving orthogonality in functional principal component analysis.
method Sequential adaptive priors for orthogonal functions using hierarchical conditionally normal distributions.
result Proposed prior leads to nearly orthogonal posterior estimates.
Method constructs orthogonal curvilinear coordinates in constant curvature spaces.
problem Creating orthogonal coordinates in spaces of constant curvature.
method Modification of Krichever's method for Euclidean space, applied to constant curvature spaces.
result Examples of orthogonal coordinate systems on the sphere and hyperbolic plane constructed.
We study the Chern-Simons partition function of orthogonal quantum group invariants, and propose a new orthogonal Labastida-Mariño-Ooguri-Vafa conjecture as well as degree conjecture for free energy associated to the orthogonal Chern-Simons partition function. We prove the degree conjecture and some interesting cases o…
The study examines how shallow neural nets converge to training samples or manifold points during diffusion.
problem Understanding when and how shallow neural nets converge to training samples or manifold points during diffusion.
method Analysis of shallow ReLU neural network denoisers trained with minimal ℓ 2 \ell^2 ℓ 2 norm, comparing score flow and diffusion flow. result Probability flow converges to training points, sums of training points, or manifold points, depending on the diffusion time scheduler.
A Clifford algebra model for M"obius geometry is presented. The notion of Ribaucour pairs of orthogonal systems in arbitrary dimensions is introduced, and the structure equations for adapted frames are derived. These equations are discretized and the geometry of the occuring discrete nets and sphere congruences is disc…
DONUT improves treatment effect estimation by enforcing orthogonality constraints.
problem Estimating treatment effects from observational data is challenging due to unobserved outcomes.
method DONUT uses a regularization framework that formalizes unconfoundedness as orthogonality, leading to deep orthogonal networks.
result DONUT outperforms state-of-the-art methods in estimating average treatment effects.
New convergence guarantees for learning with unknown nuisance parameters.
problem Learning problems with unknown nuisance parameters.
method Stochastic gradient optimization with Neyman orthogonality and approximately orthogonalized updates.
result Stochastic gradient algorithms can converge under conditions of nuisance parameters.
MuonEq improves training of matrix-valued parameters by rebalancing momentum before orthogonalization.
problem Training matrix-valued parameters with orthogonalized-update optimizers like Muon.
method MuonEq introduces three lightweight pre-orthogonalization equilibration schemes: two-sided row/column normalization (RC), row normalization (R), and column normalization (C).
result Row/column normalization acts as a zeroth-order surrogate for whitening and improves the geometry seen by orthogonalization.
Orthogonal initialization does not speed up training in ultra-wide neural networks.
problem Exploring the effect of orthogonal initialization on training speed in deep neural networks.
method Study of neural tangent kernel dynamics in FCNs and CNNs with orthogonal initialization.
result The NTK of orthogonally-initialized networks remains constant during training, suggesting no speedup in the NTK regime.
A new algorithm POGO optimizes thousands of orthogonal matrices efficiently.
problem Optimizing thousands of orthogonal constraints at scale is computationally expensive.
method Revisits Landing algorithm, uses modern adaptive optimizers, reduces hyperparameters.
result POGO optimizes thousands of orthogonal matrices in minutes, outperforming alternatives.
Batch normalization makes deep neural networks' representations increasingly orthogonal.
problem Orthogonality of deep neural network representations.
method Random linear transformations in successive batch-normalizations.
result Orthogonality of representations improves SGD performance.
The paper finds at least N orthogonal Finsler geodesic chords in a disk-like manifold.
problem Existence and multiplicity of orthogonal Finsler geodesic chords in a disk-like manifold.
method Study of Finsler geodesic chords under reversibility assumption.
result At least N orthogonal Finsler geodesic chords found in a disk-like manifold.
Optimizing over the set of orthogonal matrices is a central component in problems like sparse-PCA or tensor decomposition. Unfortunately, such optimization is hard since simple operations on orthogonal matrices easily break orthogonality, and correcting orthogonality usually costs a large amount of computation. Here we…
The paper analyzes the latent geometry of generative diffusion models.
problem The manifold overfitting phenomenon in generative models.
method Statistical physics approach to analyze the spectrum of eigenvalues and singular values of the Jacobian of the score function.
result Three distinct qualitative phases during the generative process: trivial, manifold coverage, and consolidation phases.