DONUT improves treatment effect estimation by enforcing orthogonality constraints.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A l1-norm penalized orthogonal forward regression (l1-POFR) algorithm is proposed based on the concept of leaveone- out mean square error (LOOMSE). Firstly, a new l1-norm penalized cost function is defined in the constructed orthogonal space, and each orthogonal basis is associated with an individually tunable regulari…
This paper seeks to answer the question: as the (near-) orthogonality of weights is found to be a favorable property for training deep convolutional neural networks, how can we enforce it in more effective and easy-to-use ways? We develop novel orthogonality regularizations on training deep CNNs, utilizing various adva…
One obstacle that so far prevents the introduction of machine learning models primarily in critical areas is the lack of explainability. In this work, a practicable approach of gaining explainability of deep artificial neural networks (NN) using an interpretable surrogate model based on decision trees is presented. Sim…
In this paper, we study spherical images of the modified orthogonal vector fields and Darboux vector of a regular curve which lies on the unit sphere in Euclidean 3-space.
Orthogonal deep models defend against black-box attacks by ensuring internal representations are nearly orthogonal.
Wasserstein-GANs have been introduced to address the deficiencies of generative adversarial networks (GANs) regarding the problems of vanishing gradients and mode collapse during the training, leading to improved convergence behaviour and improved image quality. However, Wasserstein-GANs require the discriminator to be…
In text classification, the problem of overfitting arises due to the high dimensionality, making regularization essential. Although classic regularizers provide sparsity, they fail to return highly accurate models. On the contrary, state-of-the-art group-lasso regularizers provide better results at the expense of low s…
TANGOS improves neural network performance on tabular data by encouraging neuron specialization.
We study conditions for the integrability of the distribution defined on a regular Poisson manifold as the orthogonal complement (with respect to some (pseudo)-Riemannian metric) to the tangent spaces of the leaves of a symplectic foliation. Examples of integrability and non-integrability of this distribution are provi…
Multi-head attention mechanism is capable of learning various representations from sequential data while paying attention to different subsequences, e.g., word-pieces or syllables in a spoken word. From the subsequences, it retrieves richer information than a single-head attention which only summarizes the whole sequen…
A parametric manifold can be viewed as the manifold of orbits of a (regular) foliation of a manifold by means of a family of curves. If the foliation is hypersurface orthogonal, the parametric manifold is equivalent to the 1-parameter family of hypersurfaces orthogonal to the curves, each of which inherits a metric and…
Study spectral flow on a warped cylinder with special boundary conditions.
The paper proves smoothness of almost-minimizers' boundaries near the free boundary.
New method relaxes PCA orthogonality constraints using explained variance of correlated components.
Through Cayley and Langlands type correspondences, we give a geometric description of the moduli spaces of real orthogonal and symplectic Higgs bundles of any signature in the regular fibres of the Hitchin fibration. As applications of our methods, we complete the concrete abelianization of real slices corresponding to…
Regularization techniques are widely used to improve the generality, robustness, and efficiency of deep convolutional neural networks (DCNNs). In this paper, we propose a novel approach of regulating DCNN convolutional kernels by a structured filter bank. Comparing with the existing regularization methods, such as $\el…
A recent theoretical analysis shows the equivalence between non-negative matrix factorization (NMF) and spectral clustering based approach to subspace clustering. As NMF and many of its variants are essentially linear, we introduce a nonlinear NMF with explicit orthogonality and derive general kernel-based orthogonal m…
The theory of slice regular functions of a quaternion variable is applied to the study of orthogonal complex structures on domains Ω of R^4. When Ω is a symmetric slice domain, the twistor transform of such a function is a holomorphic curve in the Klein quadric. The case in which Ω is the complement of a parabola is st…
PROD method improves high-dimensional regression by handling strong correlations.
OMIC improves matrix completion with orthonormal side information and nuclear-norm regularization.
We consider the problem of sampling from posterior distributions for Bayesian models where some parameters are restricted to be orthogonal matrices. Such matrices are sometimes used in neural networks models for reasons of regularization and stabilization of training procedures, and also can parameterize matrices of bo…
Stability inequalities for specific solutions in high dimensions.
Unified framework for rigidity results on -manifolds.
Subspace clustering methods based on , or nuclear norm regularization have become very popular due to their simplicity, theoretical guarantees and empirical success. However, the choice of the regularizer can greatly impact both theory and practice. For instance, regularization is guaranteed t…
We prove ultradifferentiable Chevelley restriction theorems for a wide range of ultradifferentiable classes. As a special case we find that isotropic functions, i.e., functions defined on the vector space of real symmetric matrices invariant under the action of the special orthogonal group by conjugation, possess some …
SVD training reduces DNN rank and computation load without SVD per step.
Echo state network (ESN) is viewed as a temporal non-orthogonal expansion with pseudo-random parameters. Such expansions naturally give rise to regressors of various relevance to a teacher output. We illustrate that often only a certain amount of the generated echo-regressors effectively explain the variance of the tea…
We compute approximate solutions to L0 regularized linear regression using L1 regularization, also known as the Lasso, as an initialization step. Our algorithm, the Lass-0 ("Lass-zero"), uses a computationally efficient stepwise search to determine a locally optimal L0 solution given any L1 regularization solution. We …
Deep neural networks are a promising approach towards multi-task learning because of their capability to leverage knowledge across domains and learn general purpose representations. Nevertheless, they can fail to live up to these promises as tasks often compete for a model's limited resources, potentially leading to lo…
Two-layer networks trained on low-dimensional subspaces are vulnerable to adversarial examples.
The paper shows how gradient flow on over-parametrized tensor decomposition behaves like deflation.
Study Finsler metrics with vanishing Landsberg curvature.
TGCCA analyzes higher-order tensors using orthogonal rank-R CP decomposition.
DFSOS improves sparse discriminant analysis for high-dimensional data.
We examine Higgs bundles for non-compact real forms of SO(4,C) and the isogenous complex group SL(2,C)XSL(2,C). This involves a study of non-regular fibers in the corresponding Hitchin fibrations and provides interesting examples of non-abelian spectral data.
New approach to portfolio optimization shows entropy regularization is ineffective.
Distance metric learning (DML), which learns a distance metric from labeled "similar" and "dissimilar" data pairs, is widely utilized. Recently, several works investigate orthogonality-promoting regularization (OPR), which encourages the projection vectors in DML to be close to being orthogonal, to achieve three effect…
The paper develops methods to reduce deployment risk under dynamic covariate shifts.
This paper proposes a Lasso-type estimator for a high-dimensional sparse parameter identified by a single index conditional moment restriction (CMR). In addition to this parameter, the moment function can also depend on a nuisance function, such as the propensity score or the conditional choice probability, which we es…
Multivariate Analysis (MVA) comprises a family of well-known methods for feature extraction that exploit correlations among input variables of the data representation. One important property that is enjoyed by most such methods is uncorrelation among the extracted features. Recently, regularized versions of MVA methods…
Bayesian Markowitz portfolio problem shows entropy regularization is ineffective.
We define systems of pre-extremals for the energy functional of regular rheonomic Lagrange manifolds and show how they induce well-defined Hamilton orthogonal nets. Such nets have applications in the modelling of e.g. wildfire spread under time- and space-dependent conditions. The time function inherited from such a Ha…
Study submanifolds in hyperbolic space, focusing on their boundary and Laplace operator.
Proposes nAIPW for robust ATE estimation using neural networks.
A simple regularization technique speeds up training of Neural ODEs.
New algorithm makes machine learning fairer by removing bias from data.
Overparameterized models improve performance in sequential learning tasks.