A new method improves continual learning by replaying pseudo data and using orthogonal weight modification.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
OrthoGrad improves neural calibration by constraining gradient updates orthogonally.
Method constructs orthogonal curvilinear coordinates in constant curvature spaces.
One reflection suffices for orthogonal weights, reducing GPU usage.
The popular Alternating Least Squares (ALS) algorithm for tensor decomposition is efficient and easy to implement, but often converges to poor local optima---particularly when the weights of the factors are non-uniform. We propose a modification of the ALS approach that is as efficient as standard ALS, but provably rec…
AGI modifies utility function to cooperate, conflicting with orthogonality thesis.
Introduces new curvature concept for Kähler manifolds.
A modification of the confidence screening mechanism based on adaptive weighing of every training instance at each cascade level of the Deep Forest is proposed. The idea underlying the modification is very simple and stems from the confidence screening mechanism idea proposed by Pang et al. to simplify the Deep Forest …
In the multiple linear regression setting, we propose a general framework, termed weighted orthogonal components regression (WOCR), which encompasses many known methods as special cases, including ridge regression and principal components regression. WOCR makes use of the monotonicity inherent in orthogonal components …
Pion optimizes LLMs by preserving weight matrix singular values.
Orthogonal initialization does not speed up training in ultra-wide neural networks.
Many scientific questions require estimating the effects of continuous treatments. Outcome modeling and weighted regression based on the generalized propensity score are the most commonly used methods to evaluate continuous effects. However, these techniques may be sensitive to model misspecification, extreme weights o…
New method detects RNA modifications without prior training, revealing novel sites.
Deep networks with orthogonal weights show stable fluctuations, improving generalization and training speed.
ABRF uses attention weights to improve RF performance.
MuonEq improves training of matrix-valued parameters by rebalancing momentum before orthogonalization.
ABIForest improves anomaly detection using attention weights.
Wasserstein-GANs have been introduced to address the deficiencies of generative adversarial networks (GANs) regarding the problems of vanishing gradients and mode collapse during the training, leading to improved convergence behaviour and improved image quality. However, Wasserstein-GANs require the discriminator to be…
A new classifier uses weighted orthogonal regression for robust classification with limited data.
A well-conditioned Jacobian spectrum has a vital role in preventing exploding or vanishing gradients and speeding up learning of deep neural networks. Free probability theory helps us to understand and handle the Jacobian spectrum. We rigorously show almost sure asymptotic freeness of layer-wise Jacobians of deep neura…
The paper introduces a limit version of multiple stopping options such that the holder selects dynamically a weight function that control the distribution of the payments (benefits) over time. In applications for commodities and energy trading, a control process can represent the quantity that can be purchased by a fix…
This work proves the asymptotic freeness of layerwise Jacobians in MLPs with Haar orthogonal matrices.
Recently mean field theory has been successfully used to analyze properties of wide, random neural networks. It gave rise to a prescriptive theory for initializing feed-forward neural networks with orthogonal weights, which ensures that both the forward propagated activations and the backpropagated gradients are near $…
Deep networks without non-linearities are equivalent to shallow ones.
Study connects curvature to graph theory and reveals differences.
Orthogonal initialization speeds up convergence in deep linear networks.
OPT framework improves neural network generalization by learning an orthogonal transformation.
Proposes modifications to model-based forests for HTE estimation in observational data.
In this short note we show the following result: Let () be a compact Sasaki manifold with positive transverse orthogonal bisectional curvature. Then is finite, and the universal cover of is isomorphic to a weighted Sasaki sphere. We also get some results in the case of n…
New LT-O-learners improve HLTE estimation with low overlap.
Proposes simplified SHAP for faster black-box model explanations.
Recurrent Neural Networks (RNNs) are designed to handle sequential data but suffer from vanishing or exploding gradients. Recent work on Unitary Recurrent Neural Networks (uRNNs) have been used to address this issue and in some cases, exceed the capabilities of Long Short-Term Memory networks (LSTMs). We propose a simp…
LOFT separates subspace rotation and transformation for orthogonal fine-tuning.
New method enforces orthogonality in convolutional layers for improved robustness.
We introduce a novel approach to perform first-order optimization with orthogonal and unitary constraints. This approach is based on a parametrization stemming from Lie group theory through the exponential map. The parametrization transforms the constrained optimization problem into an unconstrained one over a Euclidea…
In this paper, we introduce the algorithms of Orthogonal Deep Neural Networks (OrthDNNs) to connect with recent interest of spectrally regularized deep learning methods. OrthDNNs are theoretically motivated by generalization analysis of modern DNNs, with the aim to find solution properties of network weights that guara…
New algorithm speeds up group equivariant neural networks computations.
It is well known that Convolutional Neural Networks (CNNs) have significant redundancy in their filter weights. Various methods have been proposed in the literature to compress trained CNNs. These include techniques like pruning weights, filter quantization and representing filters in terms of a basis functions. Our ap…
Computational methods that predict differential gene expression from histone modification signals are highly desirable for understanding how histone modifications control the functional heterogeneity of cells through influencing differential gene regulation. Recent studies either failed to capture combinatorial effects…
Pulling back the weight system associated with the exceptional Lie algebra G_2 by a modification of the universal Vassiliev-Kontsevich invariant yields a link invariant; extending it to 3-nets, we derive a recursive algorithm for its evaluation.
The local linear embedding algorithm (LLE) is a non-linear dimension-reducing technique, widely used due to its computational simplicity and intuitive approach. LLE first linearly reconstructs each input point from its nearest neighbors and then preserves these neighborhood relations in the low-dimensional embedding. W…
An algorithm for efficient computation of equivariant neural network layers.
The paper proves smoothness of almost-minimizers' boundaries near the free boundary.
A weighted random survival forest is presented in the paper. It can be regarded as a modification of the random forest improving its performance. The main idea underlying the proposed model is to replace the standard procedure of averaging used for estimation of the random survival forest hazard function by weighted av…
ModHiFi identifies critical components for model modification without gradients or loss function.
The paper proves a distribution claim for neural network Jacobians.
New survival learners estimate heterogeneous treatment effects from time-to-event data.
Effective features can improve the performance of a model, which can thus help us understand the characteristics and underlying structure of complex data. Previous feature selection methods usually cannot keep more local structure information. To address the defects previously mentioned, we propose a novel supervised o…