Improves LSTM performance by initializing states via manifold learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Improves state space models' resistance to noise.
This paper studies how gradient descent in control systems can perform well on unseen data.
Study evaluates initialization strategies for infinite hidden Markov models.
HOPE improves SSMs for long-memory tasks with robust initialization and training.
Many neural networks use the tanh activation function, however when given a probability distribution as input, the problem of computing the output distribution in neural networks with tanh activation has not yet been addressed. One important example is the initialization of the echo state network in reservoir computing…
Initialization of parameters in deep neural networks has been shown to have a big impact on the performance of the networks (Mishkin & Matas, 2015). The initialization scheme devised by He et al, allowed convolution activations to carry a constrained mean which allowed deep networks to be trained effectively (He et al.…
Modeling financial time series with LSTM and trainable initial states.
Quantum states can be learned efficiently using gentle measurements.
A toy model shows how locality can emerge in the universe's Hamiltonian and initial state.
In this effort we propose a novel approach for reconstructing multivariate functions from training data, by identifying both a suitable network architecture and an initialization using polynomial-based approximations. Training deep neural networks using gradient descent can be interpreted as moving the set of network p…
In this paper, we firstly give a brief introduction of expectation maximization (EM) algorithm, and then discuss the initial value sensitivity of expectation maximization algorithm. Subsequently, we give a short proof of EM's convergence. Then, we implement experiments with the expectation maximization algorithm (We im…
Introduces a neural network-based method for efficient state and parameter estimation in complex systems.
Mixed datasets consist of both numeric and categorical attributes. Various k-means-based clustering algorithms have been developed for these datasets. Generally, these algorithms use random partition as a starting point, which tends to produce different clustering results for different runs. In this paper, we propose, …
Normalization layers are a staple in state-of-the-art deep neural network architectures. They are widely believed to stabilize training, enable higher learning rate, accelerate convergence and improve generalization, though the reason for their effectiveness is still an active research topic. In this work, we challenge…
A new method for student-initiated action advice using novelty detection.
Residual Network (ResNet) is the state-of-the-art architecture that realizes successful training of really deep neural network. It is also known that good weight initialization of neural network avoids problem of vanishing/exploding gradients. In this paper, simplified models of ResNets are analyzed. We argue that good…
Money was invented to address the difficulty in the double coincidence of wants between the supply and demand when people exchanged their goods and services. There are two information states in society: one is the initial state that people have goods and services due to division of labor; the other is the final state t…
Enhances inference of spreading processes using neural-network priors.
We present the quantum model of Bertrand duopoly and study the entanglement behavior on the profit functions of the firms. Using the concept of optimal response of each firm to the price of the opponent, we found only one Nash equilibirum point for maximally entangled initial state. The very presence of quantum entangl…
Training recurrent neural networks (RNNs) on long sequence tasks is plagued with difficulties arising from the exponential explosion or vanishing of signals as they propagate forward or backward through the network. Many techniques have been proposed to ameliorate these issues, including various algorithmic and archite…
The principle of convergence stability for geometric flows is the combination of the continuous dependence of the flow on initial conditions, with the stability of fixed points. It implies that if the flow from an initial state exists for all time and converges to a stable fixed point, then the flows of solutions…
Deep neural networks achieve state-of-the-art performance for a range of classification and inference tasks. However, the use of stochastic gradient descent combined with the nonconvexity of the underlying optimization problems renders parameter learning susceptible to initialization. To address this issue, a variety o…
Barren plateaus are not an average-case phenomenon, but a highly non-unique problem.
Two new scalable K-means initialization methods proposed for large-scale clustering.
We present a method for a certain class of Markov Decision Processes (MDPs) that can relate the optimal policy back to one or more reward sources in the environment. For a given initial state, without fully computing the value function, q-value function, or the optimal policy the algorithm can determine which rewards w…
Many optimization methods for generating black-box adversarial examples have been proposed, but the aspect of initializing said optimizers has not been considered in much detail. We show that the choice of starting points is indeed crucial, and that the performance of state-of-the-art attacks depends on it. First, we d…
In this paper, we present a novel approach for initializing deep neural networks, i.e., by turning PCA into neural layers. Usually, the initialization of the weights of a deep neural network is done in one of the three following ways: 1) with random values, 2) layer-wise, usually as Deep Belief Network or as auto-encod…
This work reduces DIM computation costs by training neural networks on single MC paths.
We consider the generic approach of using an experience memory to help exploration by adapting a restart distribution. That is, given the capacity to reset the state with those corresponding to the agent's past observations, we help exploration by promoting faster state-space coverage via restarting the agent from a mo…
We propose an algorithm for deterministic continuous Markov Decision Processes with sparse rewards that computes the optimal policy exactly with no dependency on the size of the state space. The algorithm has time complexity of and memory complexity of , where is the…
New method improves robustness in partially observable domains by training against latent distribution shifts.
Re-initializing neural networks improves generalization but not as much as other techniques.
New method improves deep learning model robustness and accuracy for long sequences.
New method for robust fixed-point smoothing without state augmentation.
The Strong Cosmic Censorship conjecture states that for generic initial data to Einstein's field equations, the maximal globally hyperbolic development is inextendible. We prove this conjecture in the class of orthogonal Bianchi class B perfect fluids and vacuum spacetimes, by showing that unboundedness of certain curv…
Quantum method prices options by evolving a state in imaginary time.
The goal of this paper is to clarify when a stochastic partial differential equation with an affine realization admits affine state processes. This includes a characterization of the set of initial points of the realization. Several examples, as the HJMM equation from mathematical finance, illustrate our results.
Due to the iterative nature of most nonnegative matrix factorization (\textsc{NMF}) algorithms, initialization is a key aspect as it significantly influences both the convergence and the final solution obtained. Many initialization schemes have been proposed for NMF, among which one of the most popular class of methods…
The paper solves pentagon equations using triangulations and edge transformations.
Constructs initial data leading to apparent horizons and tests Penrose Inequality.
Improved object segmentation and tracking in video using optical flow and initial state conditioning.
PINNs solve neuronal parameter and state estimation problems with limited data.
Optimizes deep neural network initialization variance for better performance.
The study tests inferences about neural network optimization from linear interpolation of loss landscapes.
GDB bridges geometric states with improved accuracy and generality.
In 2002, Isenberg-Mazzeo-Pollack (IMP) constructed a series of vacuum initial data sets via a gluing construction. In this paper, we investigate some local geometry of these initial data sets as well as implications regarding their spacetime developments. In particular, we state conditions for the existence of outer tr…
We study singularities of Lagrangian mean curvature flow in $\C^n$ when the initial condition is a zero-Maslov class Lagrangian. We start by showing that, in this setting, singularities are unavoidable. More precisely, we construct Lagrangians with arbitrarily small Lagrangian angle and Lagrangians which are Hamiltonia…