Initializing the weights and the biases is a key part of the training process of a neural network. Unlike the subsequent optimization phase, however, the initialization phase has gained only limited attention in the literature. In this paper we discuss some consequences of commonly used initialization strategies for va…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study evaluates initialization strategies for infinite hidden Markov models.
Recurrent Neural Networks (RNNs) can be seriously impacted by the initial parameters assignment, which may result in poor generalization performances on new unseen data. With the objective to tackle this crucial issue, in the context of RNN based classification, we propose a new supervised layer-wise pretraining strate…
This paper investigates multilevel initialization strategies for training very deep neural networks with a layer-parallel multigrid solver. The scheme is based on the continuous interpretation of the training problem as a problem of optimal control, in which neural networks are represented as discretizations of time-de…
Residual networks (ResNet) and weight normalization play an important role in various deep learning applications. However, parameter initialization strategies have not been studied previously for weight normalized networks and, in practice, initialization methods designed for un-normalized networks are used as a proxy.…
Maxout networks study gradients and propose initialization strategies.
Barren plateaus are not an average-case phenomenon, but a highly non-unique problem.
New method trains shallow neural networks with subquadratic width scaling.
Improved LLM pre-training performance through better weight and variance control.
We propose to study market efficiency from a computational viewpoint. Borrowing from theoretical computer science, we define a market to be \emph{efficient with respect to resources } (e.g., time, memory) if no strategy using resources can make a profit. As a first step, we consider memory- strategies whose a…
Stochastic variational inference is an established way to carry out approximate Bayesian inference for deep models. While there have been effective proposals for good initializations for loss minimization in deep learning, far less attention has been devoted to the issue of initialization of stochastic variational infe…
Persistent neurons improve neural network optimization by leveraging previous solutions.
Study spectral learning for odeco tensors, addressing initialization bottlenecks.
Adaptive populations such as those in financial markets and distributed control can be modeled by the Minority Game. We consider how their dynamics depends on the agents' initial preferences of strategies, when the agents use linear or quadratic payoff functions to evaluate their strategies. We find that the fluctuatio…
Study examines volatility-based strategy for Chinese ETF options, improving returns in volatile markets.
New method initializes MLPs for tabular data with tree-based feature interactions.
In an adaptive population which models financial markets and distributed control, we consider how the dynamics depends on the diversity of the agents' initial preferences of strategies. When the diversity decreases, more agents tend to adapt their strategies together. This change in the environment results in dynamical…
Paper proposes a new strategy to improve initial performance of federated models.
Whether you trade futures for yourself or a hedge fund, your strategy is counted. Long and short position limits make the number of unique strategies finite. Formulas of the numbers of strategies, transactions, do nothing actions are derived. A discrete distribution of actions, corresponding probability mass, cumulativ…
Quantum circuit models learn better with specific initialization strategies.
New method learns good initialization for gradient descent from past solutions.
XGL uses global explanations to guide human supervision in machine learning.
New method solves continuous time mean-variance model for consistent investment strategy.
Structured CNN designed using the prior information of problems potentially improves efficiency over conventional CNNs in various tasks in solving PDEs and inverse problems in signal processing. This paper introduces BNet2, a simplified Butterfly-Net and inline with the conventional CNN. Moreover, a Fourier transform i…
Parkinson's disease patients develop different speech impairments that affect their communication capabilities. The automatic assessment of the speech of the patients allows the development of computer aided tools to support the diagnosis and the evaluation of the disease severity. This paper introduces a methodology t…
Deep Hedging learns optimal strategies for various risk levels.
Two new scalable K-means initialization methods proposed for large-scale clustering.
The hypothesis that sub-network initializations (lottery) exist within the initializations of over-parameterized networks, which when trained in isolation produce highly generalizable models, has led to crucial insights into network initialization and has enabled efficient inferencing. Supervised models with uncalibrat…
New method improves adversarial training efficiency and robustness.
We consider a trader who wants to direct his portfolio towards a set of acceptable wealths given by a convex risk measure. We propose a black-box algorithm, whose inputs are the joint law of stock prices and the convex risk measure, and whose outputs are the numerical values of initial capital requirement and the funct…
We consider the problem of utility maximization for small traders on incomplete financial markets. As opposed to most of the papers dealing with this subject, the investors' trading strategies we allow underly constraints described by closed, but not necessarily convex, sets. The final wealths obtained by trading under…
We consider the martingale optimal transport duality for càdlàg processes with given initial and terminal laws. Strong duality and existence of dual optimizers (robust semi-static superhedging strategies) are proved for a class of payoffs that includes American, Asian, Bermudan, and European options with intermediate m…
An arbitrage strategy allows a financial agent to make certain profit out of nothing, i.e., out of zero initial investment. This has to be disallowed on economic basis if the market is in equilibrium state, as opportunities for riskless profit would result in an instantaneous movement of prices of certain financial ins…
Meta-strategy learns tuning parameters for online learning methods.
We study exponential Levy models with change-point which is a random variable, independent from initial Levy processes. On canonical space with initially enlarged filtration we describe all equivalent martingale measures for change-point model and we give the conditions for the existence of f-divergence minimal equival…
Optimizes dividend policies in a Brownian model with controlled rates.
Small initialization improves tensor recovery from noisy data.
Optimal trading is a recent field of research which was initiated by Almgren, Chriss, Bertsimas and Lo in the late 90's. Its main application is slicing large trading orders, in the interest of minimizing trading costs and potential perturbations of price dynamics due to liquidity shocks. The initial optimization frame…
Topic models can provide us with an insight into the underlying latent structure of a large corpus of documents. A range of methods have been proposed in the literature, including probabilistic topic models and techniques based on matrix factorization. However, in both cases, standard implementations rely on stochastic…
Consider an American option that pays G(X^*_t) when exercised at time t, where G is a positive increasing function, X^*_t := \sup_{s\le t}X_s, and X_s is the price of the underlying security at time s. Assuming zero interest rates, we show that the seller of this option can hedge his position by trading in the underlyi…
Enhances currency strategy Sharpe ratio by 30% using context-aware Learning to Rank.
Sensors which use electromagnetic induction (EMI) to excite a response in conducting bodies have long been investigated for subsurface explosive hazard detection. In particular, EMI sensors have been used to discriminate between different types of objects, and to detect objects with low metal content. One successful, p…
New model uses pretrained biochemical language models to generate drug compounds.
The performance of gradient-based optimization strategies depends heavily on the initial weights of the parametric model. Recent works show that there exist weight initializations from which optimization procedures can find the task-specific parameters faster than from uniformly random initializations and that such a w…
In this work we study the optimal execution problem with multiplicative price impact in algorithm trading, when an agent holds an initial position of shares of a financial asset. The inter-selling-decision times are modelled by the arrival times of a Poisson process. The criterion to be optimised consists in maximising…
New method for community detection in graphs faster than DCBM inference.
Motivated by recent advances in the spectral theory of auto-covariance matrices, we are led to revisit a reformulation of Markowitz' mean-variance portfolio optimization approach in the time domain. In its simplest incarnation it applies to a single traded asset and allows to find an optimal trading strategy which - fo…
The Internet is known to have had a powerful impact on on-line retailer strategies in markets characterised by long-tail distribution of sales. Such retailers can exploit the long tail of the market, since they are effectively without physical limit on the number of choices on offer. Here we examine two extensions of t…