Deep nets exhibit 'Neural Collapse' during training's final phase, simplifying decision-making.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper proposes a deep RL method for hedging variable annuities, outperforming misspecified models.
Optimal asset allocation strategy outperforms stochastic benchmark.
New loss function improves classification for imbalanced and sensitive groups.
This paper simplifies fine-tuning for small LLMs, reducing barriers for developers.
In this paper, the `Approximate Message Passing' (AMP) algorithm, initially developed for compressed sensing of signals under i.i.d. Gaussian measurement matrices, has been extended to a multi-terminal setting (MAMP algorithm). It has been shown that similar to its single terminal counterpart, the behavior of MAMP algo…
We introduce two simple models of forward-backward stochastic differential equations with a singular terminal condition and we explain how and why they appear naturally as models for the valuation of CO2 emission allowances. Single phase cap-and-trade schemes lead readily to terminal conditions given by indicator funct…
We consider an interest rate model with log-normally distributed rates in the terminal measure in discrete time. Such models are used in financial practice as parametric versions of the Markov functional model, or as approximations to the log-normal Libor market model. We show that the model has two distinct regimes, a…
We consider the class of short rate interest rate models for which the short rate is proportional to the exponential of a Gaussian Markov process x(t) in the terminal measure r(t) = a(t) exp(x(t)). These models include the Black, Derman, Toy and Black, Karasinski models in the terminal measure. We show that such intere…
In this work, we consider the problem of autonomously discovering behavioral abstractions, or options, for reinforcement learning agents. We propose an algorithm that focuses on the termination condition, as opposed to -- as is common -- the policy. The termination condition is usually trained to optimize a control obj…
Applicability of the concept of financial log-periodicity is discussed and encouragingly verified for various phases of the world stock markets development in the period 2000-2010. In particular, a speculative forecasting scenario designed in the end of 2004, that properly predicted the world stock market increases in …
Paper explains neural collapse in neural networks using a new model.
Humans tend to learn complex abstract concepts faster if examples are presented in a structured manner. For instance, when learning how to play a board game, usually one of the first concepts learned is how the game ends, i.e. the actions that lead to a terminal state (win, lose or draw). The advantage of learning end-…
TVM improves generative modeling by matching terminal velocities.
Survive method improves model-based RL by avoiding terminal states, reducing sample complexity.
The automatic classification of applications and services is an invaluable feature for new generation mobile networks. Here, we propose and validate algorithms to perform this task, at runtime, from the raw physical channel of an operative mobile network, without having to decode and/or decrypt the transmitted flows. T…
Adaptive algorithm identifies best arm with abstention, showing phase transition from polynomial to exponential error probability.
In this work we investigate approaches to reconstruct generator models from measurements available at the generator terminal bus using machine learning (ML) techniques. The goal is to develop an emulator which is trained online and is capable of fast predictive computations. The training is illustrated on synthetic dat…
Gradient descent dynamics in quadratic regression models are analyzed, revealing five phases: monotonic, catapult, periodic, chaotic, and divergent.
This paper extends neural collapse to class-imbalanced datasets using an unconstrained ReLU feature model.
A time schedule simplifies learning in flow-based models for high-dimensional data.
Resource allocation improved using machine learning from terminal positions.
Catapult phase in neural nets shows exponential loss growth before quick decrease.
We propose the Insertion-Deletion Transformer, a novel transformer-based neural architecture and training method for sequence generation. The model consists of two phases that are executed iteratively, 1) an insertion phase and 2) a deletion phase. The insertion phase parameterizes a distribution of insertions on the c…
Majority of the modern meta-learning methods for few-shot classification tasks operate in two phases: a meta-training phase where the meta-learner learns a generic representation by solving multiple few-shot tasks sampled from a large dataset and a testing phase, where the meta-learner leverages its learnt internal rep…
To improve the efficient frontier of the classical mean-variance model in continuous time, we propose a varying terminal time mean-variance model with a constraint on the mean value of the portfolio asset, which moves with the varying terminal time. Using the embedding technique from stochastic optimal control in conti…
Study shows reverberant phase is not essential for weakly-supervised dereverberation.
In this paper, we propose a phase shift deep neural network (PhaseDNN) which provides a wideband convergence in approximating a high dimensional function during its training of the network. The PhaseDNN utilizes the fact that many DNN achieves convergence in the low frequency range first, thus, a series of moderately-s…
Our paper explains deep neural collapse in multiple layers.
Reward collapse occurs when ranking-based reward models yield uniform rewards for different prompts.
Generative model prices basket options efficiently.
Improved simulation of phase transitions using hierarchical autoregressive networks.
New model explains deep learning performance at large learning rates.
Bayesian theory explains abrupt emergence of copy subcircuit in attention.
This work investigates how neural collapse improves transfer learning for large-scale models.
Large GD stepsizes improve margins and speed up training for non-homogeneous networks.
Study of two-layer ReLU neural network phase diagram at infinite-width limit.
It is typical for a machine learning system to have numerous hyperparameters that affect its learning rate and prediction quality. Finding a good combination of the hyperparameters is, however, a challenging job. This is mainly because evaluation of each combination is extremely expensive computationally; indeed, train…
DP-SGD can update fewer coordinates while maintaining privacy.
We find optimal learning rate schedules for a random feature model.
The training phases of Deep neural network~(DNN) consumes enormous processing time and energy. Compression techniques utilizing the sparsity of DNNs can effectively accelerate the inference phase of DNNs. However, it can be hardly used in the training phase because the training phase involves dense matrix-multiplicatio…
Minimizing non-convex and high-dimensional objective functions is challenging, especially when training modern deep neural networks. In this paper, a novel approach is proposed which divides the training process into two consecutive phases to obtain better generalization performance: Bayesian sampling and stochastic op…
Unsupervised learning is a discipline of machine learning which aims at discovering patterns in big data sets or classifying the data into several categories without being trained explicitly. We show that unsupervised learning techniques can be readily used to identify phases and phases transitions of many body systems…
Models for predicting aircraft motion are an important component of modern aeronautical systems. These models help aircraft plan collision avoidance maneuvers and help conduct offline performance and safety analyses. In this article, we develop a method for learning a probabilistic generative model of aircraft motion i…
Margin enlargement over training data has been an important strategy since perceptrons in machine learning for the purpose of boosting the robustness of classifiers toward a good generalization ability. Yet Breiman (1999) showed a dilemma that a uniform improvement on margin distribution does NOT necessarily reduces ge…
Deep learning has become an area of interest in most scientific areas, including physical sciences. Modern networks apply real-valued transformations on the data. Particularly, convolutions in convolutional neural networks discard phase information entirely. Many deterministic signals, such as seismic data or electrica…
Quadratic models explain neural network behavior during training.
Methodology that recently lead us to predict to an amazing accuracy the date (July 11, 2008) of reverse of the oil price up trend is briefly summarized and some further aspects of the related oil price dynamics elaborated. This methodology is based on the concept of discrete scale invariance whose finance-prediction-or…