LayerNorm transformers have dead directions that can be read from their parameters alone.
problem Locating dead directions in LayerNorm transformers
method Using the inverse-scale direction of LayerNorm affine parameters
result Predicted dead direction matches measured bottom singular direction
We give an explicit algorithm and source code for extracting equity risk factors from dead (a.k.a. "flatlined" or "hockey-stick") alphas and using them to improve performance characteristics of good (tradable) alphas. In a nutshell, we use dead alphas to extract directions in the space of stock returns along which ther…
Dead-Direction Signatures (DDS) provide a cheap, closed-form spectral reading of a network's singular complexity.
problem Estimating the complexity of deep networks through their loss singularities.
method DDS replaces the SGLD posterior chain with spectral linear algebra.
result DDS observables rank-track the network's singular complexity at the framework-predicted sign.
A new geometric concept, the dead direction, bridges singular learning theory and information geometry.
problem The gap between singular learning theory and information geometry.
method Introducing the dead direction, a unit vector along degenerating Fisher metric, and showing its KL order can be recovered.
result The KL order of the dead direction can be recovered as the decay rate of the directional Fisher curvature, providing a handle on singular geometry.
Study on financial impacts of zombie outbreak on economy.
problem Financial and economic consequences of a zombie epidemic.
method Epidemiological modeling and financial computation.
result GDP losses of 23.44% and financial market drop of 29.30% in a major industrialized nation.
CoDeQ simplifies joint model compression by integrating pruning and quantization.
problem Joint pruning and quantization methods are complex and require additional procedures.
method CoDeQ uses a dead-zone quantizer to directly induce sparsity and learn quantization parameters.
result CoDeQ achieves high sparsity and low-precision accuracy with minimal bit operations.
It is proved that the curve graph C1(Σ) of a surface Σg,n has a local pathology that had not been identified as such: there are vertices α,β in C1(Σ) such that β is a dead end of every geodesic joining α to β. It also has double dead-ends. Every dead end has depth 1.
A new optimizer DDC improves deep learning models by respecting symmetries.
problem Deep networks' loss is invariant to continuous symmetries, leading to optimization issues.
method DDC builds a Dead-Direction Conditioner that lifts a base optimizer into a G-equivariant one, preserving the quotient geometry.
result DDCAdam and DDCMuon outperform standard optimizers in various tasks, improving validation-train loss gaps and learning dynamics.
In this paper we propose a novel accurate method for dead-reckoning of wheeled vehicles based only on an Inertial Measurement Unit (IMU). In the context of intelligent vehicles, robust and accurate dead-reckoning based on the IMU may prove useful to correlate feeds from imaging sensors, to safely navigate through obstr…
A framework identifies worst-case decision points in safety-critical scenarios, improving risk assessment by 10 hours.
problem Identifying worst-case outcomes in safety-critical decision-making under uncertainty.
method Explicitly estimating distributions of expected return to identify dead-ends, tuning based on risk tolerance.
result Significantly improves risk assessment, providing indications 10 hours earlier and increasing detection by 20%.
Study on curvature in finitely generated groups, showing positive curvature in specific cases.
problem Understanding curvature in finitely generated groups.
method Analyzing dead-end elements and related elements to find curvature, studying effect of radius.
result Examples of positive curvature for arbitrary radius in lamplighter and Houghton's group.
A new method lifts training of input-convex neural networks to avoid dead weights and plateaued loss.
problem Training input-convex neural networks with non-negative weights.
method Introduces a hypernetwork that emits non-negative weights from a summary of the input batch, adding stochasticity to soften the loss landscape.
result The lift method achieves lower test loss than projected gradient descent and direct softplus reparametrization.
We introduce the notion of connection thickness of spheres in a Cayley graph, related to dead-ends and their retreat depth. It was well-known that connection thickness is bounded for finitely presented one-ended groups. We compute that for natural generating sets of lamplighter groups on a line or on a tree, connection…
The paper describes the K-theory of C∗-algebras of locally finite graphs.
problem Computing the K-theory of C∗-algebras of locally finite graphs. method Using a directed graph representation and Cuntz-Krieger algebra, the paper computes the K-theory of C∗(Γ). result The K-theory of C∗(Γ) is determined by the graph's genus, number of ends, and dead-ends. Deep Neural Networks are highly over-parameterized and the size of the neural networks can be reduced significantly after training without any decrease in performance. One can clearly see this phenomenon in a wide range of architectures trained for various problems. Weight/channel pruning, distillation, quantization, m…
In this paper, we study the trainability of rectified linear unit (ReLU) networks. A ReLU neuron is said to be dead if it only outputs a constant for any input. Two death states of neurons are introduced; tentative and permanent death. A network is then said to be trainable if the number of permanently dead neurons is …
Firms miscount their customers who stop buying without saying goodbye.
problem Counting non-contractual customers accurately.
method Estimating repeat purchase probabilities and extrapolating to infinite time.
result The count of alive customers is only partially identified, with a wide range of estimates.
Paper proposes a deep learning method for better IMU gyroscope data.
problem Improving accuracy of IMU gyroscope data for robot orientation estimation.
method Dilated convolution neural network, proper loss function, key points identification.
result Algorithm outperforms state-of-the-art on unseen test sequences.
The electrocardiogram (ECG) is a widely-used medical test, typically consisting of 12 voltage versus time traces collected from surface recordings over the heart. Here we hypothesize that a deep neural network can predict an important future clinical event (one-year all-cause mortality) from ECG voltage-time traces. We…
We describe a bottom-up framework, based on the identification of appropriate order parameters and determination of phase diagrams, for understanding progressively refined agent-based models and simulations of financial markets. We illustrate this framework by starting with a deterministic toy model, whereby N indepe…
Proposes a flexible neural model for multi-state survival analysis.
problem Limited applicability of Cox models for multi-state and competing events.
method Uses neural ordinary differential equations to solve Kolmogorov forward equations.
result Demonstrates state-of-the-art performance and interpretability.
Storage has become a constrained resource on smartphones. Gaming is a popular activity on mobile devices and the explosive growth in the number of games coupled with their growing size contributes to the storage crunch. Even where storage is plentiful, it takes a long time to download and install a heavy app before it …
Optimal dynamic fees for AMMs: A stochastic control approach
problem Fee policy of a liquidity provider in AMM
method Ergodic control problem
result Optimal fee is independent of wealth and constant relative risk aversion
A simple method flags images as out-of-distribution based on their distance to nearest neighbors.
problem Detecting images not aligned with a trained model's in-distribution data.
method Flag images as OOD if their average distance to K nearest neighbors is large in the classifier's representation space.
result Simple methods can outperform more complex ones when considering learned representations.
In this study we investigate the potential for using synthetic aperture radar (SAR) data to provide high resolution defoliation and regrowth mapping of trees in the tundra-forest ecotone. Using aerial photographs, four areas with live forest and four areas with dead trees were identified. Quad-polarimetric SAR data fro…
Gradient descent on LSE objectives implicitly performs EM, leading to collapse without volume control.
problem Gradient collapse in autoencoders without volume control.
method Introduced a single-layer encoder with an LSE objective and InfoMax regularization for volume control.
result Gradient--responsibility identity holds exactly; LSE alone collapses; variance prevents dead components; decorrelation prevents redundancy.
A deep learning model improves pedestrian tracking accuracy.
problem Pedestrian tracking accuracy is low, especially with inertial measurement unit.
method Deep learning model using IMU and LIDAR data, attention mechanism.
result Preliminary results show improved accuracy.
Applicability of the concept of financial log-periodicity is discussed and encouragingly verified for various phases of the world stock markets development in the period 2000-2010. In particular, a speculative forecasting scenario designed in the end of 2004, that properly predicted the world stock market increases in …
Model compares altruism and individualism in wealth dynamics.
problem Comparing altruism and individualism in wealth dynamics.
method Minimalist dynamical model of wealth evolution and sharing among N agents.
result Altruism leads to more global median wealth at early times but individualists accumulate most wealth in the long run.
Study quantifies gender bias in language models across 7 languages.
problem Measuring gender bias in language models across multiple languages.
method Curated dataset of politicians, multilingual language models, probing language models.
result Larger language models do not show significant gender bias compared to smaller ones.
Generative models learn from unlabeled videos via object segmentation and scene modeling.
problem Learning generative models from unlabelled videos.
method Decomposed into three subtasks: motion segmentation, background and foreground modeling, and scene sampling.
result Approach allows learning models that generalize beyond occlusions and represent scenes in a modular fashion.
TASFAR adapts regression models without labeled source data.
problem Lack of labeled source data for domain adaptation.
method Uses prediction confidence to estimate target label distribution and calibrate source model.
result Substantially reduces errors in various regression tasks.
BC-ACI corrects time series forecast bias, improving prediction intervals.
problem Persistent bias in time series forecasts leads to overly conservative prediction intervals.
method Augments ACI with an EWM estimate of forecast bias to correct nonconformity scores and re-center intervals.
result Reduces Winkler interval scores by 13-17% under distribution shifts, improving calibration.
Researchers develop a neural network that learns like humans, overcoming forgetting and structure issues.
problem Catastrophic forgetting and structure limitations in neural networks.
method Memory playback strategy and dynamic structure extension using conditional variational autoencoder (CVAE).
result The method effectively prevents forgetting and allows for dynamic network growth.
PoPCoin aims to create a more equitable cryptocurrency.
problem Inequality in traditional money systems.
method Develops two rules for PoPCoin: equal distribution and demurrage.
result PoPCoin can limit monetary inequality and incentivize rapid growth.
AEN-SAEs address feature starvation in sparse autoencoders by stabilizing the geometric alignment of sparse coding.
problem Feature starvation in sparse autoencoders, leading to unstable and misaligned representations.
method Adaptive Elastic Net SAEs (AEN-SAEs) combine ℓ2 and ℓ1 terms to stabilize the sparse coding map and control feature interactions. result AEN-SAEs mitigate feature starvation without heuristic resampling, maintaining competitive reconstruction abilities.
Bayesian deep learning predicts satellite collisions.
problem Space debris poses planetary risk.
method Bayesian deep learning with LSTM networks.
result Predicts conjunction event evolution with uncertainties.
T-KAN improves HFT LOB forecasting with learnable splines.
problem Alpha decay in HFT LOB forecasting models.
method T-KAN uses learnable B-spline activation functions to model market signals.
result 19.1% relative improvement in F1-score at k = 100 horizon.
GOPO optimizes large models in Hilbert space, avoiding Kullback-Leibler's curvature.
problem Optimizing large language models with Kullback-Leibler divergence's curvature issues.
method GOPO uses Hilbert space L2(pi_k) with orthogonality constraints and a work-dissipation functional.
result GOPO achieves competitive generalization with stable gradient dynamics and entropy preservation.
Logical neural networks solve mazes by filling dead ends, but not all methods generalize well.
problem Understanding how logical neural networks extrapolate solutions to mazes.
method Examined recurrent and implicit neural networks trained on maze-solving tasks.
result Models fail to generalize well to diverse maze sizes, suggesting limitations in learning scalable algorithms.
Given a set of points that sample a shape, the Rips complex of the data points is often used in machine-learning to provide an approximation of the shape easily-computed. It has been proved recently that the Rips complex captures the homotopy type of the shape assuming the vertices of the complex meet some mild samplin…
It is well known that the problem of vanishing/exploding gradients is a challenge when training deep networks. In this paper, we describe another phenomenon, called vanishing nodes, that also increases the difficulty of training deep neural networks. As the depth of a neural network increases, the network's hidden node…
The refugee crisis is perhaps the single most challenging problem for Europe today. Hundreds of thousands of people have already traveled across dangerous sea passages from Turkish shores to Greek islands, resulting in thousands of dead and missing, despite the best rescue efforts from both sides. One of the main reaso…
Proposes a new deep learning framework for financial stock trading.
problem Lack of effective techniques to fuse multi-channel financial time-series data.
method Inspired by convolution transform learning, SDCF processes channels through 1-D convolutions, fuses outputs with fully-connected layers, and applies softmax classification.
result Proposed framework yields better results than state-of-the-art techniques for stock trading.
Develops neural network for directed hypergraphs for node classification.
problem Irregular data structure, particularly directed graphs.
method Directed hypergraph neural network and semi-supervised learning method.
result Novel directed hypergraph neural network achieves highest accuracies on node classification tasks.
We introduce the notion of directed diagrammatic reducibility which is a relative version of diagrammatic reducibility. Directed diagrammatic reducibility has strong group theoretic and topological consequences. A multi-relator version of the Freiheitssatz in the presence of directed diagrammatic reducibility is given.…
The paper studies kernel smoothing and mean shift for directional data, deriving convergence rates and mode estimation.
problem Statistical and computational problems of kernel smoothing for directional data.
method Generalization of mean shift to directional data, derivation of convergence rates, and investigation of mode estimation.
result Statistical convergence rates of directional KDE and its derivatives, ascending property of directional mean shift, and mode estimation.
During the last two decades, we easilly see that the World Wide Web's link structure is modeled as the directed graph. In this paper, we will model the World Wide Web's link structure as the directed hypergraph. Moreover, we will develop the PageRank algorithm for this directed hypergraph. Due to the lack of the World …