Why does Deep Learning work? What representations does it capture? How do higher-order representations emerge? We study these questions from the perspective of group theory, thereby opening a new approach towards a theory of Deep learning. One factor behind the recent resurgence of the subject is a key algorithmic step…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Why does Deep Learning work? What representations does it capture? How do higher-order representations emerge? We study these questions from the perspective of group theory, thereby opening a new approach towards a theory of Deep learning. One factor behind the recent resurgence of the subject is a key algorithmic step…
A new neural network initialization method is proposed for faster and more accurate training.
We present a general framework for training deep neural networks without backpropagation. This substantially decreases training time and also allows for construction of deep networks with many sorts of learners, including networks whose layers are defined by functions that are not easily differentiated, like decision t…
A deep convolutional fuzzy system (DCFS) on a high-dimensional input space is a multi-layer connection of many low-dimensional fuzzy systems, where the input variables to the low-dimensional fuzzy systems are selected through a moving window across the input spaces of the layers. To design the DCFS based on input-outpu…
Extreme learning machine (ELM) is a new single hidden layer feedback neural network. The weights of the input layer and the biases of neurons in hidden layer are randomly generated, the weights of the output layer can be analytically determined. ELM has been achieved good results for a large number of classification ta…
The success of deep neural networks has inspired many to wonder whether other learners could benefit from deep, layered architectures. We present a general framework called forward thinking for deep learning that generalizes the architectural flexibility and sophistication of deep neural networks while also allowing fo…
We propose a neural architecture search (NAS) algorithm, Petridish, to iteratively add shortcut connections to existing network layers. The added shortcut connections effectively perform gradient boosting on the augmented layers. The proposed algorithm is motivated by the feature selection algorithm forward stage-wise …
Differentiable cutting-plane layers solve parametric mixed-integer linear optimization problems.
DLBricks automates DL benchmarking on CPUs, reducing effort and time.
In this paper, we made an extension to the convergence analysis of the dynamics of two-layered bias-free networks with one output. We took into consideration two popular regularization terms: the and norm of the parameter vector , and added it to the square loss function with coefficient $λ/…
This paper finds a new way to compress CNN weights, improving on pruning and quantization.
The performance of graph neural nets (GNNs) is known to gradually decrease with increasing number of layers. This decay is partly attributed to oversmoothing, where repeated graph convolutions eventually make node embeddings indistinguishable. We take a closer look at two different interpretations, aiming to quantify o…
A usual reinsurance policy for insurance companies admits one or two layers of the payment deductions. Under optimal criterion of minimizing the conditional tail expectation (CTE) risk measure of the insurer's total risk, this article generalized an optimal stop-loss reinsurance policy to an optimal multi-layer reinsur…
FSNet improves online time series forecasting by balancing fast adaptation and old knowledge.
LIDS assesses LLM summaries with interpretable key words.
This study analyzes why attention layers in neural networks can cause signal loss and proposes a solution.
Optimal strategies are found for a repeated betting game using diffusion approximation.
Convolutional networks struggle with repeating patterns in ECGs.
Study explores algorithmic collusion in repeated games using various learning dynamics.
We study two systems of tangle equations that arise when modeling the action of the Integrase family of proteins on DNA. These two systems--direct and inverted repeats--correspond to two different possibilities for the initial DNA sequence. We present one new class of solutions to the tangle equations. In the case of i…
Scorio.jl ranks systems from repeated tasks using various methods.
We investigate the question of when distinct branched surfaces in the complement of a 2-bridge knot support essential surfaces with identical boundary slopes. We determine all instances in which this occurs and identify an infinite family of knots for which no boundary slopes are repeated.
LAFF algorithm balances adaptability and non-exploitability in repeated games.
Link concordance and Whitney towers linked to Milnor invariants.
Risk is part of the fabric of every business; surprisingly, there is little work on establishing best practices for systematic, repeatable risk identification, arguably the first step of any risk management process. In this paper, we present a proposal that constitutes a more holistic risk management approach, a method…
It has long been known that a Milnor invariant with no repeated index is an invariant of link homotopy. We show that Milnor's invariants with repeated indices are invariants not only of isotopy, but also of self C_k-moves. A self C_k-move is a natural generalization of link homotopy based on certain degree k clasper su…
RSO uses random weight perturbations to train deep networks without gradients.
We investigate some geometric properties of the real algebraic variety of symmetric matrices with repeated eigenvalues. We explicitly compute the volume of its intersection with the sphere and prove a Eckart-Young-Mirsky-type theorem for the distance function from a generic matrix to points in . We exhibit conne…
This paper tackles the reduction of redundant repeating generation that is often observed in RNN-based encoder-decoder models. Our basic idea is to jointly estimate the upper-bound frequency of each target vocabulary in the encoder and control the output words based on the estimation in the decoder. Our method shows si…
Machine learning models for repeated measurements are limited. Using topological data analysis (TDA), we present a classifier for repeated measurements which samples from the data space and builds a network graph based on the data topology. When applying this to two case studies, accuracy exceeds alternative models wit…
Study optimizes product assortment for retailers with repeated exposures and patience costs.
Graph Neural Networks (GNNs) are based on repeated aggregations of information across nodes' neighbors in a graph. However, because common neighbors are shared between different nodes, this leads to repeated and inefficient computations. We propose Hierarchically Aggregated computation Graphs (HAGs), a new GNN graph re…
Generative compression technique reduces neural network size and improves performance on microcontrollers.
We present Sequential Attend, Infer, Repeat (SQAIR), an interpretable deep generative model for videos of moving objects. It can reliably discover and track objects throughout the sequence of frames, and can also generate future frames conditioning on the current frame, thereby simulating expected motion of objects. Th…
Study analyzes broker's gain from trade in repeated context-based trading.
A new method combines machine learning with mixed-effects models for better repeated measurement analysis.
We propose a random convolutional neural network to generate a feature space in which we study image classification and retrieval performance. Put briefly we apply random convolutional blocks followed by global average pooling to generate a new feature, and we repeat this k times to produce a k-dimensional feature spac…
New neural network processes 3D volumes with improved equivariance.
This paper examines transitions in sniping behavior among algorithmic traders, finding new profitable strategies.
Repeated self-distillation improves model performance significantly.
Study compares deep learning models for volatility prediction using multivariate data.
XSPNs combine SPNs and MEVMs for efficient inference in data with repeated parts.
Ribbon: Scalable Approximation and Robust Uncertainty Quantification
Serverless cloud computing speeds up double machine learning model estimation.
The paper uses machine learning to optimize rework policies in semiconductor manufacturing.
Origami uses SGX enclaves and blinding to protect deep neural network inference privacy.
This research extends the Pareto/NBD model using neural networks for better out-of-sample predictions.