Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

6.3%12.5%18.8%25.0% · Oct 199319922001200920172026
48 results for hidden capacity

New method reduces overfitting in deep neural networks by measuring and regulating hidden unit diversity.

problem Overfitting in deep neural networks.
method Introduces a new redundancy measure based on mutual information to improve generalization.
result Reduction of redundancy improves generalization capacity, reducing overfitting.

Wide hidden layer TCM nets capacity analyzed using RDT and fl RDT.

problem Capacity analysis of wide hidden layer TCM nets.
method Employed Fully Lifted Random Duality Theory (fl RDT) for capacity characterization.
result Explicit, closed form capacity characterizations for a generic class of hidden layer activations.

New analysis shows capacity of treelike neural networks with various activations.

problem Analyzing the capacity of treelike neural networks with diverse activations.
method Utilized Random Duality Theory and its partially lifted version to handle various activations.
result The capacity of treelike neural networks decreases for large network width but converges to a constant value.

Optimal hidden-target learning for online inventory optimization on general convex sets.

problem Online inventory optimization (OIO) on arbitrary bounded convex capacity sets.
method Maintaining a hidden target and projecting it onto the feasible order-up-to set.
result The method improves the best known regret guarantee for OIO on general convex sets from inverse to inverse-square-root dependence on the common-demand probability.

New bounds on neural network capacity for treelike sign perceptrons using RDT.

problem Determining the capacity of treelike sign perceptrons neural networks.
method Random Duality Theory (RDT) to establish upper bounds.
result Mathematically rigorous bounds on network capacity for any number of neurons.

Improved neural network capacity analysis using simplified RDT.

problem Analyzing the memorization capabilities of sign perceptron neural networks.
method Developed a simplified, partially lifted Random Duality Theory (fl RDT) approach.
result Concrete capacity bounds universally improve over previous best known ones.

Neural networks learn more efficiently with hidden factorial structures.

problem Challenges in high-dimensional statistical learning.
method Controlled experimental framework to test neural networks' ability to exploit hidden factorial structures.
result Neural networks can leverage hidden factorial structures to learn discrete distributions more efficiently.

Recurrent neural networks are powerful models for processing sequential data, but they are generally plagued by vanishing and exploding gradient problems. Unitary recurrent neural networks (uRNNs), which use unitary recurrence matrices, have recently been proposed as a means to avoid these issues. However, in previous …

2016-10-31abs ↗pdf ↗

We explore a framework called boosted Markov networks to combine the learning capacity of boosting and the rich modeling semantics of Markov networks and applying the framework for video-based activity recognition. Importantly, we extend the framework to incorporate hidden variables. We show how the framework can be ap…

2014-08-06abs ↗pdf ↗

Two potential bottlenecks on the expressiveness of recurrent neural networks (RNNs) are their ability to store information about the task in their parameters, and to store information about the input history in their units. We show experimentally that all common RNN architectures achieve nearly the same per-task and pe…

2016-11-29abs ↗pdf ↗

Chemical networks outperform spiking neural networks in classification tasks.

problem Learning tasks with spiking neural networks require hidden layers, which are computationally expensive.
method Used deterministic mass-action kinetics to prove chemical reaction networks without hidden layers can solve tasks previously solved by spiking neural networks.
result A chemical reaction network without hidden layers outperforms a spiking neural network with hidden layers in a handwritten digit classification task.

Normalization layers control deep neural network capacity, improving stability and generalization.

problem Excessive capacity in deep neural networks leads to overfitting and poor generalization.
method Developed a theoretical framework to explain normalization's role in capacity control.
result Normalization layers reduce the Lipschitz constant exponentially, smoothing the loss landscape and enhancing generalization.

This work presents entropic constraints from DAGs with hidden variables.

problem Characterizing causal relations in systems with hidden variables.
method Entropic inequality constraints derived from ee-separation relations.
result These constraints can learn about true causal models from observed data.

Combining Bayesian nonparametrics and a forward model selection strategy, we construct parsimonious Bayesian deep networks (PBDNs) that infer capacity-regularized network architectures from the data and require neither cross-validation nor fine-tuning when training the model. One of the two essential components of a PB…

2018-05-22abs ↗pdf ↗

Analyzes how ReLU affects neural network capacity and solution space geometry.

problem Understanding the effects of ReLU on neural network capacity and solution space.
method Analytical and numerical studies of two-layer neural networks with ReLU activations.
result The capacity of ReLU networks remains finite as the number of hidden neurons increases, unlike for threshold units.

Matryoshka hides secret models in a carrier model, achieving high capacity and robustness.

problem Stealing functionality of private ML data by hiding models in a carrier model.
method Parameter sharing approach exploiting the learning capacity of the carrier model.
result Hides a 26x larger secret model or 8 secret models in the carrier model.

A general Boltzmann machine with continuous visible and discrete integer valued hidden states is introduced. Under mild assumptions about the connection matrices, the probability density function of the visible units can be solved for analytically, yielding a novel parametric density function involving a ratio of Riema…

2017-12-20abs ↗pdf ↗

Differential forms and symmetric tensors show contrasting singular behaviors in a specific geometric setting.

problem Exploring differential forms and symmetric tensors on a specific geometric setting.
method Analyzing differential forms and symmetric tensors on the quadrant C2C_2 with subset diffeology.
result Symmetric tensors exhibit singularities that accumulate, while differential forms are smooth.

Reservoir Computing (RC) refers to a Recurrent Neural Networks (RNNs) framework, frequently used for sequence learning and time series prediction. The RC system consists of a random fixed-weight RNN (the input-hidden reservoir layer) and a classifier (the hidden-output readout layer). Here we focus on the sequence lear…

2017-06-24abs ↗pdf ↗

Deep dynamic generative models are developed to learn sequential dependencies in time-series data. The multi-layered model is designed by constructing a hierarchy of temporal sigmoid belief networks (TSBNs), defined as a sequential stack of sigmoid belief networks (SBNs). Each SBN has a contextual hidden state, inherit…

2015-09-23abs ↗pdf ↗

Understanding the representational power of Restricted Boltzmann Machines (RBMs) with multiple layers is an ill-understood problem and is an area of active research. Motivated from the approach of \emph{Inherent Structure formalism} (Stillinger & Weber, 1982), extensively used in analysing Spin Glasses, we propose a no…

2018-06-12abs ↗pdf ↗

We develop a novel method for training of GANs for unsupervised and class conditional generation of images, called Linear Discriminant GAN (LD-GAN). The discriminator of an LD-GAN is trained to maximize the linear separability between distributions of hidden representations of generated and targeted samples, while the …

2017-07-25abs ↗pdf ↗

Deep ReLU networks show that 4 layers suffice for unique input recovery.

problem Injectivity capacity of deep ReLU networks.
method Developed a program connecting deep ReLU injectivity to an ll-extension of the 0\ell_0 spherical perceptrons, using random duality theory.
result Only 4 layers are needed for unique input recovery, showing expansion saturation effect.

Wide neural networks can degrade performance, contrary to conventional wisdom.

problem Understanding the limitations of increasing network width in neural networks.
method Using Deep Gaussian Processes to decouple capacity and width, analyzing their effects on representational power and non-Gaussianity.
result Wide neural networks can become less adaptable and more Gaussian, leading to performance degradation.

Improved classifier accuracy by using more of the class-specific structure in trained models.

problem Softmax ignores valuable information encoded in the full array of class response distributions.
method Developed a hybrid classifier (Softmax-Pooling Hybrid, SPHSPH) that uses Softmax on high-scoring samples and a log-likelihood method on low-scoring samples.
result Reduces test set error by 6% to 23% using the exact same trained model.

Proposes M-CHMM for robust modeling of multivariate healthcare time series.

problem Challenges in analyzing multivariate healthcare time series data.
method Mixture of coupled hidden Markov models (M-CHMM) with two sampling algorithms.
result Improves data fit, handles missing and noisy measurements, and enhances prediction accuracy.

Neural networks outperform kernels by learning features better.

problem Current theories of feature learning do not adequately assess feature quality.
method Introduced feature quality metric and examined existing theories empirically.
result Current theories of feature learning do not provide a sufficient foundation for neural network generalization.

Mean field Gaussian inference limits mutual information to regularize neural networks.

problem Understanding and quantifying the regularization effect of mean field Gaussian inference.
method Empirically observed and theoretically quantified mutual information limitation through noise.
result Bounding mutual information between parameters and data effectively regularizes neural networks.

A new deep learning framework captures multi-scale spatio-temporal dependencies.

problem Designing and analyzing deep learning models for complex spatio-temporal analytics.
method Developed an I2^2DRNN model with three modules for integrating and learning multi-scale spatio-temporal data.
result The I2^2DRNN model outperforms classical and state-of-the-art models in capturing meaningful multi-scale spatio-temporal dependencies.

We study the problem of large scale, multi-label visual recognition with a large number of possible classes. We propose a method for augmenting a trained neural network classifier with auxiliary capacity in a manner designed to significantly improve upon an already well-performing model, while minimally impacting its c…

2014-12-20abs ↗pdf ↗

New neural network approach solves Poisson equations efficiently.

problem Approximating solutions to Poisson equations with Dirichlet boundary conditions.
method Using shallow ReLUα-networks to solve Laplace operator equations.
result Neural networks can approximate solutions to the Laplace operator with Dirichlet boundary conditions efficiently.

NADINE builds MLPs from streaming data, overcoming forgetting issues.

problem Building deep neural networks from streaming data efficiently and avoiding forgetting.
method NADINE uses a fully open MLP structure that dynamically evolves its depth and width online, resolving catastrophic forgetting through soft-forgetting and adaptive memory.
result NADINE outperforms existing methods in nine data stream classification and regression problems.

UWM-JEPA predicts future scenarios in belief space, improving accuracy in partially observed environments.

problem Predicting future scenarios in partially observed environments with uncertainty.
method Introduces UWM-JEPA, a JEPA world model with a density-matrix latent and learned unitary predictor.
result UWM-JEPA achieves 0.77 accuracy on a hidden-velocity indicator task, outperforming LSTM-JEPA.

Paper presents a method for probabilistic load forecasting using adaptive online learning.

problem Inability to assess intrinsic uncertainties and capture dynamic changes in consumption patterns.
method Adaptive online learning of hidden Markov models for recursive parameter updates and sequential prediction.
result Significant improvement in performance compared to existing techniques across various scenarios.