Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

12.5%25.0%37.5%50.0% · Oct 199319922001200920182026
48 results for Universal Capacity Control

New theory predicts which large DNNs will have best test accuracy.

problem Predicting which large pre-trained DNNs will have the best test accuracy.
method Heavy-Tailed Self-Regularization (HT-SR) and Universal capacity control metric based on power law exponents.
result Universal capacity control metric correlates well with reported test accuracies of large-scale DNNs.

There are (at least) three approaches to quantifying information. The first, algorithmic information or Kolmogorov complexity, takes events as strings and, given a universal Turing machine, quantifies the information content of a string as the length of the shortest program producing it. The second, Shannon information…

2011-10-17abs ↗pdf ↗

Paper tackles inventory management with deep learning, improving performance and adherence to constraints.

problem Managing inventory with limited resources and constraints.
method Proposes a novel method to sample from a distribution of possible constraint paths, extends exo-IDP formulation, introduces neural coordinator, and uses modified DirectBackprop algorithm.
result Deep reinforcement learning policies with a neural coordinator outperform classic baselines in terms of performance and adherence to constraints.

This paper controls the capacity of weight-normalized deep neural networks using rectified linear units.

problem Capacity control of weight-normalized deep neural networks.
method Establishes upper bounds on Rademacher complexities and analyzes approximation properties of Lp,qL_{p,q} weight normalized networks.
result For L1,L_{1,\infty} weight normalized networks, the approximation error is controlled by the L1L_1 norm of the output layer, and generalization error depends on the square root of depth.

Normalization layers control deep neural network capacity, improving stability and generalization.

problem Excessive capacity in deep neural networks leads to overfitting and poor generalization.
method Developed a theoretical framework to explain normalization's role in capacity control.
result Normalization layers reduce the Lipschitz constant exponentially, smoothing the loss landscape and enhancing generalization.

Adding noise controls capacity of function compositions.

problem Large capacity of function compositions with bounded capacity classes.
method Adding Gaussian noise to the output of F\mathcal{F} before composing with H\mathcal{H}.
result Noise effectively controls the capacity of HF\mathcal{H} \circ \mathcal{F}, offering a general recipe for modular design.

Economic growth is unpredictable unless demand is quantified. We solve this problem by introducing the demand for unpaid spare time and a user quantity named human capacity. It organizes and amplifies spare time required for enjoying affluence like physical capital, the technical infrastructure for production, organize…

2012-06-12abs ↗pdf ↗

Improved neural network capacity analysis using simplified RDT.

problem Analyzing the memorization capabilities of sign perceptron neural networks.
method Developed a simplified, partially lifted Random Duality Theory (fl RDT) approach.
result Concrete capacity bounds universally improve over previous best known ones.

TVS-FNNs can approximate any continuous function on expanded input spaces.

problem Processing a broader range of inputs like sequences and matrices.
method Proving a universal approximation theorem for TVS-FNNs.
result TVS-FNNs can approximate any continuous function on expanded input spaces.

Study capacity constraints in continual learning with a simple model.

problem Understanding optimal resource allocation for agents with limited memory and compute resources.
method Analyzes a capacity-constrained linear-quadratic-Gaussian (LQG) sequential prediction problem and demonstrates optimal capacity allocation strategies.
result Derives a solution to the capacity-constrained LQG sequential prediction problem and shows how to optimally allocate capacity across sub-problems in the steady state.

Normalizing flows are shown to be equivalent to Bayesian networks, revealing new insights.

problem Understanding the limitations and capabilities of normalizing flows.
method Revisiting normalizing flows as probabilistic graphical models and analyzing their structure.
result Normalizing flows can be reduced to Bayesian networks, revealing new insights into their structure and capabilities.

Sparse codes improve optimal control tasks with correlated inputs.

problem Optimal control tasks with correlated feature inputs.
method Used a sparse code to represent natural images in an optimal control task solved with neuro-dynamic programming.
result An over-complete sparse code increases memory capacity and learning speed beyond a complete code.

Study optimizes pricing under uncertainty and capacity constraints.

problem Optimizing pricing decisions under demand uncertainty and capacity constraints.
method Analyzes linear demand, stochastic noise, and finite capacity; uses certified demand forecasts and control variates.
result Certified demand forecasts reduce regret from O(T)O(\sqrt{T}) to O(logT)O(\log T) under certain conditions.

This paper uses unsupervised learning and dimensionality reduction to optimize sewer system control.

problem Challenges in controlling large sewer systems efficiently.
method Divide sewer system into subcatchments, collect factors, cluster, apply PCA, simulate control scenarios.
result Priority control measures applied to clusters with similar hydraulic characteristics yield the best overflow reduction.

Deep residual networks can approximate any continuous function using control theory.

problem Universal approximation capabilities of deep residual neural networks.
method Relating residual networks to control systems and using Lie algebraic techniques.
result Deep residual networks with adequately deep layers can approximate any continuous function on a compact set.

Generalizes memory and forecasting capacities for nonlinear recurrent networks with dependent inputs.

problem Understanding memory and forecasting capabilities in networks with dependent inputs.
method Formulated bounds for memory and forecasting capacities in terms of network size and input properties.
result Proved that memory capacity for linear recurrent networks with independent inputs is given by the rank of the controllability matrix.

HardNet adds hard constraints to neural networks without sacrificing performance.

problem Ensuring adherence to input-dependent constraints in neural networks.
method Appends a differentiable enforcement layer to neural networks for end-to-end training with hard constraint guarantees.
result HardNet retains neural networks' universal approximation capabilities and enables efficient optimization.

Deep learning models can generalize well even when they fit training data perfectly.

problem Generalization in over-parameterized deep learning models.
method Combining empirical risk minimization with capacity control, exploring inductive biases and smooth empirical risk minimizers.
result Double descent phenomenon: test error can decrease after interpolation point.

Remove symmetries to improve model optimization and performance.

problem Symmetries in loss functions trap models in low-capacity states, hindering training and optimization.
method Proposes syre, a simple algorithm to remove symmetries in neural networks.
result Removing symmetries correlates well with improved optimization and performance.

The paper compares isoperimetric quotients and capacities in weighted manifolds.

problem Comparing isoperimetric quotients and capacities in weighted manifolds.
method Analysis of weighted Laplacian of the distance function and techniques for non-compact submanifolds.
result Parabolicity and hyperbolicity criteria for weighted manifolds.

We consider the problem of finding optimal strategies that maximize the average growth-rate of multiplicative stochastic processes. For a geometric Brownian motion the problem is solved through the so-called Kelly criterion, according to which the optimal growth rate is achieved by investing a constant given fraction o…

2015-10-17abs ↗pdf ↗

Modeling alignment as resource-limited cognitive processes, researchers derive performance bounds.

problem Systematic deviations in feedback-based alignment of large language models.
method Modeling alignment as a two-stage cascade UoHoYU o H o Y given SS, with cognitive and total capacities.
result Capacity-coupled Alignment Performance Interval derived from Fano and PAC-Bayes bounds.

Graph neural networks struggle to distinguish certain graph structures.

problem Difficulty in distinguishing graphs with graph neural networks.
method Analysis of communication capacity in message-passing model of graph neural networks.
result Capacity of MPNN needs to grow linearly for trees and quadratically for general connected graphs.

Paper applies theorem to find optimal investment boundary in stochastic capacity expansion.

problem Finding optimal investment boundary in a stochastic, time-inhomogeneous capacity expansion problem.
method Applies Bank and El Karoui Representation Theorem to solve first order conditions involving a non-integral term.
result Existence of base capacity ly(t)l^{\star}_y(t), showing optimal investment process becomes active at this level.

The paper tackles imbalanced classification under operational constraints, proposing a framework to maximize sensitivity.

problem Detecting minority class observations under severe class imbalance and operational constraints.
method Formal classification framework under capacity constraints, maximizing sensitivity while respecting a user-defined label limit.
result The optimal classifier under capacity constraints is equivalent to the Bayes classifier with reweighted prior probabilities.

Truncated Singular Value Decomposition (SVD) calculates the closest rank-kk approximation of a given input matrix. Selecting the appropriate rank kk defines a critical model order choice in most applications of SVD. To obtain a principled cut-off criterion for the spectrum, we convert the underlying optimization prob…

2011-02-15abs ↗pdf ↗

New study shows how model complexity affects test risk, challenging classical theory.

problem Understanding how test risk scales with model complexity for large over-parametrized deep networks.
method Developed norm-based capacity measures for random features based estimators, providing precise characterization of estimator's norm concentration and test error.
result Predicted learning curve shows a phase transition from under- to over-parameterization, confirming classical U-shaped behavior with appropriate capacity measures.

Transformers learn to cluster Gaussian mixtures as well as the EM algorithm.

problem Learning guarantees of Transformers in multi-class clustering of Gaussian mixtures.
method Developed a theory connecting Transformer's Softmax Attention layers to the EM algorithm's workflow.
result Transformers achieve minimax optimal rate for clustering Gaussian mixtures with sufficient training samples and initialization.

Optimal trading strategy in Proof-of-Stake blockchain using continuous-time control.

problem Finding the optimal balance between stake utility and consumption utility in Proof-of-Stake blockchain.
method Continuous-time control approach, dynamic programming, Hamilton-Jacobi-Bellman (HJB) equations.
result Close-form solutions for linear and convex utility functions, optimal strategies identified.

Hybrid controller combines model-based and policy-based reinforcement learning.

problem Combining model-based and policy-based reinforcement learning for stability and robustness.
method Designs a hybrid controller that interpolates a model-based linear controller and a differentiable policy.
result Proven to maintain stability and universal approximation properties.

We information-theoretically reformulate two measures of capacity from statistical learning theory: empirical VC-entropy and empirical Rademacher complexity. We show these capacity measures count the number of hypotheses about a dataset that a learning algorithm falsifies when it finds the classifier in its repertoire …

2011-11-23abs ↗pdf ↗

Proposes PIC and POIC for measuring task difficulty in RL.

problem Lack of metrics to measure task difficulty in RL.
method Introduces policy information capacity (PIC) and policy-optimal information capacity (POIC) as metrics based on mutual information.
result Empirically shows PIC and POIC correlate with task solvability better than alternatives.

RAF model explains neural networks' dual rule learning and fact memorization.

problem Understanding how neural networks learn rules and memorize facts simultaneously.
method Introduces the Rules-and-Facts (RAF) model to bridge generalization and memorization.
result Characterizes conditions for simultaneous rule learning and fact memorization in neural networks.

A new framework uses deep reinforcement learning to improve aircraft separation in busy airspace.

problem Improving aircraft separation in high-density, dynamic airspace constrained by human controllers.
method Proximal Policy Optimization with an attention network for distributed vehicle autonomy.
result The framework significantly reduces offline training time and increases performance.

The memory capacity of linear echo state networks is accurately calculated using new numerical methods.

problem Numerical evaluations of memory capacity in recurrent neural networks often contradict theoretical bounds.
method Developed robust numerical approaches exploiting MC neutrality with respect to the input mask matrix.
result Memory curves fully agree with theory when using the proposed methods.