Study excess capacity in neural networks using Rademacher complexity.
problem Understanding how much capacity deep networks have beyond what's needed for classification.
method Unified Rademacher complexity bounds for function composition and convolutional layers, considering Lipschitz constants and initialization norms.
result There is substantial excess capacity per task, and capacity can be kept similar across different tasks.
Paper develops an online learning algorithm for functional data models.
problem Recovering slope functions or predictors in functional data models.
method Online regularized learning algorithm in reproducing kernel Hilbert spaces with polynomially decaying step-size.
result Established fast convergence rates for estimation error without capacity assumption.
Exchanges acquire excess processing capacity to accommodate trading activity surges associated with zero-sum high-frequency trader (HFT) "duels." The idle capacity's opportunity cost is an externality of low-latency trading. We build a model of decentralized exchanges (DEX) with flexible capacity. On DEX, HFTs acquire …
Investor-driven information diffusion affects excess comovement in China and the U.S. markets.
problem Investor-driven information diffusion and its impact on excess comovement.
method Cross-sectional analysis of 4,533 Chinese and 4,517 U.S. stocks from 2010 to 2022.
result Retail-driven information diffusion significantly drives excess comovement in China, while institution-driven diffusion is the primary driver in the U.S.
Normalization layers control deep neural network capacity, improving stability and generalization.
problem Excessive capacity in deep neural networks leads to overfitting and poor generalization.
method Developed a theoretical framework to explain normalization's role in capacity control.
result Normalization layers reduce the Lipschitz constant exponentially, smoothing the loss landscape and enhancing generalization.
This study uses NLP to predict stock performance based on analyst reports.
problem Predicting stock performance using textual information from analyst reports.
method Natural language processing (NLP) and a customized BERT deep learning model for Chinese text.
result Strong positive sentiment in analyst reports increases excess return and intraday volatility, while strong negative sentiment increases volatility and trading volume but decreases excess return.
RAF model explains neural networks' dual rule learning and fact memorization.
problem Understanding how neural networks learn rules and memorize facts simultaneously.
method Introduces the Rules-and-Facts (RAF) model to bridge generalization and memorization.
result Characterizes conditions for simultaneous rule learning and fact memorization in neural networks.
New method targets sparsity to prevent overfitting in deep nets.
problem Overfitting in deep neural networks with small datasets.
method Targeted sparsity regularization to visualize and counteract overfitting.
result Significant increase in image classification performance without overfitting.
Statistical learning approach for spatial data prediction.
problem Predicting values at unknown locations from spatial data with complex dependence.
method Nonparametric finite-sample predictive analysis, kernel ridge regression.
result Non-asymptotic bounds for excess risk in isotropic stationary Gaussian processes.
GD outperforms ridge regression and SGD in linear regression problems.
problem Comparing the risks of GD, ridge regression, and SGD in linear regression problems.
method Instance-wise finite-sample risk analysis of GD, ridge regression, and SGD.
result GD outperforms ridge regression and is incomparable with SGD in some cases.
The paper analyzes the generalization of deep neural networks for metric and similarity learning.
problem Lack of rigorous understanding of generalization performance in metric and similarity learning.
method Derive explicit form of true metric, construct structured deep ReLU neural network, establish excess risk bounds.
result Explicit excess risk bounds for metric and similarity learning are derived.
Sparse portfolio strategy from mutual funds' favorite stocks in China A share market.
problem Building a sparse portfolio from mutual funds' favorite stocks in a market with limited fund information.
method Analyzed mutual fund favorite stocks, used portfolio optimizer with constraints, and compared different methods.
result Sparse portfolios consistently outperform the benchmark index 930950.CSI.
Researchers analyze backdoor data poisoning attacks and identify a memorization capacity parameter.
problem Understanding and mitigating backdoor data poisoning attacks in machine learning models.
method Formal theoretical framework, statistical and computational analysis, explicit constructions, and algorithm design.
result Identified a memorization capacity parameter to assess vulnerability to backdoor attacks and developed algorithms to detect and mitigate them.
Study on KRR with power-law data, showing better sample complexity.
problem High-dimensional kernel ridge regression with anisotropic power-law covariance.
method Explicit characterization of kernel spectrum and asymptotic analysis of excess risk.
result Sample complexity is governed by effective dimension, not ambient dimension.
Secret neural networks hidden within trained models.
problem Excess capacity in neural networks allows embedding secret models.
method Novel framework for hiding secret neural networks within carrier networks.
result Detection of hidden networks is computationally infeasible.
This paper extends financial theory to measure learnable market structure under computational constraints.
problem Understanding learnable market structure under bounded computational capacity.
method Introduces financial epiplexity as a measure of learnable market structure, extending classical information theory.
result Proves that equal entropy does not imply equal epiplexity and derives thresholds for useful regimes.
This review is about the convenience, the benefits, as well as the destructive capacities of money. It deals with various aspects of money creation, with its value, and its appropriation. All sorts of money tend to get corrupted by eventually creating too much of them. In the long run, this renders money worthless and …
This work uses PAC-Bayes for structured prediction with ILE, yielding insights and algorithms.
problem Structured prediction with interdependent outputs and implicit loss embeddings.
method PAC-Bayes perspective applied to ILE framework, deriving generalization bounds and learning algorithms.
result Two learning algorithms derived from PAC-Bayes bounds, analyzed and implemented.
Overprocuring reserves can improve network efficiency by using excess reserves for congestion management.
problem Optimizing energy and reserve allocation between zones to minimize costs and ensure deliverability.
method Developed allocation models for co-allocating traded energy and reserve products, considering both deterministic and stochastic flows.
result Excess reserve supplies can be used for congestion management, leading to additional network benefits.
Mathematical study of excess growth rate connects info theory with finance.
problem Understanding the excess growth rate in portfolio theory.
method Axiomatic characterization theorems of excess growth rate in terms of relative entropy, Jensen's inequality gap, and logarithmic divergence.
result Established rich connections between information theory and finance.
In statistical learning theory, convex surrogates of the 0-1 loss are highly preferred because of the computational and theoretical virtues that convexity brings in. This is of more importance if we consider smooth surrogates as witnessed by the fact that the smoothness is further beneficial both computationally- by at…
Paper improves learning rates for GSC loss functions using iterated Tikhonov regularization.
problem Improving learning rates for GSC loss functions.
method Iterated Tikhonov regularization using proximal point method.
result Achieves fast and optimal rates for GSC loss functions.
New tool detects 'fleeting modes' causing excess risk in financial markets.
problem Detecting portfolios with statistically significant excess risk in financial markets.
method Random Matrix Theory to identify 'fleeting modes' independent of underlying correlation structure.
result Fleeting modes exist in both futures and equity markets, and momentum is a source of excess risk.
The paper explores the information-theoretic nature of excess risk in machine learning.
problem Understanding the excess risk in machine learning models.
method Formulates the minimax excess risk as a zero-sum game and modifies it to allow swapping of the order of play.
result Proves that under certain conditions, the duality gap is zero, allowing for the application of Bayesian results to provide bounds on minimax excess risk.
We study various capacities on compact Kähler manifolds which generalize the Bedford-Taylor Monge-Ampère capacity. We then use these capacities to study the existence and the regularity of solutions of complex Monge-Ampère equations.
A new formula reveals symmetries between mean excess and ES functions.
problem Optimizing risk measures in financial models.
method Established a reverse ES optimization formula.
result Reveals elegant symmetries and relationships between mean excess and ES functions.
The paper analyzes the excess risk of PCA and provides a precise characterization.
problem Understanding the excess risk of principal component analysis (PCA).
method Established a central limit theorem for PCA error and derived the excess risk distribution.
result Obtained a non-asymptotic upper bound on the excess risk of PCA.
Solves a discrete logarithmic Minkowski problem for electrostatic p-capacity.
problem Characterize measures generated by electrostatic p-capacity.
method Solves the discrete logarithmic Minkowski problem for 1 < p < n.
result Solves the discrete logarithmic Minkowski problem for measures in general position.
Study excess risk in statistical inference with transformations.
problem Excess risk in estimating random variables from feature vectors and transformations.
method Characterize lossless transformations, develop test statistics, and information-theoretic bounds.
result Strongly consistent partitioning test statistic for lossless transformations.
CapOptix uses options theory to price capacity in electricity markets.
problem Traditional capacity market designs fail to account for risk and price shocks.
method Interprets capacity commitments as reliability options and uses Markov Regime Switching Process.
result CapOptix provides more accurate pricing of capacity premia compared to existing mechanisms.
Over-parameterized models reduce Out-of-Distribution (OOD) generalization loss.
problem Understanding how over-parameterized models handle non-trivial distributional shifts.
method Investigating random feature models and examining non-trivial natural distributional shifts.
result Increasing model parameterization reduces OOD loss.
In this article, we propose the notion of the general p-affine capacity and prove some basic properties for the general p-affine capacity, such as affine invariance and monotonicity. The newly proposed general p-affine capacity is compared with several classical geometric quantities, e.g., the volume, the p-var…
While symplectic manifolds have no local invariants, they do admit many global numerical invariants. Prominent among them are the so-called symplectic capacities. Different capacities are defined in different ways, and so relations between capacities often lead to surprising relations between different aspects of sympl…
Study rigidity by logarithmic capacity and related functions.
problem Rigidity phenomena in kernel functions and capacities.
method Exploration of Bergman kernel, logarithmic capacity, Green's function, and Euclidean distance/volume.
result Established rigidity theorems by logarithmic capacity.
Study binary perceptrons' capacity using random duality theory.
problem Characterize the capacity of binary perceptrons with general thresholds.
method Utilized fully lifted random duality theory (fl RDT) to characterize the capacity.
result Characterizations match replica symmetry breaking predictions and uncover the capacity for zero-threshold scenario.
Study capacity constraints in continual learning with a simple model.
problem Understanding optimal resource allocation for agents with limited memory and compute resources.
method Analyzes a capacity-constrained linear-quadratic-Gaussian (LQG) sequential prediction problem and demonstrates optimal capacity allocation strategies.
result Derives a solution to the capacity-constrained LQG sequential prediction problem and shows how to optimally allocate capacity across sub-problems in the steady state.
New complete panel dataset for LMICs helps analyze innovation and development.
problem Lack of complete data for empirical analyses in LMICs.
method Predictive Mean Matching multiple imputation technique.
result Created a large dataset of 47 variables for 82 LMICs from 2005-2019.
Upper bounds for Lagrangian capacities of Liouville domains
problem Lagrangian capacity of Liouville domains
method Using S1-equivariant techniques result Extremal Lagrangian torus on the boundary of ellipsoid
Memory capacity of DAM scales exponentially with feature separation, unaffected by correlations.
problem Understanding how feature correlations impact DAM's capacity.
method Developed an empirical framework to analyze DAM's capacity under varying feature correlations and pattern separations.
result Memory capacity scales exponentially with feature separation, unaffected by correlations.
Proves local maximizers for higher Ekeland-Hofer capacities in 4D star-shaped domains.
problem Finding local maximizers for higher Ekeland-Hofer capacities in specific domains.
method Analogous to 4D local Viterbo conjecture, proving maximizers for rational ellipsoids.
result Local maximizers of the k-th Ekeland-Hofer capacities are symplectomorphic to rational ellipsoids.
We use a continuous-time random walk (CTRW) to model market fluctuation data from times when traders experience excessive losses or excessive profits. We analytically derive "superstatistics" that accurately model empirical market activity data (supplied by Bogachev, Ludescher, Tsallis, and Bunde)that exhibit transitio…
Develops a theory for mth order p-affine capacity for convex bodies containing the origin.
problem Defines and studies the mth order p-affine capacity for convex bodies containing the origin.
method Provides equivalent definitions, proves properties, and establishes inequalities.
result Establishes inequalities comparing to other geometric measures.
Derives an empirical capacity model for self-attention neural networks.
problem Theoretical capacity of large transformer models is not fully utilized by current optimization algorithms.
method Analyzes memory capacity of transformers using synthetic training data and common training algorithms.
result Derives an empirical capacity model (ECM) for a generic transformer.
Improves online learning algorithms for functional models with capacity assumptions.
problem Convergence rates of online stochastic gradient descent algorithms for functional linear models.
method Characterizations of slope function regularity, kernel space capacity, and sampling process covariance operator.
result Capacity assumptions can alleviate saturation of convergence rates as function regularity increases.
We introduce the concept of pseudo symplectic capacities which is a mild generalization of that of symplectic capacities. As a generalization of the Hofer-Zehnder capacity we construct a Hofer-Zehnder type pseudo symplectic capacity and estimate it in terms of Gromov-Witten invariants. The (pseudo) symplectic capacitie…
Study relates symplectic homology capacity to periodic orbits in Liouville domains.
problem Relating symplectic homology capacity to periodic orbits in Liouville domains.
method Uses positive symplectic homology and Hofer-Zehnder capacity to establish bounds and existence of periodic points.
result Non-zero positive symplectic homology implies finite upper bound for Hofer-Zehnder capacity relative to skeleton and Hamiltonian diffeomorphisms.
Learning capacity measures model complexity, correlating with test loss and sample size.
problem Understanding model complexity and its relation to test performance.
method Formal correspondence between thermodynamics and inference; learning capacity as a measure of effective dimensionality.
result Learning capacity correlates with test loss and is a small fraction of model parameters.
Proponents of behavioral finance have identified several "puzzles" in the market that are inconsistent with rational finance theory. One such puzzle is the "excess volatility puzzle". Changes in equity prices are too large given changes in the fundamentals that are expected to change equity prices. In this paper, we of…