Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

60120180240 · May 202619922001200920172026
48 results for fixed residues

MLP residual networks implement a selective coarse-graining procedure governed by the spectral structure of the input distribution.

problem Understanding the coarse-graining procedure in MLP residual networks
method Analyzing a pure MLP residual stack on synthetic Markov chain sequences
result MLP residual networks implement a selective coarse-graining procedure governed by the spectral structure of the input distribution

REMAL: Residual Equilibrium Manifold Active Learning for Surrogate-Based Multidisciplinary Design Analysis

problem Multidisciplinary design analysis of coupled engineering systems requires solving equilibrium states where all disciplinary coupling variables are consistent.
method Residual manifold surrogate modeling framework for coupled systems.
result REMAL learns a surrogate model of the joint residual manifold via multitask Gaussian process models.

Using the notion of equivariant Kirwan map, as defined by Goldin, we prove that -- in the case of Hamiltonian torus actions with isolated fixed points -- Tolman and Weitsman's description of the kernel of the Kirwan map can be deduced directly from the residue theorem of Jeffrey and Kirwan. A characterization of the ke…

2002-11-06abs ↗pdf ↗

Study conjugacy classes of parabolic diffeomorphisms fixing the origin.

problem Understanding conjugacy classes of parabolic diffeomorphisms fixing the origin.
method Establish results on differentiability classes and order of tangency, focusing on the invariance of residues under low-regular conjugacies.
result Sharp results on invariance of residues under low-regular conjugacies, extending previous work on Schwarzian derivatives.

The paper establishes principles for initializing and designing GNNs with ReLU activations to avoid oversmoothing and correlation collapse.

problem Oversmoothing and correlation collapse in deep ReLU GNNs.
method The paper derives and validates three principles for initialization and architecture selection in finite width graph neural networks with ReLU activations.
result Correct initialization, residual aggregation operators, and residual connections significantly improve early training dynamics in deep ReLU GNNs.

Improved stochastic approximation method reduces residual error.

problem Reducing residual error in stochastic approximation algorithms.
method Fixed-schedule one-quarter barrier and bias-corrected acceleration.
result Achieves T1/2+o(1)T^{-1/2+o(1)} residual reduction with O(1)O(1) primitive samples.

Consider the space RΔR_Δ of rational functions of several variables with poles on a fixed arrangement ΔΔ of hyperplanes. We obtain a decomposition of RΔR_Δ as a module over the ring of differential operators with constant coefficients. We generalize to the space RΔR_Δ the notions of principal part and of residue, and …

1999-03-30abs ↗pdf ↗

Residual connections significantly boost the performance of deep neural networks. However, there are few theoretical results that address the influence of residuals on the hypothesis complexity and the generalization ability of deep neural networks. This paper studies the influence of residual connections on the hypoth…

2019-04-02abs ↗pdf ↗

In this effort, we propose a new deep architecture utilizing residual blocks inspired by implicit discretization schemes. As opposed to the standard feed-forward networks, the outputs of the proposed implicit residual blocks are defined as the fixed points of the appropriately chosen nonlinear transformations. We show …

2019-05-24abs ↗pdf ↗

Normalization layers are a staple in state-of-the-art deep neural network architectures. They are widely believed to stabilize training, enable higher learning rate, accelerate convergence and improve generalization, though the reason for their effectiveness is still an active research topic. In this work, we challenge…

2019-01-27abs ↗pdf ↗

An automorphism αα of a group GG is normal if it fixes every normal subgroup of GG setwise. We give an algebraic description of normal automorphisms of relatively hyperbolic groups. In particular, we prove that for any relatively hyperbolic group GG, Inn(G)Inn(G) has finite index in the subgroup Autn(G)Aut_n(G) of normal au…

2008-09-14abs ↗pdf ↗

CRC improves multivariate forecasting accuracy without risking performance degradation.

problem Systematic errors and lack of guarantees in multivariate forecasters.
method CRC uses a causality-inspired encoder and hybrid corrector with a safety mechanism.
result CRC consistently improves accuracy and ensures high non-degradation rates.

Study infinite-depth limits of neural networks with fixed width.

problem Understanding the behavior of neural networks as depth increases with fixed width.
method Analyzing finite-width residual networks with random Gaussian weights, focusing on the infinite-depth limit.
result The pre-activations converge to a zero-drift diffusion process, differing from the infinite-width limit.

The paper introduces a diagnostic method to detect grokking transitions in models before test accuracy improves.

problem Detecting the transition from training to generalization in machine learning models.
method Summarize task-dependent observables as empirical distributions, map them to Wasserstein/quantile coordinates, and analyze using Hankel dynamic mode decomposition.
result The diagnostic method achieves AUROC \(\approx\) 0.93 for grokking-vs-non-grokking discrimination at the run level.

MGDL refines deep neural networks by training grades sequentially, improving stability.

problem Training deep neural networks is challenging due to nonconvex optimization landscapes.
method MGDL trains deep networks grade by grade, freezing previously learned grades and training new ones to fit residuals.
result MGDL guarantees vanishing error in a fixed-width multigrade ReLU architecture.

Study Transformer layers under cross-entropy training using mean field control.

problem Understanding the behavior of Transformer layers in cross-entropy training.
method Continuous-depth mean field control analysis, treating depth as time and layer parameters as controls.
result Derivation of a Pontryagin condition for the limiting population problem, involving the softmax residual.

This paper improves bond market making by adjusting hit-ratios for client flow quality.

problem Economic misleading of raw hit-ratios in corporate bond market making.
method Stochastic-control framework with residual-quality-adjusted hit-ratio.
result Optimal quotes decompose into various components, improving service/economics frontier.

We show that any smooth bi-Lipschitz hh can be represented exactly as a composition hm...h1h_m \circ ... \circ h_1 of functions h1,...,hmh_1,...,h_m that are close to the identity in the sense that each (hiId)\left(h_i-\mathrm{Id}\right) is Lipschitz, and the Lipschitz constant decreases inversely with the number mm of functions com…

2018-04-13abs ↗pdf ↗

On a complex curve, we establish a correspondence between integrable connections with irregular singularities, and Higgs bundles such that the Higgs field is meromorphic with poles of any order. The moduli spaces of these objects are obtained by fixing at each singularity the polar part of the connection. We prove that…

2001-11-08abs ↗pdf ↗

Unified learning-rate scale for CNNs and ResNets, avoiding depth imbalance.

problem Challenges in choosing an appropriate learning rate for deep networks, especially as depth increases.
method Introduces Arithmetic-Mean μμP (AM-μμP), constraining network-wide average pre-activation second moment to a constant scale, combined with residual-aware He fan-in initialization.
result Demonstrates a 3/2-3/2 scaling law for learning rates across depths, enabling zero-shot learning-rate transfer.

In this paper we investigate panel regression models with interactive fixed effects. We propose two new estimation methods that are based on minimizing convex objective functions. The first method minimizes the sum of squared residuals with a nuclear (trace) norm regularization. The second method minimizes the nuclear …

2018-10-25abs ↗pdf ↗

VR-GHAL method solves stochastic fixed-point equations with high probability.

problem Solving stochastic fixed-point equations in normed spaces with nonexpansive or contractive operators.
method VR-GHAL, a variance-reduced gradual Halpern method for quadratically smoothable Banach spaces, using clipped stochastic differences.
result The method achieves a high-probability residual bound, reducing the residual nearly geometrically across epochs.

We present in this article a family of new combinatorial identities via purely differential/complex geometry methods, which include as a speical case a unified and explicit formula for Chern numbers of all complex flag manifolds. Our strategy is to construct concrete circle actions with isolated fixed points on these m…

2017-02-06abs ↗pdf ↗

Deep linear ResNets converge globally with certain transformations.

problem Global convergence of training deep linear ResNets.
method Gradient descent and stochastic gradient descent for training LL-hidden-layer linear ResNets.
result GD and SGD can converge to global minimum for deep linear ResNets with specific transformations.

We prove a functorial correspondence between a category of logarithmic sl2\mathfrak{sl}_2-connections on a curve XX with fixed generic residues and a category of abelian logarithmic connections on an appropriate spectral double cover π:ΣXπ: Σ\to X. The proof is by constructing a pair of inverse functors $π^{\text{ab}}, π…

2019-02-09abs ↗pdf ↗

Improved stochastic Halpern iteration for fixed-point approximation in normed spaces.

problem Approximating fixed-points of nonexpansive and contractive operators in normed finite-dimensional spaces.
method Stochastic Halpern iteration with minibatch, analyzing oracle complexity.
result Improved oracle complexity for nonexpansive operators, with a lower bound of Ω(ε3)Ω(\varepsilon^{-3}).

Let MM be a compact, connected symplectic manifold with a Hamiltonian action of a compact nn-dimensional torus G=TnG=T^n. Suppose that σσ is an anti-symplectic involution compatible with the GG-action. The real locus of MM is XX, the fixed point set of σσ. Duistermaat uses Morse theory to give a description of the…

2001-07-20abs ↗pdf ↗

PROBE optimizes best-arm identification with cheap proxies, improving sample complexity.

problem Fixed-confidence best-arm identification with costly rewards and correlated cheap proxies.
method PROBE uses control-variate adjustment and phase elimination to learn residual variance online.
result PROBE achieves oracle sample complexity up to a constant factor and additive calibration cost.

A method for finding most influential sets reduces a complex problem to a sequence of simpler top-kk problems.

problem Identifying most influential subsets in complex models.
method Reduces the problem to a sequence of top-kk problems using Dinkelbach's method.
result The method returns a globally optimal set for the univariate ratio objective, including partial linear models.

This paper contains several results concerning circle action on almost-complex and smooth manifolds. More precisely, we show that, for an almost-complex manifold M2mnM^{2mn}(resp. a smooth manifold N4mnN^{4mn}), if there exists a partition λ=(λ1,...,λu)λ=(λ_{1},...,λ_{u}) of weight mm such that the Chern number $(c_{λ_{1}}... c_{λ_{…

2010-08-28abs ↗pdf ↗

Residual finiteness is known to be an important property of groups appearing in combinatorial group theory and low dimensional topology. In a recent work [2] residual finiteness of quandles was introduced, and it was proved that free quandles and knot quandles are residually finite. In this paper, we extend these resul…

2019-02-08abs ↗pdf ↗