New algorithm ensures global convergence in deep neural networks beyond NTK regime.
problem Existing global convergence guarantees do not apply to practical deep networks.
method Proposes an algorithm with global convergence guarantees under the expressivity condition.
result Algorithm ensures global convergence in practical settings beyond NTK regime.
BLAE solves batched linear bandits with optimal regret and practical performance.
problem Batched linear bandit problem with limited adaptivity.
method Integrates arm elimination with regularized G-optimal design, achieving minimax optimal regret.
result Achieves minimax optimal regret in both large- K K K and small- K K K regimes with O ( log log T ) O(\log\log T) O ( log log T ) batches. A practical algorithm improves approximate OT distances using quantization.
problem Substantial computational burden in computing OT distances for large samples.
method Introduces a quantization step to estimate OT distances between measures.
result The quantization step improves the performance of approximate solvers for entropy-regularized transport.
This paper studies nonlinear representation learning dynamics beyond the NTK regime.
problem Efficient reasoning and inference in raw sensory data representations.
method Identifies common model structure assumption and data-architecture alignment condition for global convergence and optimality.
result Theoretical framework explains network size effects and provides practical model structure guidelines.
New framework estimates treatment effects in extreme data.
problem Hindered by unavailability of counterfactual outcomes and rarity of extreme data.
method Proposes a new framework based on extreme value theory.
result Quantifies treatment effects using tail decay rates of potential outcomes.
NTK theory fails to predict practical behavior of large-width neural networks.
problem Theoretical limits of NTK do not match practical neural network architectures.
method Empirical investigation of NTK's applicability to large-width architectures.
result Practically relevant behavior of large-width architectures differs from NTK theory.
SRRM improves recursive transport surrogates in the small-discrepancy regime.
problem Insufficient understanding of recursive partitioning methods' statistical behavior and resolution in the small-discrepancy regime.
method Introduced Selective Recursive Rank Matching (SRRM) to improve the resolution of Recursive Rank Matching (RRM).
result SRRM yields a higher-fidelity practical surrogate for the Wasserstein distance at moderate additional computational cost.
Quantum algorithms for financial derivatives and credit risk.
problem Estimating credit risk and option pricing in realistic financial models.
method Developed a regime switching volatility model for financial markets, using a Markov chain to determine volatility parameters.
result Quantum algorithms can be applied to realistic financial models, bringing quantum computing closer to practical applications.
The paper tackles adaptive targeting in networks with interference effects.
problem Adaptive targeting under network interference in a bandit setting.
method Linear model in a sparse regime, analyzing different levels of knowledge of the interference structure.
result Unified view of how knowledge of the interference structure affects online learning efficiency.
A TTA framework improves forecasting accuracy in non-stationary time series.
problem Improving forecasting accuracy in non-stationary time series.
method Normalization-based test-time adaptation for causal timeseries forecasting and direction classification.
result Normalization-based TTA improves forecasting error in synthetic gradual drift and can even hurt in aggressive norm-only adaptation in financial markets.
Investigates JM for reducing downside risk in market regimes.
problem Mitigating downside risk during market downturns.
method Statistical jump model for identifying market regimes, optimizing penalty for state transitions.
result JM-guided strategies outperform traditional models in reducing risk and enhancing returns.
Large learning rates work surprisingly well in standard parameterization, contrary to theory.
problem Theoretical limits of large learning rates do not match practical network behavior.
method Fine-grained analysis of learning rates and network behavior under cross-entropy loss.
result There are two distinct sub-regimes of unstable learning rates, with a controlled divergence regime where features continue to evolve.
Measures strategy durability through minimum regime performance, revealing trade-offs between efficiency and resilience.
problem Systematic investing strategies are vulnerable to regime changes, affecting their effectiveness and performance.
method Introduces minimum regime performance (MRP) to quantify the durability of systematic strategies, capturing how performance deteriorates under changing market conditions.
result Higher long-term Sharpe ratios do not always correlate with higher MRP, highlighting a new dimension of portfolio fragility.
We develop and evaluate tolerance interval methods for dynamic treatment regimes (DTRs) that can provide more detailed prognostic information to patients who will follow an estimated optimal regime. Although the problem of constructing confidence intervals for DTRs has been extensively studied, prediction and tolerance…
Balls-and-Bins sampling improves DP-SGD privacy and utility.
problem Improving privacy and utility in DP-SGD implementations.
method Introducing Balls-and-Bins sampling as an alternative to shuffling in DP-SGD.
result Balls-and-Bins sampling achieves utility comparable to shuffling while offering better privacy amplification.
Paper derives analytical formulas for NLD-CEV moments with regime switching.
problem Analytical tractability of NLD-CEV models under stochastic regimes.
method Hybrid system approach using Feynman-Kac formula for solving interconnected PDEs.
result Exact closed-form expressions for fractional-order conditional moments.
Deep neural networks can interpolate any dataset in the overparametrized regime.
problem Interpolating any dataset with deep neural networks in the overparametrized regime.
method Proving universal approximations and interpolating any dataset with deep neural networks, considering specific conditions on activation functions.
result Interpolation of any dataset is possible in the overparametrized regime with deep neural networks.
Develops a method to estimate personalized treatment regimes from summary statistics.
problem Estimating optimal treatment regimes for a target population when individual-level data is unavailable.
method A weighting framework that tailors a treatment regime for the target population using summary statistics.
result Consistent and asymptotically normal estimator for optimal treatment regimes.
Framework models multiscale dynamics with Bayesian learning for regime changes.
problem Analyzing complex interactions between fast and slow processes.
method Hierarchical state-space modeling with Sequential Monte Carlo.
result Bayesian approach accurately tracks state transitions and identifies switching dynamics.
New insights into how neural networks learn features, especially when they are very wide.
problem Understanding how gradient flow in wide neural networks selects solutions, especially in the feature-learning regime.
method Axiomatizing the canonical regularizer as a function-space energy and lift, and deriving geodesic ridge for the feature-learning regime.
result Gradient flow in feature-learning networks biases towards ridge regularization, distorting the inductive bias and damaging pretrained networks.
Volatility forecasting and return prediction in high-frequency Chinese equity markets.
problem Improving statistical forecasting performance and economic strategy outcomes in equity markets.
method Developing a sequential two-stage framework combining realized volatility modeling and XGBoost return prediction.
result Regime-aware volatility forecasting outperforms baseline models.
FTPL policy achieves best-of-both-worlds regret in decoupled bandits with reduced computational cost.
problem Decoupled multi-armed bandit problem with observed and unobserved losses.
method Follow-the-Perturbed-Leader (FTPL) policy that avoids convex optimization and resampling.
result Achieves constant regret in stochastic regime and optimal O ( K T ) O(\sqrt{KT}) O ( K T ) regret in adversarial regime. This study examines how ChiNext IPOs' initial returns are influenced by regulation regime changes.
problem Investors' behavior and pricing of ChiNext IPOs under different regulation regimes.
method Analysis of three time periods with two different regulation regimes and three sets of listing day trading restrictions.
result Regulation regime changes significantly impact ChiNext IPO pricing and overreaction.
New bounds for high-dimensional sparse linear bandits, balancing information and regret.
problem Stochastic linear bandits with high-dimensional sparse features.
method Derivation of minimax regret lower and upper bounds for explore-then-commit algorithm.
result Optimal rate of Θ ( n 2 / 3 ) Θ(n^{2/3}) Θ ( n 2/3 ) for data-poor regime, complemented by O ( n ) O(\sqrt{n}) O ( n ) under signal magnitude assumption. New model clusters mixed-type data with missing values, improving air quality analysis.
problem Clustering mixed-type data with missing values and regime persistence.
method Statistical jump model incorporating regime persistence and handling missing data.
result Superior performance in inferring persistent air quality regimes compared to traditional methods.
Proposes pT-Learning for optimal dynamic treatment regimes in mHealth.
problem Challenges in learning optimal dynamic treatment regimes with large intervention options and infinite time horizon.
method Proximal Temporal consistency Learning (pT-Learning) framework for adaptively adjusting between deterministic and stochastic policies.
result Minimax estimator avoids double sampling issue and can incorporate off-policy data.
How initialization and loss function affect the learning of a deep neural network (DNN), specifically its generalization error, is an important problem in practice. In this work, by exploiting the linearity of DNN training dynamics in the NTK regime \citep{jacot2018neural,lee2019wide}, we provide an explicit and quanti…
Study on rich regime training in deep learning, finding active parameters in bottom layers.
problem Understanding the practical success of deep learning models.
method Empirical study on rich regime training with benchmark datasets, re-initialization analysis, and probabilistic Layer-Wise Sparse SGD.
result Probabilistic Layer-Wise Sparse SGD matches vanilla SGD's generalization performance with improved efficiency.
Study examines how training regime affects neural networks' forgetting.
problem Catastrophic forgetting in neural networks when learning multiple tasks sequentially.
method Analyzes the impact of different training regimes (learning rate, batch size, regularization) on forgetting.
result Training regimes that widen tasks' local minima help prevent catastrophic forgetting.
HireVAE adapts to market regimes for online stock prediction.
problem Building an online and adaptive factor model for stock prediction.
method HireVAE uses a hierarchical latent space to estimate latent factors from historical market information.
result HireVAE outperforms previous methods in active returns across benchmarks.
New analysis shows optimal embedding learning rate depends on vocabulary size, not just model width.
problem Optimal learning rate for language model embeddings is not well understood, especially with large vocabularies.
method Theoretical analysis of training dynamics, interpolation between μ μ μ P and LV regimes. result Optimal embedding learning rate scales as Θ ( w i d t h ) Θ(\sqrt{width}) Θ ( w i d t h ) in the LV regime, not Θ ( w i d t h ) Θ(width) Θ ( w i d t h ) as μ μ μ P predicts. Linearized attention fails to converge to NTK limit even at large widths.
problem Understanding the convergence of attention mechanisms to the kernel regime.
method Analyzes linearized attention and its relationship to the NTK limit, considering practical widths and conditions.
result Linearized attention does not converge to its NTK limit at any practical width, revealing a fundamental trade-off.
Develops methods for personalized treatment decisions in the presence of unmeasured factors.
problem Personalized treatment decisions in the presence of unmeasured confounding.
method Proximal learning approaches to estimate optimal individualized treatment regimes (ITRs).
result Established identification results for different classes of ITRs, improving decision-making value function.
Gradient descent can find better tensor decompositions than lazy training in over-parameterized settings.
problem Finding better tensor decompositions in over-parameterized settings.
method Gradient descent on over-parameterized tensor decomposition problems.
result Gradient descent can find an approximate tensor decomposition with rank m = O ∗ ( r 2.5 l log d ) m = O^*(r^{2.5l}\log d) m = O ∗ ( r 2.5 l log d ) , while lazy training requires m = Ω ( d l − 1 ) m = Ω(d^{l-1}) m = Ω ( d l − 1 ) . Paper proposes a new framework to compare trading strategies by accounting for market conditions.
problem Lack of information on how trading strategy performance varies with market conditions.
method Uses a GAMLSS/ZAGA framework to model the Adjusted Information Ratio ( I R ∗ IR^{\ast} I R ∗ ) for a SVMP and BH strategy across 146 folds of the S&P 500. result Dominance of SVMP over BH is conditional on market regime, as shown by differences in expected I R ∗ IR^{\ast} I R ∗ and its variance. This work investigates square loss in overparametrized neural networks, revealing its advantages in robustness and calibration.
problem Theoretical understanding of square loss in overparametrized neural networks.
method Systematic investigation of square loss in the NTK regime for both separable and non-separable classes.
result Square loss shows fast convergence rates and robustness guarantees for overparametrized neural networks.
Log-normal continuous random cascades form a class of multifractal processes that has already been successfully used in various fields. Several statistical issues related to this model are studied. We first make a quick but extensive review of their main properties and show that most of these properties can be analytic…
Framework uses hindsight regret to audit marketing budget allocations.
problem Lack of principled way to assess strategic budget allocations.
method Hindsight regret framework based on constraint-faithful benchmark.
result Identifies practical trade-off between allocation flexibility and detectability.
A hybrid framework for American option pricing under time-varying rough volatility.
problem Pricing American options under time-varying rough volatility.
method Signature method combined with gradient-boosted ensemble for Hurst parameter estimation, regime switch, and Random Fourier Features for acceleration.
result The proposed hybrid framework improves performance over fixed-roughness baselines and reduces duality gaps in some regimes.
New method reveals true causal functions in nonlinear time series, not just scores.
problem Causal discovery in nonlinear time series often uses scalar edge scores, which hide true function-valued causal influence.
method Formalized function-valued causal influence for additive, contribution-decomposable architectures. Introduced a practical framework based on ICE for estimating causal response functions directly from trained models.
result Edges with indistinguishable scalar scores can exhibit qualitatively different functional behaviors.
We study the valuation and hedging problem of European options in a market subject to liquidity shocks. Working within a Markovian regime-switching setting, we model illiquidity as the inability to trade. To isolate the impact of such liquidity constraints, we focus on the case where the market is completely static in …
Backtests of structured strategies lose much of their predictive power in live trading.
problem Uncertainty in how marketed backtests predict live performance of structured strategies.
method Analysis of 1,726 structured strategies from ten global institutions.
result Raw backtests have limited portability into live trading and deteriorate sharply.
Deep Bayesian models estimate causal effects for dynamic treatment regimes over long follow-up times.
problem Challenges in causal effect estimation for dynamic treatment regimes with long follow-up times.
method Combining outcome regression models with deep Bayesian models for high-dimensional features.
result Stable and accurate dynamic causal effect estimation from observational data, especially with long-term follow-up.
A fast method for decentralized non-convex optimization over networks.
problem Decentralized non-convex optimization problems over a network of nodes.
method GT-SAGA, a randomized incremental gradient method that evaluates one component gradient per node per iteration.
result GT-SAGA achieves almost sure and mean-squared convergence to a first-order stationary point for general smooth non-convex problems.
New analysis reveals batch size effects on stochastic conditional gradient methods.
problem Understanding the role of batch size in stochastic conditional gradient methods.
method Deriving a new analysis focusing on momentum-based stochastic conditional gradient algorithms (e.g., Scion).
result Increasing batch size initially improves optimization accuracy but can degrade performance beyond a critical threshold.
DIVI clusters noisy high-dimensional data with stable feature gating.
problem Challenging clustering in high-dimensional noisy data.
method Data-informed variational clustering framework combining global feature gating and adaptive structure growth.
result DIVI performs competitively under severe feature noise and remains computationally feasible.
Momentum affects optimization differently at small vs large batch sizes near instability.
problem Understanding how momentum impacts optimization near the edge of stability.
method Demonstrated through batch-size dependent behavior of SGD with momentum.
result Momentum operates in two distinct regimes: amplifying stochastic fluctuations at small batch sizes and stabilizing at large batch sizes.
Study identifies two borrowing patterns in UK payday loan users.
problem Financial vulnerability of payday loan users.
method Two-state hidden Markov model (HMM) using Open Banking data.
result 36.4% of borrowers experience high-intensity exposure for 12 weeks or more.