In this note we study logarithmic transformations in the sense of differential topology on two fibers of the Hopf surface. It is known that such transformations are susceptible to yield exotic smooth structures on four-manifolds. We will show here that this is not the case for the Hopf surface, all integer homology Hop…
Develops a new option pricing model under G-expectation framework.
problem Modeling uncertainty in financial markets and robust valuation under model uncertainty.
method G-expectation framework, logarithmic transformation, finite difference schemes.
result Unified risk-neutral valuation approach yielding G-Black-Scholes equation.
Solutions near infinity to special Lagrangian equations are asymptotic to quadratic polynomials with logarithmic terms.
problem Solving special Lagrangian equations near infinity with specific conditions.
method Modified Kelvin transforms to characterize remainders in asymptotic expansions.
result Remainders in asymptotic expansions are characterized by a single smooth function in even dimensions and Cn−1,α in odd dimensions. Applying logarithmic transformations along 2-tori, we construct a generalized complex structure J_n with n type changing luci for every n≥0 on genus 1-Lefschetz fibrations with a cusp neighborhood, which include elliptic surfaces with non-zero euler characteristic. Applying a technique of broken Lefschetz fibrati…
Transformers capture combinatorial tasks with bounded error and logarithmic sample dependence.
problem Capturing complex combinatorial tasks with bounded error and sample efficiency.
method Formal definition of algorithmic capture, empirical analysis of infinite-width transformers, upper bounds on computational complexity.
result Transformers exhibit an inductive bias favoring simpler algorithmic procedures over higher complexity ones.
This paper compares Transformers and RNNs in various tasks, showing size differences.
problem Comparing representational capabilities of Transformers and RNNs across tasks.
method Analysis of differences in tasks like index lookup, nearest neighbor, and string equality.
result Size differences in Transformers and RNNs for various tasks.
The paper models financial asset prices with jumps and evaluates European option prices using numerical methods.
problem Modeling and pricing European options with jumps in delayed stochastic systems.
method Existence, uniqueness, and positivity of solutions to delayed stochastic differential equations with jumps. Application of Fourier transformation for analytical pricing and Monte-Carlo simulation with a logarithmic Euler-Maruyama scheme for numerical approximation.
result The logarithmic Euler-Maruyama scheme provides a positive and convergent method for approximating the solution to the delayed stochastic differential equations with jumps.
Round handles are affiliated with smooth 4-manifolds in two major ways: 5-dimensional round handles appear extensively as the building blocks in cobordisms between 4-manifolds, whereas 4-dimensional round handles are the building blocks of broken Lefschetz fibrations on them. The purpose of this article is to shed more…
We study the logarithmic L(α)-divergence which extrapolates the Bregman divergence and corresponds to solutions to novel optimal transport problems. We show that this logarithmic divergence is equivalent to a conformal transformation of the Bregman divergence, and, via an explicit affine immersion, is equivalent t…
The paper establishes a Poisson Poincaré-Dulac theorem for Poisson-flat connections.
problem Analyzing Poisson-flat connections with logarithmic poles.
method Defining an Euler-Poisson principal part and residue theory, establishing a Poisson Poincaré-Dulac theorem.
result Any logarithmic Poisson-flat connection is holomorphically gauge equivalent to a pure Euler-Poisson normal form.
The paper derives formulas for pricing geometric Asian options in the Volterra-Heston model.
problem Pricing geometric Asian options in the Volterra-Heston model.
method Derives semi-closed formulas using Fourier transforms and Riccati-Volterra equations.
result Derives formulas for pricing geometric Asian options with fixed and floating strikes.
Optimal unimodal fitting for linear loss functions in a sequential, efficient manner.
problem Optimal unimodal transformation of univariate model scores under linear loss functions.
method Proposes a sequential approach to estimate the optimal rectangular fit for observed samples with each new sample.
result Sequential approach achieves optimal efficiency with logarithmic time complexity per iteration.
Paper analyzes risk bounds for in-context learning in multiclass classification.
problem Risk bounds for in-context learning in multiclass classification.
method Formalizes tasks as sequences of labeled examples and queries, estimates conditional class probabilities, establishes oracle inequality for KL divergence.
result ICL achieves minimax optimal rate for conditional probability estimation.
Transformers can learn noisy linear systems with depth and IID data.
problem Learning noisy linear dynamical systems with transformers.
method Theoretical analysis of multi-layer and single-layer transformers with respect to L2-testing loss. result Single-layer transformers have a non-diminishing lower bound on approximation error, suggesting depth separation.
Logarithmic-time schedules boost large-scale language model training efficiency.
problem Improving performance and efficiency in large-scale language model training.
method Designing time-varying hyperparameters (β1,β2,λ) for AdamW, specifically logarithmic-time scheduling with damping mechanisms. result ADANA optimizer achieves up to 40% compute efficiency compared to tuned AdamW, with gains persisting as model scale increases.
We present a logarithmic-scale efficient convolutional neural network architecture for edge devices, named WaveletNet. Our model is based on the well-known depthwise convolution, and on two new layers, which we introduce in this work: a wavelet convolution and a depthwise fast wavelet transform. By breaking the symmetr…
Transformers can approximate Newton's method for logistic regression.
problem Implementing higher order optimization methods in Transformers.
method Linear attention Transformers with ReLU layers approximating second order optimization algorithms.
result Transformers can implement a single step of Newton's iteration for matrix inversion.
We introduce a property of mutation loops, called the sign stability, with a focus on an asymptotic behavior of the iteration of the tropical X-transformation. A sign-stable mutation loop has a numerical invariant which we call the cluster stretch factor, in analogy with that of a pseudo-Anosov mapping clas…
The study compares differencing methods for financial data and finds fractional differencing improves model performance.
problem Improving financial time series forecasting models using appropriate data transformation techniques.
method Comparative analysis of traditional logarithmic returns and fractional differencing methods, including tempered extensions.
result Fractional differencing methods improve model forecasting performance and trading strategy effectiveness.
Paper presents a new framework for covariance matrix estimation with geometric insights.
problem Challenges in covariance matrix estimation, especially in finding suitable models and efficient estimation methods.
method General framework for linear restrictions on different transformations of the covariance matrix, including matrix logarithm and its inverse.
result Yields an M-estimator with M-estimation allowing for straightforward asymptotic and finite sample analysis. In this article we introduce an approach for studying the geodesic X-ray transform and related geometric inverse problems by using Carleman estimates. The main result states that on compact negatively curved manifolds (resp. nonpositively curved simple or Anosov manifolds), the geodesic vector field satisfies a Carlema…
Paper proves efficiency of MARL with transformers, addressing agent complexity.
problem Theoretical understanding of MARL with many agents and limited relational reasoning.
method Set Transformer for relational reasoning, model-free and model-based MARL algorithms.
result Provable efficiency of MARL algorithms, suboptimality gaps independent of number of agents.
Self-attention prefers sparse functions of input sequences, reducing sample complexity.
problem Understanding the inductive biases of self-attention in modeling long-range dependencies.
method Theoretical analysis and synthetic experiments to probe sample complexity of learning sparse functions with Transformers.
result Bounded-norm Transformer networks can represent sparse functions of the input sequence with logarithmic sample complexity.
Solve Painleve VI to relate instanton bundles.
problem Relate instanton bundles to Painleve VI solutions.
method Generalize Hitchin's logarithmic connection to vector bundles with SL2 action.
result Identify Okamoto transformations as creation operators.
Study on minimal surfaces in Heisenberg group with duality formula.
problem Understanding minimal surfaces in Heisenberg group.
method Introducing transformation surfaces and using logarithmic derivative of moving frame.
result Derivation of Sym formula for dual minimal surface.
Transformers show strengths and weaknesses in complexity analysis.
problem Understanding the strengths and limitations of attention layers in transformers.
method Analysis of representation power through complexity parameters and task-specific constructions.
result Transformers can solve sparse averaging tasks with logarithmic complexity, but triple detection tasks require linear complexity.
Neural networks with ReLU^k approximate Sobolev functions efficiently via Radon transform.
problem Approximating functions from Sobolev spaces using shallow ReLU^k neural networks.
method Utilizing the Radon transform and discrepancy theory, we provide nearly optimal approximation rates.
result Optimal approximation rates for smoothness up to order s = k + (d+1)/2.
This work concerns testing the number of parameters in one hidden layer multilayer perceptron (MLP). For this purpose we assume that we have identifiable models, up to a finite group of transformations on the weights, this is for example the case when the number of hidden units is know. In this framework, we show that …
Investor optimizes worst-case portfolio in uncertain markets.
problem Optimizing investment in markets with potential crashes.
method Enhanced martingale approach via BSDEs and PDEs.
result Characterized indifference optimal strategies for various models.
We introduce blow-up and blow-down operations for generalized complex 4-manifolds. Combining these with a surgery analogous to the logarithmic transform, we then construct generalized complex structures on nCP2 # m \bar{CP2} for n odd, a family of 4-manifolds which admit neither complex nor symplectic structures unless…
To model categorical response variables given their covariates, we propose a permuted and augmented stick-breaking (paSB) construction that one-to-one maps the observed categories to randomly permuted latent sticks. This new construction transforms multinomial regression into regression analysis of stick-specific binar…
Study anisotropic obstacle problem for minimal surfaces using Cahn-Hoffman transform.
problem Anisotropic obstacle problem for minimal surfaces.
method Cahn-Hoffman transform to convert to isotropic problem with generalized Robin boundary condition.
result Optimal regularity of the solution and C1,1 regularity of the free boundary. Method identifies low-dimensional structure in high-dimensional probability measures.
problem Identifying low-dimensional structure in high-dimensional probability measures.
method Extends prior work on minimizing majorizations of the Kullback-Leibler divergence to identify optimal approximations within a specific class of measures.
result Connection between dimensional logarithmic Sobolev inequality and approximations with the ansatz.
We present the Insertion Transformer, an iterative, partially autoregressive model for sequence generation based on insertion operations. Unlike typical autoregressive models which rely on a fixed, often left-to-right ordering of the output, our approach accommodates arbitrary orderings by allowing for tokens to be ins…
The study explains transformer scaling laws using statistical and approximation theories.
problem Understanding why transformer scaling laws exist for large models trained on low-dimensional data.
method Established statistical estimation and mathematical approximation theories for transformers on low-dimensional manifolds.
result Predicted a power law between generalization error and model and data sizes, with power depending on intrinsic data dimension.
We study Bogomolny equations on R2×S1. Although they do not admit nontrivial finite-energy solutions, we show that there are interesting infinite-energy solutions with Higgs field growing logarithmically at infinity. We call these solutions periodic monopoles. Using Nahm transform, we show that periodic monop…
The paper provides theoretical guarantees for transformation-based models in variational inference.
problem Theoretical justification for transformation-based models in variational inference.
method Theoretical analysis of non-linear latent variable models and Gaussian process priors.
result Theoretical guarantees for implicit variational inference, achieving optimal risk bounds and approximating the true posterior.
The Cauchy problem for the homogeneous (real and complex) Monge-Ampere equation (HRMA/HCMA) arises from the initial value problem for geodesics in the space of Kahler metrics. It is an ill-posed problem. We conjecture that, in its lifespan, the solution can be obtained by Toeplitz quantizing the Hamiltonian flow define…
Activation functions play a key role in providing remarkable performance in deep neural networks, and the rectified linear unit (ReLU) is one of the most widely used activation functions. Various new activation functions and improvements on ReLU have been proposed, but each carry performance drawbacks. In this paper, w…
Study large deviations in fractional volatility models with non-Gaussian volatility.
problem Large deviations in fractional volatility models with non-Gaussian volatility.
method Established a small-noise large deviation principle for log-price.
result Logarithmic call price asymptotics for large strikes in a special case.
New bounds on NTK's smallest eigenvalue for arbitrary data without distributional assumptions.
problem Existing bounds on NTK's smallest eigenvalue require distributional assumptions and high-dimensional data.
method Novel application of the hemisphere transform.
result Bounds on NTK's smallest eigenvalue hold with high probability even for constant input dimension.
In this paper, we introduce the notions of logarithmic Poisson structure and logarithmic principal Poisson structure; we prove that the latter induces a representation by logarithmic derivation of the module of logarithmic Kahler differentials; therefore, it induces a differential complex from which we derive the notio…
This paper tackles causal representation learning with linear and general transformations.
problem Identify and recover latent causal variables and graphs under unknown transformations.
method Score-based algorithms that use gradients of log-density functions for identifiability and achievability.
result Two stochastic hard interventions per node are sufficient for identifiability of general transformations.
ZDP detects drift in large language models without labels, proving key theorems and metrics.
problem Detecting drift in large language models without task labels or output evaluations.
method Zero-Direction Probing (ZDP) framework based on null directions of transformer activations, proving theoretical guarantees.
result Proves the Variance--Leak Theorem, Fisher Null-Conservation, Rank--Leak bound, and logarithmic-regret guarantee.
This study analyzes how one-layer transformers learn regular language recognition tasks.
problem Understanding how one-layer transformers solve regular language recognition tasks like even pairs and parity check.
method Theoretical analysis of training dynamics and gradient descent for a one-layer transformer.
result A one-layer transformer can solve even pairs directly but needs CoT for parity check. Training phases show rapid growth in attention layer followed by logarithmic growth in linear layer.
We consider a system of coupled free boundary problems for pricing American put options with regime-switching. To solve this system, we first employ the logarithmic transformation to map the free boundary for each regime to multi-fixed intervals and then eliminate the first-order derivative in the transformed model by …
Study phase transitions in noisy transformer dynamics on spheres.
problem Understanding phase transitions in noisy transformer dynamics on spheres.
method Sharp Beckner--Onofri/logarithmic HLS inequality, Funk--Hecke/Bessel coefficients, degree-two quartic obstruction.
result Sharp global-minimizer dichotomy and phase transitions in noisy transformer dynamics in arbitrary dimension.
Skeinformer accelerates self-attention for long sequences with linear complexity.
problem Efficiency of Transformer models in processing long sequences.
method Matrix sketching and column sampling to reduce quadratic complexity to linear.
result Skeinformer outperforms alternatives with smaller time/space footprint.