PPT optimizes transformer behavior by steering its latent posterior using prior samples.
problem Eliciting desired behavior from transformers without backpropagation.
method Posterior Prefix Tuning (PPT) uses predictive Monte Carlo (PMC) samples and importance sampling to optimize the latent posterior.
result PPT optimizes transformer behavior without backpropagation, achieving high utility across different utility functions.
Study behavior of curvatures near singular points of frontals.
problem Understanding frontals near singular points.
method Investigate principal curvatures and vectors near singular points of frontals.
result Extend Ribaucour transformations to frontals with singular points.
Transformers infer tasks from context via two modes, geometrically shaped task vectors explain their behavior.
problem Understanding how transformers infer tasks from context and the geometric properties of task vectors.
method Synthetic setting to train small transformers, mathematical characterization of task-vector geometry and inference modes.
result Task-vector geometry shapes in-distribution and out-of-distribution behavior of transformers.
Improved probabilistic forecasts using behavioral transformations.
problem Improving accuracy and consistency of probabilistic asset price forecasts.
method Behavioral transformation of fundamental expectations to disentangle sentiment-induced biases.
result Substantial forecast gains across various models and risk-preferences.
Develops first robustness verification for complex Transformers.
problem Certify prediction behavior of Transformers with complex self-attention layers.
method Resolves challenges of cross-nonlinearity and cross-position dependency in Transformers.
result Certified robustness bounds are significantly tighter than Interval Bound Propagation.
Unique Poincaré type cscK metric with singularity at smooth divisor is unique up to holomorphic transformations.
problem Proving uniqueness of Poincaré type cscK metric with singularity at a smooth divisor.
method Holomorphic transformations, asymptotic behavior analysis, fixed point problem.
result Unique Poincaré type cscK metric with singularity at a smooth divisor is unique up to holomorphic transformations.
We begin an exploration of parametric Backlund transformations for hyperbolic Monge-Ampere systems. We compute invariants for such transformations and explore the behavior of four examples regarding their invariants, symmetries, and conservation laws. We prove some preliminary results and indicate directions for furthe…
For a complete Riemannian metric, a pointwise conformal transformation may lead to a complete or incomplete transformed Riemannian metric, depending on the behavior of the conformal factor. We establish conditions on the growth of the conformal factor towards the infinity of the Riemannian metric, such that the conform…
Sparse Transformers degrade semantic information first, with early layers encoding more.
problem Understanding how sparse Transformers affect learned representations and semantic information.
method Probed Transformers with progressively pruned weights to observe changes in semantic information and model behavior.
result Complex semantic information is first to degrade in sparse Transformers, with early layers encoding more.
Transformers approximate mean-field dynamics of indistinguishable particles.
problem Approximating the dynamics of indistinguishable particles in complex systems.
method Using transformers to model the mean-field dynamics of interacting particle systems.
result Theoretical bounds on the distance between true and transformer-obtained mean-field dynamics.
The price of financial assets are, since Bachelier, considered to be described by a (discrete or continuous) time sequence of random variables, i.e a stochastic process. Sharp scaling exponents or unifractal behavior of such processes has been reported in several works. In this letter we investigate the question of sca…
New bounds show transformers need longer training for length generalization.
problem Understanding when transformers can generalize to longer inputs.
method Analyzing different settings of transformers, providing quantitative bounds.
result Transformers need training data longer than previously thought for length generalization.
Transformers can learn Markov processes with constant depth, surprising results.
problem Understanding how transformers learn context in Markov processes.
method Empirical study and theoretical analysis of attention-based transformers on Markov data.
result Transformers with constant depth can achieve low test loss on Markov sequences, matching empirical and theoretical findings.
We study strict local martingales via h-transforms, a method which first appeared in Delbaen-Schachermayer. We show that strict local martingales arise whenever there is a consistent family of change of measures where the two measures are not equivalent to one another. Several old and new strict local martingales are i…
This work studies clustering in transformer models, proving exponential convergence to a single token state.
problem Understanding the long-term behavior of tokens in transformer models.
method Investigates mean-field transformer models under specific conditions to prove exponential convergence to a single state.
result Transformer models synchronize exponentially fast to a single token state with explicit rates.
Deep learning analyzes healthcare provider actions and patient outcomes.
problem Understanding provider behavior in non-randomized healthcare settings.
method Deep causal behavioral policy learning (DC-BPL) using transformer architecture.
result Optimal provider policies identified for specific patient types.
Learned data models based on sparsity are widely used in signal processing and imaging applications. A variety of methods for learning synthesis dictionaries, sparsifying transforms, etc., have been proposed in recent years, often imposing useful structures or properties on the models. In this work, we focus on sparsif…
We construct the Nahm transform from finite energy instantons on the product of a real line and a three dimensional torus to Dirac-type singular monopoles on the dual torus. Moreover, we show the correspondence between the data which handle the asymptotic behavior of instantons at infinity and one of monopoles at singu…
Study investigates ruin probability with random premiums and risky investments.
problem Ruin probability with random premiums and risky investments.
method Laplace transform applied to a model with geometric Brownian motion.
result Asymptotic behavior of ruin probability for large initial capital values.
How can deep learning systems flexibly reuse their knowledge? Toward this goal, we propose a new class of challenges, and a class of architectures that can solve them. The challenges are meta-mappings, which involve systematically transforming task behaviors to adapt to new tasks zero-shot. The key to achieving these c…
Transformers achieve near-optimal dynamic regret in non-stationary reinforcement learning.
problem Understanding and handling non-stationary environments in reinforcement learning.
method Demonstrated that transformers can achieve nearly optimal dynamic regret bounds in non-stationary settings.
result Transformers can approximate and learn strategies for non-stationary environments, matching or outperforming existing expert algorithms.
Transformers learn to use induction heads or shortcuts based on data diversity.
problem How data diversity influences the behavior of transformers.
method Gradient-based training of a single-layer transformer on a minimal task.
result Data diversity steers transformers toward induction heads or shortcuts.
Study of immersions with Willmore energy leading to spherical and catenoid bubbles.
problem Classifying immersions with specific energy properties.
method Analyzing sequences of weak immersions with diverging conformal classes, applying Möbius transformations, and strong Wloc2,2-limits. result Obtaining spherical and catenoid bubbles as limits of immersions.
Transformers learn sparse Boolean functions through RL and SFT, revealing distinct learning behaviors.
problem Learning sparse Boolean functions with Transformers.
method Reinforcement Learning (RL) with process rewards and Supervised Fine-Tuning (SFT).
result RL learns the whole CoT chain simultaneously, while SFT learns step by step.
A robust method for decomposing spectral peaks robust to distortion and interference.
problem Decomposing spectral peaks in the presence of distortion and interference.
method Optimizing a nonparametric approach using pseudo-symmetric functions with nonincreasing behavior.
result Decomposed spectral peaks show pseudo-orthogonal behavior and power preserving equality.
Econometric framework integrates heavy-tailed distributions with behavioral probability weighting for better asset pricing.
problem Underestimation of Value-at-Risk by traditional models in asset pricing.
method Developed an econometric framework combining heavy-tailed Student's t distributions with behavioral probability weighting. result Student's t specifications outperform Gaussian models in 88.4% of cases, reducing underestimation of Value-at-Risk by 16.5 percentage points. TabPFN learns to approximate functions on tabular data.
problem TabPFN tackles function approximation on tabular data.
method Treated as a black-box function approximator generator, observed behavior on varied datasets.
result Observed behavior that is both brilliant and baffling.
LISBET automates social behavior analysis using machine learning.
problem Manual annotation of social behaviors is time-consuming, biased, and misses subtle interactions.
method Self-supervised learning on body tracking data.
result Automated detection and segmentation of social interactions.
Improves MCMC performance with adaptive affine transformations.
problem Improving the performance of Markov Chain Monte Carlo samplers.
method Adaptive learning of bijective affine transformations during sampling.
result Adaptive affine transformations improve the quality of samples at low computational cost.
New insights into X-ray transform on hyperbolic disk, with functional relations and range characterizations.
problem Understanding the X-ray transform on hyperbolic geometry.
method Derived new singular value decompositions, range characterizations, and intertwining relations with wedge-type differential operators.
result Sharp understanding of boundary behavior and invertibility settings for the X-ray transform.
Empirical time series of inter-event or waiting times are investigated using a modified Multifractal Detrended Fluctuation Analysis operating on fluctuations of mean detrended dynamics. The core of the extended multifractal analysis is the non-monotonic behavior of the generalized Hurst exponent h(q) -- the fundament…
New theory explains signal propagation in normalization-free transformers.
problem Understanding signal propagation in normalization-free transformers.
method Deriving recurrence relations for activation statistics and APJNs across layers.
result Transformers with elementwise tanh-like nonlinearities exhibit subcritical signal propagation.
Transformers cluster meaningless words around leaders for sentiment analysis.
problem Capturing context in sentiment analysis using transformers.
method Characterized transformers with hardmax self-attention and normalization, showing asymptotic convergence to clustered equilibrium.
result Transformers can effectively capture context by clustering meaningless words around leader words.
Ridge regression shows different behaviors in binary classification with noisy labels.
problem Binary classification with noisy labels and anisotropic cluster distributions.
method Investigation of ridge regression behavior in overparameterized settings with label noise.
result Ridge regression exhibits qualitatively different behavior based on the scale of cluster mean vectors and covariance matrices.
Study evaluates five LLMs for financial report analysis, revealing performance differences and variability.
problem Lack of understanding in reliability, consistency, and transparency of LLMs in financial analysis.
method Human evaluation, automated similarity metrics, and behavioral diagnostics applied to five transformer-based LLMs over U.S. 10-K filings.
result No single LLM consistently dominates across all evaluation perspectives, highlighting variability and need for interpretability.
New Transformers maintain Lipschitz continuity for robustness.
problem Ensuring robustness in Transformers for safety-sensitive applications.
method Introducing gradient-descent-type in-context Transformers with explicit Euler steps of negative gradient flows.
result Universal approximation theorem for Lipschitz continuous Transformers.
Pion optimizes LLMs by preserving weight matrix singular values.
problem Training large language models (LLMs) with standard optimizers leads to unstable weight matrices.
method Pion uses orthogonal transformations to update weight matrices, preserving their singular values.
result Pion offers a stable alternative to standard optimizers for LLM pretraining and finetuning.
This paper is devoted to study of transformations on metric spaces. It is done in an effort to produce qualitative version of quasi-isometries which takes into account the asymptotic behavior of the Gromov product in hyperbolic spaces. We characterize a quotient semigroup of such transformations on Teichmüller space by…
Transformers trained on random classification tasks generalize well and can overfit without error.
problem Understanding how transformers generalize and overfit in-context.
method Analysis of implicit regularization during gradient descent training.
result Transformers can overfit without error and still generalize well.
Efficient algorithms find optimal monotone transforms for calibration under strictly convex losses.
problem Calibrating estimations to improve performance with monotone transforms.
method Proposed linear-time and space algorithm for finding optimal monotone transforms for specific loss functions. Also proposed an anytime algorithm with linear space and pseudo-linearithmic time complexity.
result Optimal monotone transforms are unique and can be found efficiently for various strictly convex loss functions.
We apply neural nets with ReLU gates in online reinforcement learning. Our goal is to train these networks in an incremental manner, without the computationally expensive experience replay. By studying how individual neural nodes behave in online training, we recognize that the global nature of ReLU gates can cause und…
We propose the first qualitative hypothesis characterizing the behavior of visual transformation based self-supervision, called the VTSS hypothesis. Given a dataset upon which a self-supervised task is performed while predicting instantiations of a transformation, the hypothesis states that if the predicted instantiati…
Natural gradient descent is a robust optimization method for machine learning.
problem Training poorly parameterized networks efficiently.
method Optimization algorithms with natural transformation properties.
result Optimization algorithms with natural transformation properties are more efficient for poorly parameterized networks.
This paper introduces a new task to better understand Transformers in quantitative contexts.
problem Understanding Transformers in high-stakes quantitative and scientific applications.
method Introduces a novel contextual counting task and analyzes it with causal and non-causal Transformer architectures.
result Causal attention is better suited for the contextual counting task, and no positional embeddings lead to the best accuracy.
Transformers struggle to learn Markovian dynamics, showing NP-hard optimization challenges.
problem Understanding transformers' limitations in learning Markovian dynamical functions.
method Investigated through a structured ICL setup, analyzing loss landscapes and parameter optimization.
result Recovering optimal transformer parameters for Markovian functions is NP-hard.
Transformers learn to integrate information from past positions incrementally, specializing heads in distinct patterns.
problem How transformers learn to integrate information from multiple past positions with varying statistical significance.
method High-order Markov chain task, incremental learning, sparse attention patterns, simplified differential equations, stage-wise convergence, early stopping as regularizer.
result Transformers learn to specialize heads in distinct patterns, shifting from competitive to cooperative learning dynamics.
We investigate large changes, bursts, of the continuous stochastic signals, when the exponent of multiplicativity is higher than one. Earlier we have proposed a general nonlinear stochastic model which can be transformed into Bessel process with known first hitting (first passage) time statistics. Using these results w…
Investigates optimal parameter allocation in Transformers for efficiency and expressivity.
problem Balancing expressivity and efficiency in Transformer model parameters.
method Mathematical analysis and theoretical characterization of attention heads and head dimensions.
result Later layers can operate more efficiently with reduced parameters due to saturation of softmax activations.