Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

100201301401 · Jun 202019922001200920182026
48 results for scale mechanism

The paper studies scaling laws for associative memory mechanisms.

problem Understanding and optimizing learning and memorization processes.
method High-dimensional matrices of outer products of embeddings, relating to transformer models. Derived scaling laws with sample and parameter sizes. Extensive numerical experiments.
result Precise scaling laws and statistical efficiency of estimators.

Paper proposes a new hierarchical attention mechanism for multi-scale data.

problem Challenges in applying neural attention to multi-scale, multi-modal data.
method Developed a mathematical framework for multi-modal, multi-scale data and derived optimal neural attention mechanics.
result Proposed hierarchical attention mechanism improves transformer performance in multi-scale, multi-modal settings.

Exact 1-Wasserstein distance between location-scale distributions derived, with privacy effects studied.

problem Calculating the 1-Wasserstein distance between location-scale distributions and its impact on differential privacy.
method Exact expressions and special functions for 1-Wasserstein distance, new upper bounds, and asymptotic analysis.
result New linear upper bound and detailed asymptotic bounds for Gaussian case, effect of differential privacy studied.

This paper makes complex neural graphs nearly convex through iterative decomposition and scale mechanism.

problem Non-convexity in complex neural architectures limits their performance in convex optimization.
method Decompose neural graph into operators, algorithms, and functions; iteratively propagate along edges; introduce scale mechanism to transform non-convex properties.
result Proves neural graph is nearly convex in each variable when others are fixed, validating the scale mechanism.

Differentially private log-location-scale regression models improve privacy in statistical analysis.

problem Ensuring privacy in statistical regression models while maintaining accuracy.
method Integrates differential privacy into LLS regression using the functional mechanism.
result Proposed DP-LLS models satisfy ε-differential privacy and perform well under various conditions.

Neural HMM with AGA captures multi-scale dynamics in financial markets.

problem Capturing multi-scale temporal dynamics in financial markets.
method Parallel multi-resolution encoders, adaptive gating, and multi-head attention.
result Outperforms fixed-resolution baselines in predicting price movements and liquidity shocks.

A new network learns to prioritize messages for efficient multi-robot path planning.

problem Efficient path planning and coordination for large-scale multi-robot systems.
method Message-Aware Graph Attention Network (MAGAT) incorporating attention mechanisms.
result MAGAT achieves performance close to a coupled centralized expert algorithm.

New privacy mechanism reduces error in query results.

problem Achieving privacy while minimizing noise in query results.
method Extended sufficient and necessary condition for (ε,δ)(ε, δ)-differential privacy for symmetric and log-concave noise densities.
result Significantly lower mean squared errors than Laplace and Gaussian mechanisms.

AAANE embeds networks by learning attention weights for multi-scale structure.

problem Existing methods ignore the role of different scales in network embedding.
method AAANE uses an attention-based adversarial autoencoder to learn robust representations.
result AAANE outperforms existing methods on real-world networks.

This work bridges two views of feature learning in neural networks.

problem The relationship between kernel scale changes and data-adaptive feature learning in neural networks remains unresolved.
method Using statistical mechanics, the work derives analytical expressions for network output statistics across scaling regimes.
result Kernel adaptation can be reduced to an effective kernel rescaling, but multi-scale adaptive approach provides richer insights.

New approach for large-scale distributed learning systems that improve generalization performance.

problem Transitioning from centralized to distributed AI systems for complex learning tasks.
method Self-organizing hierarchical structuring mechanism based on agglomerative clustering, hierarchical generalization, and personalized learning.
result Demonstrates better generalization performance compared to conventional federated learning algorithms.

Fault diagnosis method for refrigerant leaks using scaling law.

problem Difficulty in improving fault-detection model generalization due to complex system configuration and insufficient data.
method Derive a scaling law based on physical modeling and control mechanism, apply to other systems without modification.
result Scaling exponents of different air-conditioning systems are equivalent, indicating the proposed method's applicability.

Geometrically transforms nonconservative dynamics to linearize Kepler and Manev systems.

problem Regularizing and linearizing nonconservative central force dynamics.
method Projective transformation and conformal scaling in configuration and phase spaces.
result Full linearization of Kepler and Manev dynamics in any finite dimension.

STAS selects optimal spatio-temporal scales for bias correction in precipitation forecasts.

problem Limited prior data and fixed ST scale in existing BCoPs lead to biases in numerical weather predictions.
method End-to-end deep-learning BCoP model STAS with SFM/TFM to automatically adjust spatial and temporal scales.
result STAS outperforms 8 published BCoP methods on threat scores (TS).

Method uses ANN to estimate incentive salience from large behavioral data.

problem Estimating incentive salience in naturalistic settings.
method Artificial Neural Networks (ANNs) for latent state approximation.
result ANNs produce better representations for predicting future behaviour.

New framework explains fast transfer of hyperparameters across model scales.

problem Understanding and optimizing hyperparameters for large-scale models.
method Developed a conceptual framework for HP transfer across scale, showing fast transfer is equivalent to useful transfer for compute-optimal grid search.
result Fast transfer of hyperparameters is equivalent to useful transfer for compute-optimal grid search, offering asymptotic computational advantage.

The abstract discusses financial irreversibility using quantum mechanics and projective geometry.

problem Financial irreversibility and its limitations in trading strategies.
method Projective geometry and Taylor expansion of directed distance in quantum systems.
result Fundamental asymmetry under state exchange is a key factor in financial irreversibility.

Proposes a text perturbation method using a Mahalanobis metric to balance privacy and utility.

problem Low utility of text analysis when using spherical noise for privacy-preserving text embedding.
method Regularized Mahalanobis metric to add elliptical noise, accounting for embedding space density.
result Improves privacy statistics while maintaining utility, outperforming Laplace mechanism.

Deep neural networks near edge of chaos show universal scaling laws.

problem Understanding the behavior of deep neural networks near critical points.
method Analogy to absorbing phase transitions in statistical mechanics, deterministic propagation dynamics, mean-field and directed percolation universality classes.
result Deep neural networks exhibit universal scaling laws near the edge of chaos.

Modern neural networks are often augmented with an attention mechanism, which tells the network where to focus within the input. We propose in this paper a new framework for sparse and structured attention, building upon a smoothed max operator. We show that the gradient of this operator defines a mapping from real val…

2017-05-22abs ↗pdf ↗

A multi-scale model predicts atomic-scale properties using both local and long-range information.

problem Inability of machine-learning schemes to capture long-range physical effects.
method Combines local and non-local information in a multipole expansion framework.
result Demonstrates the ability to model electrostatics, polarization, and dispersion.

Preformer improves Transformer for long-term time series forecasting.

problem Transformer's quadratic complexity and lack of context-awareness for long-term forecasting.
method Introduces Multi-Scale Segment-Correlation mechanism for efficient time series segmentation and context-aware attention.
result Preformer outperforms other Transformer-based methods in long-term time series forecasting.

CogScale benchmarks AI architectures for sequential processing.

problem Evaluating AI architectures' ability to process sequential information efficiently.
method 14 scalable synthetic tasks designed to isolate cognitive and memory abilities at different scales.
result Attention mechanisms and modern state-space models consistently maintain high performance as task difficulty scales.

Hyperfitting improves LLM generation quality by enhancing diversity, contrary to simple temperature scaling.

problem Improving open-ended generation quality of LLMs with minimal fine-tuning effort.
method Demonstrates that hyperfitting, a phenomenon where LLMs are fine-tuned to near-zero training loss, enhances generation quality and mitigates repetition.
result Hyperfitting is distinct from temperature scaling and involves a dynamic, context-dependent rank reordering mechanism in the final transformer block.

D-LinOSS models learn to dissipate energy, improving performance on long-range tasks.

problem Representational limitations of LinOSS models in long-range reasoning.
method Introducing Damped Linear Oscillatory State-Space models (D-LinOSS) that learn to dissipate latent state energy on arbitrary time scales.
result D-LinOSS consistently outperforms previous LinOSS methods on long-range learning tasks, achieving faster convergence and reducing hyperparameter search space.

This work analyzes neural scaling laws using power-law data spectra and derives analytical expressions for generalization error.

problem Understanding how neural network performance scales with key factors like data size and model complexity.
method Statistical mechanics techniques applied to one-pass stochastic gradient descent in a student-teacher framework.
result Derivation of analytical expressions for generalization error under power-law data spectra and identification of conditions for power-law scaling.

We introduce a prototype model in an attempt to capture some aspects of market dynamics simulating a trading mechanism. The model description starts with a discrete-space, continuous-time Markov process describing arrival and movement of orders with different prices. We then perform a re-scaling procedure leading to a …

2012-01-22abs ↗pdf ↗

Logarithmic-time schedules boost large-scale language model training efficiency.

problem Improving performance and efficiency in large-scale language model training.
method Designing time-varying hyperparameters (β1,β2,λ)(β_1, β_2, λ) for AdamW, specifically logarithmic-time scheduling with damping mechanisms.
result ADANA optimizer achieves up to 40% compute efficiency compared to tuned AdamW, with gains persisting as model scale increases.

Uniform scaling limits in AdamW-trained transformers converge to ODEs.

problem Understanding the dynamics of large-depth transformers trained with AdamW.
method Modeling transformer dynamics as an interacting particle system coupled through attention, proving convergence to ODEs.
result The joint dynamics of hidden states and backpropagated variables converge uniformly to an ODE system.

Improved weather forecasting with gridded pseudo-token TNPs.

problem Handling large-scale, unstructured spatio-temporal data in weather forecasting.
method Introducing gridded pseudo-token transformer neural processes (TNPs) with efficient attention mechanisms.
result Consistently outperforms baselines on various synthetic and real-world regression tasks involving large-scale data.

Study reveals how correlation matrix eigenvalues change with time scale in U.S. stocks.

problem Understanding how correlation structure of securities changes with time scale.
method Aggregated one-minute returns of 533 U.S. stocks at different time scales, estimated correlation matrix, lead-lag factor model.
result Emergence of several dominant eigenvalues as time scale increases.

SANs estimate feature importance in neural models, identifying interactions and high-ranked features.

problem Understanding and interpreting black-box neural network models.
method Self-Attention Network (SAN) architecture for feature importance estimation.
result SANs identify similar high-ranked features and feature interactions as other methods, improving predictive performance.

We present a general framework, the coupled compound Poisson factorization (CCPF), to capture the missing-data mechanism in extremely sparse data sets by coupling a hierarchical Poisson factorization with an arbitrary data-generating model. We derive a stochastic variational inference algorithm for the resulting model …

2017-01-09abs ↗pdf ↗