Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

4998146195 · Jun 202019922001200920182026
48 results for lower latency

New algorithms decode Markov chains with near-optimal performance, even with small latency.

problem Online decoding of nthn^{th} order ergodic Markov chains with latency constraints.
method Deterministic and randomized algorithms using dynamic programs, with lower bounds established.
result Near-optimal performance of algorithms with minimal latency, outperforming existing methods.

Optimizes latency and false alarm probability in change detection problems.

problem Balancing latency and false alarms in non-stationary environments.
method Develops order-optimal change detectors under specified latency and false alarm levels.
result Derives a universal lower bound on latency and develops order-optimal detectors.

Benanza speeds up DL model optimization by automatically generating micro-benchmarks and identifying inefficiencies.

problem Slow characterization/optimization cycles for DL models on GPUs.
method Benanza includes a model processor, benchmark generator, database, and analyzer.
result Benanza identifies optimizations in parallel layer execution, cuDNN, framework inefficiency, layer fusion, and Tensor Cores.

Collage inference uses redundancy to reduce cloud image classification latency variance.

problem Reducing latency variance in cloud image classification.
method Integrates collage-cnn for low-cost redundancy in multi-image classification.
result Significant reduction in 99th percentile tail latency and inference latency variation.

Density-Softmax improves uncertainty estimation and robustness without sampling, reducing model size and latency.

problem Sampling-based uncertainty estimation methods suffer from large model size and high latency.
method Combines a Lipschitz-constrained feature extractor with the softmax layer to create a sampling-free deterministic framework.
result Density-Softmax reduces over-confidence under distribution shifts and achieves competitive results in uncertainty and robustness.

Collage-CNN reduces cloud inference latency by 1.47X with 9X reduced latency variation.

problem Reducing latency variance in cloud machine learning inference.
method Proposes a novel Collage-CNN model that combines multiple images for classification, providing redundancy and cost efficiency.
result Significant reduction in 99th percentile tail latency and variation in inference latency.

LITE models improve query-document relevance with learnable late interactions.

problem Improving query-document relevance with lower latency and storage.
method Proposes learnable late-interaction models (LITE) that use factorized query and document embeddings followed by a learnable scorer.
result Empirically, LITE outperforms previous late-interaction models in re-ranking tasks.

Paper presents a new method to train deep neural networks with reduced memory access.

problem High computational and storage complexity of deep neural networks.
method Boolean logic minimization to remove memory access and reduce resource usage.
result Significantly lower latency and two orders of magnitude fewer computing resources.

A power-law fit to the empirical inference-compute frontier in LOB prediction suggests a scaling-law-style frontier.

problem Limit order book prediction
method Using a suite of models ranging from small decision trees to neural LOB architectures
result A power-law fit to the low- and mid-compute non-MLPLOB frontier extrapolates across multiple orders of magnitude and attains R2=0.941R^2=0.941 on the excluded high-compute MLPLOB target frontier.

A modified VDCNN model reduces size and latency for mobile platforms.

problem Memory and processing constraints on mobile platforms.
method Temporal Depthwise Separable Convolutions and Global Average Pooling.
result The squeezed model (SVDCNN) is 10x-20x smaller with minimal accuracy loss.

Improved multilingual speech recognition with low latency for nine Indic languages.

problem Imbalance in training data across languages and low latency in interactive applications.
method Conditioning on a language vector and training language-specific adapter layers.
result Lower word error rate than monolingual E2E models and conventional systems.

DIET-SNN optimizes SNNs for faster, lower-energy image classification.

problem High inference latency and inefficient input encoding in SNNs.
method End-to-end backpropagation to optimize membrane leak and firing threshold.
result Achieves top-1 accuracy of 69% on ImageNet with 5 timesteps and 12x less compute energy.

Gradient codes adapt to varying straggler counts in distributed learning.

problem Mitigating slow worker nodes (stragglers) in distributed machine learning.
method Proposes a flexible gradient coding scheme that concatenates codes for different straggler tolerances, adapting to actual straggler counts.
result Significantly lower latency compared to fixed-tolerance gradient codes.

Research examines how strategic latency manipulation impacts Ethereum's network efficiency and decentralization.

problem Impact of artificial latency on Ethereum's network efficiency and decentralization.
method Comprehensive analysis of MEV-Boost auction system and empirical validation with a pilot.
result Increased profitability for node operators and significant systemic challenges like heightened network inefficiencies and centralization risks.

Research optimizes C++ patterns for HFT, reducing latency and improving profitability.

problem Optimizing latency-critical code for high-frequency trading systems.
method Creation of a Low-Latency Programming Repository, optimisation of trading strategy, implementation of Disruptor pattern.
result Significant performance improvements in speed and profitability.

Efficiently improves non-autoregressive sequence models for better translation performance.

problem Heavy inference latency and inconsistent output sentences in non-autoregressive models.
method Incorporates a structured inference module with an efficient CRF approximation and dynamic transition technique.
result Significantly better translation performance (BLEU score 26.80) compared to previous non-autoregressive models.

MASnet enhances speech on mobile devices with low latency.

problem Efficiently enhancing speech on mobile devices with low latency.
method MASnet processes linear-scale spectrograms, using ratio masks to enhance noisy frames, and operates in low-latency incremental inference mode.
result MASnet achieves efficient speech enhancement with low latency, reducing FMA/s operations.

Proposes deep-RL with GANs for ultra-reliable low-latency communication.

problem Resource allocation for URLLC with high reliability and low latency.
method Experienced deep-RL framework using GANs to pre-train and optimize resource allocation.
result Deep-RL framework achieves near-optimal reliability and latency under URLLC constraints.

PHAZE framework uses zkML and hashing for fast, verifiable LHC trigger decisions.

problem Inefficient inference on large machine learning models for LHC trigger performance.
method Cryptographic techniques like hashing and zkML for low latency, certifiable inference.
result Achieves nanosecond-order latency for LHC triggers, enabling dynamic low-level triggers.

Two solutions improve privacy-preserving inference with reduced latency and wider neural network support.

problem Balancing accuracy, security, and computational complexity in machine learning with sensitive data.
method Combining Homomorphic Encryption with transfer learning and novel data representation methods.
result More than 10x improvement in latency with wider neural network support.

The possibility of latency arbitrage in financial markets has led to the deployment of high-speed communication links between distant financial centers. These links are noisy and so there is a need for coding. In this paper, we develop a gametheoretic model of trading behavior where two traders compete to capture laten…

2015-04-27abs ↗pdf ↗

New strategy improves liquidity takers' performance in markets with latency.

problem Latency affects liquidity takers' ability to execute limit orders effectively.
method Modelled LOB and MLOs as a marked point process, used variational analysis and FBSDEs to find optimal price limits.
result Optimal trading strategy improves marksmanship in markets with latency.

TinyML models detect RF and cyber threats in spacecraft with low latency.

problem Detecting cyber-RF threats in autonomous spacecraft with low latency.
method Analysis of classical models (RF, LR, SVM, MLP) for latency-accuracy trade-offs.
result Logistic Regression achieves microsecond-level inference with minimal accuracy loss.

Speed bumps reduce but do not fully eliminate investment in fast trading technology.

problem Limiting low-latency trading to curb investment in fast trading technology.
method Built an experimental trading platform to test the effects of speed bumps on investment in fast trading technology.
result Asymmetric speed bumps reduce investment in speed by only 20%, and increasing the magnitude further reduces investment by 8.33%. Symmetric speed bumps have no effect on investment levels.

Blockchain trading faces limits due to time-consuming settlement, exposing arbitrageurs to price risk.

problem Time-consuming settlement in blockchain trading limits arbitrage opportunities.
method Analysis of Bitcoin network and order book data.
result Cross-exchange price differences coincide with high settlement latency and low default risk.

This study analyzes satellite communication latency using a stochastic geometry model.

problem Latency analysis of LEO satellite relay communication systems.
method Stochastic geometry framework with spherical BPP models, suboptimal satellite relay selection strategy.
result Derives distance distributions and analytical expressions for transmission delays.

This paper proposes an EM approach to reduce inference latency in NAR sequence generation.

problem High inference latency in NAR models due to multi-modality in sequence generation.
method A unified EM framework that jointly optimizes AR and NAR models, with iterative refinement.
result The proposed approach achieves competitive performance with existing NAR models and significantly reduces inference latency.

Hierarchical FL reduces latency in HCNs by sharing model updates.

problem Latency and privacy issues in federated learning across heterogeneous cellular networks.
method Hierarchical federated learning, gradient sparsification, periodic averaging.
result Significant reduction in communication latency without compromising model accuracy.

Improves text-to-speech speed by interleaving character reading and audio synthesis.

problem Latency in text-to-speech models limits their use in time-sensitive tasks.
method Reinforcement learning to train an agent to choose the order of character reading and audio synthesis.
result The proposed method successfully balances latency and audio quality.