MASnet enhances speech on mobile devices with low latency.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study validates low latency's impact on trading profits.
PHAZE framework uses zkML and hashing for fast, verifiable LHC trigger decisions.
Research optimizes C++ patterns for HFT, reducing latency and improving profitability.
Automated tool reduces FPGA inference latency to 5 μs for deep neural networks.
Exchanges implement intentional trade delays to limit the harmful impact of low-latency trading. Do such "speed bumps" curb investment in fast trading technology? Data is scarce since trading technologies are proprietary. We build an experimental trading platform where participants face speed bumps and can invest in fa…
Paper presents FPGA implementation for efficient recurrent neural networks.
The usability and practicality of any machine learning (ML) applications are largely influenced by two critical but hard-to-attain factors: low latency and low cost. Unfortunately, achieving low latency and low cost is very challenging when ML depends on real-world data that are highly distributed and rapidly growing (…
CoinTossX is a low-latency, open-source matching engine for financial trading.
MLaaS (ML-as-a-Service) offerings by cloud computing platforms are becoming increasingly popular. Hosting pre-trained machine learning models in the cloud enables elastic scalability as the demand grows. But providing low latency and reducing the latency variance is a key requirement. Variance is harder to control in a…
Optimizes latency and false alarm probability in change detection problems.
FlashIV solves Black-Scholes implied volatility efficiently and accurately.
In this paper, a novel experienced deep reinforcement learning (deep-RL) framework is proposed to provide model-free resource allocation for ultra reliable low latency communication (URLLC). The proposed, experienced deep-RL framework can guarantee high end-to-end reliability and low end-to-end latency, under explicit …
The possibility of latency arbitrage in financial markets has led to the deployment of high-speed communication links between distant financial centers. These links are noisy and so there is a need for coding. In this paper, we develop a gametheoretic model of trading behavior where two traders compete to capture laten…
PolyLUT uses polynomials to reduce FPGA latency.
Reducing the latency variance in machine learning inference is a key requirement in many applications. Variance is harder to control in a cloud deployment in the presence of stragglers. In spite of this challenge, inference is increasingly being done in the cloud, due to the advent of affordable machine learning as a s…
Paper predicts interference for better LA in URLLC.
This paper tackles URLLC in 6G networks with deep learning.
We develop a supervised machine learning model that detects anomalies in systems in real time. Our model processes unbounded streams of data into time series which then form the basis of a low-latency anomaly detection model. Moreover, we extend our preliminary goal of just anomaly detection to simultaneous anomaly pre…
Kolmogorov-Arnold Networks enable ultrafast online learning with fixed-point quantization.
Select-DC reduces GFLOPS for uncertainty estimation in neural networks.
DCT-SNN uses DCT to reduce inference latency in SNNs.
This study analyzes satellite communication latency using a stochastic geometry model.
Recent results at the Large Hadron Collider (LHC) have pointed to enhanced physics capabilities through the improvement of the real-time event processing techniques. Machine learning methods are ubiquitous and have proven to be very powerful in LHC physics, and particle physics as a whole. However, exploration of the u…
Power-SMC reduces inference latency for training-free LLM reasoning.
TinyML models detect RF and cyber threats in spacecraft with low latency.
A blockchain replaces central counterparties with time-consuming consensus protocols to record the transfer of ownership. This settlement latency slows cross-exchange trading, exposing arbitrageurs to price risk. Off-chain settlement, instead, exposes arbitrageurs to costly default risk. We show with Bitcoin network an…
Exchanges acquire excess processing capacity to accommodate trading activity surges associated with zero-sum high-frequency trader (HFT) "duels." The idle capacity's opportunity cost is an externality of low-latency trading. We build a model of decentralized exchanges (DEX) with flexible capacity. On DEX, HFTs acquire …
When applying machine learning to sensitive data, one has to find a balance between accuracy, information security, and computational-complexity. Recent studies combined Homomorphic Encryption with neural networks to make inferences while protecting against information leakage. However, these methods are limited by the…
A new method reduces inference cost for FwFM by allowing it to scale with item fields only.
In this paper, we present iPrescribe, a scalable low-latency architecture for recommending 'next-best-offers' in an online setting. The paper presents the design of iPrescribe and compares its performance for implementations using different real-time streaming technology stacks. iPrescribe uses an ensemble of deep lear…
Optimizing distributed learning systems is an art of balancing between computation and communication. There have been two lines of research that try to deal with slower networks: {\em communication compression} for low bandwidth networks, and {\em decentralization} for high latency networks. In this paper, We explore a…
GPU-accelerates multiuser detection for 5G URLLC systems.
A power-law fit to the empirical inference-compute frontier in LOB prediction suggests a scaling-law-style frontier.
We present a systematic analysis on the performance of a phonetic recogniser when the window of input features is not symmetric with respect to the current frame. The recogniser is based on Context Dependent Deep Neural Networks (CD-DNNs) and Hidden Markov Models (HMMs). The objective is to reduce the latency of the sy…
In this paper, we study how to solve resource allocation problems in ultra-reliable and low-latency communications by unsupervised deep learning, which often yield functional optimization problems with quality-of-service (QoS) constraints. We take a joint power and bandwidth allocation problem as an example, which mini…
DIET-SNN optimizes SNNs for faster, lower-energy image classification.
New video compression method outperforms traditional approaches.
GOBO compresses 99.9% of BERT model parameters to 3 bits, improving inference efficiency.
Innovative neural networks reduce memory usage for efficient, accurate segmentation.
Density-Softmax improves uncertainty estimation and robustness without sampling, reducing model size and latency.
A blockchain-based federated learning system with latency analysis.
Two new methods reduce random forest latency and improve accuracy.
EQ-Net combines LLR estimation and quantization using deep learning.
Paper proposes new Bayesian neural network models for efficient learning.
Deep learning models have become state of the art for natural language processing (NLP) tasks, however deploying these models in production system poses significant memory constraints. Existing compression methods are either lossy or introduce significant latency. We propose a compression method that leverages low rank…
Optimized neural networks for Edge TPU achieve high accuracy in real-time image classification.
Machine Learning models are often composed of pipelines of transformations. While this design allows to efficiently execute single model components at training time, prediction serving has different requirements such as low latency, high throughput and graceful performance degradation under heavy load. Current predicti…