In this paper, a novel experienced deep reinforcement learning (deep-RL) framework is proposed to provide model-free resource allocation for ultra reliable low latency communication (URLLC). The proposed, experienced deep-RL framework can guarantee high end-to-end reliability and low end-to-end latency, under explicit …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper tackles URLLC in 6G networks with deep learning.
End-to-end Text-to-speech (TTS) system can greatly improve the quality of synthesised speech. But it usually suffers form high time latency due to its auto-regressive structure. And the synthesised speech may also suffer from some error modes, e.g. repeated words, mispronunciations, and skipped words. In this paper, we…
LanguaShrink compresses prompts using psycholinguistic principles to reduce costs.
DIET-SNN optimizes SNNs for faster, lower-energy image classification.
Machine Learning models are often composed of pipelines of transformations. While this design allows to efficiently execute single model components at training time, prediction serving has different requirements such as low latency, high throughput and graceful performance degradation under heavy load. Current predicti…
HOLMES improves real-time model serving for ICU patients, balancing accuracy and speed.
Multilingual end-to-end (E2E) models have shown great promise in expansion of automatic speech recognition (ASR) coverage of the world's languages. They have shown improvement over monolingual systems, and have simplified training and serving by eliminating language-specific acoustic, pronunciation, and language models…
In this paper, we present a new open source toolkit for automatic speech recognition (ASR), named CAT (CRF-based ASR Toolkit). A key feature of CAT is discriminative training in the framework of conditional random field (CRF), particularly with connectionist temporal classification (CTC) inspired state topology. CAT co…
Automated multi-task learning algorithm that optimizes network topology.
A blockchain-based federated learning system with latency analysis.
The rising popularity of intelligent mobile devices and the daunting computational cost of deep learning-based models call for efficient and accurate on-device inference schemes. We propose a quantization scheme that allows inference to be carried out using integer-only arithmetic, which can be implemented more efficie…
Efficient neural network for audio source separation.
Recently, Transformer has gained success in automatic speech recognition (ASR) field. However, it is challenging to deploy a Transformer-based end-to-end (E2E) model for online speech recognition. In this paper, we propose the Transformer-based online CTC/attention E2E ASR architecture, which contains the chunk self-at…
Optimizes latency and false alarm probability in change detection problems.
Study validates low latency's impact on trading profits.
Research examines how strategic latency manipulation impacts Ethereum's network efficiency and decentralization.
The Tick library simulates and learns Hawkes processes with latency effects.
Research optimizes C++ patterns for HFT, reducing latency and improving profitability.
We present a new algorithm for video coding, learned end-to-end for the low-latency mode. In this setting, our approach outperforms all existing video codecs across nearly the entire bitrate range. To our knowledge, this is the first ML-based method to do so. We evaluate our approach on standard video compression test …
This paper presents a novel end-to-end methodology for enabling the deployment of low-error deep networks on microcontrollers. To fit the memory and computational limitations of resource-constrained edge-devices, we exploit mixed low-bitwidth compression, featuring 8, 4 or 2-bit uniform quantization, and we model the i…
MASnet enhances speech on mobile devices with low latency.
CryptoNAS improves PI accuracy by 3.4% with 2.4x less latency.
PHAZE framework uses zkML and hashing for fast, verifiable LHC trigger decisions.
This paper studies optimal market making for large-tick assets in the presence of latency. We consider a random walk model for the asset price, and formulate the market maker's optimization problem using Markov Decision Processes (MDP). We characterize the value of an order and show that it plays the role of one-period…
BRP-NAS uses GCNs to predict neural network performance for more efficient NAS.
The possibility of latency arbitrage in financial markets has led to the deployment of high-speed communication links between distant financial centers. These links are noisy and so there is a need for coding. In this paper, we develop a gametheoretic model of trading behavior where two traders compete to capture laten…
Automated tool reduces FPGA inference latency to 5 μs for deep neural networks.
TinyML models detect RF and cyber threats in spacecraft with low latency.
A new algorithm for cryo-EM data collection that balances reward and latency.
High frequency trading has led to widespread efforts to reduce information propagation delays between physically distant exchanges. Using relativistically correct millisecond-resolution tick data, we document a 3-millisecond decrease in one-way communication time between the Chicago and New York areas that has occurred…
S3NAS finds high-accuracy CNN architectures for NPUs in 3 hours.
DCT-SNN uses DCT to reduce inference latency in SNNs.
PolyLUT uses polynomials to reduce FPGA latency.
When applying machine learning to sensitive data, one has to find a balance between accuracy, information security, and computational-complexity. Recent studies combined Homomorphic Encryption with neural networks to make inferences while protecting against information leakage. However, these methods are limited by the…
This study analyzes satellite communication latency using a stochastic geometry model.
Paper analyzes how latency affects optimal order execution in markets.
Paper presents FPGA implementation for efficient recurrent neural networks.
New attacks exploit neural network energy and latency, increasing costs by 10-200x.
Exchanges implement intentional trade delays to limit the harmful impact of low-latency trading. Do such "speed bumps" curb investment in fast trading technology? Data is scarce since trading technologies are proprietary. We build an experimental trading platform where participants face speed bumps and can invest in fa…
This paper proposes an EM approach to reduce inference latency in NAR sequence generation.
Improves text-to-speech speed by interleaving character reading and audio synthesis.
A blockchain replaces central counterparties with time-consuming consensus protocols to record the transfer of ownership. This settlement latency slows cross-exchange trading, exposing arbitrageurs to price risk. Off-chain settlement, instead, exposes arbitrageurs to costly default risk. We show with Bitcoin network an…
Freezing intermediate layers reduces DNN inference latency.
Reinforcement learning approaches have long appealed to the data management community due to their ability to learn to control dynamic behavior from raw system performance. Recent successes in combining deep neural networks with reinforcement learning have sparked significant new interest in this domain. However, pract…
NeuraLUT maps neural networks to lookup tables, reducing latency and improving expressivity.
We resolve the fundamental problem of online decoding with general order ergodic Markov chain models. Specifically, we provide deterministic and randomized algorithms whose performance is close to that of the optimal offline algorithm even when latency is small. Our algorithms admit efficient implementation vi…
Paper predicts interference for better LA in URLLC.