Improves matrix multiplication throughput for asymmetric bit-width operands.
problem Matrix multiplications between asymmetric bit-width operands, especially 8- and 4-bit, are not efficiently handled by existing SIMD instructions.
method Proposes a new SIMD matrix multiplication instruction that uses mixed precision on inputs (8- and 4-bit) and accumulates into 16-bit output, improving throughput.
result Offers 2x improvement in throughput compared to existing symmetric-operand-size instructions, with negligible overflow.
Enhances high-throughput imaging of microtubule networks, improving clarity and consistency.
problem Fluorescence noise obscures microtubule structures in high-throughput imaging.
method CycleGAN learning to enhance low-resolution images of microtubule networks.
result CycleGAN effectively identifies microtubules with high accuracy (0.93+ AUC-ROC).
Bayesian method improves hit identification in compound screening.
problem Identifying candidate hits from thousands to millions of compounds.
method Bayesian nonparametric modeling for cross-plate correlation and statistical strength.
result Significant improvements in hit identification sensitivity and specificity.
Deep learning improves decoding of constrained sequence codes, reducing errors and increasing throughput.
problem Errors during transmission of constrained sequence codes.
method Deep learning, specifically MLP and CNN networks.
result Achieved low bit error rates close to MAP decoding and improved system throughput.
EB improves asset pricing by mining large strategies without lookahead bias.
problem Lack of unbiased asset pricing models with out-of-sample performance.
method Empirical Bayes applied to 136,000 long-short strategies.
result EB provides unbiased predictions with transparent intuition.
This paper evaluates quantization techniques for deep learning inference.
problem Reducing the size and improving inference of deep neural networks.
method Review and empirical evaluation of quantization parameters for various neural network models.
result 8-bit quantization maintains accuracy within 1% of floating-point models.
We propose a method for downlink coordinated multipoint (DL CoMP) in heterogeneous fifth generation New Radio (NR) networks. The primary contribution of our paper is an algorithm to enhance the trigger of DL CoMP using online machine learning. We use support vector machine (SVM) classifiers to enhance the user downlink…
New algorithm boosts deep learning training speed.
problem Efficiently train deep neural networks on large clusters.
method Synchronous distributed SGD with amortized inference model.
result Dynamic cutoff improves convergence and training time.
Proposes a model for identifying 4G cells with network throughput problems.
problem Challenges in identifying 4G cells with network throughput issues due to network complexity and privacy concerns.
method Data-driven model using clustering and Deep Neural Networks (DNNs). Model parameters are learned from a small number of expert-labeled data. Multiple clustering models capture common features for problematic cells.
result The proposed model outperforms a simple classifier in identifying cells with network throughput problems.
Ithemal predicts processor throughput from instructions, outperforming existing tools.
problem Accurately predicting processor throughput from instructions is challenging and time-consuming.
method Uses a hierarchical LSTM-based approach to predict throughput based on opcodes and operands of instructions in a basic block.
result Ithemal predicts throughput more accurately and faster than existing tools.
This work improves wireless network learning by using side-information about interference.
problem Improving online learning algorithms in wireless networks.
method Exploiting side-information like interference levels to improve learning algorithms.
result Improved learning algorithms achieve higher throughput with fewer samples.
Hydra boosts efficiency for long-context reasoning in resource-constrained settings.
problem Quadratic complexity of transformers limits long-context reasoning in resource-constrained systems.
method Hydra uses a modular architecture with adaptive routing between sparse global attention, mixture-of-experts, and dual memories.
result Hydra achieves significant throughput and accuracy improvements for long-context reasoning.
Automatically computes reference ranges for UK Biobank cardiac data.
problem Improving healthcare by discovering patterns in large-scale population data.
method Fully automatic pipeline for 3D cardiac MR image analysis.
result Statistically significant agreement between manual and automatic indexes.
Deep reinforcement learning boosts throughput in RF-powered cognitive radio networks.
problem Maximizing throughput in large-scale, decentralized RF-powered cognitive radio networks.
method Proposes deep reinforcement learning to find optimal policies for network throughput maximization.
result Deep reinforcement learning outperforms existing techniques in large-scale RF-CRN environments.
The paper optimizes LLM inference systems through queueing theory.
problem Efficient LLM inference for AI agents under various routing topologies.
method Developed a fluid-limit framework for multi-class batched processing networks under K-FCFS scheduling.
result Proved that work-conserving scheduling algorithms maximize throughput for LLM inference.
Hybrid BFP-FP improves DNN training accuracy with 8.5x higher throughput.
problem Limited dynamic range of fixed-point arithmetic for DNN training convergence.
method Introducing HBFP, a hybrid BFP-FP approach.
result HBFP matches floating point's accuracy while delivering up to 8.5x higher throughput.
IMPACT improves RL training speed without sacrificing sample efficiency.
problem Limited sample efficiency in scalable RL architectures.
method Proposes IMPACT, extending IMPALA with target networks, circular buffers, and truncated importance sampling.
result IMPACT achieves higher rewards and significantly reduces training time compared to IMPALA.
Active learning with Gaussian processes improves crop phenotype data collection.
problem Scalability issue in high throughput phenotyping for crop improvement.
method Active learning algorithm with Gaussian Process model.
result Superior performance compared to current practices on sorghum data.
Hi-RES framework extracts medical relations from articles and EHRs.
problem Manual annotation bottleneck in relation extraction.
method Labeling sentences, creating improved negative samples, using pretrained language models, and combining EHR embeddings.
result Significant accuracy increases in relation extraction, up to 0.998 for disorder-location relations.
A new method selects inducing points to optimize high-throughput Bayesian optimisation.
problem Current inducing point selection methods sacrifice high-fidelity modeling of promising regions.
method Information-theoretic criterion to select inducing points maximizing global and maximum value uncertainties.
result Surrogate models support high-precision high-throughput Bayesian optimisation.
A central problem in neuroscience is reconstructing neuronal circuits on the synapse level. Due to a wide range of scales in brain architecture such reconstruction requires imaging that is both high-resolution and high-throughput. Existing electron microscopy (EM) techniques possess required resolution in the lateral p…
High-speed model accurately simulates neuromorphic devices.
problem Accurately modeling stochastic synapses in large-scale neuromorphic systems.
method Generative vector autoregressive model based on resistive memory cell data.
result Fast, high-throughput model reproduces synaptic parameters and correlations.
New neural network models predict molecular properties without 3D geometry, speeding up high-throughput screening.
problem Predicting molecular properties for large, complex molecules without computationally expensive 3D geometry.
method Message-passing neural networks trained with and without 3D structural information.
result Message-passing neural networks achieve similar accuracy to state-of-the-art methods without 3D geometry.
BCAE-2D compresses 3D data from a time projection chamber at high speed.
problem Compressing high-speed, sparse 3D data from a time projection chamber.
method 2D Bicephalous Convolutional Autoencoder (BCAE-2D) approach.
result 3x speedup in compression throughput with improved reconstruction accuracy.
High-throughput 3D control training system achieves 100,000 FPS.
problem Lack of efficient, single-machine reinforcement learning systems.
method Sample Factory combines asynchronous sampling and off-policy correction.
result Achieves 100,000 FPS on 3D control problems without sacrificing sample efficiency.
DRL-DPT improves energy efficiency in wireless networks with deterministic power control.
problem Severe performance degradation in traditional ICIC schemes with complex interference patterns.
method Deep Reinforcement Learning with Deterministic Policy and Target (DRL-DPT) framework.
result Consistently outperforms existing schemes in terms of energy efficiency and throughput.
New algorithm tackles resource allocation in multi-armed bandits to balance speed and throughput.
problem Balancing speed and throughput in stochastic multi-armed bandits with limited resources.
method Proposes an algorithm that trades off between information accumulation and throughput.
result Upper bounds the time taken to find the best arm with a given target success probability.
CodeX improves DNN acceleration on FPGAs by encoding and customizing bitwidth.
problem Efficiently accelerating deep neural networks on FPGAs with limited memory.
method Nonlinear encoding, automated bitwidth customization, FPGA streaming buffers.
result Average 4.65x throughput improvement on MNIST, SVHN, CIFAR-10.
Cactus improves auto-regressive decoding speed without sacrificing quality.
problem Accelerating auto-regressive decoding while maintaining output quality.
method Formalizes speculative sampling as constrained optimization and proposes Cactus for controlled divergence from the verifier distribution.
result Empirically validated effectiveness across various benchmarks.
PRETZEL optimizes machine learning prediction serving systems for better performance.
problem Low latency, high throughput, and graceful performance degradation under heavy load in prediction serving systems.
method Introducing a novel white box architecture enabling both end-to-end and multi-model optimizations.
result Average 5.5x reduction in 99th percentile latency, 25x reduction in memory footprint, and 4.7x increase in throughput compared to state-of-the-art approaches.
Paper introduces OARF benchmark suite for federated learning systems.
problem Limited diversity in federated learning benchmarks.
method Characterizes OARF benchmark suite with diverse data and applications.
result Federated learning can effectively increase end-to-end throughput.
Many modern data sets are sampled with error from complex high-dimensional surfaces. Methods such as tensor product splines or Gaussian processes are effective/well suited for characterizing a surface in two or three dimensions but may suffer from difficulties when representing higher dimensional surfaces. Motivated by…
A tutorial on optimizing complex functions with partial knowledge.
problem Optimizing functions with limited or partial information.
method Grey-box Bayesian optimization, blending black-box and white-box approaches.
result Improves optimization performance by leveraging internal information.
Chemical space is so large that brute force searches for new interesting molecules are infeasible. High-throughput virtual screening via computer cluster simulations can speed up the discovery process by collecting very large amounts of data in parallel, e.g., up to hundreds or thousands of parallel measurements. Bayes…
Small batch training improves deep neural network performance and stability.
problem Improving deep neural network performance and stability with limited computational resources.
method Experimental comparison of test performance for different mini-batch sizes, focusing on learning rate scaling and training duration.
result Best performance achieved for mini-batch sizes between 2 and 32, contrasting recent work advocating larger batch sizes.
CoinTossX is a low-latency, open-source matching engine for financial trading.
problem Efficiently matching orders in financial markets with low latency and high throughput.
method Developed in Java, orders submitted via UDP SBE, low-latency message transport (Aeron Media Driver). Separates order generation and matching.
result Demonstrated low-latency, high-throughput performance in various deployment scenarios.
Quantization scheme improves inference efficiency for deep learning models.
problem Limited computational resources and energy constraints in deploying deep learning models.
method Proposes a quantization scheme to reduce precision and improve efficiency without sacrificing accuracy.
result End-to-end post-quantization accuracies comparable to reference model achieved using a single inference batch calibration.
New algorithms speed up learning from large screens of proteins.
problem Lack of scaled data hampers biological machine learning.
method Optimized high throughput screens and generative models.
result Maximized information gain with consistent estimates of p(y∣x). MLPerf benchmarks ML training to drive performance improvements.
problem Unique challenges in ML training benchmarks.
method Developed MLPerf to overcome ML training's specific challenges.
result Quantitatively evaluated MLPerf's effectiveness.
We present Caffe con Troll (CcT), a fully compatible end-to-end version of the popular framework Caffe with rebuilt internals. We built CcT to examine the performance characteristics of training and deploying general-purpose convolutional neural networks across different hardware architectures. We find that, by employi…
New method converts ANN gates to SNNs with AMOS neurons for improved image classification.
problem Efficiently converting ANN gates to SNNs for neuromorphic hardware.
method Introducing AMOS conversion for gates in ANNs, improving accuracy and throughput.
result Improved accuracy of SNNs for ImageNet from 74.60% to 80.97%.
TUNet improves protein classification in cell images.
problem Classifying specific proteins in human cells using microscopy images.
method TUNet model incorporating segmentation maps for improved classification.
result TUNet achieves competitive performance in protein classification.
This paper improves neural network training performance by optimizing concurrency and operation scheduling.
problem Managing and scheduling fine-grained operations in neural network training for high performance.
method Extending TensorFlow runtime to enable automatic concurrency control and scheduling, using performance modeling.
result Achieved 33% average performance improvement on neural network models, up to 49%.
As Machine Learning (ML) applications increase in data size and model complexity, practitioners turn to distributed clusters to satisfy the increased computational and memory demands. Unfortunately, effective use of clusters for ML requires considerable expertise in writing distributed code, while highly-abstracted fra…
BSBO optimizes constraints for high-throughput experiments.
problem Optimizing high-throughput experiments with combinatorial constraints.
method Stochastic Bayesian optimization with submodular decomposition.
result BSBO outperforms heuristics in real-world protein datasets.
We introduce a new sequential Monte Carlo algorithm we call the particle cascade. The particle cascade is an asynchronous, anytime alternative to traditional particle filtering algorithms. It uses no barrier synchronizations which leads to improved particle throughput and memory efficiency. It is an anytime algorithm i…
Guidelines for deploying deep learning models on smartphones are developed.
problem Lack of unified guidelines for real-time deployment of deep learning solutions on smartphones.
method Unified flow of implementation for Android and iOS, use of multi-threading, benchmarking framework.
result Developed deployment approach allows easy conversion of deep learning models into real-time smartphone apps.
Con-TS optimizes wireless link throughput with latency constraints.
problem Optimizing rate selection for wireless links with latency constraints.
method Proposes Con-TS, a constrained Thompson sampling algorithm for stochastic MAB problems.
result Con-TS achieves upper bounds on expected constraint violations and throughput loss.