The study corrects misconceptions in GBDT speed benchmarks.
problem Misleading speed benchmarks of GBDT algorithms.
method Explained and criticized several straightforward benchmarking methods, outlined fair benchmark requirements.
result A fair GBDT speed benchmark requires specific conditions.
SaaS uses training speed to infer unknown labels in semi-supervised learning.
problem Inference of unknown labels in semi-supervised learning.
method Uses learning speed during stochastic gradient descent to infer unknown labels.
result Achieves state-of-the-art results in semi-supervised learning benchmarks.
SG-DNI outperforms standard neural interfaces in speed and accuracy.
problem Inflexible structure of neural networks limits parallelization.
method Introduces synthetic gradients with decoupled neural interfaces (SG-DNI).
result SG-DNI is over 3-fold faster with comparable accuracy.
Non-autoregressive method speeds up protein folding prediction 23 times.
problem Generating protein sequences with higher order interactions.
method Discrete diffusion conditioned on 3D structure using ProteinMPNN.
result 23 times speed up in inference without performance loss.
LaRT models LLMs' response accuracy and CoT length to evaluate reasoning ability and speed.
problem Valid evaluation of Large Language Models (LLMs) via response accuracy and chain-of-thought length.
method Introduces Latency-Response Theory (LaRT) to jointly model response accuracy and CoT length using latent ability and latent speed.
result LaRT yields higher estimation accuracy and shorter confidence intervals for latent traits compared to IRT.
Grouped Gaussian Processes improve solar power and wind speed forecasting.
problem Forecasting distributed solar power and wind speed at multiple sites.
method Coupled Gaussian process priors over groups of node and weight functions.
result Our approach maintains or improves point-prediction accuracy and provides better quantification of predictive uncertainties.
This work introduces a new benchmark to compare neural network training algorithms.
problem Lack of reliable benchmarks to compare training algorithms effectively.
method Developed a new benchmark called AlgoPerf: Training Algorithms benchmark.
result Demonstrated the feasibility of the benchmark and set a provisional state-of-the-art.
Model predicts traffic speed using urban incidents.
problem Accurately predicting traffic speed in urban areas.
method Deep Incident-Aware Graph Convolutional Network (DIGC-Net).
result Model outperforms competing benchmarks in traffic speed prediction.
Benanza speeds up DL model optimization by automatically generating micro-benchmarks and identifying inefficiencies.
problem Slow characterization/optimization cycles for DL models on GPUs.
method Benanza includes a model processor, benchmark generator, database, and analyzer.
result Benanza identifies optimizations in parallel layer execution, cuDNN, framework inefficiency, layer fusion, and Tensor Cores.
Two new SR methods improve image quality and speed.
problem Improving HR image quality and computational speed.
method VDSR-ResNeXt and SRCGAN.
result Both methods outperformed existing techniques in benchmark tests.
New interior-point method tackles Wasserstein barycenter problem efficiently.
problem Computing Wasserstein barycenter for large support measures.
method Adapted interior-point method exploiting problem's matrix structure.
result Achieves a well-balanced tradeoff between accuracy and speed.
Unconstrained MLIPs outperform constrained ones in accuracy and speed.
problem Improving the efficiency and accuracy of machine-learned interatomic potentials.
method Investigated unconstrained models trained on large datasets compared to physically constrained models.
result Unconstrained MLIPs can be superior in accuracy and speed compared to physically constrained models.
Improved GAN training speed and quality with FastGAN.
problem Slower convergence and less expressive models in GAN training.
method Adversarial training with FastGAN algorithm.
result Better generation quality with less training time.
MultiRocket boosts TSC speed and accuracy with pooling and transformations.
problem Efficient time series classification with high accuracy.
method Multiple pooling operators and transformations applied to raw and differenced series.
result MultiRocket outperforms MiniRocket and is competitive with state-of-the-art methods in terms of accuracy and speed.
A new interpolation method speeds up neural ODE training.
problem Efficiently approximating gradients in neural ODEs.
method Interpolation-based technique to approximate gradients.
result Our method trains neural ODEs faster than the reverse dynamic method.
Reservoir Memory Machines solve benchmark tasks faster than Neural Turing Machines.
problem Training Neural Turing Machines is hard and limits their applicability.
method Proposes Reservoir Memory Machines, combining neural network flexibility with Turing machine capabilities, but with faster training via alignment and linear regression.
result Reservoir Memory Machines solve benchmark tasks as well as Neural Turing Machines but are much faster to train.
In this paper we propose an algorithm that builds sparse decision DAGs (directed acyclic graphs) from a list of base classifiers provided by an external learning method such as AdaBoost. The basic idea is to cast the DAG design task as a Markov decision process. Each instance can decide to use or to skip each base clas…
This paper speeds up large-scale deep learning training.
problem Training large-scale deep architectures is slow and resource-intensive.
method Systematic approach to identify bottlenecks, develop guidelines, and derive lemmas.
result Developed procedures and lemmas for setting minibatch size, choosing algorithms, and determining component quantities.
Integrating wind power into the grid is challenging because of its random nature. Integration is facilitated with accurate short-term forecasts of wind power. The paper presents a spatio-temporal wind speed forecasting algorithm that incorporates the time series data of a target station and data of surrounding stations…
Tensor trains speed up option pricing for multi-asset options.
problem Speeding up option pricing for multi-asset options.
method Tensor train learning algorithms to compress functions with parameter dependence.
result The proposed method outperforms Monte Carlo-based pricing in computational complexity.
CodedFedL speeds up federated learning in MEC networks by 15x.
problem Slow convergence in federated learning due to heterogeneity and stochastic fluctuations.
method Injects structured coding redundancy into federated learning to mitigate stragglers and speed up training.
result CodedFedL speeds up the training procedure by up to 15x compared to benchmark schemes.
We speed up Gaussian process cross-validation calculations and improve model diagnostics.
problem Efficiently calculating cross-validation residuals and their covariances in Gaussian processes.
method Generalized fast Gaussian process leave-one-out formulae to multiple-fold cross-validation, highlighting covariance structures.
result Correcting for residual covariances in cross-validation improves back to Maximum Likelihood Estimation.
Two methods using Chebyshev tensors improve accuracy and speed in computing Dynamic Initial Margin.
problem Computing Dynamic Initial Margin (DIM) with high accuracy and speed.
method Two methods based on Chebyshev tensors implemented in Monte Carlo engine.
result Better accuracy, speed, and implementation efforts compared to benchmarks.
Meta-Surrogate model speeds up HPO benchmarking.
problem Limited and expensive real-world HPO benchmarks.
method Meta-surrogate model trained on off-line data.
result More coherent and statistically significant conclusions faster.
PyTorch combines usability and speed in deep learning.
problem Combining usability and speed in deep learning frameworks.
method Imperative programming style, Pythonic interface, hardware acceleration support.
result PyTorch achieves both usability and high performance.
Research optimizes C++ patterns for HFT, reducing latency and improving profitability.
problem Optimizing latency-critical code for high-frequency trading systems.
method Creation of a Low-Latency Programming Repository, optimisation of trading strategy, implementation of Disruptor pattern.
result Significant performance improvements in speed and profitability.
A new federated learning method speeds up training by selecting faster nodes first.
problem System heterogeneity and stragglers slow down federated learning.
method Adaptive selection of nodes based on data statistical characteristics.
result Significant speedup in wall-clock time compared to standard federated learning.
New benchmarks provide full training data for NAS research.
problem Limited training data on popular benchmarks restricts multi-fidelity techniques.
method SVD and noise modeling to create surrogate benchmarks with full training info.
result Learning curve extrapolation framework improves single-fidelity algorithms.
Develops a new trading strategy for statistical arbitrage with path-dependent signals.
problem Optimal execution in statistical arbitrage strategies with dynamic predictive signals.
method Signature-based framework modeling alpha and trading speed as linear functionals of truncated signature of market path.
result Fitted policy achieves higher return on turnover compared to a z-score benchmark.
DeepOBS benchmarks deep learning optimizers.
problem Quantitative evaluation of deep learning optimizers.
method Python package with realistic benchmarks and back-ends.
result Automates benchmarking of stochastic optimizers.
Replicated Softmax model, a well-known undirected topic model, is powerful in extracting semantic representations of documents. Traditional learning strategies such as Contrastive Divergence are very inefficient. This paper provides a novel estimator to speed up the learning based on Noise Contrastive Estimate, extende…
Fastest video anomaly detection via teacher-student distillation.
problem Anomaly detection in video at high speed.
method Adversarial knowledge distillation from object-level teacher models.
result 7-62 times faster than state-of-the-art methods.
A deep learning model speeds up computation of numerous implied volatilities.
problem Frequent computation of numerous implied volatilities using iteration methods like Newton-Raphson reaches processing speed limits.
method Emulated Newton-Raphson method using PyTorch and optimized with TensorRT.
result Up to 1,000 times faster than a benchmark implementation of Newton-Raphson.
Faster deep neural networks converge and generalize better.
problem Slower training and insufficient data for deep networks.
method Optimization algorithm based on generalized-optimal updates.
result Two orders of magnitude speed up over traditional back-propagation.
Warm-start strategies speed up GP inference by 19x.
problem Efficient sequential inference in Gaussian processes.
method Three warm-start strategies exploiting smaller linear systems.
result Warm-starting achieves up to 19x speed-up in convergence.
AcceleratedLiNGAM speeds up causal discovery methods for large datasets.
problem Slow causal discovery methods for large-scale datasets.
method Parallelized LiNGAM method with GPU acceleration.
result Up to 32-fold speed-up on benchmark datasets.
RL for image captions improved with a language prior.
problem Learning biases and large sample space issues in RL image captioning.
method Added a language prior to constrain the action space.
result RL with the language prior module performs better in readability and speed.
A simple baseline for extreme multi-label classification using random projections.
problem Automatically annotating data points with relevant labels from a large label vocabulary.
method On-the-fly global embedding using random projections, with an ensemble of learners.
result Competitive accuracy compared to existing methods, with significant speed-up and model-size reduction.
TiDE uses MLP for fast, simple long-term time-series forecasting.
problem Long-term time-series forecasting challenges.
method Time-series Dense Encoder (TiDE) based on MLP.
result TiDE matches or outperforms Transformer models while being 5-10x faster.
Loihi neuromorphic chip outperforms conventional hardware in keyword spotting efficiency.
problem Benchmarking keyword spotting efficiency on neuromorphic hardware.
method Comparative analysis of a two-layer neural network trained to recognize a single phrase on Intel's Loihi neuromorphic chip and conventional hardware devices.
result Loihi outperforms conventional hardware on energy cost per inference for this keyword spotting application.
New method pools graphs with edge features for molecular data.
problem Pooling graphs with edge features for molecular data.
method Proposes two types of pooling layers compatible with edge-feature graph-convolutional architecture.
result Significantly outperforms previous benchmarks on three out of four MoleculeNet datasets.
SRSGD improves DNN training speed and accuracy.
problem Training deep neural networks is computationally expensive and slow.
method Scheduled Restart SGD (SRSGD) combines NAG momentum with momentum reset.
result SRSGD achieves better error rates with fewer training epochs.
FAVANO improves federated learning for resource-constrained environments.
problem Asynchronous communication in federated learning leads to bias and scalability issues.
method FAVANO is a novel asynchronous federated learning framework for resource-constrained environments.
result FAVANO outperforms existing methods on standard benchmarks.
This paper improves NAQ method for faster convergence on Tensorflow.
problem Non-convex optimization problems.
method Modified Nesterov's Accelerated Quasi-Newton (NAQ) method on Tensorflow.
result mNAQ converges better and faster than first and second order optimizers.
PointPillars improves object detection speed and accuracy in point clouds.
problem Encoding point clouds for efficient object detection.
method PointPillars uses PointNets to learn pillar representations of point clouds, combined with a lean downstream network.
result PointPillars outperforms previous encoders in both speed and accuracy.
A hardware-based reservoir computing system predicts time series with high speed and accuracy.
problem Processing time-dependent signals with high speed and accuracy.
method A hardware-based reservoir computing system using a field-programmable gate array (FPGA) for both the reservoir and output layers.
result Achieves comparable accuracy to software approaches but with a superior real-time prediction rate up to 160 MHz.
Bayesian neural networks speed up numerical integration.
problem Scalability of Bayesian quadrature methods.
method Bayesian Stein networks using neural networks and Laplace approximation.
result Orders of magnitude speed-up on benchmark functions and real-world problems.
Queue-based resampling tackles online class imbalance learning with selective resampling of past examples.
problem Online class imbalance learning under class imbalance and concept drift.
method Queue-based resampling algorithm that selectively includes past examples in the training set.
result Queue-based resampling outperforms state-of-the-art methods in terms of learning speed and quality.