Secure and efficient distributed learning on devices with limited communication.
problem Limited communication and security in distributed on-device learning.
method Proposes SLSGD, a robust distributed optimization algorithm with efficient communication and attack tolerance.
result Stabilizes convergence and tolerates data poisoning on a small number of workers.
Survey on-device ML challenges and future directions.
problem Training machine learning models on-device with limited resources.
method Reformulated as resource constrained learning, comparing techniques from various AI areas.
result Identification of open challenges and future research directions.
FD and FAug reduce communication in on-device ML with non-IID data.
problem Minimize communication overhead in on-device ML with non-IID data.
method Federated distillation (FD) and federated augmentation (FAug).
result FD with FAug reduces communication by 26x while maintaining high accuracy.
On-device federated learning updates edge models by exchanging trained results.
problem Limited training data at edge devices due to model drift.
method OS-ELM for sequential training and autoencoder for anomaly detection, combined with federated learning.
result The proposed approach produces a merged model as accurately as traditional methods with lower costs.
This paper tackles energy-efficient machine learning on low-power devices.
problem Energy consumption in machine learning due to data communication.
method Dynamic averaging for integer exponential families on low-power processors.
result Achieves comparable model quality with significantly less communication and energy.
FedZKT enables resource-constrained devices to participate in federated learning with heterogeneous models.
problem Inequality in resource allocation hinders participation from resource-constrained devices in federated learning.
method Zero-shot knowledge transfer through a server-assigned distillation process.
result FedZKT effectively transfers knowledge across heterogeneous on-device models without requiring comparable local training efforts.
A lightweight FPGA-based reinforcement learning approach for edge devices.
problem Resource constraints and inefficiency of DQN on edge devices.
method OS-ELM based training algorithm and L2 regularization for stability.
result 29.77x and 89.40x faster than conventional DQN-based approach for CartPole-v0 task.
The paper optimizes neural network inference on mobile GPUs.
problem Limited computing power and thermal constraints on mobile CPUs.
method Leverage mobile GPUs for neural network inference.
result Real-time inference of deep neural networks on Android and iOS devices.
Paper tackles intermittent learning for energy-constrained machine learning tasks.
problem Energy-constrained machine learning tasks on intermittently powered systems.
method Developed an algorithm and heuristics for efficient learning under energy constraints.
result Improves energy efficiency by up to 100% and reduces learning examples by up to 50%.
Personalized stress model using transfer learning from 20 participants.
problem Limited generalizability of machine learning models due to individual physiological differences.
method Transfer learning from a base model trained on 20 participants' physiological data collected in real-time.
result Improved model personalization and cross-domain performance.
A hybrid framework reduces ML complexity on edge devices.
problem Limited memory and energy on edge devices.
method Compressed data collection and tailored deep learning network.
result Significant reduction in computational complexity and memory.
Paper proposes low-rank gradient approximation to save memory for deep neural network training.
problem Memory limitation on mobile devices for deep neural network training.
method Approximating gradient matrices using low-rank parameterization.
result Reduces training memory by about 33.0% for Adam optimization and 4.5% relative lower word error rate on ASR personalization task.
Federated learning framework extended for personalizing global models.
problem Evaluate personalization strategies for global models on-device.
method Extend federated learning framework, develop tools for analysis and evaluation.
result Personalization yields significant benefits for a large user population.
FixyNN splits CNN models into fixed and trainable parts for efficient on-device inference.
problem Energy inefficiency in on-device CNN inference for real-time computer vision.
method Co-designed hardware accelerator platform with transfer learning for training.
result Achieved nearly 2x better energy efficiency than a conventional accelerator.
Paper proposes federated learning for SNNs to enable low-power, online training.
problem Limited data at each device for on-device SNN training.
method Federated Learning (FL) for cooperative SNN training, leveraging local and global feedback.
result FL-SNN achieves significant advantages over separate training and offers a flexible trade-off between accuracy and communication load.
User-specific KWS system learns new keywords on-device.
problem Out-of-vocabulary problem in traditional KWS systems.
method Query-by-example enrollment and testing, using phonetic posteriors and FST.
result Promising performance on two keywords, preserving simplicity.
Decentralized machine learning is a promising emerging paradigm in view of global challenges of data ownership and privacy. We consider learning of linear classification and regression models, in the setting where the training data is decentralized over many user devices, and the learning algorithm must run on-device, …
New method FedEx accelerates federated hyperparameter tuning.
problem Federated hyperparameter tuning challenges in distributed learning.
method FedEx method connecting to weight-sharing, adapted for federated optimization.
result FedEx outperforms natural baselines on various benchmarks.
The paper uses quaternions to model quantum learning on devices.
problem Designing adaption and optimization techniques for quantum learning machines.
method Division algebra of quaternions to model computation and measurement on qubits, developing a training framework.
result Established quantum information processing units similar to neurons in classical approaches.
This work reduces computation cost for on-device CNN training.
problem High computation cost during on-device CNN training.
method Self-supervised instance filtering and error map pruning.
result Substantial computation saving without significant accuracy loss.
WEST compresses word embeddings and softmax layers for memory efficiency.
problem Memory constraints in large vocabulary models.
method WEST encodes words with sequences of sub-units, improving compression without performance loss.
result WEST achieves significant compression without sacrificing performance.
Paper proposes an AutoML framework for efficient device-edge co-inference.
problem Finding optimal hyper-parameters for model sparsity and feature compression.
method Sequential decision problem solved using deep reinforcement learning (DRL).
result Achieves better communication-computation trade-off and significant speedup.
FedML aims to improve FL research by providing a library and benchmark.
problem Inconsistent FL algorithm development and performance comparison.
method FedML offers an open research library and benchmark supporting diverse computing paradigms and flexible API design.
result FedML facilitates fair algorithm comparison and development in federated learning.
EdgeSpeechNets improve speech recognition on mobile devices.
problem Deploying deep learning for speech recognition on edge devices is challenging.
method Human-machine collaboration for designing efficient DNN architectures.
result EdgeSpeechNets achieve higher accuracy with smaller network size and lower computational cost.
Transfer learning improves activity recognition accuracy on smartwatches.
problem Inconsistent activity recognition accuracy across different users.
method On-device calibration of GBM by tuning parameters without retraining.
result Significant improvement in user-based accuracy for activity recognition.
A new quantization strategy reduces Transformer model size and inference time.
problem Heavy computation load and memory overhead in Transformer models for mobile devices.
method Mixed precision quantization with varying bits per word in embedding blocks.
result 11.8x smaller model size and 3.5x speed up for on-device NMT.
ONLAD Core detects anomalies in edge devices with fast learning and low power.
problem Anomaly detection in edge devices with concept drift and data transfers.
method Highly optimized neural network-based anomaly detection on edge devices.
result ONLAD Core achieves fast anomaly detection and low power consumption.
ActiveHARNet improves resource efficiency in deep learning for HAR and fall detection.
problem Resource efficiency and real-time learning for HAR models.
method Deep ensembled model with incremental learning and active learning.
result Significant efficiency boost during inference and reduction in acquired pool points.
This paper optimizes AI inference on edge devices with reduced communication and computation costs.
problem Efficiently performing AI inference on resource-constrained edge devices with reduced communication and computation costs.
method A three-step framework for effective inference: model split point selection, communication-aware model compression, and task-oriented encoding of intermediate features.
result Our proposed framework achieves a better trade-off and significantly reduces inference latency compared to baseline methods.
Two new methods reduce communication costs in federated learning.
problem Heavy communication costs in federated learning, especially with non-IID data.
method FedMMD uses MMD constraint for two-stream model training; FedFusion aggregates local and global features.
result FedMMD and FedFusion reduce communication costs by 20% and 60% respectively.
The rising popularity of intelligent mobile devices and the daunting computational cost of deep learning-based models call for efficient and accurate on-device inference schemes. We propose a quantization scheme that allows inference to be carried out using integer-only arithmetic, which can be implemented more efficie…
MDLdroid improves mobile deep learning for personal sensing with faster training.
problem Continuous local changes and resource constraints in personal mobile sensing affect global model performance.
method ChainSGD-reduce approach to reduce overhead and balance resources.
result 2x to 3.5x faster training on off-the-shelf mobile devices compared to single-device training.
Mobile training improves speech recognition for users with unique speech characteristics.
problem Limited generalization of speaker-independent speech recognition models for users with very different speech characteristics.
method Securely training personalized end-to-end speech recognition models on mobile devices, splitting gradient computation to reduce memory usage.
result On-device personalization achieved 58.1% relative word error rate reduction compared to 63.7% in a server environment, with 18.7% performance degradation.
PyTorch adds tools for pruning neural networks.
problem Model size and resource constraints in machine learning.
method Pruning techniques to reduce model size and capacity.
result Facilitates adoption of pruning in PyTorch.
VoiceFilter-Lite separates speech from background in real-time for on-device speech recognition.
problem Separate speech from background in real-time for on-device speech recognition.
method Asymmetric loss, adaptive runtime suppression, quantization to 8-bit.
result VoiceFilter-Lite achieves real-time speech separation and maintains speech recognition performance.
New pruning methods improve energy efficiency of neural networks.
problem Energy-efficient neural networks for devices with limited resources.
method Magnitude and Gradient based pruning at initialization and training of sparse architectures.
result Proposed novel pruning methods prevent full layer pruning and improve training.
PairNets optimize AI models for fast IoT applications.
problem Slow training and high memory usage of deep neural networks.
method Developed Pairwise Neural Networks (PairNets) with low memory and fast training.
result PairNets achieve faster training (one epoch) and lower prediction errors.
Paper removes sensitive data from IoT and Big Data for privacy.
problem Privacy concerns in IoT and Big Data.
method Develops new supervised and adversarial learning methods to remove sensitive data.
result Models maintain predictive model utility while making sensitive predictions ineffective.
A framework for large-scale federated learning with non-IID data.
problem Stability of models trained on non-IID data in federated learning.
method Generation of non-IID datasets and modular evaluation framework.
result Open-source benchmark for large-scale federated learning research.
Much of the focus in the design of deep neural networks has been on improving accuracy, leading to more powerful yet highly complex network architectures that are difficult to deploy in practical scenarios, particularly on edge devices such as mobile and other consumer devices given their high computational and memory …
Framework prevents data leakage in mobile cloud DNNs.
problem Data leakage from cloud DNNs poses privacy risks.
method Privacy-preserving reinforcement learning framework.
result Framework successfully defends against various privacy attacks.
Improves E2E ASR performance on numeric sequences with additional training data and denormalization.
problem Challenges in recognizing numeric sequences out-of-vocabulary in ASR systems.
method Uses text-to-speech for additional numeric training data and a small-footprint neural network for denormalization.
result Reduction of WER by up to a factor of 8 in the longest numeric sequences.
Faster convergence in federated learning for non-convex problems.
problem Accelerating convergence in federated learning for non-convex models.
method Reformulated federated learning as gradient-based method with biased gradients, proving convergence for non-convex problems and proposing an accelerated algorithm.
result Proved federated averaging algorithm converges for non-convex problems and proposed an accelerated federated learning algorithm with convergence guarantee.
Federated learning improves by unbiased gradient aggregation and controllable meta updating.
problem Gradient biases and inconsistency between target and optimization objectives in federated averaging.
method Unbiased gradient aggregation with keep-trace gradient descent and gradient evaluation strategy, controllable meta updating with small data samples.
result Faster convergence and higher accuracy with different network architectures in various FL settings.
Study presents a low-cost local motion planner for vineyard navigation.
problem Autonomous navigation in vineyards with limited resources.
method RGB-D camera, dual layer control algorithm, deep learning synergy.
result Robust motion planning for vineyard navigation achieved.
Entity linking is the task of mapping potentially ambiguous terms in text to their constituent entities in a knowledge base like Wikipedia. This is useful for organizing content, extracting structured data from textual documents, and in machine learning relevance applications like semantic search, knowledge graph const…
This paper proposes a communication-efficient deep anomaly detection framework for industrial IoT.
problem Accurately detecting anomalies in time-series data from edge devices in industrial IoT.
method A federated learning-based approach with an Attention Mechanism-based Convolutional Neural Network-Long Short Term Memory (AMCNN-LSTM) model and gradient compression.
result The proposed framework accurately and timely detects anomalies with reduced communication overhead.
Federated MTL learns personalized models under mixed distributions.
problem Heterogeneity of local data distributions leads to poor global model performance.
method Proposes federated MTL under mixture of distributions, using penalized optimization and federated EM-like algorithms.
result Models with higher accuracy and fairness than state-of-the-art methods.