FedZKT enables resource-constrained devices to participate in federated learning with heterogeneous models.
problem Inequality in resource allocation hinders participation from resource-constrained devices in federated learning.
method Zero-shot knowledge transfer through a server-assigned distillation process.
result FedZKT effectively transfers knowledge across heterogeneous on-device models without requiring comparable local training efforts.
On-device federated learning updates edge models by exchanging trained results.
problem Limited training data at edge devices due to model drift.
method OS-ELM for sequential training and autoencoder for anomaly detection, combined with federated learning.
result The proposed approach produces a merged model as accurately as traditional methods with lower costs.
This paper optimizes how deep learning models are distributed across different devices.
problem Optimizing how large, complex neural networks are split across multiple devices.
method Identified and solved an optimization problem for device placement of DNN operators.
result Automated algorithms that solve the device placement problem for modern pipelined settings.
RNNs learn device models from input/output data.
problem Learning complex device models from limited data.
method Empirical study using RNNs to model six different devices.
result RNNs can generate functional software-only models of hardware devices.
Survey on-device ML challenges and future directions.
problem Training machine learning models on-device with limited resources.
method Reformulated as resource constrained learning, comparing techniques from various AI areas.
result Identification of open challenges and future research directions.
This paper analyzes mobile device training of deep learning models.
problem Performance characterization of training deep learning models on mobile devices.
method Experiments on NVIDIA TX2, benchmark suite, and tools for performance analysis.
result Interesting performance problems and opportunities revealed.
SplitEasy trains ML models on mobile devices without server data transfer.
problem Training complex DL models on resource-limited mobile devices.
method Split learning approach where sensitive layers are trained locally, computationally intensive layers on server.
result SplitEasy trains models on mobile devices with minimal data transfer, near-constant time per sample.
This paper optimizes AI inference on edge devices with reduced communication and computation costs.
problem Efficiently performing AI inference on resource-constrained edge devices with reduced communication and computation costs.
method A three-step framework for effective inference: model split point selection, communication-aware model compression, and task-oriented encoding of intermediate features.
result Our proposed framework achieves a better trade-off and significantly reduces inference latency compared to baseline methods.
Semi-decentralized federated learning combines device-to-server and device-to-device communications for faster convergence.
problem Faster convergence in federated learning with decentralized model training.
method Two timescale hybrid federated learning (TT-HF) with cooperative D2D model aggregations.
result Achieves sublinear convergence rate of O(1/t) with adaptive control algorithm.
Paper tackles efficient allocation of multiple devices to users for AutoML services.
problem Allocating multiple devices to multiple users for AutoML services efficiently.
method Develops a multi-device, multi-tenant algorithm for GP-EI, achieving near-linear speedup.
result Achieves near-linear speedup when users are many more than devices.
AMS improves video inference on edge devices by adapting a small model with online knowledge distillation.
problem High computation cost of Deep Neural Networks for real-time video inference on edge devices.
method AMS uses a remote server to continually train and adapt a small model on edge devices, using online knowledge distillation from a large model.
result 0.4--17.8 percent mean Intersection-over-Union improvement in video semantic segmentation.
Survey of knowledge distillation for resource-limited devices.
problem Deploying large deep learning models on resource-limited devices.
method Knowledge distillation using a smaller model trained with information from a larger model.
result A new metric (distillation metric) for comparing different knowledge distillation algorithms.
DarkneTZ protects edge devices from DNN model leaks using TEE and model partitioning.
problem Privacy risks of pre-trained DNNs on edge devices through membership inference attacks.
method Model partitioning into sensitive and untrusted parts, leveraging TEE.
result DarkneTZ provides reliable model privacy with minimal performance overhead.
FD and FAug reduce communication in on-device ML with non-IID data.
problem Minimize communication overhead in on-device ML with non-IID data.
method Federated distillation (FD) and federated augmentation (FAug).
result FD with FAug reduces communication by 26x while maintaining high accuracy.
Federated learning enables private model training across devices.
problem Private and collaborative machine learning across multiple devices.
method Designing scalable, privacy-preserving FL systems using graph-based optimization.
result Personalized models for each device while maintaining data privacy.
Fog learning distributes ML model training across heterogeneous devices and networks.
problem Challenges with conventional federated learning in heterogeneous networks.
method Intelligent distribution of ML model training across nodes from edge devices to cloud servers.
result Enhanced federated learning with multi-layer hybrid framework considering network, heterogeneity, and proximity.
TOCO framework compresses neural networks based on tolerance analysis.
problem Deploying large neural networks on edge devices with limited resources.
method TOCO uses tolerance analysis to perform fine-grained compression, allowing flexibility to hardware changes.
result Fine-grained compression of neural networks on edge devices.
Paper optimizes neural architectures for multiple device constraints.
problem NAS ignores device constraints like latency and energy.
method Developed MONAS and DPP-Net for multi-objective optimization.
result Found Pareto-optimal architectures for various devices.
Paper proposes low-rank gradient approximation to save memory for deep neural network training.
problem Memory limitation on mobile devices for deep neural network training.
method Approximating gradient matrices using low-rank parameterization.
result Reduces training memory by about 33.0% for Adam optimization and 4.5% relative lower word error rate on ASR personalization task.
POET enables large neural network training on tiny devices with reduced energy.
problem Training large neural networks on memory-limited edge devices.
method Jointly optimizes rematerialization and paging for memory reduction, formulating an MILP for energy-efficient training.
result POET trains ResNet-18 and BERT within Cortex-M memory constraints, outperforming current methods in energy efficiency.
Federated learning algorithm reduces global model size by combining local and global representations.
problem Scalability issues in training large models on private data distributed over multiple devices.
method Proposes a federated learning algorithm that jointly learns compact local representations and a global model.
result The global model can be smaller since it only operates on local representations, reducing the number of communicated parameters.
The paper optimizes neural network inference on mobile GPUs.
problem Limited computing power and thermal constraints on mobile CPUs.
method Leverage mobile GPUs for neural network inference.
result Real-time inference of deep neural networks on Android and iOS devices.
Mobile training improves speech recognition for users with unique speech characteristics.
problem Limited generalization of speaker-independent speech recognition models for users with very different speech characteristics.
method Securely training personalized end-to-end speech recognition models on mobile devices, splitting gradient computation to reduce memory usage.
result On-device personalization achieved 58.1% relative word error rate reduction compared to 63.7% in a server environment, with 18.7% performance degradation.
New ML approach for edge devices tackles deployment challenges.
problem Challenges in deploying ML models on edge devices.
method Specialized ML development and deployment approach for edge devices.
result Prototype demonstrates efficient and high-quality solutions.
Deep learning identifies unknown IoT devices in network traffic.
problem Unauthorized IoT devices pose security risks in BYOD environments.
method Deep learning applied to network traffic images for IoT device identification.
result Over 99% accuracy in identifying 10 IoT devices and non-white-listed devices.
Paper proposes an AutoML framework for efficient device-edge co-inference.
problem Finding optimal hyper-parameters for model sparsity and feature compression.
method Sequential decision problem solved using deep reinforcement learning (DRL).
result Achieves better communication-computation trade-off and significant speedup.
A new federated learning framework for handling device heterogeneity.
problem Handling device heterogeneity in federated learning.
method Superquantile-based objective with parameterized levels of conformity, optimized using secure aggregation.
result The optimization algorithm converges to a stationary point.
Federated learning on edge devices achieves high accuracy with minimal data exchange.
problem Training deep neural networks on edge devices while maintaining user privacy.
method Training CNN, LSTM, and MLP on MNIST data using federated learning on edge devices (Raspberry Pi4s). Experimentally tested on IID and non-IID samples.
result Up to 85% test accuracy achieved with 2 minutes of training time and <10 MB data exchange per device.
A method for trust evaluation of devices in human-device coexistence systems.
problem Efficient trust evaluation of devices in systems with diverse physical and social attributes.
method Canonical correlation analysis-enhanced hypergraph self-supervised learning (HSLCCA).
result The proposed HSLCCA method significantly outperforms baseline algorithms in identifying trusted devices.
PruneFL reduces FL training time on edge devices by pruning model size.
problem Limited computation and communication resources on edge devices in FL.
method Adaptive and distributed parameter pruning during FL process.
result Pruned model converges to similar accuracy as original model with reduced training time.
DoCoFL compresses model updates for cross-device federated learning.
problem Downlink compression for cross-device federated learning where clients may appear only once.
method Proposes DoCoFL framework for downlink compression in cross-device federated learning.
result Significant bi-directional bandwidth reduction with competitive accuracy.
EDCompress optimizes energy efficiency of CNN models on edge devices.
problem Low energy consumption for edge devices with diverse dataflow types.
method Energy-aware model compression using reinforcement learning.
result Improves energy efficiency by 20X, 17X, 37X in various networks.
Two approaches scale up DNN optimization for diverse edge devices.
problem Optimizing DNNs for edge devices with varying performance requirements.
method Reuse performance predictors on proxy devices and build scalable predictors.
result Optimized DNN designs for many different edge devices without lengthy optimization.
DirNet compresses RNNs for mobile devices with minimal accuracy loss.
problem High computational and memory demands of RNNs on mobile devices.
method Dynamic dictionary learning for adaptive sparsity and compression rate.
result Significant accuracy improvement with eight times model size reduction.
Federated Learning leaks user-specific information, making devices deanonymizable.
problem Federated Learning leaks user-specific information, making devices deanonymizable.
method Identified subtle variations in model updates that encode user-specific data. Proposed data-augmentation strategies to mitigate deanonymization.
result Data-augmentation strategies offer substantial protection against deanonymization threats with little effect on utility.
PPGnet model estimates heart rate from PPG signals without motion artifacts.
problem Wearable PPG devices struggle with motion artifacts.
method End-to-end deep learning model using 8-second PPG signals.
result Achieved mean absolute error of 3.36+-4.1 BPM on IEEE SPC 2015 dataset.
This paper tackles energy-efficient machine learning on low-power devices.
problem Energy consumption in machine learning due to data communication.
method Dynamic averaging for integer exponential families on low-power processors.
result Achieves comparable model quality with significantly less communication and energy.
Flower framework simplifies federated learning experiments on edge devices.
problem Realistic implementation of Federated Learning on edge devices is challenging.
method Developed a comprehensive federated learning framework, Flower, supporting large-scale experiments on heterogeneous devices.
result Flower enables federated learning experiments with up to 15M client size using only two high-end GPUs.
CheckNet verifies neural network inference on untrusted devices.
problem Ensuring secure and tamper-proof inference on untrusted devices.
method A checksum-based approach for neural network inference verification.
result Excellent attack detection and success bounds on various models.
Federated learning struggles with non-IID data, but a strategy improves model accuracy.
problem Federated learning accuracy drops significantly with non-IID data.
method Identified weight divergence as the cause, quantified by EMD, and proposed a solution of sharing a subset of globally shared data.
result Accuracy can be increased by 30% for CIFAR-10 with only 5% globally shared data.
RONA compresses complex models while ensuring privacy.
problem Deploying complex deep neural networks on mobile devices poses privacy risks and computational constraints.
method RONA uses knowledge distillation, hint learning, and self learning to train a compact neural network with differential privacy guarantees.
result RONA achieves 20x compression and 19x speed-up with 0.97% accuracy loss on SVHN while maintaining strong privacy.
High-speed model accurately simulates neuromorphic devices.
problem Accurately modeling stochastic synapses in large-scale neuromorphic systems.
method Generative vector autoregressive model based on resistive memory cell data.
result Fast, high-throughput model reproduces synaptic parameters and correlations.
WaveletNet improves edge device efficiency with logarithmic convolution.
problem Efficiency and performance on edge devices for CNNs.
method Introduces WaveletNet architecture with wavelet convolution and depthwise fast wavelet transform.
result WaveletNet achieves superior and comparable performance to state-of-the-art models on CIFAR-10 and ImageNet.
Paper proposes a multi-phase pruning pipeline for deep ensemble learning on IIoT devices.
problem Computational limitations of IoT devices for deep learning models.
method Generates diverse pruned models, applies integer quantization, and uses clustering-based pruning.
result Significant reduction in model size (up to 90%) and improved performance (up to 7%) on IIoT devices.
Distributed learning adapts to diverse devices, improving performance.
problem Training neural networks on devices with varying capabilities and resources.
method Each device trains a customized neural network, sharing parameters with others.
result Achieves higher rewards on more powerful devices without sacrificing weaker ones.
MetaDVFS uses device and application metadata to improve DVFS efficiency.
problem Improving energy efficiency in mobile platforms with diverse applications and hardware.
method Formulates DVFS as a multi-task reinforcement learning problem and introduces MetaDVFS, leveraging metadata for knowledge transfer.
result MetaDVFS achieves up to 26% improvement in Quality of Experience and up to 17% improvement in Performance-Power Ratio.
Efficient quantization scheme for neural networks using integer arithmetic.
problem Efficient inference on mobile devices with limited computational resources.
method Quantization and co-design training procedure for integer-only inference.
result Improves accuracy-latency tradeoff on various models and hardware.
Flexible device participation improves federated learning convergence.
problem Strict device participation limits federated learning reach.
method Analytical results and new aggregation scheme for flexible participation.
result Convergence improved with flexible device participation.