Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

14294357 · Jun 202019922001200920182026
48 results for server failures

DC-Prophet predicts catastrophic server failures in datacenters.

problem Forecasting catastrophic machine failures in datacenters.
method Two-stage framework based on One-Class Support Vector Machine and Random Forest.
result DC-Prophet achieves an AUC of 0.93 and F3-score of 0.88 in predicting the next machine failure.

LiuBei is a resilient ML algorithm that tolerates Byzantine workers and servers without trusting any component.

problem Byzantine failures in distributed ML solutions.
method Byzantine-resilient ML algorithm that aggregates gradients and replicates parameter servers, using a filtering mechanism and scatter/gather protocol.
result LiuBei achieves Byzantine resilience to both servers and workers and guarantees convergence, with an accuracy loss of around 5% and a 24% convergence overhead.

BrainTorrent uses peer-to-peer FL for medical image segmentation without a central server.

problem Lack of sufficient annotated data for personalized medical models.
method Peer-to-peer federated learning framework without a central server.
result BrainTorrent outperforms traditional server-based FL and achieves similar performance to pooled data training.

Predicts hard disk failure with LSTM networks.

problem Improving system reliability of backend data servers.
method Deep LSTM Network for pattern recognition, hyper-parameter tuning, feature extraction, online predictions.
result Average precision of 0.8435 in predicting disk failure 10 days in advance.

BRIDGE improves decentralized learning resilience against Byzantine failures.

problem Decentralized learning in distributed systems is vulnerable to Byzantine failures.
method Introduces BRIDGE, a scalable Byzantine-resilient decentralized gradient descent algorithm.
result Proves algorithmic and statistical convergence guarantees for BRIDGE.

DRACO improves robustness in distributed training by using redundant gradients.

problem Byzantine failures and adversarial compute nodes corrupting model training.
method DRACO uses redundant gradients to eliminate adversarial updates, leveraging coding theory.
result DRACO provides problem-independent robustness guarantees and is significantly faster than median-based approaches.

New federated learning protocols resist Byzantine failures and offer privacy guarantees.

problem Resisting Byzantine failures in federated learning.
method Proposes robust federated learning protocols with optimal statistical rates and privacy guarantees.
result Achieves nearly optimal statistical rates and tight rate in terms of all parameters for strongly convex losses.

This paper introduces a new framework for collective online learning of Gaussian processes in massive multi-agent systems.

problem The inefficiency of centralized communication in distributed machine learning systems.
method A novel Collective Online Learning Gaussian Process framework that allows each agent to build its local model and exchange it with others via peer-to-peer communication.
result Empirical results demonstrate the efficiency of the framework on both synthetic and real-world datasets.

We introduce a methodology for efficient monitoring of processes running on hosts in a corporate network. The methodology is based on collecting streams of system calls produced by all or selected processes on the hosts, and sending them over the network to a monitoring server, where machine learning algorithms are use…

2017-07-12abs ↗pdf ↗

Lapse improves parameter servers by dynamically allocating parameters, achieving near-linear scaling.

problem Efficiently managing distributed training with reduced communication overhead.
method Integrate dynamic parameter allocation into parameter servers, proposing Lapse.
result Lapse provides near-linear scaling and can be orders of magnitude faster than existing parameter servers.

A decentralized approach for multi-source domain adaptation.

problem Transfer knowledge from multiple related domains to an unlabeled target domain.
method Federated Dataset Dictionary Learning (FedDaDiL) framework, eliminating central server, using Wasserstein barycenters.
result Our decentralized approach effectively adapts source domains to an unlabeled target domain.

Algorithm stabilizes queues in asymmetric systems with unknown service rates.

problem Stabilizing queues in multi-class multi-server systems with unknown service rates.
method Proposes UCB and Thompson Sampling algorithms to stabilize queues while learning service rates.
result Achieves system stability with an average queue length bound of \(O(\min\{N,K\}/ε)\) for large time horizon \(T\).

Corella protects client data privacy in multi-server learning with correlated queries.

problem Protecting client data privacy in multi-server machine learning.
method Proposes a private multi-server learning approach using correlated queries and strong noise.
result Mitigates client data leakage with high accuracy and minimal computational effort.

MDLdroid improves mobile deep learning for personal sensing with faster training.

problem Continuous local changes and resource constraints in personal mobile sensing affect global model performance.
method ChainSGD-reduce approach to reduce overhead and balance resources.
result 2x to 3.5x faster training on off-the-shelf mobile devices compared to single-device training.

Paper explores differential privacy in high-dimensional federated learning, tackling server trustworthiness and estimation.

problem Maintaining privacy in distributed environments with high-dimensional data.
method Investigates scenarios with untrusted and trusted central servers, introduces novel federated estimation algorithms for linear regression models.
result Tight minimax rates depend on high-dimensionality even with sparsity assumptions, and novel algorithms handle slight variations among distributed models.

A smart device optimizes server selection for energy and latency in dynamic networks.

problem Optimizing server selection for mobile devices in edge computing with uncertainty and dynamic changes.
method Formulated as a budget-limited multi-armed bandit problem, a policy is proposed to minimize regret.
result The proposed method outperforms existing solutions in terms of energy and latency.

CSE-FSL reduces communication and storage costs in federated learning.

problem High communication and storage costs in federated learning.
method CSE-FSL uses an auxiliary network to locally update client models and sends only selected epochs' smashed data.
result Significant communication reduction with state-of-the-art convergence and model accuracy.

Enhanced federated learning reduces communication costs and improves model accuracy.

problem Reducing communication costs in federated learning.
method Asynchronous model update and temporally weighted aggregation.
result The proposed algorithm outperforms baseline in terms of communication cost and model accuracy.

Detects backdoors in outsourced models by replicating training steps across multiple servers.

problem Detecting backdoors in models trained on cloud providers without prior knowledge.
method Replicate training steps across multiple servers to identify deviations and malicious updates.
result 99.6% accuracy in identifying backdoored models out of 50% malicious providers.

Stanza separates convolutional and fully connected layers for faster deep learning training.

problem Heavy data transfer between workers and servers in distributed deep learning.
method Layer separation: most nodes train convolutional layers, others train fully connected layers only.
result Significant acceleration of training time (1.34x--13.9x) over current systems.

This paper improves privacy in federated learning without a trusted server.

problem Privacy in federated learning with silos that distrust each other.
method Introduces Inter-Silo Record-Level Differential Privacy (ISRL-DP) and accelerated algorithms for convex and smooth losses.
result Achieves optimal privacy and accuracy tradeoffs in federated learning.

This work provides bounds on generalization error and privacy leakage in federated learning.

problem Bounding generalization error and privacy leakage in federated learning.
method Information-theoretic framework for classical, distributed, and federated learning.
result Upper and lower bounds on generalization error and privacy leakage.

Paper tackles federated linear bandit learning with AirComp for noisy channels.

problem Minimize cumulative regret in federated linear bandit learning.
method Proposes a federated linear bandits scheme using over-the-air computation (AirComp) over noisy fading channels.
result Determines the regret bound of the proposed scheme.

Client-based machine learning uses mobile devices for computation, improving privacy and reducing data upload.

problem Exploiting mobile devices for machine learning tasks to protect privacy and reduce data upload.
method Leveraging local hardware and data on mobile devices for computation-intensive tasks, only uploading results.
result Client-based machine learning can relieve server burdens and protect user privacy.

Economics tool predicts failure times in reliability systems.

problem Predicting optimal failure times in weighted k-out-of-n reliability systems with heterogeneous component failure.
method Using rational expectations to analyze and predict failure times in reliability systems with heterogeneous component failure.
result Different measures are optimal for predicting system failure depending on component failure distributions.

Paper proposes a GPU-based system for training massive deep learning models in ads systems.

problem Training massive deep learning models with terabyte-scale parameters in ads systems.
method Hierarchical GPU parameter server with 3-layer storage (GPU High-Bandwidth Memory, CPU main memory, SSD).
result 4-node hierarchical GPU parameter server trains a model 2X faster than a 150-node in-memory system.

DFedAvgM is a decentralized FedAvg with momentum for privacy and communication efficiency.

problem Efficiently train models with privacy and communication efficiency in federated learning.
method Decentralized Federated Averaging with Momentum (DFedAvgM) on clients connected by an undirected graph, using stochastic gradient descent with momentum and quantization.
result DFedAvgM converges under trivial assumptions and can be improved with the PŁ property, numerically verified.

Paper improves communication in distributed optimization, reducing worker-to-server data exchanges.

problem Efficiency in server-to-worker communication in distributed optimization.
method MARINA-P, a novel downlink compression method using correlated compressors; M3, combining MARINA-P with uplink compression.
result MARINA-P achieves provably superior server-to-worker communication complexity with increasing number of workers.

Unified model predicts multi-mode failure with multi-sensor data.

problem Independent failure mode and RUL prediction ignores inherent relationship.
method Hierarchical Bayesian framework with Cox model, Gaussian process, and multinomial distributions.
result Robust uncertainty quantification and accurate prediction of multi-mode failure.