Mobile big data contains vast statistical features in various dimensions, including spatial, temporal, and the underlying social domain. Understanding and exploiting the features of mobile data from a social network perspective will be extremely beneficial to wireless networks, from planning, operation, and maintenance…
The paper explores machine learning in mobile big data analysis.
problem Challenges in mobile big data analysis.
method Discussion and review of existing methods.
result Identification of main challenges and future directions.
Predicts gender and age from mobile phone data for marketing.
problem Enhance marketing offers by predicting customer demographics.
method Machine learning algorithms applied to CDRs, CRM, and billing info.
result 85.6% accuracy in gender prediction, 65.5% in age prediction.
This research detects anomalies and predicts traffic using CDR data.
problem Detecting and predicting anomalies in mobile network traffic.
method Utilized CDR data, k-means clustering for anomaly detection, neural network for anomaly-free data, and ARIMA for traffic prediction.
result Anomaly-free data leads to better model generalization and prediction performance.
Currently, the world is witnessing a mounting avalanche of data due to the increasing number of mobile network subscribers, Internet websites, and online services. This trend is continuing to develop in a quick and diverse manner in the form of big data. Big data analytics can process large amounts of raw data and extr…
Deep-MAPS uses machine learning for mobile air pollution sensing in Beijing.
problem Ubiquitous sensing of urban air quality.
method Machine learning framework (Deep-MAPS) based on mobile and fixed sensors.
result Deep-MAPS achieves high spatial-temporal resolution (1km-by-1km and 1 hour) with over 85% accuracy.
Survey of big data in cyber-physical systems, including data security and green challenges.
problem Managing vast amounts of data in cyber-physical systems.
method Taxonomy and overview of data collection, storage, access, processing, and analysis.
result First panoramic survey on big data for CPS, addressing cybersecurity and green challenges.
Paper removes sensitive data from IoT and Big Data for privacy.
problem Privacy concerns in IoT and Big Data.
method Develops new supervised and adversarial learning methods to remove sensitive data.
result Models maintain predictive model utility while making sensitive predictions ineffective.
The emergence of mobile games has caused a paradigm shift in the video-game industry. Game developers now have at their disposal a plethora of information on their players, and thus can take advantage of reliable models that can accurately predict player behavior and scale to huge datasets. Churn prediction, a challeng…
Big data from phone calls improves credit scoring models and profits.
problem Improving credit scoring models to enhance financial inclusion.
method Combining call-detail records and traditional data to build scorecards using social network analytics.
result Combining call-detail records with traditional data significantly increases model performance and profit.
To accommodate heterogeneous tasks in Internet of Things (IoT), a new communication and computing paradigm termed mobile edge computing emerges that extends computing services from the cloud to edge, but at the same time exposes new challenges on security. The present paper studies online security-aware edge computing …
Survey of mobility studies using mobile phone data.
problem Understanding human mobility patterns.
method Data Science techniques applied to mobile phone datasets.
result Applications in urban planning, data traffic prediction, etc.
GeneCAI optimizes DNN compression hyper-parameters for mobile devices.
problem Efficient deployment of complex DNNs on resource-limited devices.
method GeneCAI uses genetic algorithm to learn optimal hyper-parameters.
result GeneCAI finds models with better accuracy-complexity trade-off.
New algorithm clusters streaming data efficiently.
problem Challenges in clustering streaming data.
method Online clustering algorithm for unknown number of clusters.
result Produces partitions close to full data clustering.
The paper presents a probabilistic method to discover daily human mobility patterns from mobile data.
problem Discovering daily human mobility patterns from mobile data.
method A non-parameter Bayesian modeling method, Infinite Gaussian Mixture Model, combined with Kullback-Leibler divergence for automatic clustering.
result The IGMM-based algorithm outperforms the GMM-based algorithm in discovering mobility patterns.
FLAME auto-labels mobile data efficiently on diverse processors.
problem Accurately and efficiently labeling mobile data with unknown labels on heterogeneous processors.
method Self-adaptive auto-labeling system Flame that schedules and executes workloads on mobile processors.
result Flame achieves high labeling accuracy and performance on heterogeneous mobile processors.
Model predicts urban population using mobile data traffic.
problem Estimating urban population dynamics from mobile data.
method Data-driven approach combining NetMob 2023 and ENACT datasets.
result NetMob 2023 data can estimate urban population with XGBoost models.
Graphs model human mobility patterns, reducing errors in data matching.
problem Lack of high-quality data and computational resources for graph-based mobility analysis.
method Embedding graphs into a continuous space to address matching, modeling, and visualization challenges.
result Approx 40% decrease in error on average in matched graphs vs unmatched ones.
Gait patterns reveal emotions, offering a non-invasive method for automated recognition.
problem Automated emotion recognition from gait patterns.
method Data collection, preprocessing, and classification techniques.
result Gait patterns can indicate different emotion states, making them a promising source for emotion detection.
Mobile sensing is an emerging technology that utilizes agent-participatory data for decision making or state estimation, including multimedia applications. This article investigates the structure of mobile sensing schemes and introduces crowdsourcing methods for mobile sensing. Inspired by social network, one can estab…
Extracts patterns from mobile network data for better resource management.
problem Improving network efficiency and resource allocation for mobile users.
method Spatiotemporal analysis of internet activity records (IARs) data.
result Developed a mobile traffic partitioning scheme.
Method discovers user habits from mobile data.
problem Understanding human mobility patterns and habits.
method Density-based clustering for spatio-temporal data and Gaussian Mixture Model (GMM).
result Many unique habits were identified from the datasets.
TraLFM models human mobility patterns from traffic trajectories.
problem Understanding human mobility patterns from traffic data.
method Latent factor modeling of sequential, personal, and temporal factors.
result TraLFM significantly outperforms state-of-the-art methods in latent factor analysis and next location prediction.
Understanding the spatiotemporal distribution of people within a city is crucial to many planning applications. Obtaining data to create required knowledge, currently involves costly survey methods. At the same time ubiquitous mobile sensors from personal GPS devices to mobile phones are collecting massive amounts of d…
Paper infers human mobility from sparse trajectories.
problem Modeling and inferring human mobility from sparse trajectory data.
method Proposes a single trajectory inference algorithm and a deep learning architecture for multiple trajectories.
result Deep learning model achieves 2x overall accuracy improvement on sparse trajectories.
Study predicts traffic congestion based on population mobility data.
problem Predicting traffic congestion in multimodal transport networks.
method Machine learning methods applied to population mobility data.
result Likely prediction of congestion based on population movements.
Federated Learning improves mobile data privacy by training classifiers without sharing raw data.
problem Privacy concerns in packet classification due to sensitive data sharing.
method Apply Federated Learning to mobile packet classification tasks, training models without raw data sharing.
result Demonstrated effectiveness of the approach in terms of performance, cost, and privacy.
SplitEasy trains ML models on mobile devices without server data transfer.
problem Training complex DL models on resource-limited mobile devices.
method Split learning approach where sensitive layers are trained locally, computationally intensive layers on server.
result SplitEasy trains models on mobile devices with minimal data transfer, near-constant time per sample.
Study uses mobile phone data to map Chagas disease risk zones.
problem Identifying geographical spread of Chagas disease.
method Analyzing geolocalized call records and public health information.
result Generated risk maps for public health campaigns.
Paper develops fine-grain spatiotemporal risk scores using high-resolution mobility data.
problem Developing reliable spatiotemporal risk scores for safe economic reopening.
method Hawkes process-based technique leveraging high-resolution cell-phone location signals.
result Fine-grain spatiotemporal risk scores based on high-resolution mobility data provide useful insights for safe re-opening.
We present and test a sequential learning algorithm for the short-term prediction of human mobility. This novel approach pairs the Exponential Weights forecaster with a very large ensemble of experts. The experts are individual sequence prediction algorithms constructed from the mobility traces of 10 million roaming mo…
The study introduces measures of collective mobility from aggregated OD data.
problem Understanding large-scale mobility patterns from aggregated data.
method Developed a framework using synthetic and real data to interpret network-level mobility.
result Aggregated mobility measures reveal network structure and flow constraints.
Paper presents a robust model to improve prediction accuracy for real-life mobile phone data.
problem Noisy instances in real-life mobile phone data affect model accuracy.
method Identify and eliminate noisy instances using naive Bayes and Laplace estimators, then build a decision tree model.
result The robust model improves prediction accuracy as shown by experimental results.
ARDEN improves deep learning performance on mobile devices by protecting privacy in the cloud.
problem Balancing privacy and performance in mobile deep learning with limited device capacity.
method ARDEN partitions DNN across mobile devices and cloud, using data transformation and noise addition for privacy, and noisy training for robustness.
result ARDEN enhances inference performance on cloud while maintaining strong privacy.
This paper develops synthetic mobility datasets to protect privacy while maintaining realism.
problem Privacy concerns restrict sharing real mobility datasets, leading to lack of reproducibility.
method The paper benchmarks RNNs, GANs, and copulas to generate realistic synthetic trajectories.
result The approach generates trajectories that are statistically and semantically similar to real-world data.
This paper tackles computational bottlenecks in federated learning on mobile devices.
problem Computationally heterogeneous mobile devices hinder federated learning efficiency.
method Proposes efficient algorithms to schedule mobile devices based on data heterogeneity.
result Achieves up to 100x speedup and 7% accuracy gain in federated learning.
MDLdroid improves mobile deep learning for personal sensing with faster training.
problem Continuous local changes and resource constraints in personal mobile sensing affect global model performance.
method ChainSGD-reduce approach to reduce overhead and balance resources.
result 2x to 3.5x faster training on off-the-shelf mobile devices compared to single-device training.
The tremendous growth of positioning technologies and GPS enabled devices has produced huge volumes of tracking data during the recent years. This source of information constitutes a rich input for data analytics processes, either offline (e.g. cluster analysis, hot motion discovery) or online (e.g. short-term forecast…
Machine learning methods are used to discover complex nonlinear relationships in biological and medical data. However, sophisticated learning models are computationally unfeasible for data with millions of features. Here we introduce the first feature selection method for nonlinear learning problems that can scale up t…
New models predict mobility flows as well as complex machine learning but are simpler and interpretable.
problem Incomplete understanding and modeling of human mobility flows.
method Developed simple machine-learned, closed-form models of mobility.
result These models predict mobility flows more accurately than gravity or complex machine/deep learning models.
Hidden Markov Models analyze mobile health data to identify APNS states.
problem Subjective self-report measures of APNS lead to errors and biases.
method Exploratory hidden Markov factor models and Stabilized Expectation-Maximization algorithm.
result Identified homogeneous APNS states and dynamic transitions.
UrbanRhythm reveals urban dynamics from mobility data.
problem Understanding changing urban activities over time.
method Extracting staying, leaving, arriving attributes; using Saak transform; clustering for city states; motif analysis for short-term regularity.
result Characterized urban dynamics as city state transformations over time.
Paper presents a new time-series segmentation technique for mobile phone user behavior.
problem Current segmentation techniques do not accurately capture individual user behavior over time.
method Behavior-Oriented Time Segmentation (BOTS) technique that considers temporal coverage and number of incidences.
result BOTS technique better captures user behavior at various times of day and week.
Paper proposes a new classifier for gender detection in mobile telematics.
problem Detecting gender through mobile telematics data.
method Choquet fuzzy integral vertical bagging classifier combining random forest and rough set theory.
result Choquet fuzzy integral vertical bagging classifier outperforms other classifiers.
Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…
In this study, we present a machine learning approach to infer the worker and student mobility flows on daily basis from static censuses. The rapid urbanization has made the estimation of the human mobility flows a critical task for transportation and urban planners. The primary objective of this paper is to complete i…
Paper introduces IGMM-GAN for multimodal anomaly detection in mobility data.
problem Lack of ground truth data and dependence on pre-processing for anomaly detection in human mobility.
method Coupled IGMM-GAN for generating realistic synthetic datasets and multimodal anomaly detection.
result IGMM-GAN improves anomaly detection performance over existing GAN methods.
Study of urban lifestyles from mobility data of 1.2M people in 11 U.S. cities.
problem Lack of interpretability in digital mobility data for understanding urban lifestyles.
method Privacy-enhanced dataset of mobility visitation patterns, latent activity behavior decomposition.
result Detected 12 latent activity behaviors that describe urban lifestyles, not single lifestyles.