Recently, air pollution is one of the most concerns for big cities. Predicting air quality for any regions and at any time is a critical requirement of urban citizens. However, air pollution prediction for the whole city is a challenging problem. The reason is, there are many spatiotemporal factors affecting air pollut…
This paper analyzes air pollution trends in Rwanda using low-cost sensors and machine learning.
problem Lack of reliable air pollution data in Rwanda due to high costs of equipment.
method Analysis of existing data and development of forecasting models using low-cost sensors and machine learning.
result Proposes forecasting models for air pollution data collected by low-cost sensors.
DGPs improve air quality inference from sparse data.
problem Accurate air quality monitoring in unmonitored areas.
method Deep Gaussian Processes with Doubly Stochastic Variational Inference.
result DGPs outperform state-of-the-art models in AQ inference.
High levels of air pollution may seriously affect people's living environment and even endanger their lives. In order to reduce air pollution concentrations, and warn the public before the occurrence of hazardous air pollutants, it is urgent to design an accurate and reliable air pollutant forecasting model. However, m…
Researchers develop a new method to assess variable importance in spatial machine learning models for air pollution exposure prediction.
problem Understanding the mechanism captured by machine learning models in air pollution studies, especially with spatial correlation.
method Leave-one-out approach for variable importance measure applicable to models with separable mean and covariance components.
result The new method highlights differences in model mechanisms even for similar prediction accuracies.
GraphSVR forecasts urban air pollution robustly across stations and seasons.
problem Nonlinear, nonstationary, spatiotemporally dependent urban air pollution forecasting challenges.
method Combines graph convolutional learning and support vector regression.
result GraphSVR improves predictive accuracy and maintains stable performance across seasons and outlier-prone episodes.
Study shows reducing anthropogenic emissions significantly lowers PM2.5 levels but has little effect on O3 in Delhi.
problem Understanding and mitigating the effects of anthropogenic emissions on air pollution in Delhi.
method Predictive modeling, causal inference, Gaussian Process modeling, Granger causality analysis.
result Reductions in anthropogenic emissions lead to significant decreases in PM2.5 levels but have little effect on O3. Paper uses Gaussian Processes to monitor air quality in Kampala.
problem Monitoring air pollution in cities with limited sensor coverage.
method Gaussian Processes for nowcasting and forecasting air pollution.
result Demonstrates the effectiveness of Gaussian Processes in air quality monitoring.
Weather2vec learns representations to adjust for non-local confounding in air pollution studies.
problem Non-local confounding in evaluating environmental policies and climate events on health outcomes.
method weather2vec framework using balancing scores to learn representations of non-local information.
result The framework effectively adjusts for confounding in air pollution studies.
China's rapid economic growth resulted in serious air pollution, which caused substantial losses to economic development and residents' health. In particular, the road transport sector has been blamed to be one of the major emitters. During the past decades, fluctuation in the international oil prices has imposed signi…
Air quality is closely related to public health. Health issues such as cardiovascular diseases and respiratory diseases, may have connection with long exposure to highly polluted environment. Therefore, accurate air quality forecasts are extremely important to those who are vulnerable. To estimate the variation of seve…
Tackling air pollution is an imperative problem in South Korea, especially in urban areas, over the last few years. More specially, South Korea has joined the ranks of the world's most polluted countries alongside with other Asian capitals, such as Beijing or Delhi. Much research is being conducted in environmental sci…
The dynamic nature of air quality chemistry and transport makes it difficult to identify the mixture of air pollutants for a region. In this study of air quality in the Houston metropolitan area we apply dynamic principal component analysis (DPCA) to a normalized multivariate time series of daily concentration measurem…
Mobile and ubiquitous sensing of urban air quality has received increased attention as an economically and operationally viable means to survey atmospheric environment with high spatial-temporal resolution. This paper proposes a machine learning based mobile air pollution sensing framework, called Deep-MAPS, and demons…
Inferring air quality from a limited number of observations is an essential task for monitoring and controlling air pollution. Existing inference methods typically use low spatial resolution data collected by fixed monitoring stations and infer the concentration of air pollutants using additional types of data, e.g., m…
Artificial Neural Network predicts PM2.5 pollution with low-cost sensors.
problem Costly and bulky PM2.5 monitoring instruments limit real-time, high-resolution data.
method Analytical equations derived using Artificial Neural Network (ANN).
result RMSE of 1.7973 ug/m3 and R2 of 0.9986 for eight predictors; 7.5372 ug/m3 and 0.9708 for three predictors.
The study corrects measurement error in evaluating health effects of multiple pollutants.
problem Bias in estimating health effects of air pollution constituents due to mismeasurement.
method Used a linear regression calibration model and extended DML approach to correct for measurement error.
result Identified two PM2.5 constituents (Br and Mn) that show a negative causal effect on cognitive function after correction.
Interpretable additive models outperform complex DL and hybrid pipelines for air quality forecasting.
problem Accurate forecasting of urban air pollution for public health and policy guidance.
method Investigated lightweight additive models (FBP, NP) vs. deep learning and hybrid pipelines on Beijing PM2.5 and PM10 data.
result Facebook Prophet consistently outperformed NeuralProphet and traditional models, achieving high R2 values. Air quality forecasting has been regarded as the key problem of air pollution early warning and control management. In this paper, we propose a novel deep learning model for air quality (mainly PM2.5) forecasting, which learns the spatial-temporal correlation features and interdependence of multivariate air quality rel…
Fine particulate matter (PM2.5) is one of the criteria air pollutants regulated by the Environmental Protection Agency in the United States. There is strong evidence that ambient exposure to (PM2.5) increases risk of mortality and hospitalization. Large scale epidemiological studies on the health effects of P…
The significance of air pollution and the problems associated with it are fueling deployments of air quality monitoring stations worldwide. The most common approach for air quality monitoring is to rely on environmental monitoring stations, which unfortunately are very expensive both to acquire and to maintain. Hence e…
Study improves PM concentration forecasting using MCCR loss.
problem Forecasting particulate matter concentration in South Korea.
method Used MCCR loss for regression analysis of air pollution and weather data.
result MCCR loss is more effective for extreme value forecasting.
MapLUR uses deep learning on map images to estimate NO2 pollution, outperforming traditional methods.
problem Limited availability of data for traditional LUR models makes them hard to adapt to new areas.
method Data-driven, open-source approach using convolutional neural networks trained on map data.
result MapLUR significantly outperforms traditional LUR models, including those with manually engineered features.
Predicts local AQI using mobile sensor data, improving accuracy by 71.654 MSE.
problem Inaccurate AQI data from sparse sensors in developing countries.
method Spatio-temporal GNNs for fine-grained AQI forecasting.
result Significant improvement in AQI prediction accuracy (71.654 MSE reduction).
Modeling air pollutants using data-driven techniques and sparse identification of nonlinear dynamics.
problem Predicting concentrations of air pollutants using hidden physical laws.
method Sparse identification of nonlinear dynamics (SINDy) for parsimonious systems of ordinary differential equations.
result More than half of the critical points are saddle points, indicating system instability.
Data collection in economically constrained countries often necessitates using approximate and biased measurements due to the low-cost of the sensors used. This leads to potentially invalid predictions and poor policies or decision making. This is especially an issue if methods from resource-rich regions are applied wi…
DCK improves air quality index prediction with probabilistic spatial models.
problem Non-Gaussian, complex spatial structure of air quality index.
method Deep classifier kriging (DCK) for non-Gaussian, nonlinear spatial prediction.
result DCK outperforms conventional methods in predictive accuracy and uncertainty quantification.
AirRL uses RL to infer urban air quality from selected stations.
problem Inferring fine-grained urban air quality from limited monitoring stations.
method Reinforcement learning model with a dynamic station selector and air quality regressor.
result AirRL achieves highest performance in air quality inference experiments.
RESPIRE calibrates low-cost air-quality sensors for CO levels, resistant to outliers.
problem Calibrating LCAQ sensors against regulatory-grade monitors is expensive and time-consuming.
method PROvably outlier-resistant semi-parametric regression technique.
result RESPIRE offers improved prediction in cross-site, cross-season, and cross-sensor settings.
Across numerous applications, forecasting relies on numerical solvers for partial differential equations (PDEs). Although the use of deep-learning techniques has been proposed, actual applications have been restricted by the fact the training data are obtained using traditional PDE solvers. Thereby, the uses of deep-le…
Quantile gradient boosted trees outperform other models in predicting NO2 concentration distributions.
problem Forecasting high NO2 concentration episodes for effective air quality management.
method Compared 10 probabilistic forecasting models for NO2 concentration prediction.
result Quantile gradient boosted trees model outperformed others in predicting NO2 concentration distributions.
Engine predicts real-time air quality with high resolution.
problem Real-time prediction of air pollutants for health monitoring.
method Combines official data, models, land cover, traffic data for high-resolution predictions.
result Engine produces predictions with resolution of a few dozen meters.
Engine forecasts NO2, O3, PM2.5, PM10 with high accuracy.
problem Accurate long-term air quality forecasting.
method Convolutional LSTM network trained on grid data.
result 4-day forecasts significantly outperform simple benchmarks.
Double machine learning improves causal effect estimation by relaxing assumptions.
problem Estimating causal effects with observational data.
method Double/debiased machine learning (DML) framework.
result DML improves adjustment for nonlinear confounding relationships.
Work addresses long-term accuracy issues in IoT air quality sensors.
problem Limited accuracy of IoT air quality sensors in long-term field deployments.
method Adaptive machine learning strategies for network calibration.
result Prolongs the validity of multisensor calibration models for continuous learning.
CRE discovers interpretable subgroups with heterogeneous treatment effects.
problem Identifying subgroups with notable treatment effect heterogeneity.
method Causal Rule Ensemble (CRE) using an ensemble-of-trees approach.
result CRE offers interpretable decision rules and high stability in subgroup discovery.
This paper establishes that so-called instrumental variables enable the identification and the estimation of a fully nonparametric regression model with Berkson-type measurement error in the regressors. An estimator is proposed and proven to be consistent. Its practical performance and feasibility are investigated via …
Framework detects shape shifts in functional profiles using Fréchet mean and shape invariant model.
problem Detecting shape shifts in functional profiles.
method Combining Fréchet mean and shape invariant model for interpretable parameterization of profile deviations.
result Potential shifts in shape deformation process distinguished by significant shifts in amplitude and/or phase.
Study compares geostatistical and machine learning models for PM2.5 prediction.
problem Improving accuracy of hourly PM2.5 maps across California.
method Traditional geostatistical methods (kriging, land use regression) and machine learning models (neural networks, random forests, support vector machines) were evaluated.
result Ensemble model enhanced predictive accuracy of PM2.5 concentration by correcting PurpleAir data bias.
We consider evidence integration from potentially dependent observation processes under varying spatio-temporal sampling resolutions and noise levels. We develop a multi-resolution multi-task (MRGP) framework while allowing for both inter-task and intra-task multi-resolution and multi-fidelity. We develop shallow Gauss…
Study prenatal PM2.5 exposure and 4th grade reading scores, identifying critical windows of susceptibility.
problem Understanding the impact of prenatal PM2.5 exposure on educational outcomes.
method Developed a locally adaptive Bayesian regression model with B-spline basis expansion and dynamic shrinkage priors.
result Prenatal PM2.5 exposure during early and late pregnancy is most adverse for 4th grade reading scores.
A new method for separating mixed signals in space and time.
problem Nonlinear and nonstationary spatio-temporal data challenges.
method Identifiable autoregressive variational autoencoder.
result The method outperforms existing techniques in blind source separation and spatio-temporal prediction.
Ensemble learning is a standard approach to building machine learning systems that capture complex phenomena in real-world data. An important aspect of these systems is the complete and valid quantification of model uncertainty. We introduce a Bayesian nonparametric ensemble (BNE) approach that augments an existing ens…
New algorithm balances spatial data approximation and prediction accuracy.
problem Lack of methods considering spatial correlation and downstream modeling in dimension reduction.
method Formalizes approximation and modeling utility as metrics, proposes a balanced algorithm.
result Optimal trade-off between approximation accuracy and downstream modeling utility.
Aggregated data is commonplace in areas such as epidemiology and demography. For example, census data for a population is usually given as averages defined over time periods or spatial resolutions (cities, regions or countries). In this paper, we present a novel multi-task learning model based on Gaussian processes for…
A new method learns priors for Bayesian optimisation to improve performance.
problem Bayesian optimisation tasks often assume strong similarity, which is violated in many cases.
method Replace strong similarity assumption with shape similarity, learn priors for hyperparameters.
result PLeBO and prior transfer find good inputs in fewer evaluations.
This study examines whether PCA can effectively identify nitrogen pollution sources in rivers.
problem Identifying pollution sources in rivers for effective environmental management.
method Principal Component Analysis and its modifications, along with Independent Component Analysis and Factor Analysis, are applied to nitrogen pollution source identification.
result PCA and related techniques can be powerful tools for uncovering nitrogen pollution sources in rivers.
Online algorithms for identifying river pollution sources.
problem Real-time estimation of river pollution sources from downstream data.
method Gradient-based online learning algorithms with adaptive step sizes and escaping from saddle points module.
result High estimation accuracy in three dimensions, superior to existing methods.