Machine learning predicts plant phenotypes from soil microbiome data.
problem Predicting plant phenotypes from soil microbiome data.
method Two models (random forest and Bayesian neural network) were used to predict plant phenotypes from soil properties and microbial population density.
result Human decisions and normalization strategies significantly impact model performance.
Data mining involves the systematic analysis of large data sets, and data mining in agricultural soil datasets is exciting and modern research area. The productive capacity of a soil depends on soil fertility. Achieving and maintaining appropriate levels of soil fertility, is of utmost importance if agricultural land i…
In precision agriculture (PA), soil sampling and testing operation is prior to planting any new crop. It is an expensive operation since there are many soil characteristics to take into account. This paper gives an overview of soil characteristics and their relationships with crop yield and soil profiling. We propose a…
Improves classification of microbiome data using mixture distributions.
problem Challenges in classifying sparse and heterogeneous microbiome count data.
method Distance-based classification using mixture distributions.
result The method outperforms existing distance-based classifiers and machine learning approaches.
Agricultural research has been profited by technical advances such as automation, data mining. Today, data mining is used in a vast areas and many off-the-shelf data mining system products and domain specific data mining application soft wares are available, but data mining in agricultural soil datasets is a relatively…
AI enhances microbiology and microbiome research through machine learning.
problem Understanding microbial life and its impact on health and the environment.
method AI-driven approaches including machine learning and deep learning.
result Transformative role in enhancing microbial life understanding.
We model microbiome interactions as graphs to interpret complex dynamics.
problem Understanding the differences in microbiome profiles between healthy and ill individuals.
method Developed a method to learn low-dimensional graph representations of time-evolving microbiome interactions.
result Extracted graph features that highlight microbes and interactions strongly correlated with clinical diseases.
Study uses machine learning to predict soil organic carbon content in northern Iran.
problem Estimating soil organic carbon content for understanding soil functions.
method Applied machine learning algorithms including DNN, SVM, ANN, etc., with genetic algorithm feature selection.
result DNN model showed highest accuracy with low prediction error and uncertainty.
Cameras are an essential part of sensor suite in autonomous driving. Surround-view cameras are directly exposed to external environment and are vulnerable to get soiled. Cameras have a much higher degradation in performance due to soiling compared to other sensors. Thus it is critical to accurately detect soiling on th…
Data augmentation improves microbiome disease prediction.
problem Improving predictive models for microbiome data.
method Defined novel data augmentation strategies for simplex-valued data.
result Set new state-of-the-art for disease prediction tasks.
GraphKKE learns fixed-length feature vectors from time-evolving graphs of human microbiome data.
problem Understanding dynamic changes in human microbiome graphs over time.
method Spectral analysis of transfer operators and graph kernels.
result GraphKKE captures temporal changes in human microbiome graphs.
Soil moisture is an important variable that determines floods, vegetation health, agriculture productivity, and land surface feedbacks to the atmosphere, etc. Accurately modeling soil moisture has important implications in both weather and climate models. The recently available satellite-based observations give us a un…
Unified framework combines dependent microbiome tests.
problem Combining dependent microbiome association tests.
method Generalized meta-analysis framework for dependent tests.
result The vanilla Cauchy combination is a special case.
In construction projects, estimation of the settlement of fine-grained soils is of critical importance, and yet is a challenging task. The coefficient of consolidation for the compression index (Cc) is a key parameter in modeling the settlement of fine-grained soil layers. However, the estimation of this parameter is c…
Improved prediction of soil parameters using Multi-target Stacked Generalisation on EDXRF spectra.
problem Challenges in predicting multiple soil parameters accurately from EDXRF spectra.
method Multi-target Stacked Generalisation (MTSG) method combining multiple regression models.
result MTSG significantly improved prediction accuracy for multiple soil parameters, reducing average error from 0.67 to 0.64.
In this paper, we investigate the potential of estimating the soil-moisture content based on VNIR hyperspectral data combined with LWIR data. Measurements from a multi-sensor field campaign represent the benchmark dataset which contains measured hyperspectral, LWIR, and soil-moisture data conducted on grassland site. W…
Develops methods for causal inference in compositional data using instrumental variables.
problem Interpreting summary statistics like diversity indices as causal effects in compositional data.
method Statistical data transformations and regression techniques tailored for compositional data.
result Advantages and limitations of the proposed methods demonstrated on synthetic and real microbiome data.
KernelBiome tackles microbiome research by improving predictive performance and interpretability.
problem Challenges in analyzing high-throughput sequencing data, especially in microbiome research.
method KernelBiome is a kernel-based nonparametric regression and classification framework for compositional data, incorporating prior knowledge and capturing complex signals.
result Improved predictive performance compared to state-of-the-art machine learning methods, with two novel quantities for interpretability.
Bayesian model fuses diverse microbiome data types.
problem Challenges in fusing different types of microbiome data.
method Flexible multinomial-Gaussian generative model with variational EM algorithm.
result Inferred latent variables provide common dimensionality reduction and predictive posterior distribution.
Bayesian method for imputing missing values in tensor data.
problem Missing data in multi-way arrays (tensors) in biomedical studies.
method Flexible Bayesian framework with conjugate priors for CP factorization.
result Accurately captures uncertainty in microbiome profiles at missing timepoints.
New model identifies microbial subcommunities robustly, accounting for cross-sample heterogeneity.
problem Inference in LDA is sensitive to the number of subcommunities and often creates artificial ones.
method Incorporates logistic-tree normal (LTN) model into LDA to account for cross-sample heterogeneity.
result Restores robustness of inference and identifies meaningful subcommunities.
PHIBP models complex microbiome data with shared parameters.
problem Complex, sparse count data in microbiome analysis.
method Bayesian nonparametric framework with shared species parameters.
result Flexible multivariate count model with tractable inference.
New models for analyzing microbiome data with interactions.
problem Analyzing compositional data with interactions.
method Exponential family models with generalized score matching.
result Effective estimation methods for compositional data with interactions.
New geometric approach for analyzing compositional data like gut microbiomes.
problem Analyzing non-negative compositional data with relative values only.
method Reinterpret compositional data as quotient topology of a sphere, using spherical harmonics and reflection group actions.
result Construction of Reproducing Kernel Hilbert Space (RKHS) for compositional data.
New method for analyzing compositional data, addressing biases in summary statistics.
problem Inadequate effect measures for compositional data, especially in high-dimensionality and sparsity.
method Perturbation-based effect measures, average perturbation effects.
result Proposed estimators efficiently estimate average perturbation effects, outperforming existing techniques.
Soil texture is important for many environmental processes. In this paper, we study the classification of soil texture based on hyperspectral data. We develop and implement three 1-dimensional (1D) convolutional neural networks (CNN): the LucasCNN, the LucasResNet which contains an identity block as residual network, a…
The Soil Moisture Active Passive (SMAP) mission has delivered valuable sensing of surface soil moisture since 2015. However, it has a short time span and irregular revisit schedule. Utilizing a state-of-the-art time-series deep learning neural network, Long Short-Term Memory (LSTM), we created a system that predicts SM…
A hybrid model combines machine learning with a land surface model to improve soil moisture predictions.
problem Improving soil moisture predictions in climatological situations.
method Noah land-surface model integrated with Gaussian Processes, using autoregressive model for out-of-sample results.
result 3-fold reduction in RMSE using one-year leave-one-out cross-validation.
Scientific investigations that incorporate next generation sequencing involve analyses of high-dimensional data where the need to organize, collate and interpret the outcomes are pressingly important. Currently, data can be collected at the microbiome level leading to the possibility of personalized medicine whereby tr…
Tree-based variational inference improves PLN model for hierarchical count data.
problem Limited applicability of PLN model in ecosystems due to lack of hierarchical tree structures.
method Introduced PLN-Tree model integrating structured variational inference techniques.
result Enhanced generative improvements and practical interpretability in microbiome modeling.
Develops a method for estimating networks and covariate associations in compositional data.
problem Estimating network interactions and covariate associations for compositional data.
method Hierarchical Bayesian model with spike-and-slab priors for edge and covariate selection, variational EM for inference.
result The proposed method outperforms existing methods in network recovery accuracy.
Geotechnics adopts data-driven methods from materials informatics.
problem Soil complexity and lack of comprehensive data.
method Leveraging deep learning and transfer learning for feature extraction.
result Revolutionary potential of advanced computational tools in geotechnics.
Generative model identifies temporal count data components with regime-dependent contributions.
problem Modeling temporal count data with regime-dependent dynamics.
method Generative framework combining regime-adaptive dynamics with Poisson log-normal emissions.
result Established identifiability of the model and revealed co-variation patterns and regime shifts.
SwiGAN generates drought scenarios for climate risk management.
problem Natural catastrophes and droughts increase insurance costs.
method Conditional GANs for generating spatio-temporal SWI maps.
result Simulates drought patterns up to 2050 for French regions.
Soil organic carbon (SOC) plays a major role in the global carbon budget. It can act as a source or a sink of atmospheric carbon, thereby possibly influencing the course of climate change. Improving the tools that model the spatial distributions of SOC stocks at national scales is a priority, both for monitoring change…
PolyILR: A Tree-Structured Orthonormal Decomposition of Compositional Data
problem Representing compositional data with hierarchical structure
method PolyILR: A canonical orthonormal decomposition of the Aitchison tangent space aligned with any tree topology
result PolyILR yields stable, interpretable features and enables inference at multiscale tree resolution
The performance of land surface models (LSMs) significantly affects the understanding of atmospheric and related processes. Many of the LSMs' soil and vegetation parameters were unknown so that it is crucially important to efficiently optimize them. Here I present a globally applicable and computationally efficient met…
Generative model for tabular data density regression.
problem Estimating conditional distribution of outcomes given covariates.
method Tree-based flow model for efficient sampling and likelihood evaluation.
result Our method achieves comparable or superior performance with reduced training and sampling costs.
This paper addresses measurement errors in high-dimensional compositional data using a log-contrast model calibration approach.
problem Measurement errors in high-dimensional regression models involving compositional covariates.
method Calibration approach for the linear log-contrast model under lenient sparsity conditions.
result Established asymptotic normality of the estimator for inference.
Study uses machine learning to identify IBD biomarkers from gut microbiota.
problem Identifying biomarkers for Inflammatory Bowel Disease (IBD) from gut microbiota.
method Ensemble feature selection methods (CMIM, FCBF, mRMR, XGBoost) applied to IBD-associated metagenomics dataset.
result XGBoost minimizes microbiota used for IBD diagnosis, improving classification accuracy.
Hybrid models combine domain knowledge and data-driven learning for Earth observation.
problem Challenges in modelling Earth observation data with either purely mechanistic or data-driven methods.
method Gaussian process convolution models, specifically latent force models (LFMs), integrating physical knowledge into multioutput GP models.
result Model automatically estimates soil moisture persistence and discovers latent forces related to precipitation.
Paper tackles multi-source domain adaptation for regression.
problem Predicting HDL cholesterol levels using gut microbiome data.
method Two-step procedure: 1) Extend a flexible single-source DA algorithm for classification to regression. 2) Augment with ensemble learning for multi-source DA.
result Consistent improvement in HDL cholesterol level prediction performance over existing methods.
Buried landmines and unexploded remnants of war are a constant threat for the population of many countries that have been hit by wars in the past years. The huge amount of human lives lost due to this phenomenon has been a strong motivation for the research community toward the development of safe and robust techniques…
New ZIPLN model accounts for zero-inflation in multivariate count data.
problem Zero-inflation in multivariate count data.
method Introduced Zero-Inflated PLN (ZIPLN) model with variational inference.
result ZIPLN significantly improves log-likelihood and reduces dispersion.
Compositional data have two unique characteristics compared to typical multivariate data: the observed values are nonnegative and their summand is exactly one. To reflect these characteristics, a specific regularized regression model with linear constraints is commonly used. However, linear constraints incur additional…
DL models can outperform regionalized models in hydrology by pooling diverse data.
problem Traditional wisdom in hydrology suggests regionalization improves model performance, but DL models can unify data for better performance.
method Used DL models on pooled data from different regions, showing improved performance compared to regionalized models.
result DL models can improve performance by pooling diverse data, highlighting the 'data synergy' effect.
We develop a cross-sectional research design to identify causal effects in the presence of unobservable heterogeneity without instruments. When units are dense in physical space, it may be sufficient to regress the "spatial first differences" (SFD) of the outcome on the treatment and omit all covariates. The identifyin…
GeoLifeCLEF 2020 dataset pairs species observations with environmental data.
problem Understanding geographic species distribution.
method Presented a dataset of species observations with environmental features.
result Advances in location-based species recommendation.