Paper improves nuclear threat detection by better modeling background radiation.
problem Improving nuclear threat detection sensitivity and specificity.
method Applied Poisson Principal Component Analysis (PCA) to model background radiation.
result Poisson PCA outperforms standard Gaussian PCA in modeling background radiation.
New methods model gamma-ray data to better understand Galactic emissions.
problem Uncertain diffuse Galactic gamma-ray emissions bias data interpretation.
method Gaussian processes and variational inference for flexible modeling.
result More robust interpretation of gamma-ray sky, especially dark matter signals.
Deep learning improves gamma-ray energy estimation and event selection.
problem Improving gamma-ray event selection and energy estimation.
method Adapted convolutional neural networks (CNN) for gamma-ray astronomy.
result Significant improvement in gamma-ray energy estimation and event selection.
The paper confirms two groups of gamma-ray bursts using a new nonparametric metric.
problem Determining the number of inherent groups in gamma-ray bursts.
method A new nonparametric interpoint distance-based measure, combined with clustering methods.
result Confirms two groups of short and long gamma-ray bursts.
CNNs improve particle identification in ground-based gamma-ray astronomy.
problem Identifying particles in gamma-ray astronomy images.
method Used convolutional neural networks (CNNs) with PyTorch and TensorFlow.
result Improved accuracy in identifying gamma-rays and background particles.
VAE improves MSI data analysis for tissue sub-types.
problem Analyzing MSI data from unprocessed samples.
method Applied Variational Autoencoders for data reduction and pattern detection.
result VAEs outperform standard methods in detecting tissue sub-types.
Develops a hybrid MtFA approach for high-dimensional data clustering.
problem Scalability issues in traditional MtFA estimation methods for high-dimensional data.
method Integrates profile likelihood method into EM framework for efficient parameter estimation.
result Demonstrates superior computational efficiency and clustering accuracy compared to existing methods.
Deep Relevance Regularization improves neural network performance in tumor typing.
problem Confounding factors hinder neural network performance in multi-laboratory imaging mass spectrometry data.
method Introduces Deep Relevance Regularization to restrict neural network focus.
result Deep Relevance Regularization robustifies neural networks and improves interpretability.
Deep learning improves tumor classification in IMS data.
problem Automated feature extraction and classification for tumor classification in IMS data.
method Adapted deep convolutional network architecture and sensitivity analysis.
result Competitive performance and biologically interpretable results.
Neural network predicts electron-ionization mass spectra quickly.
problem Identifying unknown molecules not in existing libraries.
method Lightweight neural network model for predicting mass spectra.
result High accuracy predictions of small molecule mass spectra.
TiK-means extends K-means for skewed groups, revealing structured clusters.
problem Clustering skewed groups using traditional K-means.
method Introduces TiK-means, a modified K-means algorithm that estimates skewness-transformation parameters.
result Reveals structured clusters that explain the skewness of groups.
New algorithm clusters GRBs into two groups: short and long duration.
problem Determining the number of clusters in gamma-ray bursts.
method Completely parameter-free clustering algorithm.
result Clusters indicate two main groups: short and long duration GRBs.
DeepNovoV2 improves de novo peptide sequencing from mass spectrometry data.
problem De novo peptide sequencing from mass spectrometry data for personalized cancer vaccines.
method DeepNovoV2 combines T-Net and recurrent neural networks for end-to-end training and prediction.
result DeepNovoV2 achieves 13.01-23.95\% higher accuracy than previous methods.
Two classes of gamma-ray bursts (GRBs), short and long, have been determined without any doubts, and are usually ascribed to different progenitors, yet these classes overlap for a variety of descriptive parameters. A subsample of 46 long and 22 short Fermi GRBs with estimated Hurst Exponents (HEs), complemented by mi…
Background: High-throughput proteomics techniques, such as mass spectrometry (MS)-based approaches, produce very high-dimensional data-sets. In a clinical setting one is often interested in how mass spectra differ between patients of different classes, for example spectra from healthy patients vs. spectra from patients…
We study a simple modification to the conventional time of flight mass spectrometry (TOFMS) where a \emph{variable} and (pseudo)-\emph{random} pulsing rate is used which allows for traces from different pulses to overlap. This modification requires little alteration to the currently employed hardware. However, it requi…
Improved peptide identification from mass spectrometry data.
problem Lack of large ground truth datasets for protein identification.
method Deep neural networks trained on imperfect hand-coded models.
result 43% improvement over standard matching methods.
Developed HIquant to quantify proteoforms accurately without biases.
problem Peptide-specific biases in quantifying proteoforms by MS methods.
method First-principles model (HIquant) for quantifying proteoform stoichiometries.
result High accuracy in quantifying fractional PTM occupancy without external standards.
Improved protein identification in mass spectrometry data.
problem Expanding peptide scoring capabilities in tandem mass spectrometry.
method Deriving concave emission distributions for dynamic Bayesian networks.
result Efficiently learned scoring function outperforms state-of-the-art.
Neural networks predict substructures from mass spectra to identify chemical threats.
problem Identifying unknown chemical threats from mass spectra and formulas.
method Data-driven approach using neural networks to rank and match substructures.
result Substructure classifiers achieve over 90% micro F1-score and correctly identify structures in 88-71% of cases.
New framework distinguishes lung cancer subtypes using MALDI mass spectrometry.
problem Distinguishing between adenocarcinoma and squamous cell carcinoma subtypes in lung cancer.
method Supervised topological data analysis on MALDI mass spectrometry imaging data.
result The proposed framework successfully classifies lung cancer subtypes with competitive results.
Mass spectrometry (MS) is an important technique for chemical profiling which calculates for a sample a high dimensional histogram-like spectrum. A crucial step of MS data processing is the peak picking which selects peaks containing information about molecules with high concentrations which are of interest in an MS in…
A new algorithm ThreeSieves maximizes submodular functions efficiently in streaming data.
problem Maximizing submodular functions in high-dimensional data streams efficiently.
method ThreeSieves algorithm, designed for streaming data, ignoring worst-case guarantees for high probability solutions.
result ThreeSieves outperforms state-of-the-art methods in terms of performance and resource usage.
A flexible machine learning model infers the morphology of the Galactic Center Excess.
problem Inferring the unknown morphology of the Galactic Center Excess using Fermi gamma-ray data.
method Used a Gaussian process (GP) to model the Galactic Center Excess (GCE) as a flexible, non-parametric machine learning model.
result The best-fit GP contains morphological features not typically associated with traditional GCE studies, such as a localized bright source and a diagonal arm.
Efficiently clusters incomplete data without imputation or full EM, faster and more accurate.
problem Clustering partially recorded data efficiently.
method Model-based approach using multivariate t-distributions, considering only observed values.
result Approach is more accurate and computationally efficient than alternatives.
Syncytial clustering merges groups from standard algorithms to reveal complex data structures.
problem Challenges in finding clusters with irregular structures.
method Estimates nonparametric overlap between clusters and merges groups with high overlap.
result Always a top performer in identifying groups with regular and irregular structures.
New method recovers relative rates in spatial compositional data from IMS.
problem Challenges in analyzing spatial data from IMS due to competitive sampling.
method Hierarchical Variational Graph Fused Lasso using heavy-tailed graphical lasso prior and automatic differentiation variational inference.
result Our method outperforms state-of-the-practice point estimate methodologies in IMS and has superior posterior coverage.
The study uses DPGMM to analyze pulsar families in parameter space.
problem Classifying and understanding different types of pulsars.
method Unsupervised machine learning with Dirichlet process Gaussian mixture model (DPGMM).
result DPGMM provides insights into pulsar families and their relations.
Deep learning aligns GC-MS peaks for biomarker discovery.
problem Aligning retention times of GC-MS peaks across different samples.
method ChromAlignNet, a deep learning model for peak alignment.
result ChromAlignNet outperforms existing methods on complex data sets.
Microbial identification is a central issue in microbiology, in particular in the fields of infectious diseases diagnosis and industrial quality control. The concept of species is tightly linked to the concept of biological and clinical classification where the proximity between species is generally measured in terms o…
New framework selects features from noisy data with hidden variables.
problem Feature selection from noisy data with hidden variables.
method Developed a mathematical framework for feature selection from real-world data with non-linear observations.
result Successful variable selection possible even without knowing model parameters.
Paper presents a method for identifying isotope envelopes in MALDI-ToF data.
problem Deisotoping of isotopic peaks in MALDI-ToF molecular imaging data.
method Uses Mamdani-Assilan fuzzy system and spatial maps of molecular distribution to identify isotope envelopes.
result Proposed method detects overlapping envelopes and analyzes large data sets.
ForestDSH hashes improve nearest neighbor search in high-dimensional data.
problem High-dimensional classification and nearest neighbor search.
method Distribution-sensitive hashing using a forest of decision trees.
result ForestDSH hashes outperform LSH and state-of-the-art methods in speed and accuracy.
Develops a test for strict stationarity in stochastic processes.
problem Testing strict stationarity of discrete time stochastic processes.
method Window averaged sample estimate of second order cumulant spectrum, asymptotic complex standard normal distribution test.
result Test statistic derived and demonstrated with 137Cs gamma ray decay data.
Study proposes SVDD framework for classifying water saturation in imbalanced geological datasets.
problem Classification of petrophysical properties from imbalanced datasets with nonlinear and heterogeneous subsurface properties.
method Support Vector Data Description (SVDD) for one class classification of water saturation.
result Proposed SVDD framework outperforms other classifiers in terms of g metric means and execution time.
Clustering in high-dimensional spaces is nowadays a recurrent problem in many scientific domains but remains a difficult task from both the clustering accuracy and the result understanding points of view. This paper presents a discriminative latent mixture (DLM) model which fits the data in a latent orthonormal discrim…
Random small feature subsets outperform FS in diverse datasets.
problem The significance of selected features in high-dimensional datasets is questionable.
method Analysis of 28 diverse datasets (microarray, RNA-Seq, etc.).
result Any arbitrary set of features performs as well as or better than selected features across datasets.
New scalable algorithm for non-negative linear regression with entropy-regularized OT loss.
problem Generalizing task-specific linear models to broader applications.
method Sinkhorn-like scaling iterations for convex penalty and datafit terms.
result Simple multiplicative updates for various penalty and datafit terms.
SMM improves signal recovery from noisy data.
problem Estimating signals from noisy and scaled observations.
method Spiked mixture model (SMM) with EM algorithm.
result SMM outperforms GMM in signal recovery.
Generative model gradients enhance MS/MS peptide identification.
problem Improving peptide identification from MS/MS spectra.
method Leverage log-likelihood gradients of generative models in a kernel-based classifier.
result Fisher kernel outperforms other methods on MS/MS datasets.
Probabilistic PARAFAC2 improves robustness to noise.
problem Improving robustness of PARAFAC2 to noise and determining the number of factors.
method Developed two probabilistic formulations of PARAFAC2 with variational procedures for inference.
result Probabilistic PARAFAC2 is more robust to noise and model order misspecification.
Novel process model for metabolomics data analysis.
problem Analyzing complex metabolomics data.
method Data-driven and hypothesis-driven data mining approaches using various techniques.
result Demonstrated applicability and strengths of MeKDDaM model.
We present the first public release of our generic neural network training algorithm, called SkyNet. This efficient and robust machine learning tool is able to train large and deep feed-forward neural networks, including autoencoders, for use in a wide range of supervised and unsupervised learning applications, such as…
To draw inferences about gamma-ray burst (GRB) source populations based on Swift observations, it is essential to understand the detection efficiency of the Swift burst alert telescope (BAT). This study considers the problem of modeling the Swift/BAT triggering algorithm for long GRBs, a computationally expensive proce…
This study benchmarks anomaly detection methods in a functional setup.
problem Efficient anomaly detection in high-frequency sensor data.
method Functional analysis approach for multivariate data.
result Comparison of anomaly detection methods on real datasets.
Selective prediction framework reduces errors in molecular structure identification from MS/MS.
problem High-stakes applications require reliable molecular structure identification from MS/MS data.
method Selective prediction framework using risk-coverage tradeoff and uncertainty quantification.
result First-order confidence measures and retrieval-level aleatoric uncertainty achieve strong risk-coverage tradeoffs.
Hybrid system uses SVM, ANFIS, and expert knowledge for oil saturation prediction.
problem Predicting oil saturation from well logs in a noisy dataset.
method Two-stage DKFIS: SVM for classification, ANFIS for prediction, expert knowledge refinement.
result DKFIS improves prediction accuracy compared to ANFIS alone.
Nonparametric empirical Bayes denoising on Riemannian manifolds
problem Denoising measurements on compact Riemannian manifolds
method Using a surrogate oracle denoiser based on the marginal distribution of measurements
result Achieving nearly the Bayes risk in a low-noise regime