Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

2795588361,115 · Jun 202019922001200920172026
48 results for NASA data

The paper classifies U.S. crop types using hyperspectral satellite imagery.

problem Classifying crop types from hyperspectral satellite imagery.
method Gaussian Bayesian models and neural networks applied to NASA data.
result Bayesian methods outperform standard LDA and QDA.

Procedural terrain generation for video games has been traditionally been done with smartly designed but handcrafted algorithms that generate heightmaps. We propose a first step toward the learning and synthesis of these using recent advances in deep generative modelling with openly available satellite imagery from NAS…

2017-07-11abs ↗pdf ↗

As the Industrial Internet of Things (IIoT) grows, systems are increasingly being monitored by arrays of sensors returning time-series data at ever-increasing 'volume, velocity and variety' (i.e. Industrial Big Data). An obvious use for these data is real-time systems condition monitoring and prognostic time to failure…

2018-04-09abs ↗pdf ↗

Reanalysis datasets combining numerical physics models and limited observations to generate a synthesised estimate of variables in an Earth system, are prone to biases against ground truth. Biases identified with the NASA Modern-Era Retrospective Analysis for Research and Applications, Version 2 (MERRA-2) aerosol optic…

2019-10-14abs ↗pdf ↗

Entropy-based GP adaptive design improves failure probability estimation.

problem Limited accuracy in failure probability estimation due to model evaluation costs.
method Entropy-based Gaussian process (GP) adaptive design combined with multifidelity importance sampling (MFIS).
result More accurate failure probability estimates and higher confidence.

Remaining Useful Life (RUL) of an equipment or one of its components is defined as the time left until the equipment or component reaches its end of useful life. Accurate RUL estimation is exceptionally beneficial to Predictive Maintenance, and Prognostics and Health Management (PHM). Data driven approaches which lever…

2019-04-12abs ↗pdf ↗

Deep learning improves oilfield equipment maintenance and reduces downtime.

problem Predicting equipment failure in oilrigs to minimize downtime.
method Developed and tested neural networks on oilfield datasets, using data processing and feature extraction.
result Deep learning can predict oilfield equipment failure with reduced downtime.

This paper improves federated learning for industrial predictive analytics by accommodating client heterogeneity.

problem Traditional federated models assume homogeneity in degradation processes, which doesn't apply to industrial settings.
method Personalized federated prognostic model using proximal gradient descent algorithm for joint parameter estimation.
result The proposed model enhances performance and provides comprehensive failure time distributions.

The study finds the best elliptical trajectory for planets using a variation of the hodograph theorem.

problem Finding the best elliptical trajectory for planets.
method Using a variation of the circular hodograph theorem, the study finds the best fitting ellipse for planetary trajectories by minimizing the sum of square distances from the points to the plane.
result The study finds that the best fitting ellipse for planetary trajectories minimizes the sum of square distances from the points to the plane.

Proposes a federated learning approach for industrial asset failure prediction.

problem Lack of data and privacy concerns in industrial prognostics.
method Two-stage federated learning: dimension reduction and parameter estimation.
result Validated the approach using simulated and real data.

Framework predicts remaining useful life of DSH subsystems under unknown failure modes.

problem Predicting remaining useful life of DSH subsystems with unknown failure modes.
method Unsupervised framework using mixture of Gaussian regressions and Expectation-Maximization algorithm.
result Improved prediction accuracy and interpretability of RUL.

Recent anomaly detection benchmarks are flawed, potentially misleading progress.

problem Flawed benchmark datasets create misleading progress reports.
method Identified four flaws in benchmark datasets and introduced a new archive.
result Published comparisons may be unreliable due to flaws in benchmark datasets.

TadGAN detects anomalies in time series data using GANs and LSTM.

problem Challenges in detecting anomalies in time series data, especially without labeled data.
method TadGAN uses Generative Adversarial Networks (GANs) with LSTM Recurrent Neural Networks to capture temporal correlations and compute anomaly scores.
result TadGAN outperforms 8 baseline methods in most cases, achieving the highest averaged F1 score.

New method disentangles sources of different timescales in planetary seismic data.

problem Unsupervised source separation of multi-scale seismic data from planetary missions.
method Wavelet scattering spectra for multi-scale clustering and variational autoencoder for source separation.
result Disentangles sources with different timescales in InSight mission seismic data.

Paper compares neural networks and time-series models for weather derivative pricing.

problem Pricing accuracy and regime adaptation for temperature and precipitation weather derivatives.
method Benchmarked harmonic-regression/ARMA vs. feed-forward neural network for temperature. Used CNN for precipitation, adapting to seasonal heterogeneity.
result CNN yields more accurate pricing, especially for regime-adapted seasonal data.

Generative model downgrades coarse satellite images to fine resolution.

problem Reconstructing fine resolution satellite images from coarse scale inputs.
method Combines U-Net transfer encoder with diffusion-based generative model.
result Excellent performance (R2 = 0.65 to 0.94) across seasonal regional splits.

Study shows visual feedback and monetary incentives reduce plugload energy consumption in commercial buildings.

problem Mitigating energy consumption in commercial buildings through occupant plugload control.
method Field experiments with visual feedback and monetary incentives in government and university buildings.
result Mean energy reduction of ~9.52% in office environments and ~21.61% in university environments with visual feedback.

EGFs use ergodicity to simplify generative flows for easier training and imitation learning.

problem Challenges in training generative flows, especially in continuous settings and for imitation learning.
method EGFs leverage ergodicity to build simple flows with universality guarantees and tractable FM loss. They introduce a KL-weakFM loss for IL training without a separate reward model.
result EGFs simplify generative flow training and enable effective imitation learning.

AER combines auto-encoder and LSTM for better time series anomaly detection.

problem Anomaly detection in time series data with limited labeled data and ambiguous definitions.
method AER (Auto-Encoder with Regression) integrates auto-encoder and LSTM for joint predictions and reconstructions.
result AER achieves the highest F1 score across 12 datasets with comparable runtime.

New CH covariance class improves spatial statistics by balancing differentiability and tail behavior.

problem Lack of control over mean-square differentiability and tail behavior in Matérn covariance functions.
method Developed a new Confluent Hypergeometric (CH) covariance class using a scale mixture of Matérn and polynomial covariances.
result The CH class offers improved theoretical properties and better performance in extrapolative settings.

Paper tackles cyber threats to PHM systems using adversarial examples.

problem Vulnerability of IoT sensors and DL algorithms to cyber attacks.
method Adopted adversarial example crafting techniques from computer vision to PHM domain.
result PHM models are vulnerable to adversarial attacks, leading to inaccurate remaining useful life estimation.

Accelerates pulsar light curve inference with learned representations and optimization.

problem Computational expense of Markov chain Monte Carlo methods for posterior inference.
method Combining U-Net latent representations with local simulator-guided optimization.
result 120x reduction in inference time (24 hours to 12 minutes) with accuracy preserved.

New method speeds up galaxy analysis from hours to seconds.

problem Infeasibility of state-of-the-art SED analyses for large surveys.
method Amortized Neural Posterior Estimation (ANPE) for scalable Bayesian inference.
result Posterior distributions of 12 model parameters estimated in seconds per galaxy.

This study evaluates uncertainty quantification methods for deep learning in predictive maintenance.

problem Uncertainty quantification for reliable decision-making in predictive maintenance.
method State-of-the-art variational inference algorithms for Bayesian neural networks (BNN), Monte Carlo Dropout (MCD), deep ensembles (DE), and heteroscedastic neural networks (HNN) were tested.
result No method clearly outperforms others in all situations, but DE and MCD provide more conservative uncertainty estimates.

Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…

2017-12-12abs ↗pdf ↗

Prevents sensitive data generation in diffusion models using labeled and unlabeled data.

problem Generating sensitive data in diffusion models using unlabeled data.
method Positive-Unlabeled Diffusion Models, approximating ELBO with labeled and unlabeled data.
result Prevents the generation of sensitive data without compromising image quality.

Study reveals Data Shapley's inconsistent performance in data selection tasks.

problem Inconsistency of Data Shapley's performance in data selection across different settings.
method Hypothesis testing framework and identification of utility functions.
result Data Shapley's performance is no better than random selection without specific constraints.

PRRO generates synthetic tabular data that improves SL performance and class distribution.

problem Low SL utility of synthetic data due to class imbalance and overlooked data relationships.
method Data pruning and column reordering to optimize SL utility.
result Synthetic data generated with PRRO enhances predictive performance and class distribution.

Defines data science as a natural ecosystem with challenges and missions.

problem Challenges and missions in data science due to 5D complexities and data life cycle phases.
method Systemic and data-centric view of data science as a fusion of data universe and its challenges, formalizing a general-purpose architecture.
result Essential data science as a natural ecosystem integrating specific disciplines and high-impact applications.

Proposes using probabilistic models for privacy-preserving synthetic data.

problem Designing high-quality synthetic data for privacy preservation.
method Formulate the problem through probabilistic modelling, choosing a model for the data.
result Statistical discoveries can be reliably reproduced from synthetic data.

Paper creates fair synthetic data ensuring equal predictions across sensitive attributes.

problem Ensuring fair predictions across sensitive attributes in synthetic data.
method Equalizing target probability distributions across sensitive attributes in synthetic data generation.
result Synthetic data provides strong fair predictions, equal across all thresholds.