Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

2795588361,115 · Jun 202019922001200920172026
48 results for reanalysis data

New method reconstructs past foehn occurrences using unsupervised and supervised learning.

problem Reconstructing past foehn occurrences due to lack of direct measurement.
method Combining unsupervised and supervised learning methods to infer foehn occurrences from reanalysis data.
result Accurate hourly reconstructions of past foehn occurrences for 83 years.

Deep learning model predicts European weather parameters.

problem Ensemble weather prediction using deep learning.
method Conditional deep convolutional generative adversarial network (GAN) and Monte-Carlo dropout.
result Forecast skill for geopotential height and two-meter temperature is good, but precipitation is challenging.

Reanalysis datasets combining numerical physics models and limited observations to generate a synthesised estimate of variables in an Earth system, are prone to biases against ground truth. Biases identified with the NASA Modern-Era Retrospective Analysis for Research and Applications, Version 2 (MERRA-2) aerosol optic…

2019-10-14abs ↗pdf ↗

Machine learning predicts Atlantic blocking using limited data.

problem Underestimation of blocking event duration in climate models.
method Transfer Learning and Explainable AI (SHAP analysis)
result High-pressure anomalies in specific regions contribute to blocking events.

The paper evaluates the probability distributions of analog-to-target distances for multiple analogs.

problem Understanding the performance of analog applications through the distribution of distances to target states.
method Theoretical analysis and numerical experiments using dynamical systems theory.
result The size of the catalog and dimensionality affect the probability distributions of the K-best analogs.

We present a case-study demonstrating the usefulness of Bayesian hierarchical mixture modelling for investigating cognitive processes. In sentence comprehension, it is widely assumed that the distance between linguistic co-dependents affects the latency of dependency resolution: the longer the distance, the longer the …

2017-02-02abs ↗pdf ↗

Hierarchical causal models help understand cause and effect in nested data.

problem Learning cause and effect from nested hierarchical data.
method Extend structural causal models and causal graphical models with inner plates, develop graphical identification technique and estimation methods.
result Hierarchical data can enable causal identification even when non-hierarchical data cannot.

CE improves climate uncertainty quantification using GCM ensembles and observational data.

problem Uncertainty in climate projections due to model inadequacies and variability.
method Conformal ensembles integrating GCM ensembles and observational data.
result CE generates statistically rigorous, easy-to-interpret uncertainty estimates.

Softmax is an output activation function for modeling categorical probability distributions in many applications of deep learning. However, a recent study revealed that softmax can be a bottleneck of representational capacity of neural networks in language modeling (the softmax bottleneck). In this paper, we propose an…

2018-05-28abs ↗pdf ↗

We harness the power of Bayesian emulation techniques, designed to aid the analysis of complex computer models, to examine the structure of complex Bayesian analyses themselves. These techniques facilitate robust Bayesian analyses and/or sensitivity analyses of complex problems, and hence allow global exploration of th…

2017-03-03abs ↗pdf ↗

Generative AI predicts Arctic sea ice dynamics over decades.

problem Reproducing realistic sea ice dynamics from days to decades is computationally challenging.
method Introduced GenSIM, a generative AI model trained on 20 years of sea-ice-ocean simulation data.
result Generative AI predicts realistic sea ice evolution for 30 years, capturing long-term trends and physical consistency.

OceanForecastBench offers a comprehensive benchmark for data-driven ocean forecasting models.

problem Lack of open-source, standardized benchmarks for data-driven ocean forecasting models.
method Proposes OceanForecastBench, a benchmark with high-quality data and evaluation pipeline.
result Offers the most comprehensive benchmarking framework for data-driven ocean forecasting.

Signature kernel scoring rule improves weather forecasting by capturing temporal and spatial dependencies.

problem Lack of suitable scoring rules for probabilistic weather forecasting.
method Reframe weather variables as continuous paths using iterated integrals (signature kernels) to capture temporal and spatial dependencies.
result Signature kernel scoring rule outperforms conventional methods in weather forecasting, especially for long-term forecasts.

M-CaStLe discovers causal structures in multivariate space-time data.

problem Challenges in causal graph discovery for high-dimensional gridded data.
method Generalizes CaStLe to multivariate analyses, using local embeddings and pooling spatial replicates.
result More accurately recovers multivariate causal structure and identifies physical dynamics.

NN-GPR improves climate model predictions by preserving fine-scale spatial information.

problem Dilution of fine-scale spatial information and bias in model averaging.
method Gaussian process regression with an infinitely wide deep neural network.
result NN-GPR produces more accurate and detailed climate projections.

Hybrid framework predicts Arctic permafrost decline, risks infrastructure, and provides tools.

problem Tackles permafrost decline and infrastructure risk assessment in Arctic territories.
method Hybrid physics-machine learning framework integrating 2.9 million observations.
result Projects mean permafrost fraction decline of -20.3 pp under RCP8.5 forcing, with high-risk zones identified.

The paper investigates overfitting in hyperparameter optimization.

problem Overfitting in hyperparameter optimization (overtuning).
method Formal definition, large-scale reanalysis of HPO benchmark data, analysis of factors affecting overtuning.
result Overtuning is more common than previously assumed, leading to worse generalization error in 10% of cases.

New interpretation of RNN forget gate improves learnability for long-term sequential data.

problem Improving learnability of recurrent neural networks for long-term temporal dependencies.
method Generalized theory of gated RNNs, focusing on gradient behavior over time.
result Existing RNNs satisfy the gradient condition for initial training, suggesting validity of forget gate interpretation.

Improves trial efficiency by adjusting for historical prognostic scores.

problem Reducing statistical uncertainty in randomized trial estimates.
method Linear covariate adjustment using a prognostic model trained on historical data.
result Prognostic covariate adjustment achieves minimum variance and reduces mean-squared error.

Temporal difference (TD) learning is a popular algorithm for policy evaluation in reinforcement learning, but the vanilla TD can substantially suffer from the inherent optimization variance. A variance reduced TD (VRTD) algorithm was proposed by Korda and La (2015), which applies the variance reduction technique direct…

2020-01-07abs ↗pdf ↗

FCNv2 robustness tested under noise and random initial conditions.

problem Assessing AI weather forecasting model robustness to input noise.
method Two experiments with varying noise levels and random initial conditions.
result FCNv2 preserves hurricane features under low to moderate noise, but underestimates intensity and persistence.

Generative model downgrades coarse satellite images to fine resolution.

problem Reconstructing fine resolution satellite images from coarse scale inputs.
method Combines U-Net transfer encoder with diffusion-based generative model.
result Excellent performance (R2 = 0.65 to 0.94) across seasonal regional splits.

EnKBS smoothes complex systems with future observations for causal inference.

problem Improving state estimation in complex systems with rapid dynamics.
method Continuous-time ensemble Kalman-Bucy smoother for nonlinear dynamical systems.
result EnKBS provides derivative-free framework with high skill in various scientific problems.

Study combines variational inference and transformers for seasonal climate predictions.

problem Lack of robust seasonal predictions due to limited historical records and computational constraints.
method Combines variational inference with transformer models trained on climate model output.
result Method provides skilful predictions beyond climate change-induced trends in various regions.

Driven by climatic processes, wind power generation is inherently variable. Long-term simulated wind power time series are therefore an essential component for understanding the temporal availability of wind power and its integration into future renewable energy systems. In the recent past, mainly power curve based mod…

2019-12-09abs ↗pdf ↗

A new method uses deep learning to evaluate causal theories without strict assumptions.

problem Evaluating causal theories represented as DAGs requires arbitrary assumptions that can bias results.
method Causal-graphical normalizing flows (cGNFs) use deep neural networks to empirically evaluate DAGs without functional form assumptions.
result cGNFs allow flexible, semi-parametric estimation of causal effects from DAGs.

Study assesses risk of upward lightning at wind turbines using direct measurements and machine learning.

problem Risk underestimation of upward lightning at wind turbines due to limited detection by current standards.
method Direct UL measurements linked to meteorological reanalysis data using random forests.
result Risk maps based on case study events show high probabilities coincide with actual UL events.

We propose a novel method to forecast the future from the present using time-reversed data.

problem Forecasting the future from past data, exploiting temporal asymmetry.
method Retrodictive forecasting via inverse MAP optimization over a Conditional Variational Autoencoder (CVAE).
result The method successfully predicts future events in time-reversible and irreversible processes.

New method predicts wind farm power and wakes using weather patterns.

problem Inefficient and computationally intensive wind energy resource assessment.
method Unsupervised clustering of ERA5 data on wind velocity, WRF simulations at cluster centers, and post-processing.
result Accurate long-term predictions of power and wakes with reduced computational time.

Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…

2017-12-12abs ↗pdf ↗

Prevents sensitive data generation in diffusion models using labeled and unlabeled data.

problem Generating sensitive data in diffusion models using unlabeled data.
method Positive-Unlabeled Diffusion Models, approximating ELBO with labeled and unlabeled data.
result Prevents the generation of sensitive data without compromising image quality.

Study reveals Data Shapley's inconsistent performance in data selection tasks.

problem Inconsistency of Data Shapley's performance in data selection across different settings.
method Hypothesis testing framework and identification of utility functions.
result Data Shapley's performance is no better than random selection without specific constraints.

PRRO generates synthetic tabular data that improves SL performance and class distribution.

problem Low SL utility of synthetic data due to class imbalance and overlooked data relationships.
method Data pruning and column reordering to optimize SL utility.
result Synthetic data generated with PRRO enhances predictive performance and class distribution.

Defines data science as a natural ecosystem with challenges and missions.

problem Challenges and missions in data science due to 5D complexities and data life cycle phases.
method Systemic and data-centric view of data science as a fusion of data universe and its challenges, formalizing a general-purpose architecture.
result Essential data science as a natural ecosystem integrating specific disciplines and high-impact applications.

Differential privacy allows quantifying privacy loss resulting from accessing sensitive personal data. Repeated accesses to underlying data incur increasing loss. Releasing data as privacy-preserving synthetic data would avoid this limitation, but would leave open the problem of designing what kind of synthetic data. W…

2019-12-10abs ↗pdf ↗

Paper creates fair synthetic data ensuring equal predictions across sensitive attributes.

problem Ensuring fair predictions across sensitive attributes in synthetic data.
method Equalizing target probability distributions across sensitive attributes in synthetic data generation.
result Synthetic data provides strong fair predictions, equal across all thresholds.

Efficient synthetic data generation improves model performance on tabular data.

problem Improving model robustness and performance with scarce or low-quality data.
method Hardness characterization to identify high-value training points, generating synthetic data only from these points.
result Synthetic data generated from hardest points outperforms non-targeted methods on tabular datasets.