The study introduces a holdout-based framework to assess synthetic data fidelity and privacy.
problem Evaluating the quality and privacy of synthetic data solutions for mixed-type tabular data.
method Holdout-based empirical assessment framework measuring fidelity and privacy risk.
result Synthetic data samples are as close to the training as to the holdout data, indicating generalization and independence from individual records.
Enhances multi-fidelity modeling with DGPs for different input domains.
problem Improving prediction accuracy with multi-fidelity models using different input domains.
method Extends Deep Gaussian Processes (DGPs) to handle different input domains for high and low-fidelity models.
result Demonstrates improved performance on real-world physical problems.
Fidel-TS creates a new benchmark for time series forecasting models.
problem Lack of high-quality benchmarks for time series forecasting models.
method Formalized high-fidelity benchmark principles, including data sourcing integrity, leak-free design, and structural clarity. Created Fidel-TS, a new large-scale benchmark.
result Demonstrated the limitations of prior benchmarks and potential discrepancies in model evaluation.
This paper reviews Gaussian process-based multi-fidelity techniques for different fidelity relationships.
problem Combining accurate and cheap models for complex system design.
method Gaussian process-based multi-fidelity modeling techniques for varying fidelity relationships.
result Comparison of techniques on analytical and aerospace engineering problems.
This work introduces a new metric to assess the fidelity of surrogate models to the underlying data-generating signal.
problem The limitations of fidelity-based explanations in explainable AI.
method Introduces the linearity score λ(f) to quantify the extent of a regression network's linear decodability. result High-fidelity surrogates can underperform compared to simpler models and even linear baselines trained directly on the data.
This paper presents a method to efficiently estimate rare event probabilities using a combination of high and low-fidelity models.
problem Estimating the probability of failure for complex systems using high-fidelity models is expensive and inaccurate for rare events.
method The paper introduces a multi-fidelity surrogate modeling strategy using active learning and subset simulation to merge high and low-fidelity models.
result The method significantly reduces computational cost while maintaining high accuracy in estimating rare event probabilities.
In this article, we consider a stochastic numerical simulator to assess the impact of some factors on a phenomenon. The simulator is seen as a black box with inputs and outputs. The quality of a simulation, hereafter referred to as fidelity, is assumed to be tunable by means of an additional input of the simulator (e.g…
Deep neural networks (DNNs) have recently achieved state-of-the-art performance and provide significant progress in many machine learning tasks, such as image classification, speech processing, natural language processing, etc. However, recent studies have shown that DNNs are vulnerable to adversarial attacks. For inst…
SHAP Distance assesses semantic fidelity of synthetic tabular data.
problem Semantic fidelity of synthetic tabular data is not well evaluated.
method SHAP Distance, defined as cosine distance between global SHAP attribution vectors.
result SHAP Distance detects semantic discrepancies overlooked by standard measures.
GAR generalizes autoregression for efficient multi-fidelity fusion.
problem Efficiently combining low-fidelity and high-fidelity simulation results.
method Generalized autoregression (GAR) using tensor formulation and latent features.
result GAR outperforms state-of-the-art methods with a large margin in RMSE.
Enhanced multi-fidelity models improve digital twin accuracy and uncertainty quantification.
problem Lack of detailed application-specific data and inaccurate sensor data hinder surrogate model learning for digital twins.
method Proposes a multi-fidelity surrogate model framework integrating PCFE and GP, and deep-HPCFE with auto-regression schemes.
result Demonstrates improved accuracy and uncertainty quantification in digital twin systems.
Study characterizes harmful low-fidelity data sources for surrogate models.
problem Identifying which low-fidelity data sources to use in constructing surrogate models.
method Employed benchmark filtering techniques to assess harmful sources using limited data.
result Provided guidelines for using low-fidelity sources in an industrial setting.
Develops a framework for cost-efficient Bayesian optimization with constraints.
problem Optimizing designs with minimal cost in constrained search spaces.
method Constrained multi-fidelity Bayesian optimization (CMFBO) with automatic stopping criterion.
result Minimizes overall sampling costs while ensuring feasibility.
Proposes a method to improve surrogate modeling and design optimization using latent variables.
problem Improving efficiency in multi-fidelity adaptive sampling without hierarchical assumptions.
method A framework using a latent variable Gaussian process to capture correlations between different fidelity models and optimize adaptive sampling.
result Demonstrates superior performance in convergence rate and robustness compared to existing methods.
A cost-efficient method for hyperparameter tuning using multi-fidelity Bayesian optimization.
problem Expensive hyperparameter tuning with limited knowledge transfer methods.
method Amortized Auto-Tuning (AT2) framework for multi-task, multi-fidelity Bayesian optimization.
result AT2 leads to the best hyperparameter recommendation and is more cost-efficient.
New method for reliability analysis using multi-fidelity models.
problem Reliability analysis of complex systems with high computational costs.
method Adaptive Multi-fidelity Gaussian Process for Reliability Analysis (AMGPRA) with collective learning function (CLF).
result AMGPRA achieves similar or higher accuracy with reduced computational costs compared to state-of-the-art methods.
Researchers develop methods to reduce simulation costs for cardiovascular modeling.
problem High computational cost of high-fidelity simulations in cardiovascular modeling.
method Use low-fidelity approximations, neural networks, and normalizing flows to construct surrogates.
result Validated methods reduce computational cost while maintaining accuracy.
Paper proposes a method to estimate total variation distance for synthetic data fidelity.
problem Assessing the fidelity of synthetic data generated by AI.
method Discriminative approach to estimate total variation distance between two distributions.
result Estimation of total variation distance reduces to quantifying Bayes risk in classification.
Paper evaluates synthetic retail data for fidelity, utility, and privacy.
problem Ensuring accurate synthetic data in retail.
method Differentiates between continuous and discrete data, measures fidelity and utility, and uses Differential Privacy for privacy.
result Validated framework for reliable and scalable synthetic data evaluation.
Automated rock fragmentation assessment using deep learning and spatial statistics.
problem Assessing post-blast rock fragmentation in real-time.
method Fine-tuned YOLO12l-seg model for instance segmentation, followed by spatial statistics.
result Framework accurately assesses rock fragmentation patterns in real-time.
ECS evaluates synthetic CXR images' distributional fidelity.
problem Evaluating synthetic CXR images' distributional fidelity under privacy constraints.
method Characteristic function transforms of feature embeddings.
result ECS uncovers clinically relevant distributional discrepancies.
A new UNet variant reduces spectral artifacts in image transformations.
problem Spectral artifacts caused by traditional UNet upsampling layers.
method Introduced a Guided UNet (GUNet) architecture using a novel upsampling module.
result GUNet produces higher fidelity outputs in image transformations.
Saliency maps are a popular approach to creating post-hoc explanations of image classifier outputs. These methods produce estimates of the relevance of each pixel to the classification output score, which can be displayed as a saliency map that highlights important pixels. Despite a proliferation of such methods, littl…
Stable GFlowNets prevent loss spikes and mode collapse in training.
problem Unstable training of GFlowNets leading to loss spikes and mode collapse.
method Assessed sensitivity of GFlowNet objectives, derived loss-to-TV bounds, and proposed Stable GFlowNets.
result Stable GFlowNets improve training behavior and distributional fidelity.
Maximal Rate of Stepwise Uncertainty Reduction selects simulations to reduce uncertainty efficiently.
problem Efficiently estimating quantities of interest from multi-fidelity simulations.
method Bayesian sequential strategy that maximizes the ratio of expected uncertainty reduction to simulation cost.
result MR-SUR strategy unifies and provides principled approaches to develop new methods.
Managers of US National Forests must decide what policy to apply for dealing with lightning-caused wildfires. Conflicts among stakeholders (e.g., timber companies, home owners, and wildlife biologists) have often led to spirited political debates and even violent eco-terrorism. One way to transform these conflicts into…
This paper uses CausalGANs and RL with LLM to predict bond yields.
problem Challenges in financial bond yield forecasting due to data scarcity and market conditions.
method Proposes a novel framework combining CausalGANs, RL, and LLM for synthetic data generation and trading signals.
result Improves forecasting performance over existing methods with low Mean Absolute Error.
Recent studies illustrate how machine learning (ML) can be used to bypass a core challenge of molecular modeling: the tradeoff between accuracy and computational cost. Here, we assess multiple ML approaches for predicting the atomization energy of organic molecules. Our resulting models learn the difference between low…
The unconditional generation of high fidelity images is a longstanding benchmark for testing the performance of image decoders. Autoregressive image models have been able to generate small images unconditionally, but the extension of these methods to large images where fidelity can be more readily assessed has remained…
Most of existing manifold learning methods rely on Mean Squared Error (MSE) or ℓ2 norm. However, for the problem of image quality assessment, these are not promising measure. In this paper, we introduce the concept of an image structure manifold which captures image structure features and discriminates image dist…
Framework audits synthetic datasets for trustworthiness across various use cases.
problem Assessing the trustworthiness of synthetic datasets and models.
method Holistic auditing framework focusing on bias, fidelity, utility, robustness, and privacy.
result Introduces a trustworthiness index and model selection process for controllable trade-offs.
Unified four trade-off curves for assessing generative model proximity.
problem Quantitative assessment of proximity between two probability distributions.
method Unified four existing curves: PR, Lorenz, ROC, and Rényi divergence frontiers.
result Explicit relationship between PR and Lorenz curves with domain adaptation bounds.
CTBench benchmarks cryptocurrency time series generation for trading applications.
problem Lack of comprehensive benchmarks for cryptocurrency time series generation.
method Developed a comprehensive benchmark extsf{CTBench} with 13 metrics across 5 dimensions.
result Uncovered trade-offs between statistical fidelity and real-world profitability.
Proposes a multi-fidelity machine learning strategy integrating low-fidelity deterministic and high-fidelity Bayesian models.
problem Addressing the accuracy-efficiency trade-off in machine learning with scarce high-fidelity data.
method Integrates a non-probabilistic regression model for low-fidelity with a Bayesian model for high-fidelity, trained in a staggered scheme.
result Achieves comparable performance in mean and uncertainty estimation with reduced training time and effective mitigation of overfitting.
A new BO framework reduces costs by using low-fidelity data.
problem Optimizing expensive experiments with low-fidelity data.
method Developed a multi-fidelity cost-aware Bayesian optimization framework.
result Significantly outperforms state-of-the-art BO methods.
The paper proposes a new method to calibrate multiple computer models simultaneously.
problem Calibrating multiple computer models one at a time is inefficient.
method Developed a probabilistic framework using customized neural networks.
result Simultaneous calibration improves predictive accuracy but can be non-identifiable in high dimensions.
This work improves surrogate models for balancing accuracy and cost in multi-fidelity methods.
problem Balancing accuracy and computational cost in multi-fidelity methods.
method Develops context-aware surrogate models for multi-fidelity importance sampling and Bayesian inverse problems.
result Context-aware surrogate models can lead to runtime speedups of up to one order of magnitude.
This paper presents a novel signal compression algorithm based on the Blaschke unwinding adaptive Fourier decomposition (AFD). The Blaschke unwinding AFD is a newly developed signal decomposition theory. It utilizes the Nevanlinna factorization and the maximal selection principle in each decomposition step, and achieve…
Discovering novel materials can be greatly accelerated by iterative machine learning-informed proposal of candidates---active learning. However, standard \emph{global-scope error} metrics for model quality are not predictive of discovery performance, and can be misleading. We introduce the notion of \emph{Pareto shell-…
Machine learning combines high- and low-fidelity models for efficient uncertainty quantification and optimization.
problem Efficiently combining high- and low-fidelity models for uncertainty quantification and optimization.
method Machine learning-based multi-fidelity methods for uncertainty quantification and optimization.
result Unified perspective on multi-fidelity priors for optimization.
Despite the advances of deep learning in specific tasks using images, the principled assessment of image fidelity and similarity is still a critical ability to develop. As it has been shown that Mean Squared Error (MSE) is insufficient for this task, other measures have been developed with one of the most effective bei…
This paper improves surrogate modeling for noisy data.
problem Uncertainty in high-fidelity models due to noise.
method Comprehensive framework for multi-fidelity surrogate modeling.
result Estimates uncertainty in high-fidelity model predictions.
A new MCMC method combines low and high-fidelity models to reduce computation.
problem Inefficient computation of expensive target densities in scientific applications.
method Pseudo-marginal MCMC approach using a telescoping series of low-fidelity models.
result Asymptotically exact multi-fidelity MCMC algorithms for reduced computational cost.
The paper compares multi-fidelity methods for Gaussian process surrogates in physics.
problem Limited availability of data due to expensive simulations.
method Extending non-linear autoregressive methods to multi-fidelity models and incorporating delay terms.
result Multi-fidelity methods generally have smaller prediction error for the same computational cost.
This text discusses several popular explanatory methods that go beyond the error measurements and plots traditionally used to assess machine learning models. Some of the explanatory methods are accepted tools of the trade while others are rigorously derived and backed by long-standing theory. The methods, decision tree…
Efficiently predicts high-fidelity PDE solutions using multi-fidelity Gaussian processes.
problem Expensive high-fidelity solutions for PDEs on discretized domains.
method Multi-Fidelity High-Order Gaussian Process (MFHoGP) that integrates multi-fidelity examples and scales to large numbers of outputs.
result Significantly reduces the cost of high-fidelity PDE solutions through efficient Gaussian process modeling.
FNO model predicts GCS pressure fields with 81% less data, even with limited high-fidelity data.
problem Accurate prediction of complex physical behaviors in large-scale 3D geological carbon storage problems with limited data.
method Multi-fidelity Fourier Neural Operator (FNO) for efficient training with multi-fidelity datasets.
result Multi-fidelity FNO model predicts pressure fields with reasonable accuracy even with limited high-fidelity data.
Paper optimizes multi-fidelity function with fast learning rates.
problem Optimizing a locally smooth function with limited budget and varying fidelity approximations.
method Kometo algorithm that achieves simple regret rates without knowing function smoothness or fidelity assumptions.
result Kometo algorithm outperforms previous methods empirically.