New method preserves GCM spatial dependencies for better climate projections.
problem Systemic biases in GCM output and loss of spatial/temporal dependencies.
method SPECD approach using Vecchia approximation and semi-parametric quantile regression.
result SPECD preserves key marginal and joint distribution properties of precipitation and temperature.
EnScale learns to downscale climate models efficiently, capturing both spatial and temporal consistency.
problem Downscaling climate models from coarse to high-resolution data is computationally expensive and challenging.
method EnScale uses generative models and proper scoring rules to map GCM data to RCM data, reducing computational cost.
result EnScale achieves competitive performance and computational efficiency in downscaling multiple climate variables.
This paper corrects climate model biases using a factor model approach.
problem Systematic biases in GCM outputs due to unobserved confounders.
method Factor model approach to learn latent confounders from historical data and apply them to enhance bias correction.
result Significant improvements in the accuracy of precipitation outputs.
This paper constructs GCM hypersurfaces in Kerr spacetimes.
problem Extending the Kerr family stability proof to full stability.
method Concatenating a 1-parameter family of GCM spheres by solving an ODE system.
result Removes symmetry restrictions in GCM procedure.
DoWhy-GCM extends causal inference in graphical models for diverse queries.
problem Addressing diverse causal queries in graphical causal models.
method Specify cause-effect relations via a causal graph, fit causal mechanisms, pose causal queries.
result Identification of root causes, attribution of causal influences, diagnosis of causal structures.
New method defines GCM spheres in Kerr perturbations, proving their stability.
problem Stability of GCM spheres in Kerr perturbations.
method Effective uniformization theorem, canonical definition of ℓ=1 modes, intrinsic existence theorem. result Stability of GCM spheres in Kerr perturbations proven.
Paper compares feature selection methods using GCM and LOCO, showing GCM methods generally outperform LOCO.
problem Feature selection and importance estimation in model-agnostic settings.
method Comparison of feature selection methods related to GCM and LOCO under three model settings.
result GCM-related methods generally outperform LOCO under suitable regularity conditions, as shown by theoretical and empirical results.
Study combines variational inference and transformers for seasonal climate predictions.
problem Lack of robust seasonal predictions due to limited historical records and computational constraints.
method Combines variational inference with transformer models trained on climate model output.
result Method provides skilful predictions beyond climate change-induced trends in various regions.
This the first in a series of papers whose ultimate goal is to establish the full nonlinear stability of the Kerr family for ∣a∣≪m. The paper builds on the strategy laid out in \cite{KS} in the context of the nonlinear stability of Schwarzschild for axially symmetric polarized perturbations. In fact the central id…
We consider the pricing of European-style structured credit payoff in a static framework, where the underlying default times are independent given a common factor. A practical application would consist of the pricing of nth-to-default baskets under the Gaussian copula model (GCM). We provide necessary and sufficient co…
CE improves climate uncertainty quantification using GCM ensembles and observational data.
problem Uncertainty in climate projections due to model inadequacies and variability.
method Conformal ensembles integrating GCM ensembles and observational data.
result CE generates statistically rigorous, easy-to-interpret uncertainty estimates.
The methodology of community detection can be divided into two principles: imposing a network model on a given graph, or optimizing a designed objective function. The former provides guarantees on theoretical detectability but falls short when the graph is inconsistent with the underlying model. The latter is model-fre…
Generative model improves wind field downscaling from coarse climate models.
problem Limited spatial resolution and biases in GCMs for wind energy studies.
method SerpentFlow for domain alignment and conditional fine-scale learning.
result Improved spatial coherence, inter-variable consistency, robustness under climate change.
AnomalyCD discovers anomaly causes in large systems with binary flags, reducing computational burden.
problem Learning graphical causal models from large-scale binary anomaly data is computationally expensive.
method AnomalyCD uses anomaly data-aware causality testing, sparse data compression, and edge pruning.
result AnomalyCD reduces computation overhead and improves accuracy on binary anomaly datasets.
Statistical downscaling of global climate models (GCMs) allows researchers to study local climate change effects decades into the future. A wide range of statistical models have been applied to downscaling GCMs but recent advances in machine learning have not been explored. In this paper, we compare four fundamental st…
Develops geometric causal models for causal inference from dependent data.
problem Causal inference from structured, dependent data (e.g., spatial, network, molecular).
method Geometric causal models (GCMs) exploiting symmetries of data generating process, combining group theory, ergodic theory, and Bayesian inference.
result Establishes identification and estimation of causal effects from dependent data.
Turaev's shadow formula calculates the SU(2)-Reshetikhin-Turaev-Witten invariants using shadows, and its expression is somehow similar to a Euler characteristic. We give a short proof of this formula using skein theory. The formula applies to pairs (M,G) where M is a closed oriented 3-manifold and GcM is a (possibly em…
A new measure of causal influence quantifies intrinsic contributions in DAGs.
problem Quantifying intrinsic causal contributions in Directed Acyclic Graphs (DAGs).
method Recursive decomposition of node contributions, structure-preserving interventions, Shapley symmetrization.
result A measure of intrinsic causal contribution that is invariant to node relabeling.
This paper proves a canonical foliation on null infinity for Kerr-like black holes.
problem Establishing well-defined physical quantities on null infinity for Kerr-like black holes.
method Existence and uniqueness results for GCM spheres by Klainerman-Szeftel.
result Existence of a canonical foliation on future null infinity with well-defined physical quantities.
ECCIT improves conditional independence tests by calibrating for miscalibration.
problem Inaccurate frequentist guarantees in CITs, especially in small samples and misspecified models.
method Empirically Calibrated Conditional Independence Tests (ECCIT) that optimize and correct for miscalibration.
result ECCIT achieves valid FDR with higher power than existing calibration strategies.
New method selects causal features from diverse data types.
problem Discovering causal relationships from non-continuous data types.
method Transformation-Model (TRAM) based Invariant Causal Prediction (TRAM-ICP) with TRAM-GCM and TRAM-Wald tests.
result Improved power and type I error control for diverse response types.
MinShap improves feature selection accuracy and stability in complex models.
problem Challenging feature selection in unknown non-linear relationships.
method Combines Shapley values with a modified approach to feature selection.
result MinShap outperforms state-of-the-art algorithms in accuracy and stability.
The paper presents a novel approach to multi-output regression using probabilistic circuits.
problem Capturing correlations between multiple output dimensions in large-scale regression problems.
method Employing a mixture of single-output Gaussian process experts encoded via a probabilistic circuit.
result The method can capture correlations between output dimensions and often outperforms other approaches.
A scalable MOGP model with stochastic variational inference for many outputs.
problem Efficiently modeling data from multiple sources with many outputs.
method Stochastic variational inference for Latent Variable MOGP (LV-MOGP).
result Computational complexity per iteration is independent of the number of outputs.
Multi-output learning aims to simultaneously predict multiple outputs given an input. It is an important learning problem due to the pressing need for sophisticated decision making in real-world applications. Inspired by big data, the 4Vs characteristics of multi-output imposes a set of challenges to multi-output learn…
Unified framework for output analysis using Monte Carlo sampling.
problem Accurately assess the quality of estimated values in predictive models.
method Unified output analysis framework through Monte Carlo sampling, leveraging fast iterative bootstrap sampling and higher-order influence functions.
result Clear advantage in building more robust confidence intervals with higher coverage probability.
Enhances robustness of MOGP regression for multiple correlated outputs.
problem Model misspecification and outliers in MOGP regression.
method Extends RCGP framework to multi-output setting.
result Provable robust MOGP with joint correlation capture.
DeepICMGP surrogate models multiple outputs efficiently.
problem Challenges in modeling dependencies between multiple outputs using traditional multi-output GPs.
method Introduces hierarchical coregionalization structures across layers in DGPs.
result Demonstrates competitive performance and active learning strategies.
Improved code translation by preserving structure with composed fine-tuning.
problem Improving code translation accuracy with unlabeled code outputs.
method Pre-trained denoiser to capture output structure, composed fine-tuning to fine-tune predictor.
result Composed fine-tuning significantly improves generalization over standard fine-tuning.
Multi-Output Dependence (MOD) learning is a generalization of standard classification problems that allows for multiple outputs that are dependent on each other. A primary issue that arises in the context of MOD learning is that for any given input pattern there can be multiple correct output patterns. This changes the…
New method learns output embeddings for structured prediction.
problem Structured prediction with output embeddings.
method Jointly learns output embedding and regression function.
result Structured predictor is a consistent estimator with smaller complexity.
Proposes GPLFR for predicting high-dimensional outputs with few data.
problem Predicting high-dimensional outputs from limited data.
method GPLFR combines Gaussian process and linear-Gaussian decoding for high-dimensional prediction.
result GPLFR outperforms existing methods in predicting high-dimensional outputs.
New model predicts multiple outputs with missing labels.
problem Missing group labels in multi-output regression.
method Weakly-supervised multi-output model using correlated Gaussian processes.
result Model excels in multi-output settings with missing labels.
RePULSe improves language model alignment by reducing undesired outputs without sacrificing overall performance.
problem Aligning language models with human preferences while minimizing undesired outputs.
method Integrates probabilistic inference into RL training to reduce undesired outputs.
result RePULSe achieves a better balance between expected reward and undesired output probability.
Sig-PCA integrates model outputs and observations to correct model biases.
problem Improving model accuracy and reliability by correcting biases and numerical approximations.
method Sig-PCA framework that combines summary statistics from model outputs with localized observations via a neural network.
result Corrects model outputs to align closely with observational data, preserving essential statistical information.
Models such as Sequence-to-Sequence and Image-to-Sequence are widely used in real world applications. While the ability of these neural architectures to produce variable-length outputs makes them extremely effective for problems like Machine Translation and Image Captioning, it also leaves them vulnerable to failures o…
Gaussian processes (GPs), or distributions over arbitrary functions in a continuous domain, can be generalized to the multi-output case: a linear model of coregionalization (LMC) is one approach. LMCs estimate and exploit correlations across the multiple outputs. While model estimation can be performed efficiently for …
Automatically identifies geometric flat outputs for robotic systems.
problem Lack of systematic and practical means to identify flat outputs for arbitrary robotic systems.
method Casts the search for a globally valid, equivariant flat output as an optimization problem using Riemannian geometry, Lie group theory, and differential forms.
result Approximate transcription of continuum formulation to a quadratic program achieves precise agreement with known closed-form flat outputs.
Wide Boosting improves GB's performance on multivariate output tasks.
problem Lack of flexibility in fitting probabilistic multi-dimensional outputs.
method Inserts matrix multiplication between GB output and loss function.
result Wide Boosting outperforms Gradient Boosting on multivariate output tasks.
Proposes SHORE model for efficient MOR with sparsity and scalability.
problem Challenges of interpretability and scalability in MOR with high-dimensional outputs.
method Incorporates sparsity requirements and a two-stage optimization framework for efficient compression.
result Theoretical and empirical validation of the proposed framework's efficiency and accuracy.
In many applications of supervised learning, multiple classification or regression outputs have to be predicted jointly. We consider several extensions of gradient boosting to address such problems. We first propose a straightforward adaptation of gradient boosting exploiting multiple output regression trees as base le…
For many important problems the quantity of interest is an unknown function of the parameters, which is a random vector with known statistics. Since the dependence of the output on this random vector is unknown, the challenge is to identify its statistics, using the minimum number of function evaluations. This problem …
The paper discusses the role of monetary policy when potential output depends on the inflation rate. If the intention of the central bank is to maximize actual output growth, then it has to be credibly committed to a strict inflation targeting rule, and to take the MOGIR (the Maximizing Output Growth Inflation Rate) as…
Recently there has been an increasing interest in methods that deal with multiple outputs. This has been motivated partly by frameworks like multitask learning, multisensor networks or structured output data. From a Gaussian processes perspective, the problem reduces to specifying an appropriate covariance function tha…
This study compares multivariate vs univariate machine learning for multi-output regression.
problem When to use multivariate ensemble techniques over separate univariate models.
method Comparative analysis of different multivariate approaches for multi-output regression.
result Multivariate ensemble techniques outperform separate univariate models in simulations.
This paper studies the problem of learning the conditional distribution of a high-dimensional output given an input, where the output and input may belong to two different domains, e.g., the output is a photo image and the input is a sketch image. We solve this problem by cooperative training of a fast thinking initial…
Safe active learning for multi-output Gaussian processes reduces data acquisition costs and ensures safety.
problem Expensive data acquisition and safety concerns in multi-output regression problems.
method Proposes a safe active learning approach considering data informativeness and safety constraints.
result Improved convergence compared to competitors on simulated and real-world datasets.
Reduced-rank method improves least-squares regression under output regularity.
problem Least-squares regression with infinite dimensional outputs.
method Reduced-rank method for solving least-squares problems with output regularity assumptions.
result Learning bounds and improved statistical performance compared to full-rank method.