Perturbation theory improves nonparametric instrumental variable estimation accuracy.
problem Improving nonparametric instrumental variable estimation accuracy in high-dimensional settings.
method Perturbative approach based on physics perturbation theory, extending kernel ridge methods with higher-order corrections.
result First-order perturbative corrections reduce prediction error by up to 99% in high-dimensional ill-defined cases.
The majority of traditional classification ru les minimizing the expected probability of error (0-1 loss) are inappropriate if the class probability distributions are ill-defined or impossible to estimate. We argue that in such cases class domains should be used instead of class distributions or densities to construct …
We construct solutions with prescribed asymptotics to the Einstein constraint equations using a cut-off technique. Moreover, we give various examples of vacuum asymptotically flat manifolds whose center of mass and angular momentum are ill-defined.
Recover simple irreversible Finsler geometry from travel time data
problem Stable recovery of a simple irreversible Finsler geometry
method Use a Gromov-Hausdorff distance adapted to irreversible metric spaces
result Unique and Lipschitz-stable recovery
Unified framework for estimating indirect effects in observational studies with unmeasured confounding.
problem Challenges in evaluating indirect effects due to unmeasured confounding and unethical exposures.
method Developed a unified identification and estimation framework using proximal causal inference.
result Unified identification and estimation of PIIE and causal effect of an intervening variable in settings with pervasive unmeasured confounding.
New method evaluates personalized treatment in critical care, robust to death.
problem Truncation by death in critical care makes traditional DTR evaluation ineffective.
method Principal stratification-based approach, focusing on always-survivor value function, with a semiparametrically efficient, multiply robust estimator.
result Demonstrates robustness and efficiency of the method for personalized treatment optimization.
The estimation of probabilities of network edges from the observed adjacency matrix has important applications to predicting missing links and network denoising. It has usually been addressed by estimating the graphon, a function that determines the matrix of edge probabilities, but this is ill-defined without strong a…
In this paper we study the implications of contingent payments on the clearing wealth in a network model of financial contagion. We consider an extension of the Eisenberg-Noe financial contagion model in which the nominal interbank obligations depend on the wealth of the firms in the network. We first consider the prob…
A cocycle Ω:P×G→H taking values in a Lie group H for a free right action of G on P defines a principal bundle Q with the structure group H over P/G. The Chern character of a vector bundle associated to Q defines then characteristic classes on X. This observation becomes useful in the case …
A generic degenerate Lagrangian system of even and odd variables on an arbitrary smooth manifold is examined in terms of the Grassmann-graded variational bicomplex. Its Euler-Lagrange operator obeys Noether identities which need not be independent, but satisfy first-stage Noether identities, and so on. However, non-tri…
Reconstruction error is a prevalent score used to identify anomalous samples when data are modeled by generative models, such as (variational) auto-encoders or generative adversarial networks. This score relies on the assumption that normal samples are located on a manifold and all anomalous samples are located outside…
Empirical study shows Randomized Signature Methods improve portfolio optimization in financial markets.
problem Drift estimation in non-linear, non-parametric financial markets is challenging.
method Applied Randomized Signature Methods for non-linear, non-parametric drift estimation in multi-variate financial markets.
result Randomized Signature Methods provide features on the same scale and improve portfolio optimization in real-world settings.
Proposes GDTW for aligning time series on different, incomparable spaces.
problem Dynamic time warping requires comparable spaces, but time series can live on different, incomparable spaces.
method Gromov dynamic time warping (GDTW) considers intra-relational geometry to avoid comparability requirements.
result Demonstrates effectiveness of GDTW in aligning, combining, and comparing time series on incomparable spaces.
Estimating treatment effects in time series with hidden confounding.
problem Estimating treatment effects in time series with hidden confounding.
method A neural framework that learns individual-level counterfactuals and flexible matching procedures.
result Improves counterfactual estimation under latent bias.
The paper proposes a Gaussian mixture model for Hilbert-space-valued data.
problem Challenges in characterizing probability measures for infinite-dimensional random objects.
method Gaussian mixture framework based on kernel mean embeddings.
result The proposed algorithm yields a dense class of approximations in infinite-dimensional spaces.
Method ranks generative models without needing latent factor supervision.
problem Challenges in selecting generative models for qualities like disentanglement.
method Ranking generative models based on training dynamics, without requiring labels for latent factors.
result Method correlates with supervised disentanglement metrics and can predict downstream performance.
A new method for unfolding histograms without matrix inversion.
problem Matrix inversion in experimental physics, especially in high-energy particle physics.
method Sampling many distributions, folding them through the response matrix, and choosing the closest one to the data.
result Performs as well as traditional methods in well-defined inverse problems and outperforms them in ill-defined ones.
Proposes a new flow model to better represent data on manifolds.
problem Flow models struggle to represent data on lower-dimensional manifolds accurately.
method Introduces a manifold prior that leverages spread divergence to improve model performance.
result Improves both sample and representation quality, identifies manifold intrinsic dimension.
Aggregated variables can mask causal effects, turning unconfounded into confounded relations.
problem Aggregated variables can mask causal effects, leading to paradoxical confounding.
method Analysis of how aggregated variables can change the definition of causality and the feasibility of causal relations.
result Macro causal relations are defined by micro states, not just aggregated variables.
Many problems in machine learning involve calculating correspondences between sets of objects, such as point clouds or images. Discrete optimal transport provides a natural and successful approach to such tasks whenever the two sets of objects can be represented in the same space, or at least distances between them can…
MarkerMap selects key genes for cell type analysis in single-cell RNA-seq.
problem Selecting informative genes from large single-cell RNA-seq datasets is challenging and computationally intensive.
method MarkerMap is a generative model that identifies minimal gene sets explaining cell type variability.
result MarkerMap outperforms existing methods in both supervised and unsupervised marker selection.
We analyze GANs using neural tangent kernels, revealing flaws and advancing understanding.
problem Flaws in previous GAN analysis models.
method Neural Tangent Kernel framework for infinite-width discriminator.
result New insights into GAN convergence and generated distribution.
Paper introduces a new geometric homology theory and applies it to Gromov-Witten theory.
problem Developing a new homology theory for orbifolds with corners.
method Using stratification and triangulation theories of Lie groupoids and their orbit spaces, extending to Lie groupoids with corners.
result Proposes and proves the geometric homology theory (GHT), a flexible generalization of singular homology.
Method identifies latent variables from high-dimensional data with piecewise affine mixing.
problem Identifying latent variables from high-dimensional observations with dependencies and piecewise affine transformations.
method Proposes a two-stage method with sparsity and Gaussianity regularization.
result Effectively recovers ground-truth latent variables from synthetic and image data.
We introduce a new generative model where samples are produced via Langevin dynamics using gradients of the data distribution estimated with score matching. Because gradients can be ill-defined and hard to estimate when the data resides on low-dimensional manifolds, we perturb the data with different levels of Gaussian…
This work explores function-space inference using KL divergence and proposes Bayesian linear regression as a benchmark.
problem Approximating the predictive posterior distribution of Bayesian models without parameter posterior approximation.
method Employing Kullback-Leibler divergence and proposing featurized Bayesian linear regression as a benchmark.
result Minimizing KL divergence leads to an ill-defined objective function, highlighting limitations of this approach.
Active learning selects optimal measurement times for inferring continuous paths from sparse data.
problem Inferring continuous probability paths from sparse snapshots in high-fidelity domains like single-cell biology.
method Extends active experimentation to the space of measures using Linearized Optimal Transport (LOT) for probabilistic surrogate modeling.
result Empirical results show that the proposed strategy outperforms uncertainty-agnostic baselines.
Introduces intrinsic Riemannian cross-covariance for manifold-valued random objects.
problem Covariance estimation for random objects on Riemannian manifolds.
method Defines covariance and correlation via parallel transport.
result Proposed covariance is independent of coordinate choices.
New method learns kernels in nonlocal operators robustly.
problem Learning kernels in nonlocal operators is ill-posed.
method Nonparametric regression with Tikhonov regularization.
result Robust estimator of kernel yields homogenized model.
Unified framework for constrained diffusion models on nonconvex sets with efficient landing mechanism.
problem Efficiently modeling generative models under nonconvex constraints.
method Unified framework with overdamped and underdamped dynamics, landing mechanism.
result Significantly reduces computational cost while maintaining sample quality.
New dynamical torsion for contact Anosov flows connects to Reidemeister torsion.
problem Understanding contact Anosov flows and their properties.
method Introducing dynamical torsion and showing its properties.
result Locally constant ratio between dynamical and Turaev torsion.
Dropout, a stochastic regularisation technique for training of neural networks, has recently been reinterpreted as a specific type of approximate inference algorithm for Bayesian neural networks. The main contribution of the reinterpretation is in providing a theoretical framework useful for analysing and extending the…
Develops new shape metrics for high-dimensional objects.
problem Lack of single metrics to describe shape in high dimensions.
method Introduces hyper-Sphericity and hyper-Shape Proportion metrics.
result Discriminates between different shapes in high dimensions.
The problem of developing binary classifiers from positive and unlabeled data is often encountered in machine learning. A common requirement in this setting is to approximate posterior probabilities of positive and negative classes for a previously unseen data point. This problem can be decomposed into two steps: (i) t…
3D BF theory on certain 3-manifolds evaluated via residues and large k limits.
problem Singular and ill-defined partition function of 3D BF theory.
method Direct evaluation of path integral for specific 3-manifolds, using residues and large k limits of Chern-Simons matrix integrals.
result 3 definitions of the integral offer insights into the sum/integral over all flat connections.
The use of machine learning systems to support decision making in healthcare raises questions as to what extent these systems may introduce or exacerbate disparities in care for historically underrepresented and mistreated groups, due to biases implicitly embedded in observational data in electronic health records. To …
Detection of emerging topics are now receiving renewed interest motivated by the rapid growth of social networks. Conventional term-frequency-based approaches may not be appropriate in this context, because the information exchanged are not only texts but also images, URLs, and videos. We focus on the social aspects of…
Proposes a new risk model using stable laws to manage company-wide losses.
problem Managing aggregate risks and pricing policies in the presence of systematic risk.
method Develops a modified risk model using multivariate stable distributions to account for various risk phenomena.
result Computes the Tail Conditional Expectation of aggregate risks and corresponding allocations.
New criterion selects optimal number of clusters based on stability.
problem Challenges in selecting optimal number of clusters in non-parametric clustering.
method Proposes a stability-based validation criterion combining between-cluster and within-cluster stability.
result Empirically demonstrates effectiveness in selecting optimal number of clusters.
The problem of finding a reduced dimensionality representation of categorical variables while preserving their most relevant characteristics is fundamental for the analysis of complex data. Specifically, given a co-occurrence matrix of two variables, one often seeks a compact representation of one variable which preser…
A new method measures heterogeneity without needing categorical partitioning or distance measurement.
problem Measuring heterogeneity in non-categorical data requires categorical partitioning and distance measurement, limiting applicability.
method Representational Rényi heterogeneity (RRH) transforms data into a latent space where heterogeneity can be measured without these requirements.
result RRH can generalize existing indices and better responds to changes in mixture component separation and weighting.
A new method improves flow matching by dynamically weighting density estimates.
problem High-dimensional integration inefficiency in flow matching.
method Density-weighted Dynamic Stein operators.
result Significant improvement in vector field smoothness and sampling efficiency.
Unified framework for sampling from complex distributions, including discrete and mixed-variable systems.
problem Sampling from complex unnormalized distributions, especially in discrete or mixed-variable systems.
method Enforces time-reversibility using a prescribed physical transition kernel to minimize Maximum Mean Discrepancy (MMD).
result Demonstrates accurate reproduction of thermodynamic observables and mode-switching behavior across diverse systems.
MeLa learns task relations by inferring global labels for robust FSL.
problem Few-shot learning with limited global labels.
method Meta Label Learning (MeLa) and augmented pre-training.
result MeLa outperforms existing methods across diverse benchmarks.
Noise can affect the overparametrization of QNNs, enabling new directions but also suppressing sensitivity.
problem The overparametrization of QNNs in the presence of noise.
method Analyzing the Quantum Fisher Information Matrix (QFIM) to understand how noise affects the rank of QFIM.
result Noise can turn previously-zero eigenvalues of the QFIM to non-zero, enabling exploration of new directions.
Learning the true ordering between objects by aggregating a set of expert opinion rank order lists is an important and ubiquitous problem in many applications ranging from social choice theory to natural language processing and search aggregation. We study the problem of unsupervised rank aggregation where no ground tr…
This paper analyzes Mean Decrease Impurity (MDI) variable importance in random forests.
problem Lack of interpretability in random forest variable importances.
method Analysis of Mean Decrease Impurity (MDI) in random forests.
result MDI provides a variance decomposition of the output when variables are independent and there are no interactions.
The paper explores fairness, welfare, and equity in personalized pricing across various applications.
problem Interplay of fairness, welfare, and equity in personalized pricing based on customer features.
method Comprehensive literature review and observational metrics without underlying valuation distribution assumptions.
result Personalized pricing can expand access, improve welfare, and increase revenue or budget utilization.