CCs learn high-dimensional distributions from heterogeneous data.
problem Learning high-dimensional distributions from heterogeneous data.
method Introducing characteristic circuits (CCs) that learn from data and use spectral domain.
result CCs outperform state-of-the-art density estimators on common benchmark data sets.
New models capture heterogeneous network density, improving community detection.
problem Empirical networks are often globally sparse but locally dense.
method Latent Poisson models generating hidden multigraphs.
result These models improve community detection in sparse networks.
Proposes a new method to handle data heterogeneity in causal inference.
problem Challenges of collaborating between different data centers due to heterogeneity.
method Collaborative inverse propensity score weighting estimator to adjust for distribution shift.
result Significant improvements over traditional meta-analysis methods when dealing with increased heterogeneity.
Using an intuitive concept of what constitutes a meaningful community, a novel metric is formulated for detecting non-overlapping communities in undirected, weighted heterogeneous networks. This metric, modularity density, is shown to be superior to the versions of modularity density in present literature. Compared to …
Study on evolving interfaces with complex curvature and density effects.
problem Understanding the dynamics of evolving heterogeneous elastic interfaces.
method Modeling an evolving curve with a density function, analyzing the associated gradient flow evolution.
result Analysis of the preservation and asymptotic behavior of geometric properties in the evolving system.
Expands causal clustering framework with hierarchical and density-based methods.
problem Identifying heterogeneous treatment effects in unknown subgroup structure.
method Integrates hierarchical and density-based clustering algorithms into causal k-means clustering.
result Plug-in estimators for causal clustering are simple and readily implementable.
Conformal-DP improves differential privacy on manifold data by calibrating perturbations based on local densities.
problem Lack of density-awareness in existing differential privacy mechanisms for manifold data leads to biased and suboptimal privacy-utility trade-offs.
method Proposes Conformal-DP, a density-aware differential privacy mechanism using conformal transformations to calibrate perturbations based on local densities.
result Demonstrates improved privacy-utility trade-off in heterogeneous data distribution settings compared to state-of-the-art mechanisms.
The paper explores strong identifiability and parameter learning in regression models with heterogeneous responses.
problem Understanding heterogeneity in data populations through conditional distributions of a response variable.
method Investigation of strong identifiability, convergence rates, and posterior contraction behavior in finite mixture of regression models.
result Theoretical findings on conditions for strong identifiability and rates of convergence in regression mixture models.
We introduce RNADE, a new model for joint density estimation of real-valued vectors. Our model calculates the density of a datapoint as the product of one-dimensional conditionals modeled using mixture density networks with shared parameters. RNADE learns a distributed representation of the data, while having a tractab…
Algorithm aligns 3D density maps using Wasserstein distance.
problem Aligning 3D density maps in cryogenic electron microscopy.
method Minimizing 1-Wasserstein distance after rigid transformation using Bayesian optimization.
result Improved accuracy and efficiency in protein molecule alignment.
Bayesian approach learns nonparametric mixture components from heterogeneous data.
problem Realistic modeling of heterogeneous data populations with nonparametric mixture components.
method Bayesian nonparametric modeling using Dirichlet process mixture priors.
result Posterior contraction rates for component densities are nearly polynomial, improving over deconvolution methods.
HIRM models noisy, sparse, heterogeneous relational data using hierarchical clustering and Dirichlet processes.
problem Modeling noisy, sparse, and heterogeneous relational data.
method Hierarchical Chinese restaurant process and Dirichlet process mixture for clustering and modeling relation values.
result HIRM generalizes standard models and discovers relational structure in real-world datasets.
Hierarchical nucleation patterns emerge in deep neural network layers.
problem Understanding the generation of meaningful representations in deep neural networks.
method Analysis of the probability density of ImageNet dataset across hidden layers.
result Density peaks in subsequent layers mirror the semantic hierarchy of concepts, resembling nucleation process.
Federated Learning with L0 constraint improves sparsity and performance.
problem Inherent sparsity in data and models leads to dense models with poor generalizability.
method L0 constraint on model density achieved through probabilistic gates and federated stochastic gradient descent.
result Achieves target sparsity (rho) in FL with minimal loss in statistical performance.
Generative model captures complex dependence in financial data.
problem Complex dependence structure in business and financial data.
method Multivariate generative model with heterogeneous and asymmetric tail dependence.
result Novel moment learning algorithm for scalable parameter estimation.
Proposes PFWCP for multi-agent tasks with privacy and validity guarantees.
problem Challenges in uncertainty quantification for multi-agent settings.
method Personalized federated weighted conformal prediction (PFWCP) combining local density ratio weighting and weighted quantile aggregation.
result Asymptotically valid coverage guarantees for each agent in heterogeneous settings.
T-EMDE bridges the heterogeneity gap between image and text modalities.
problem Finding similarities between image and text modalities with non-related feature spaces.
method Inspired by EMDE, T-EMDE uses sketches for multimodal operations, avoiding self-attention's quadratic complexity.
result T-EMDE achieves state-of-the-art results and reduces model latency.
Bayesian X-Learner calibrates uncertainty and robustness for CATE estimation under heavy-tailed data.
problem Estimating heterogeneous treatment effects with calibrated uncertainty and robustness to heavy-tailed outcomes.
method Bayesian X-Learner using cross-fitted doubly robust pseudo-outcomes and MCMC for a full posterior over CATE.
result Bayesian X-Learner achieves robust and calibrated CATE estimation on real and contaminated data.
This paper examines a heterogeneous beliefs model in which there is a process that is only partially observed by the agents. The economy contains a risky asset producing dividends continuously in time. The dividends are observed by the agents. The dividends are assumed to be a known function of some other unobserved pr…
Bayesian framework for model uncertainty identifies complex heterogeneity without strong assumptions.
problem Identifying complex heterogeneity in factorial data with varying covariates.
method Rashomon Partition Sets (RPS) using l0 prior for robust model uncertainty.
result RPS provides a robust set of models capturing complex heterogeneity without strong assumptions.
A new method predicts electron density accurately from atom-centered models.
problem Predicting electron density accurately from atom-centered models.
method Gradient-based approach to minimize loss function in an optimized sparse feature space.
result Extremely accurate predictions of electron density and total energies.
WDL models density curves using Wasserstein distance and flexible mixture models.
problem Modeling entire distribution and non-negativity constraints.
method Wasserstein distance, Semi-parametric Conditional Gaussian Mixture Models (SCGMM), Majorization-Minimization optimization.
result WDL better characterizes nonlinear dependence of conditional densities.
Multitask Gaussian process regression reduces data generation costs for molecular property prediction.
problem Data bottleneck in training surrogate models for molecular properties.
method Multitask Gaussian process regression over heterogeneous data sources (CC and DFT).
result Predicts at CC-level accuracy with over an order of magnitude reduction in data generation cost.
The distribution of absorbed dose in radionuclide therapy with Lu177 can be approximated by convolving an image of the time-integrated activity distribution with a dose voxel kernel representing different tissue types. This fast but inaccurate approximation is unsuitable for personalised dosimetry because it negle…
Method reconstructs financial networks from aggregate data, revealing critical link density.
problem Reconstructing financial networks from aggregate data is challenging due to unreconstructability phases.
method Random graph generation with desired link density and replicated constraints.
result There is a critical link density below which networks become unreconstructable.
Paper develops efficient recursive learning for multi-channel systems with heterogeneous dynamics.
problem Accurately learning system dynamics in complex, multi-channel systems with nonlinear and noisy data.
method Formulates system as Gaussian process state-space models (GPSSMs), introduces heterogeneous multi-output kernel, and develops recursive inference framework.
result Matches SOTA offline GPSSMs in accuracy with 1/100 runtime, and outperforms SOTA online GPSSMs by 70% in accuracy under noise with 1/20 runtime.
New method uses kernel deviance measures to discover causal relationships in heterogeneous data.
problem Discovering causal relationships in complex, heterogeneous datasets.
method KIIM-HT, a novel score measure based on heterogeneous transformations of RKHS embeddings.
result KIIM-HT outperforms previous methods in causal discovery tasks.
Paper proposes a probabilistic method to handle missing data in decision trees.
problem Handling missing data in decision trees.
method At deployment time, use density estimators to compute expected predictions. At learning time, fine-tune tree parameters to minimize expected prediction loss.
result Effective compared to baselines in experiments.
Transfer learning improves loan recovery rate forecasting under data scarcity.
problem Data scarcity in loan portfolios limits RR modeling accuracy.
method Introduces FT-MDN-Transformer, a mixture-density tabular Transformer architecture for TL.
result FT-MDN-Transformer outperforms baseline models in RR forecasting, especially under covariate and conditional shifts.
Study on dynamic curves with elastic energy and spontaneous curvature.
problem Modeling and analyzing dynamic planar curves with elastic energy.
method Gradient flow of inclination angle, nonlocal quasilinear system, local well-posedness, global existence, convergence.
result Local well-posedness, global existence, convergence of the flow for weak regularity initial data.
Paper characterizes optimal graph clustering limits under a new model.
problem Graph clustering under varying edge density signals.
method Introduced Popularity-Adjusted Block Model (PABM) to address SBM and DCBM limitations.
result Cluster recovery possible even when edge density signals vanish, highlighting local connectivity differences.
Reconstructing weighted networks from partial information is necessary in many important circumstances, e.g. for a correct estimation of systemic risk. It has been shown that, in order to achieve an accurate reconstruction, it is crucial to reliably replicate the empirical degree sequence, which is however unknown in m…
The paper applies Gaussianization to analyze Earth data, simplifying complex multivariate distributions.
problem Challenges in accurately estimating information content in high-dimensional, heterogeneous Earth data.
method Multivariate Gaussianization for robust probability density estimation.
result Validates the method for estimating information-theoretic measures in Earth system data.
Outlier detection amounts to finding data points that differ significantly from the norm. Classic outlier detection methods are largely designed for single data type such as continuous or discrete. However, real world data is increasingly heterogeneous, where a data point can have both discrete and continuous attribute…
This note will extend the research presented in Brown & Rogers (2009) to the case of CRRA agents. We consider the model outlined in that paper in which agents had diverse beliefs about the dividends produced by a risky asset. We now assume that the agents all have CRRA utility, with some integer coefficient of relative…
Quantum time evolution exhibits rich physics, attributable to the interplay between the density and phase of a wave function. However, unlike classical heat diffusion, the wave nature of quantum mechanics has not yet been extensively explored in modern data analysis. We propose that the Laplace transform of quantum tra…
We study cross-country GDP losses due to financial crises in terms of frequency (number of loss events per period) and severity (loss per occurrence). We perform the Loss Distribution Approach (LDA) to estimate a multi-country aggregate GDP loss probability density function and the percentiles associated to extreme eve…
Combines Xgboost and transductive SVM for semi-supervised learning.
problem Improving semi-supervised learning performance with heterogeneous tabular data.
method Proposes an optimization-based ensemble method to adaptively combine Xgboost and transductive SVM.
result Significantly improves classification accuracy over state-of-the-art methods.
MCBP detects boundaries in high-dimensional data using curvature.
problem Boundary detection in high-dimensional data.
method MCBP uses mean curvature to model data manifold curvature.
result MCBP improves clustering performance in complex scenarios.
Counterfactual explanations can be obtained by identifying the smallest change made to a feature vector to qualitatively influence a prediction; for example, from 'loan rejected' to 'awarded' or from 'high risk of cardiovascular disease' to 'low risk'. Previous approaches often emphasized that counterfactuals should be…
We prove a law of large numbers for the loss from default and use it for approximating the distribution of the loss from default in large, potentially heterogenous portfolios. The density of the limiting measure is shown to solve a non-linear SPDE, and the moments of the limiting measure are shown to satisfy an infinit…
We offer mathematical tractability and new insights for a framework of exponential utility with non-negative consumption, a constraint often omitted in the literature giving rise to economically unviable solutions. Specifically, using the Kuhn-Tucker theorem and the notion of aggregate state price density (Malamud and …
We study a model of wealth dynamics [Bouchaud and Mézard 2000, \emph{Physica A} \textbf{282}, 536] which mimics transactions among economic agents. The outcomes of the model are shown to depend strongly on the topological properties of the underlying transaction network. The extreme cases of a fully connected and a ful…
This paper introduces an agent-based artificial financial market in which heterogeneous agents trade one single asset through a realistic trading mechanism for price formation. Agents are initially endowed with a finite amount of cash and a given finite portfolio of assets. There is no money-creation process; the total…
The Stochastic Block Model (SBM) is a widely used random graph model for networks with communities. Despite the recent burst of interest in recovering communities in the SBM from statistical and computational points of view, there are still gaps in understanding the fundamental information theoretic and computational l…
VSAE learns from missing heterogeneous data by modeling latent dependencies.
problem Learning from partially-observed heterogeneous data with missingness.
method Variational selective autoencoder (VSAE) models joint distribution of observed, unobserved, and missing data.
result VSAE improves over state-of-the-art models in data generation and imputation tasks.
We present and analyze a model for the evolution of the wealth distribution within a heterogeneous economic environment. The model considers a system of rational agents interacting in a game theoretical framework, through fairly general assumptions on the cost function. This evolution drives the dynamic of the agents i…
With the improvement of medical data capturing, vast amount of continuous patient monitoring data, e.g., electrocardiogram (ECG), real-time vital signs and medications, become available for clinical decision support at intensive care units (ICUs). However, it becomes increasingly challenging to model such data, due to …