Federated learning calibrates insurance indices from renewable energy producers' data.
problem Calibrating parametric insurance indices under heterogeneous renewable energy production losses.
method Federated learning framework using Tweedie GLMs and distributed optimization.
result Federated learning recovers comparable index coefficients under moderate heterogeneity.
This paper characterizes the equilibrium in a continuous time financial market populated by heterogeneous agents who differ in their rate of relative risk aversion and face convex portfolio constraints. The model is studied in an application to margin constraints and found to match real world observations about financi…
R2P method identifies homogeneous and heterogeneous subgroups for better treatment effect estimation.
problem Current subgroup analysis methods are weak in identifying homogeneous and heterogeneous subgroups and lack confidence estimates.
method R2P uses an arbitrary ITE estimator and quantifies uncertainty robustly.
result R2P produces more homogeneous and heterogeneous partitions than other methods.
Efficiently solves heterogeneous QPs by reducing variables using instance-specific projections.
problem Solving high-dimensional quadratic programming problems efficiently.
method Data-driven framework with a graph neural network generating projections tailored to each QP instance.
result Produces high-quality solutions with reduced computation time, outperforming existing methods.
Productivity and credit limits affect aggregate production in non-monotonic ways.
problem Understanding how aggregate production is influenced by individual characteristics and financial constraints.
method Analytical proof of non-monotonic effects of productivity and credit limits on aggregate production in a general equilibrium model.
result Equilibrium aggregate production can be non-monotonic in both individual productivity and credit limit.
Method for understanding heterogeneous treatment effects in complex causal graphs.
problem Heterogeneity and comorbidity in healthcare problems.
method Developed a new approach to characterize heterogeneous causal effects (HCEs) in graphical contexts, including heterogeneous causal graphs (HCGs) with confounders and mediators.
result Established theoretical forms and properties of HCEs in linear and nonlinear models, and developed interactive structural learning for estimation.
HeteroFL trains diverse clients with varying capabilities efficiently.
problem Training models on clients with different computational and communication capacities.
method Proposes HeteroFL framework to adaptively distribute subnetworks based on client capabilities.
result Adaptive subnetwork distribution leads to efficient computation and communication.
We develop a behavioral asset pricing model in which agents trade in a market with information friction. Profit-maximizing agents switch between trading strategies in response to dynamic market conditions. Due to noisy private information about the fundamental value, the agents form different evaluations about heteroge…
Proposes a flexible deep learning model for complex distributions.
problem Complex shapes, strong skews, and multiple modes in output variable distributions.
method Uncountable Mixture of Asymmetric Laplacians (UMAL) deep learning framework.
result UMAL can estimate heterogeneous distributions without strong assumptions.
New method for robust prediction valid under non-exchangeable data.
problem Non-exchangeability in data violates robustness of common CP methods.
method Introduces efficient CP approach for general non-exchangeable data.
result Produces provably valid confidence sets for non-exchangeable data.
Spatially-aware model improves earthquake hazard assessment accuracy.
problem Misrepresentation of seismic effects across diverse landscapes.
method Causal Bayesian network with Gaussian Processes and normalizing flows.
result Achieves up to 35.2% AUC improvement over existing methods.
Method improves robustness and generalizability of CATE estimation.
problem Lack of external validity in site-specific models for diverse populations.
method Minimax-regret framework with robust optimization.
result Interpretable closed-form solution for generalizable CATE model.
This paper surveys techniques to personalize federated learning models.
problem Personalized models outperform shared models for some clients, reducing participation.
method Surveys recent research on personalizing federated learning models.
result Personalization techniques improve model performance for individual clients.
This paper examines a heterogeneous beliefs model in which there is a process that is only partially observed by the agents. The economy contains a risky asset producing dividends continuously in time. The dividends are observed by the agents. The dividends are assumed to be a known function of some other unobserved pr…
How do macro-financial shocks affect investor behavior and market dynamics? Recent evidence on experience effects suggests a long-lasting influence of personally experienced outcomes on investor beliefs and investment, but also significant differences across older and younger generations. We formalize experience-based …
Improved forecasting of investment dynamics across heterogeneous panels using a two-stage model.
problem Forecasting investment dynamics in heterogeneous panels with varying dynamics.
method Two-stage architecture: global pooled AR(1) for shared persistence, local models for residual dynamics.
result Significant improvement in out-of-sample R2 from 0.630 to 0.677, with a gain of 0.047. This work generates synthetic EHRs with privacy guarantees for machine learning tasks.
problem Privacy concerns and heterogeneity in EHR data limit their use in machine learning.
method Generative Adversarial Networks (GANs) with differential privacy (DP) for synthetic data generation.
result Synthetic EHRs maintain performance close to real data, even with DP applied.
New model identifies microbial subcommunities robustly, accounting for cross-sample heterogeneity.
problem Inference in LDA is sensitive to the number of subcommunities and often creates artificial ones.
method Incorporates logistic-tree normal (LTN) model into LDA to account for cross-sample heterogeneity.
result Restores robustness of inference and identifies meaningful subcommunities.
SLiCE learns contextual node embeddings for link prediction in heterogeneous networks.
problem Link prediction requires specific contextual information not captured by static node embeddings.
method Self-supervised pre-training with localized attention mechanisms.
result SLiCE significantly outperforms existing methods on link prediction tasks.
This paper characterizes and explains the disagreement between two graph embedding methods.
problem Understanding why two popular graph embedding methods produce different results.
method End-to-end analysis of ASE-LSE latent subspaces, proving conditions for agreement and disagreement.
result No maximal-disagreement graph exists; disagreement is strictly below its theoretical ceiling.
We propose a novel approach for loss reserving based on deep neural networks. The approach allows for joint modeling of paid losses and claims outstanding, and incorporation of heterogeneous inputs. We validate the models on loss reserving data across lines of business, and show that they improve on the predictive accu…
Framework harmonizes EHR data across institutions for better analysis.
problem Heterogeneity of medical codes and terminologies hinder EHR data analysis.
method MASH (Multi-source Automated Structured Hierarchy) uses neural optimal transport and learned hyperbolic embeddings to align and structure EHR data.
result MASH generates interpretable hierarchical graphs for unstructured local laboratory codes.
Paper proposes a new method for fast matrix completion.
problem Challenges in matrix completion, especially for images with heterogeneous data.
method Sparse reverse of principal component analysis.
result The method efficiently reconstructs matrices with missing data.
Unsupervised image regression detects changes in satellite images.
problem Detecting changes in heterogeneous multitemporal satellite images.
method Comparison of affinity matrices and image regression.
result Image regression improves change detection accuracy.
Resource scheduling and coordination is an NP-hard optimization requiring an efficient allocation of agents to a set of tasks with upper- and lower bound temporal and resource constraints. Due to the large-scale and dynamic nature of resource coordination in hospitals and factories, human domain experts manually plan a…
CTSyn generates high-quality synthetic tabular data.
problem Challenges in generating high-quality synthetic tabular data.
method Diffusion-based generative foundation model with autoencoder and conditional latent diffusion.
result CTSyn outperforms existing table synthesizers on standard benchmarks.
Proposes a hybrid model for stock market report classification using graph neural networks.
problem Lack of unified node embeddings for heterogeneous graphs in text datasets.
method Transductive hybrid approach combining unsupervised node representation learning and supervised node classification/edge prediction.
result Demonstrates the model's ability to classify stock market technical analysis reports.
Graph embedding is a central problem in social network analysis and many other applications, aiming to learn the vector representation for each node. While most existing approaches need to specify the neighborhood and the dependence form to the neighborhood, which may significantly degrades the flexibility of represent…
Transfer learning improves loan recovery rate forecasting under data scarcity.
problem Data scarcity in loan portfolios limits RR modeling accuracy.
method Introduces FT-MDN-Transformer, a mixture-density tabular Transformer architecture for TL.
result FT-MDN-Transformer outperforms baseline models in RR forecasting, especially under covariate and conditional shifts.
Convolutional neural networks (CNN) have recently achieved state-of-the-art results in various applications. In the case of image recognition, an ideal model has to learn independently of the training data, both local dependencies between the three components (R,G,B) of a pixel, and the global relations describing edge…
SIMPOL solves complex economic models using numerical methods.
problem Optimizing consumption and savings under uncertainty.
method SIMPOL uses a modular numerical framework combining policy iteration and finite difference schemes.
result SIMPOL produces solutions consistent with economic and mathematical theory.
Paper learns interaction laws from multiple trajectories of heterogeneous systems.
problem Estimating unknown interaction laws from multiple trajectories of heterogeneous systems.
method Nonparametric learning of interaction kernels based on pairwise distances, with convergence guarantees in L2 space. result Estimators converge at optimal min-max rate for 1-dimensional nonparametric regression.
We introduce the anti-profile Support Vector Machine (apSVM) as a novel algorithm to address the anomaly classification problem, an extension of anomaly detection where the goal is to distinguish data samples from a number of anomalous and heterogeneous classes based on their pattern of deviation from a normal stable c…
PINE embeds graph nodes flexibly, capturing any neighbor dependency.
problem Learning flexible node representations from graph neighborhoods.
method PINE uses partial permutation invariant set functions to capture any possible neighbor dependencies.
result PINE outperforms state-of-the-art methods on various graph learning tasks.
Financial markets are notoriously complex environments, presenting vast amounts of noisy, yet potentially informative data. We consider the problem of forecasting financial time series from a wide range of information sources using online Gaussian Processes with Automatic Relevance Determination (ARD) kernels. We measu…
Trans-GLMC tackles source heterogeneity in transfer learning for structured clusters.
problem Source heterogeneity makes it hard to use multiple related auxiliary sources effectively.
method Trans-GLMC constructs clusters of sources, then combines global fusion, within-cluster refinement, and target debiasing.
result Improves facility-specific prediction and identifies interpretable communities of hospitals with mutual transferability.
A new method estimates treatment effects in mixed groups, improving accuracy.
problem Estimating treatment effects in mixed groups with heterogeneous responses.
method PCM (pre-cluster and merge) approach for nonparametric estimation.
result Asymptotic consistency and significant improvement in accuracy over existing methods.
New method uses untrusted data for more precise causal analysis.
problem Causal questions with limited trusted data.
method Incorporates untrusted data and trains richer models.
result Tighter, sounder prediction intervals.
New method reduces labeler costs by aggregating predictions from local classifiers.
problem Reduce labeler costs in multiclass classification.
method Model K-class classification using smaller classifiers trained on subsets of tasks. result Near-optimal scheme for designing classifier configurations reduces labeler costs.
This note will extend the research presented in Brown & Rogers (2009) to the case of CRRA agents. We consider the model outlined in that paper in which agents had diverse beliefs about the dividends produced by a risky asset. We now assume that the agents all have CRRA utility, with some integer coefficient of relative…
PEHRT harmonizes EHR data for translational research.
problem Barriers in using EHR data for translational research.
method Common pipeline including open-source code, visualization tools, and detailed documentation.
result PEHRT harmonizes EHR data to standardized ontologies and generates robust embeddings.
The paper corrects for node degree in spectral clustering using random walk Laplacian.
problem Node degree heterogeneity in spectral clustering.
method Graph spectral embedding using the random walk Laplacian.
result The embedding provides uniformly consistent estimates of degree-corrected latent positions.
We introduce a deep multitask architecture to integrate multityped representations of multimodal objects. This multitype exposition is less abstract than the multimodal characterization, but more machine-friendly, and thus is more precise to model. For example, an image can be described by multiple visual views, which …
Modeling DEX liquidity with heterogeneous LPs and MEV bots.
problem Understanding and predicting the dynamics of decentralized cryptocurrency exchanges.
method Mean-field game approach to model liquidity providers' optimal strategies and interactions.
result Calibrated model produces consistent pool exchange rate dynamics and liquidity evolution.
We propose a novel kinetic exchange model differing from previous ones in two main aspects. First, the basic dynamics is modified in order to represent economies where immediate wealth exchanges are carried out, instead of reshufflings or uni-directional movements of wealth. Such dynamics produces wealth distributions …
Stochastic gradient descent (SGD) is a popular stochastic optimization method in machine learning. Traditional parallel SGD algorithms, e.g., SimuParallel SGD, often require all nodes to have the same performance or to consume equal quantities of data. However, these requirements are difficult to satisfy when the paral…
daep learns from irregular, multimodal astronomical data.
problem Learning from irregular, multimodal astronomical sequences.
method Diffusion Autoencoder with Perceivers (daep) tokenizes, compresses, and reconstructs data.
result daep outperforms VAE and maep baselines in reconstruction and fine-scale structure preservation.
Proposes methods for online conformal prediction with nested prediction sets across multiple confidence levels.
problem Need for uncertainty quantification with multiple confidence levels in diverse applications.
method Online optimization perspective to enforce nestedness of prediction sets while controlling quantile estimation error.
result Achieves stable coverage across all levels, strictly nested prediction sets, and improved efficiency.