HERA improves PLL by integrating heterogeneous loss and sparse-low-rank regularization.
problem Learning from data with partial labels.
method Combines heterogeneous loss and sparse-low-rank regularization.
result Achieves superior performance on artificial and real-world data.
Federated learning calibrates insurance indices from renewable energy producers' data.
problem Calibrating parametric insurance indices under heterogeneous renewable energy production losses.
method Federated learning framework using Tweedie GLMs and distributed optimization.
result Federated learning recovers comparable index coefficients under moderate heterogeneity.
The goal of this paper is to study organized flocking behavior and systemic risk in heterogeneous mean-field interacting diffusions. We illustrate in a number of case studies the effect of heterogeneity in the behavior of systemic risk in the system, i.e., the risk that several agents default simultaneously as a result…
We consider the problem of estimating a low-rank matrix from a noisy observed matrix. Previous work has shown that the optimal method depends crucially on the choice of loss function. In this paper, we use a family of weighted loss functions, which arise naturally for problems such as submatrix denoising, denoising wit…
Proposes a method to ensure low losses across all subpopulations in large datasets.
problem Standard practice of minimizing average loss fails to guarantee low losses across all subpopulations in heterogeneous datasets.
method Convex procedure that controls worst-case performance over all subpopulations of a given size with finite-sample convergence guarantees.
result Empirically, the worst-case procedure learns models that do well against unseen subpopulations.
A new framework for time series forecasting that adapts to varying patterns.
problem Forecasting multivariate time series with predictive heterogeneity.
method Validation-driven clustering framework that applies specialization based on out-of-sample predictive performance.
result Improves robustness to heavy-tailed errors and local anomalies.
A new tensor model captures complex structures in heterogeneous data.
problem Complex joint and discriminative structures in heterogeneous datasets.
method Double core tensor factorization with smoothing loss functions and linearized ADMM.
result The model accurately estimates factors even with missing entries.
S2M optimizes mining for diverse data subpopulations.
problem Scalability and uniformity in training sets with many labels and diverse data.
method Doubly-stochastic mining (S2M) computes per-example and minibatch losses on hardest labels/examples.
result S2M ensures good performance across all data subpopulations.
We study the impact of contagion in a network of firms facing credit risk. We describe an intensity based model where the homogeneity assumption is broken by introducing a random environment that makes it possible to take into account the idiosyncratic characteristics of the firms. We shall see that our model goes behi…
DFFL tackles federated learning with heterogeneous objectives and constraints.
problem Federated learning with clients having different objectives and feasible regions.
method Derived heterogeneity bounds for cost-vector distances and support-function/shape-distance terms. Lifted pointwise bounds to local-versus-federated excess-risk comparison.
result Federation is beneficial when the statistical advantage of pooling exceeds a client-specific heterogeneity penalty.
Paper extends FFT-based differential privacy method to heterogeneous compositions.
problem Computing accurate differential privacy guarantees for mixed mechanisms.
method Uses Fast Fourier Transform (FFT) for error analysis and parameter selection.
result Provides tighter bounds for heterogeneous compositions compared to homogeneous cases.
New algorithm tackles heterogeneous curvature in online convex optimization.
problem Adversarial bandit convex optimization with varying curvature.
method Developed an adaptive algorithm that learns curvature on the fly.
result Achieves optimal regret bounds even with heterogeneous curvature.
We propose a novel approach for loss reserving based on deep neural networks. The approach allows for joint modeling of paid losses and claims outstanding, and incorporation of heterogeneous inputs. We validate the models on loss reserving data across lines of business, and show that they improve on the predictive accu…
Bayesian approach for estimating heterogeneous treatment effects in RDD designs.
problem Heterogeneity in treatment effects in RDD designs can lead to misleading conclusions.
method Direct Bayesian Additive Regression Trees (BART) for modeling heterogeneous treatment effects.
result Flexibly captures complicated structures of heterogeneous treatment effects as a function of covariates.
Study optimal investment decisions for diverse risk-tolerant agents.
problem Optimizing investment choices for agents with varying risk preferences.
method Characterizes optimal behavior using certainty equivalents and lognormal risks.
result Derives optimal decision menus under known and uncertain preference distributions.
We study cross-country GDP losses due to financial crises in terms of frequency (number of loss events per period) and severity (loss per occurrence). We perform the Loss Distribution Approach (LDA) to estimate a multi-country aggregate GDP loss probability density function and the percentiles associated to extreme eve…
Proposes P-learner for estimating treatment effects with proxy variables.
problem Estimating treatment effect heterogeneity in settings with unverifiable exchangeability.
method Two-stage loss function for learning heterogeneous treatment effects with proxy variables.
result P-learner satisfies an oracle bound on estimated error.
Investors suffer welfare loss despite having better information.
problem Welfare loss among investors with absolute information advantages.
method Examined financial markets with heterogenous investors and objective measures of welfare.
result Investors incur welfare loss even with better information, revealing a double loss phenomenon.
New bucketing scheme improves Byzantine robustness for heterogeneous data.
problem Byzantine attacks on federated learning with heterogeneous data.
method Bucketing scheme to adapt robust algorithms to non-iid data.
result Bucketing scheme ensures convergence against Byzantine attacks.
Proposes HetSANN for learning heterogeneous graph structures without meta-paths.
problem Learning low-dimensional vector space of heterogeneous information networks.
method Implicitly represents heterogeneous information through entity space transformation and attention mechanism.
result Significant improvements over state-of-the-art solutions on public datasets.
Proposes Deep LTMLE for estimating dynamic treatment effects in longitudinal studies.
problem Estimating counterfactual mean outcomes under dynamic treatment policies in longitudinal settings.
method Uses a transformer architecture with temporal-difference learning for initial estimation, followed by TMLE correction and statistical inference.
result Demonstrates superior performance in complex, long-term scenarios compared to existing methods.
A new method for designing accurate emulators using deep learning with interval calibration.
problem Designing accurate emulators for scientific processes with modern machine learning methods.
method Learn-by-Calibrating (LbC) approach based on interval calibration.
result Significant improvements in generalization error over widely-used loss functions.
New method resolves causal heterogeneity by defining a resolution profile.
problem Causal subgroup analyses often oversimplify heterogeneity into a small number of groups.
method Introduces a resolution profile as a functional of the causal feature law, using Bayesian-bootstrap inference.
result Shows that the resolution profile is a continuous path with discontinuities at knots, providing integer-valued subgroup numbers.
CAP adapts optimization to class attributes for better fairness.
problem Heterogeneities across classes impede classification performance.
method CAP generates class-specific learning strategies based on attributes.
result CAP improves over naive approach and is competitive with prior art.
LARP filters data to protect model performance across various learners.
problem Protecting model accuracy in public datasets with diverse learners.
method Formalizes and analyzes LARP, a robust data prefiltering method.
result LARP provides guarantees on worst-case loss over a set of learners, with some performance trade-off.
The paper introduces a privacy-preserving method for estimating treatment effects that maintains accuracy.
problem Estimating heterogeneous treatment effects in sensitive data while protecting privacy.
method A general meta-algorithm for CATE estimation with differential privacy guarantees, using sample splitting and parallel composition.
result The meta-algorithm maintains accuracy even with differential privacy, showing that most accuracy loss is due to variance increase.
We prove a law of large numbers for the loss from default and use it for approximating the distribution of the loss from default in large, potentially heterogenous portfolios. The density of the limiting measure is shown to solve a non-linear SPDE, and the moments of the limiting measure are shown to satisfy an infinit…
AEGCN uses autoencoder constraints to improve graph node classification.
problem Node classification on graph domains with reduced information loss.
method Autoencoder-constrained graph convolutional network (AEGCN).
result Adding autoencoder constraints significantly improves graph convolutional network performance.
Operational risk is the risk relative to monetary losses caused by failures of bank internal processes due to heterogeneous causes. A dynamical model including both spontaneous generation of losses and generation via interactions between different processes is presented; the efforts made by the bank to avoid the occurr…
PIE-PINN estimates elastic properties from noisy, low-res displacement data.
problem Estimating heterogeneous elastic properties from low-resolution, noisy data.
method Probabilistic Physics-Informed Neural Network (PIE-PINN) framework combining B-spline and hierarchical scale model.
result Robust estimation of Young's modulus and Poisson's ratio from noisy, low-resolution displacement data.
The paper proposes a method to estimate heterogeneous treatment effects using pretraining strategies.
problem Estimating conditional average treatment effects (CATE) in the presence of many covariates.
method The approach leverages prognostic factors that also predict treatment effect heterogeneity, using the R-learner framework.
result The proposed method improves estimation accuracy and power for detecting treatment effect heterogeneity.
Robust X-Learner improves HTE estimation in imbalanced and heavy-tailed data.
problem Estimating HTE in imbalanced and heavy-tailed data.
method Integrates γ-divergence objective and Proxy Hessian strategy into gradient boosting.
result Reduces PEHE metric by 98.6% in semi-synthetic Criteo Uplift dataset.
This paper optimizes reinsurance contracts with belief differences between insurer and reinsurer.
problem Dynamic reinsurance design with heterogeneous beliefs under mean-variance framework.
method Modeling surplus process, applying partitioned domain optimization, solving HJB system.
result Optimal reinsurance contracts with belief heterogeneity are more complex than standard contracts.
FairMixRep learns fair representations from mixed data types.
problem Representation learning in mixed numerical and categorical data with fairness constraints.
method Efficient encoder-decoder framework + fairness constraints.
result Excellent performance in preserving information and fairness in mixed data representations.
The well known domain shift issue causes model performance to degrade when deployed to a new target domain with different statistics to training. Domain adaptation techniques alleviate this, but need some instances from the target domain to drive adaptation. Domain generalisation is the recently topical problem of lear…
Estimates treatment effects with machine learning using instruments in A/B tests.
problem Estimating heterogeneous treatment effects with unobserved confounders in A/B tests.
method Develops a statistical learning approach using machine learning methods and auxiliary models.
result Shows robustness of estimated effect model to auxiliary model errors and provides asymptotic normality for parameter estimates.
We propose a dynamical model for the estimation of Operational Risk in banking institutions. Operational Risk is the risk that a financial loss occurs as the result of failed processes. Examples of operational losses are the ones generated by internal frauds, human errors or failed transactions. In order to encompass t…
FedLoRU improves FL efficiency by using low-rank updates.
problem Communication inefficiency and performance reduction in Federated Learning.
method Proposes FedLoRU, a low-rank update framework for FL, which reduces communication costs while maintaining performance.
result FedLoRU achieves convergence rates similar to FedAvg and is robust to heterogeneous and large numbers of clients.
Paper tackles end-to-end training of complex neural networks using DIP method.
problem Training complex heterogeneous neural network models end-to-end.
method Deep Innovation Protection (DIP) method using multiobjective optimization.
result End-to-end training of complex heterogeneous neural network models is possible.
The rising use of deep learning and other big-data algorithms has led to an increasing demand for hardware platforms that are computationally powerful, yet energy-efficient. Due to the amount of data parallelism in these algorithms, high-performance 3D manycore platforms that incorporate both CPUs and GPUs present a pr…
EBM reduces dimensionality for estimating heterogeneous CATEs.
problem Estimating CATEs requires many confounding variables, increasing sample complexity.
method Proposes an EBM that learns a low-dimensional representation of variables.
result EBM representations keep CATE estimates consistent and perform better than other methods.
The paper analyzes systemic risk in an insurance model with multiple business lines and heterogeneous claims.
problem Analyzing systemic risk in a multi-dimensional insurance model with heterogeneous claims.
method A multi-dimensional Lévy process-based renewal risk model with pairwise asymptotic independence (PAI).
result Asymptotic formulas for tail probabilities and systemic risk measures are derived.
Model shows how confidence feedback can lead to different crisis outcomes.
problem Characterizing the impact of economic recessions on different social strata.
method A self-reflexive DSGE model with heterogeneous households, varying parameters to analyze crisis typologies.
result Crisis propagation can be confined to high or low income households, depending on social network structure and income inequality.
A new method reduces preference distortion in LLM alignment.
problem Vulnerability of traditional LLM alignment methods to human preference heterogeneity.
method Sign Estimator: A simple, provably consistent, and efficient estimator using binary classification loss.
result Substantially reduces preference distortion over a panel of simulated personas.
New method for clustering tasks with heterogeneous data.
problem Clustered multitask learning with semiparametric and heterogeneous nuisances.
method Adaptive fused orthogonal estimator with Neyman-orthogonal losses and data-driven fusion penalties.
result Achieves exact clustering recovery and pooled parametric convergence rates.
New estimator robust to adversarial noise and data heterogeneity.
problem Sensitive to adversarial noise and poor performance with heterogeneous data.
method Distributionally robust estimator minimizing worst-case conditional expected loss over adversarial distributions.
result Efficiently finds non-parametric local estimates via convex optimization.
Personalized activity recognition improves performance for diverse users.
problem Poor performance of impersonal algorithms for individual users.
method Personalized activity recognition using deep embeddings from a fully convolutional neural network with triplet loss.
result Novel subject triplet loss provides the best performance overall.
A new Bayesian model improves forecasting for intermittent demand.
problem Sparse observations, cold-start items, and obsolescence in intermittent demand forecasting.
method Hierarchical Bayesian TSB model with partial pooling and calibrated probabilistic configuration.
result TSB-HB achieves the lowest RMSE and RMSSE on the UCI Online Retail dataset.