New method models portfolios with leptokurtic risk factors using Gram-Charlier expansions.
problem Modeling portfolios with excess kurtosis.
method GC-like expansions of the hyperbolic-secant law to account for leptokurtosis.
result Portfolio distribution with risk factors modeled as GC-like expansions of the HS law.
Probabilistic inference in graphical models is the task of computing marginal and conditional densities of interest from a factorized representation of a joint probability distribution. Inference algorithms such as variable elimination and belief propagation take advantage of constraints embedded in this factorization …
Bayesian nonparametric models for data with heterogeneous particles.
problem Deconvolving data with heterogeneous particles, like voter tallies in elections.
method Nonparametric deconvolution models (NDMs) using two tiers of Dirichlet processes.
result NDMs can recover how factor distributions vary locally for each observation.
Paper proposes a new generative model for discrete distributions using flows on submanifolds.
problem Discretization issues and complex statistical dependencies in discrete data.
method Continuous normalizing flows on factorizing discrete measures, geodesic flow matching.
result Efficient training and broad applicability demonstrated through experiments.
This paper corrects an error in [Keller-Ressel, M. and Steiner T. "Yield curve shapes and the asymptotic short rate distribution in affine one-factor models." Finance and Stochastics 12.2 (2008): 149-172]. The error concerns the correct expression for the boundary between normal and humped yield curve behavior in affin…
New method disentangles correlated factors without independence assumption.
problem Learning disentangled representations from correlated data.
method Hausdorff Factorized Support (HFS) criterion for disentanglement.
result HFS consistently improves disentanglement and recovery across various correlation settings.
Model-based clustering imposes a finite mixture modelling structure on data for clustering. Finite mixture models assume that the population is a convex combination of a finite number of densities, the distribution within each population is a basic assumption of each particular model. Among all distributions that have …
Proposes CSG model to separate semantic and variation factors for OOD prediction.
problem Out-of-distribution examples cause conventional models to mix semantic and variation factors, leading to poor performance.
method Causal Semantic Generative model (CSG) based on causal reasoning, using variational Bayes for efficient learning and prediction.
result CSG can identify semantic factor and improve OOD prediction performance.
New distribution simplifies covariance matrix inference.
problem Efficient inference for covariance matrices in large models.
method Incorporates Inverse G-Wishart distribution for variational message passing.
result Elegant and succinct expression of variational message passing fragments.
Probabilistic graphical models compactly represent joint distributions by decomposing them into factors over subsets of random variables. In Bayesian networks, the factors are conditional probability distributions. For many problems, common information exists among those factors. Adding similarity restrictions can be v…
New model analyzes dynamic correlations in stock returns.
problem Analyzing time-varying correlations in high-dimensional data.
method Dynamic factor correlation model with novel parametrization.
result Model accurately captures heterogeneous heavy-tailed distributions and dependent shocks.
While the Matrix Generalized Inverse Gaussian (MGIG) distribution arises naturally in some settings as a distribution over symmetric positive semi-definite matrices, certain key properties of the distribution and effective ways of sampling from the distribution have not been carefully studied. In this paper…
New tests for identifying the number of latent factors in short panels with small time dimensions.
problem Determining the number of latent factors in short panels with small time dimensions.
method Eigenvalue tests based on variance-covariance matrices of asset returns, with assumptions on spherical errors or instrumental variables for factor betas.
result Established asymptotic distributional results and proposed a novel statistical test for weak factors.
Paper proposes a new algorithm for graph learning with covariance constraints.
problem Graphical models and factor analysis not jointly leveraged in graph learning processes.
method Penalized maximum likelihood estimation of an elliptical distribution with Riemannian optimization.
result Effectiveness of the proposed approach demonstrated on real-world data sets.
We analyze the ideal gas like models of markets and review the different cases where a `savings' factor changes the nature and shape of the distribution of wealth. These models can produce similar distribution of wealth as observed across varied economies. We present a more realistic model where the saving factor can v…
Deep model learns complex latent codes without assuming factor structure.
problem Learning latent codes with complex, non-factorial distributions.
method Deep generative factor analysis with beta process prior and stochastic EM algorithm.
result Preliminary results show model can approximate complex distributions.
The effects of saving and spending patterns on holding time distribution of money are investigated based on the ideal gas-like models. We show the steady-state distribution obeys an exponential law when the saving factor is set uniformly, and a power law when the saving factor is set diversely. The power distribution c…
A new model optimizes portfolios by learning stock return distributions conditioned on factors.
problem Optimizing portfolios with high-dimensional asset-specific factors.
method Conditional Diffusion Transformer architecture linking each asset's return to its factor vector.
result The model outperforms benchmarks in mean-variance and mean-CVaR optimization.
Efficiently learns Single-Index Models with constant factor approximation.
problem Learning Single-Index Models under L22 loss with unknown link functions. method An efficient algorithm using alignment sharpness for optimization.
result Achieves constant factor approximation to optimal loss for various distributions and link functions.
Optimal control in latent factor models uses Tsallis entropy for exploration.
problem Optimal control in models with latent factors.
method Reward exploration with Tsallis entropy and derive q-Gaussian distribution over states. result Optimal policy derived in a model-agnostic setting.
In the standard equilibrium and/or arbitrage pricing framework, the value of any asset is uniquely specified from the belief that only the systematic risks need to be remunerated by the market. Here, we show that, even for arbitrary large economies when the distribution of the capitalization of firms is sufficiently he…
We present a class of flexible and tractable static factor models for the term structure of joint default probabilities, the factor copula models. These high-dimensional models remain parsimonious with pair-copula constructions, and nest many standard models as special cases. The loss distribution of a portfolio of con…
The behavior of many Bayesian models used in machine learning critically depends on the choice of prior distributions, controlled by some hyperparameters that are typically selected by Bayesian optimization or cross-validation. This requires repeated, costly, posterior inference. We provide an alternative for selecting…
New probabilistic model for semi-nonnegative matrix factorization using Skellam distribution.
problem Automatic clustering of semi-nonnegative data.
method Skellam-SNMF model with EM and VBEM algorithms.
result New divergence D and algorithms outperform classic SNMF. Proposes a robust factor analysis for matrix data.
problem Robust factor analysis for matrix data with heavy-tailed or contaminated data.
method Bilinear factor analysis based on the matrix-variate t distribution. result Significantly higher breakdown point than traditional methods.
A parsimonious model reduces over-parameterization in skewed matrix variate mixtures.
problem Over-parameterization in skewed matrix variate mixtures.
method Parsimonious family of 256 models using bilinear factor analyzers constrained over clusters, with AECM algorithm for estimation.
result Extensive simulations and real-world datasets (MNIST, Olivetti faces) demonstrate the method's effectiveness.
Deep learning improves Bayes factor computation for likelihood-free models.
problem Computing Bayes factors for likelihood-free models is challenging.
method Proposes a deep learning estimator of Bayes factors using simulated data.
result Establishes consistency of the Deep Bayes Factor estimator.
Framework LiLY recovers latent causal variables from time-series data under distribution shifts.
problem Learning and correcting models under unknown distribution shifts in time-series data.
method LiLY framework that recovers latent causal variables and identifies their relations from temporal data under different distribution shifts.
result The framework reliably identifies time-delayed latent causal influences from observed variables under different distribution changes.
Collaborative filtering, especially latent factor model, has been popularly used in personalized recommendation. Latent factor model aims to learn user and item latent factors from user-item historic behaviors. To apply it into real big data scenarios, efficiency becomes the first concern, including offline model train…
New portfolios outperform traditional methods by using factor weights.
problem Improving portfolio allocation in markets driven by factors.
method Factor-weighted Dirichlet portfolios outperform uniform Dirichlet portfolios.
result Factor-weighted portfolios outperform uniformly sampled portfolios in market returns.
New model improves DNA methylation data analysis.
problem Analyzing DNA methylation data with complex distributions.
method Doubly non-central beta (DNCB) distribution for non-negative matrix factorization.
result Improves predictive performance and yields meaningful latent representations.
We introduce a class of dependence structures, that we call the Multiple Risk Factor (MRF) dependence structures. On the one hand, the new constructions extend the popular CreditRisk+ approach, and as such they formally describe default risk portfolios exposed to an arbitrary number of fatal risk factors with condition…
Learned factor graphs improve inference from time sequences using neural networks.
problem Inference from time sequences with limited labeled data.
method Combines model-based algorithms and data-driven ML tools for stationary time sequences.
result Learned factor graphs can accurately infer from small training sets.
MGLM models all possible language channel factorizations for improved multilingual generation.
problem Generating multilingual text with flexibility and quality.
method Generative joint distribution model over language channels, marginalizing all possible factorizations.
result MGLM outperforms traditional models in multilingual generation tasks.
New model explains price dynamics of Bitcoin with psychological factors.
problem Understanding price variations in cryptocurrency markets with psychological factors.
method Extended agent-based model with heterogeneous psychological parameters.
result Model shows diverse dynamics based on psychological correlation.
Gradient-based methods can be biased by distributional asymmetries in bivariate categorical data.
problem Gradient-based causal discovery methods can be biased by distributional asymmetries in bivariate categorical data.
method Identified and examined two distributional biases: Marginal Distribution Asymmetry and Marginal Distribution Shift Asymmetry. Employed two simple models to demonstrate and control these biases.
result Gradient-based methods can be biased by distributional asymmetries, and these biases can be controlled.
A common approach to analyze a covariate-sample count matrix, an element of which represents how many times a covariate appears in a sample, is to factorize it under the Poisson likelihood. We show its limitation in capturing the tendency for a covariate present in a sample to both repeat itself and excite related ones…
Optimizes risk measures given known marginal distributions of two unknown factors.
problem Determining an upper bound for spectral risk measures with unknown joint distribution.
method Introduces Maximum Spectral Measure (MSP) as a worst-case risk measure, formulated as an optimization problem with a more general objective function.
result Characterizes the continuity properties of the optimal value function and optimal solution set with respect to marginal distributions.
Study finds significant premium for low-beta stocks in firm-level idiosyncratic return distributions.
problem Understanding the role of common idiosyncratic quantile factors in asset pricing.
method Quantile factor analysis to extract common idiosyncratic quantile factors with asymmetric pricing effects.
result Significant premium for innovations to the lower-tail factor: high-beta stocks outperform low-beta stocks by around 7-8% per year.
Survey of factor analysis, PCA, variational inference, and VAE.
problem Dimensionality reduction and generative modeling of data.
method Variational inference, factor analysis, probabilistic PCA, and VAE.
result Derivation and explanation of ELBO, EM, and closed-form solutions.
We discuss the ideal gas like models of a trading market. The effect of savings on the distribution have been thoroughly reviewed. The market with fixed saving factors leads to a Gamma-like distribution. In a market with quenched random saving factors for its agents we show that the steady state income (m) distributi…
The paper analyzes a five-factor capital market model and facilitates exact simulation.
problem Analyzing and simulating a five-factor capital market model.
method Using a Vasicek interest rate model, mean-reverting excess return, and realized inflation with expectation, the paper derives the necessary distributional results and describes practical methods to overcome rank deficiency.
result Exact simulation from the model can be achieved by sampling from a seven-dimensional normal distribution.
This paper improves credit risk analysis by incorporating state-dependent recovery rates into a factor model.
problem Accurate default forecasting in credit risk analysis.
method Extends a one-factor Gaussian copula model to include state-dependent recovery rates and a common factor.
result The proposed model outperforms other models in default prediction, especially during hectic periods.
Paper identifies latent factors from noisy measurements using tensor decomposition.
problem Identification of latent factors from noisy, correlated measurements.
method Tensor decomposition of third order cross moments, Kruskal theorem, Kotlarski identity, generalized Kruskal rank.
result Full distribution of latent factors and measurement errors identified without injective measurements.
The paper analyzes market risk factors for a mining company using a VAR model with stable distribution.
problem Understanding mid- and long-term dynamics of market risk factors for a mining company.
method Two-dimensional vector autoregressive (VAR) model with α-stable distribution, identifying two regimes.
result Derives dynamics of copper price in PLN, crucial for company risk exposure.
Bayesian method estimates contamination factor for unsupervised anomaly detection.
problem No good methods for estimating contamination factor in unsupervised anomaly detection.
method Bayesian approach using mixture formulation of anomaly detector outputs.
result Estimated contamination factor distribution is well-calibrated and improves anomaly detection performance.
A new method STMF improves missing value prediction using tropical semiring.
problem Limited capability of linear models to model complex relations.
method Sparse Tropical Matrix Factorization (STMF) using tropical semiring.
result STMF outperforms NMF on real data, especially in handling extreme values.
A distributed framework for reducing high-dimensional matrix-variate time series data.
problem Reducing dimensionality of high-dimensional, heterogeneous matrix-variate time series data.
method Data partitioning, distributed two-dimensional tensor PCA, aggregation, final PCA, factor matrix computation.
result Preserves latent matrix structure, improves computational efficiency and information utilization.