Gradient-based methods can be biased by distributional asymmetries in bivariate categorical data.
problem Gradient-based causal discovery methods can be biased by distributional asymmetries in bivariate categorical data.
method Identified and examined two distributional biases: Marginal Distribution Asymmetry and Marginal Distribution Shift Asymmetry. Employed two simple models to demonstrate and control these biases.
result Gradient-based methods can be biased by distributional asymmetries, and these biases can be controlled.
Proposes a new classifier for causal discovery in categorical data.
problem Causal discovery for categorical data.
method Classification with optimal label permutation (COLP) and simple learning algorithm.
result Favorable performance compared to state-of-the-art methods.
Kendall transformation converts continuous data into categorical vectors for robust information theory.
problem Handling small number of observations and preserving ranking in continuous data.
method Kendall transformation converts ordered features into categorical vectors of pairwise order relations.
result Kendall transformation makes information theory methods applicable to continuous data robustly.
Proposes bivariate DeepKriging for efficient wind field prediction.
problem Challenges in predicting large-scale bivariate wind fields with high spatial variability and heterogeneity.
method Spatially dependent deep neural network (DNN) with embedding layer using spatial radial basis functions.
result Outperforms traditional cokriging predictors and reduces computation time.
Constructs bivariate quantiles using vine copulas for multivariate analysis.
problem Need for research in multivariate quantiles, especially for bivariate responses.
method Constructs bivariate (conditional) quantiles using vine copula based bivariate regression model with a novel tree sequence graph structure.
result Avoids typical shortfalls of regression like transformations, interactions, collinearity, and quantile crossings.
New method improves speed of estimating bivariate functional data.
problem Estimating bivariate functional data at faster rates.
method Adapting to directional regularity of bivariate processes.
result Faster rates of convergence achieved through change-of-basis.
Study uses a bivariate model to price crude oil futures.
problem Pricing crude oil futures using latent factors and state-space models.
method Modelled short and long term factors as OU processes, estimated using Kalman Filter and maximised Gaussian likelihood.
result Successfully estimated model parameters and factors from WTI Crude Oil NYMEX futures data.
New methods optimize sums of bivariate functions on finite domains.
problem Optimizing functions with multiple arguments that are sums of bivariate functions.
method Measure-valued extensions, ℓ2-approximation, entropy-regularization, linear programming, coordinate ascent. result Tractable problem formulations solvable with various methods.
TRA detects causal direction from bivariate data using geometric shapes.
problem Inferring causal direction from observational data is challenging and unreliable.
method TRA compares rank-based copula-standardized residual clouds to detect causal direction.
result TRA is robust and superior in detecting causal direction across various scenarios.
We define parametrized cobordism categories and study their formal properties as bivariant theories. Bivariant transformations to a strongly excisive bivariant theory give rise to characteristic classes of smooth bundles with strong additivity properties. In the case of cobordisms between manifolds with boundary, we pr…
Several classification methods assume that the underlying distributions follow tree-structured graphical models. Indeed, trees capture statistical dependencies between pairs of variables, which may be crucial to attain low classification errors. The resulting classifier is linear in the log-transformed univariate and b…
We collect well known and less known facts about the bivariate normal distribution and translate them into copula language. In addition, we prove a very general formula for the bivariate normal copula, we compute Gini's gamma, and we provide improved bounds and approximations on the diagonal.
A new Heckman selection model uses a bivariate contaminated normal distribution for more accurate data analysis.
problem Sample selection biases in econometric data analysis.
method Introduces a Heckman selection model using a bivariate contaminated normal distribution and presents an efficient ECM algorithm for parameter estimation.
result The proposed model outperforms normal and Student's t counterparts in real data analysis and simulation studies.
Method estimates joint distribution of bivariate outcomes.
problem Modeling dependence between bivariate outcomes.
method Semiparametric distribution regression.
result Method performs similarly or better than alternatives in finite samples.
Study classifies mappings of bivariate normal densities, revealing three types with distinct geometric and statistical properties.
problem Understanding the properties of two-component bivariate normal mixtures.
method Classification via A-equivalence and statistical analysis. result Three distinct types of mappings with specific geometric and statistical properties, and upper bounds for the number of modes.
New method infers causal relationships from nonstationary time series data.
problem Challenges in inferring causal relationships from nonstationary time series data.
method Proposes a new class of restricted SCM with time-varying filters and stationary noise, leveraging asymmetry from nonstationarity.
result Demonstrates effectiveness of the proposed methodology on various synthetic and real datasets.
Worst-case bounds on the expected shortfall risk given only limited information on the distribution of the random variables has been studied extensively in the literature. In this paper, we develop a new worst-case bound on the expected shortfall when the univariate marginals are known exactly and additional expert inf…
Our goal in this paper is to propose an alternative risk measure which takes into account the fluctuations of losses and possible correlations between random variables. This new notion of risk measures, that we call Copula Conditional Tail Expectation describes the expected amount of risk that can be experienced given …
Paper proves global optimality of a simple optimization scheme for learning DAG models.
problem Learning acyclic directed graphical models from data.
method Path-following optimization scheme for bivariate setting.
result Simple optimization scheme globally converges to global minimum.
We show that gamma distributions provide models for departures from randomness since every neighbourhood of an exponential distribution contains a neighbourhood of gamma distributions, using an information theoretic metric topology. We derive also the information geometry of the 3-manifold of McKay bivariate gamma dist…
In this paper we consider a family of Dirac-type operators on fibration P→B equivariant with respect to an action of an etale groupoid. Such a family defines an element in the bivariant K theory. We compute the action of the bivariant Chern character of this element on the image of Connes' map Φ in the cyclic…
Study assesses drought and late-frost risks in Bavaria using vine copulas.
problem Assessing risks of late-frost and drought in Bavaria due to climate change.
method Used vine copula models for non-Gaussian and asymmetric dependencies, with univariate and bivariate regression analyses.
result Identified 'at-risk' regions for forest adaptation.
New method improves bivariate causal discovery by accurately estimating cause variable complexity.
problem Improper estimation of cause variable complexity in current MDL-based methods.
method Rate-distortion MDL (RDMDL) using information dimension for cause variable complexity estimation.
result RDMDL achieves competitive performance on Tübingen dataset.
The paper develops deep learning models for personalized treatment rules in survival analysis.
problem Deriving optimal treatment rules for bivariate survival outcomes in randomized trials.
method Adaptive prediction-powered learning using deep neural networks and stochastic policies.
result Maximizes joint survival probability beyond fixed time points (t1,t2). Causal inference using observational data is challenging, especially in the bivariate case. Through the minimum description length principle, we link the postulate of independence between the generating mechanisms of the cause and of the effect given the cause to quantile regression. Based on this theory, we develop Bi…
A method is developed to estimate the parameters of a Levy copula of a discretely observed bivariate compound Poisson process without knowledge of common shocks. The method is tested in a small sample simulation study. Also, the method is applied to a real data set and a goodness of fit test is developed. With the meth…
The study evaluates financial risk using copulas and statistical tests.
problem Validating bivariate forecasts in risk evaluation.
method Using copulas to characterize dependencies, applying statistical tests to validate forecasts, removing heteroskedasticity.
result A Student copula accurately describes financial time series dependencies.
UNTIE learns representations of coupled categorical data.
problem Challenges in learning from unlabeled categorical data with complex couplings.
method UNTIE approach for unsupervised representation learning of heterogeneous couplings.
result UNTIE significantly improves categorical data representations on 25 diverse datasets.
Paper introduces Categorical Normalizing Flows for better handling of categorical data.
problem Limited application of normalizing flows on categorical data due to lack of intrinsic order.
method Categorical Normalizing Flows use continuous transformations to model latent relations in categorical data, optimizing both continuous representation and model likelihood.
result GraphCNF, a permutation-invariant generative model, outperforms state-of-the-art on molecule generation.
PCA is often used in anomaly detection and statistical process control tasks. For bivariate data, we prove that the minor projection (the least varying projection) of the PCA-rotated data is the most sensitive to distributional changes, where sensitivity is defined by the Hellinger distance between distributions before…
Modeling stock returns and volatility using a bivariate gamma generalized Laplace law.
problem Analyzing stock returns and volatility using a new statistical model.
method Maximum likelihood estimation for a bivariate generalized Laplace distribution, simplifying to linear regression.
result Explicit estimators derived with nonstandard convergence rates for certain parameter configurations.
New method distinguishes cause from effect using causal velocity.
problem Inferring causal direction from bivariate data.
method Parametrization of bivariate SCMs in terms of causal velocity, using tools from measure transport.
result Method extends beyond known model classes and requires no assumptions on noise distributions.
CADM proposes a cluster-specific distance metric for categorical data clustering.
problem Inadequate distance metrics for categorical data, especially varying within clusters.
method Cluster-customized adaptive distance metric for categorical data.
result Achieved competitive performance in categorical data clustering.
Bayesian model improves categorization of explosions from sparse data.
problem Challenges in categorizing explosions from limited data.
method Bayesian update to Event Categorization Matrix model with Bayesian Decision Theory.
result Consistent gains in overall accuracy and lower false negative rates.
Researchers study the conformal geometry of bivariate Gaussian manifolds.
problem Exploring the conformal structure of Fisher-Rao metric on statistical manifolds.
method Determined invariants of the conformal structure of the Fisher-Rao metric on the bivariate Gaussian manifold.
result The conformal holonomy group is SO0(1,6) for generic random variables, but SO0(1,4) for independent ones. Study categorizes mutual funds using natural language processing from unstructured data.
problem Categorizing mutual funds using unstructured data for financial analysis.
method Used natural language processing models to classify mutual funds from their investment strategy descriptions.
result High accuracy in categorizing mutual funds using NLP from unstructured data.
Paper introduces new methods for modeling categorical data.
problem Training generative models on categorical data like text and segmentation.
method Argmax Flows and Multinomial Diffusion models.
result Models outperform existing methods in log-likelihood.
CPML efficiently learns new metrics for categorical data.
problem Metric learning for categorical data.
method CPML (categorical projected metric learning) using Schatten p-norms.
result CPML provides efficient metric learning with improved accuracy.
Long Short-Term Memory (LSTM) infers the long term dependency through a cell state maintained by the input and the forget gate structures, which models a gate output as a value in [0,1] through a sigmoid function. However, due to the graduality of the sigmoid function, the sigmoid gate is not flexible in representing m…
A new method clusters categorical data by learning their optimal order and distance.
problem Clustering categorical data lacks a well-defined metric space.
method Order distance metric learning for categorical data.
result Superior clustering accuracy on categorical and mixed datasets.
Categorical bundles provide a natural framework for gauge theories involving multiple gauge groups. Unlike the case of traditional bundles there are distinct notions of triviality, and hence also of local triviality, for categorical bundles. We study categorical principal bundles that are product bundles in the categor…
Study compares exponential and power-law kernels in modeling high-frequency trading data.
problem Modeling high-frequency trading data with specific kernel types.
method Proposes and analyzes two bivariate Hawkes processes with exponential and power-law kernels.
result Identifies strengths and limitations of exponential and power-law kernels for high-frequency trading data.
Paper proposes robust methods to detect and treat outliers in multivariate loss reserving.
problem Distortion of traditional reserving techniques by outliers in past claims data.
method Two robust bivariate chain-ladder techniques: Adjusted Outlyingness and Bagdistance.
result Improved accuracy in estimating outstanding claim liabilities through robust methods.
StructureBoost improves gradient boosting for complex categorical variables efficiently.
problem Efficiently handling complex categorical variables with known structure.
method Two methods to overcome computational obstacles in SCDT enumeration for structured categorical variables.
result StructureBoost outperforms existing packages on complex categorical problems.
Develops a new method for decision trees using categorical variable structure.
problem Lack of structure in treating categorical variables as predictors.
method Introduces a mathematical framework to represent categorical structure and generalizes decision trees to utilize this structure.
result Improves prediction accuracy on weather data using the new method.
The Bivariate Dynamic Contagion Processes (BDCP) are a broad class of bivariate point processes characterized by the intensities as a general class of piecewise deterministic Markov processes. The BDCP describes a rich dynamic structure where the system is under the influence of both external and internal factors model…
We present a technique for clustering categorical data by generating many dissimilarity matrices and averaging over them. We begin by demonstrating our technique on low dimensional categorical data and comparing it to several other techniques that have been proposed. Then we give conditions under which our method shoul…
New model uses attention for in-context learning of categorical data.
problem Learning from categorical data in context.
method Attention-based network with self-attention and cross-attention layers, using functional gradient descent.
result Model can perform multi-step inference for categorical observations.