Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

2815628421,123 · Jun 202019922001200920172026
48 results for Bivariate categorical data

Gradient-based methods can be biased by distributional asymmetries in bivariate categorical data.

problem Gradient-based causal discovery methods can be biased by distributional asymmetries in bivariate categorical data.
method Identified and examined two distributional biases: Marginal Distribution Asymmetry and Marginal Distribution Shift Asymmetry. Employed two simple models to demonstrate and control these biases.
result Gradient-based methods can be biased by distributional asymmetries, and these biases can be controlled.

Kendall transformation converts continuous data into categorical vectors for robust information theory.

problem Handling small number of observations and preserving ranking in continuous data.
method Kendall transformation converts ordered features into categorical vectors of pairwise order relations.
result Kendall transformation makes information theory methods applicable to continuous data robustly.

Proposes bivariate DeepKriging for efficient wind field prediction.

problem Challenges in predicting large-scale bivariate wind fields with high spatial variability and heterogeneity.
method Spatially dependent deep neural network (DNN) with embedding layer using spatial radial basis functions.
result Outperforms traditional cokriging predictors and reduces computation time.

Constructs bivariate quantiles using vine copulas for multivariate analysis.

problem Need for research in multivariate quantiles, especially for bivariate responses.
method Constructs bivariate (conditional) quantiles using vine copula based bivariate regression model with a novel tree sequence graph structure.
result Avoids typical shortfalls of regression like transformations, interactions, collinearity, and quantile crossings.

Study uses a bivariate model to price crude oil futures.

problem Pricing crude oil futures using latent factors and state-space models.
method Modelled short and long term factors as OU processes, estimated using Kalman Filter and maximised Gaussian likelihood.
result Successfully estimated model parameters and factors from WTI Crude Oil NYMEX futures data.

New methods optimize sums of bivariate functions on finite domains.

problem Optimizing functions with multiple arguments that are sums of bivariate functions.
method Measure-valued extensions, 2\ell^2-approximation, entropy-regularization, linear programming, coordinate ascent.
result Tractable problem formulations solvable with various methods.

Several classification methods assume that the underlying distributions follow tree-structured graphical models. Indeed, trees capture statistical dependencies between pairs of variables, which may be crucial to attain low classification errors. The resulting classifier is linear in the log-transformed univariate and b…

2018-06-06abs ↗pdf ↗

We collect well known and less known facts about the bivariate normal distribution and translate them into copula language. In addition, we prove a very general formula for the bivariate normal copula, we compute Gini's gamma, and we provide improved bounds and approximations on the diagonal.

2009-12-15abs ↗pdf ↗

A new Heckman selection model uses a bivariate contaminated normal distribution for more accurate data analysis.

problem Sample selection biases in econometric data analysis.
method Introduces a Heckman selection model using a bivariate contaminated normal distribution and presents an efficient ECM algorithm for parameter estimation.
result The proposed model outperforms normal and Student's t counterparts in real data analysis and simulation studies.

Study classifies mappings of bivariate normal densities, revealing three types with distinct geometric and statistical properties.

problem Understanding the properties of two-component bivariate normal mixtures.
method Classification via A\mathcal{A}-equivalence and statistical analysis.
result Three distinct types of mappings with specific geometric and statistical properties, and upper bounds for the number of modes.

New method infers causal relationships from nonstationary time series data.

problem Challenges in inferring causal relationships from nonstationary time series data.
method Proposes a new class of restricted SCM with time-varying filters and stationary noise, leveraging asymmetry from nonstationarity.
result Demonstrates effectiveness of the proposed methodology on various synthetic and real datasets.

Worst-case bounds on the expected shortfall risk given only limited information on the distribution of the random variables has been studied extensively in the literature. In this paper, we develop a new worst-case bound on the expected shortfall when the univariate marginals are known exactly and additional expert inf…

2017-01-16abs ↗pdf ↗

Our goal in this paper is to propose an alternative risk measure which takes into account the fluctuations of losses and possible correlations between random variables. This new notion of risk measures, that we call Copula Conditional Tail Expectation describes the expected amount of risk that can be experienced given …

2012-05-19abs ↗pdf ↗

In this paper we consider a family of Dirac-type operators on fibration PBP \to B equivariant with respect to an action of an etale groupoid. Such a family defines an element in the bivariant KK theory. We compute the action of the bivariant Chern character of this element on the image of Connes' map ΦΦ in the cyclic…

2005-04-06abs ↗pdf ↗

Study assesses drought and late-frost risks in Bavaria using vine copulas.

problem Assessing risks of late-frost and drought in Bavaria due to climate change.
method Used vine copula models for non-Gaussian and asymmetric dependencies, with univariate and bivariate regression analyses.
result Identified 'at-risk' regions for forest adaptation.

New method improves bivariate causal discovery by accurately estimating cause variable complexity.

problem Improper estimation of cause variable complexity in current MDL-based methods.
method Rate-distortion MDL (RDMDL) using information dimension for cause variable complexity estimation.
result RDMDL achieves competitive performance on Tübingen dataset.

The paper develops deep learning models for personalized treatment rules in survival analysis.

problem Deriving optimal treatment rules for bivariate survival outcomes in randomized trials.
method Adaptive prediction-powered learning using deep neural networks and stochastic policies.
result Maximizes joint survival probability beyond fixed time points (t1,t2)(t_1, t_2).

A method is developed to estimate the parameters of a Levy copula of a discretely observed bivariate compound Poisson process without knowledge of common shocks. The method is tested in a small sample simulation study. Also, the method is applied to a real data set and a goodness of fit test is developed. With the meth…

2012-12-01abs ↗pdf ↗

The study evaluates financial risk using copulas and statistical tests.

problem Validating bivariate forecasts in risk evaluation.
method Using copulas to characterize dependencies, applying statistical tests to validate forecasts, removing heteroskedasticity.
result A Student copula accurately describes financial time series dependencies.

UNTIE learns representations of coupled categorical data.

problem Challenges in learning from unlabeled categorical data with complex couplings.
method UNTIE approach for unsupervised representation learning of heterogeneous couplings.
result UNTIE significantly improves categorical data representations on 25 diverse datasets.

Paper introduces Categorical Normalizing Flows for better handling of categorical data.

problem Limited application of normalizing flows on categorical data due to lack of intrinsic order.
method Categorical Normalizing Flows use continuous transformations to model latent relations in categorical data, optimizing both continuous representation and model likelihood.
result GraphCNF, a permutation-invariant generative model, outperforms state-of-the-art on molecule generation.

Modeling stock returns and volatility using a bivariate gamma generalized Laplace law.

problem Analyzing stock returns and volatility using a new statistical model.
method Maximum likelihood estimation for a bivariate generalized Laplace distribution, simplifying to linear regression.
result Explicit estimators derived with nonstandard convergence rates for certain parameter configurations.

New method distinguishes cause from effect using causal velocity.

problem Inferring causal direction from bivariate data.
method Parametrization of bivariate SCMs in terms of causal velocity, using tools from measure transport.
result Method extends beyond known model classes and requires no assumptions on noise distributions.

Bayesian model improves categorization of explosions from sparse data.

problem Challenges in categorizing explosions from limited data.
method Bayesian update to Event Categorization Matrix model with Bayesian Decision Theory.
result Consistent gains in overall accuracy and lower false negative rates.

Researchers study the conformal geometry of bivariate Gaussian manifolds.

problem Exploring the conformal structure of Fisher-Rao metric on statistical manifolds.
method Determined invariants of the conformal structure of the Fisher-Rao metric on the bivariate Gaussian manifold.
result The conformal holonomy group is SO0(1,6)SO^{0}(1,6) for generic random variables, but SO0(1,4)SO^{0}(1,4) for independent ones.

Study categorizes mutual funds using natural language processing from unstructured data.

problem Categorizing mutual funds using unstructured data for financial analysis.
method Used natural language processing models to classify mutual funds from their investment strategy descriptions.
result High accuracy in categorizing mutual funds using NLP from unstructured data.

Long Short-Term Memory (LSTM) infers the long term dependency through a cell state maintained by the input and the forget gate structures, which models a gate output as a value in [0,1] through a sigmoid function. However, due to the graduality of the sigmoid function, the sigmoid gate is not flexible in representing m…

2019-05-25abs ↗pdf ↗

Categorical bundles provide a natural framework for gauge theories involving multiple gauge groups. Unlike the case of traditional bundles there are distinct notions of triviality, and hence also of local triviality, for categorical bundles. We study categorical principal bundles that are product bundles in the categor…

2015-06-14abs ↗pdf ↗

Study compares exponential and power-law kernels in modeling high-frequency trading data.

problem Modeling high-frequency trading data with specific kernel types.
method Proposes and analyzes two bivariate Hawkes processes with exponential and power-law kernels.
result Identifies strengths and limitations of exponential and power-law kernels for high-frequency trading data.

Paper proposes robust methods to detect and treat outliers in multivariate loss reserving.

problem Distortion of traditional reserving techniques by outliers in past claims data.
method Two robust bivariate chain-ladder techniques: Adjusted Outlyingness and Bagdistance.
result Improved accuracy in estimating outstanding claim liabilities through robust methods.

StructureBoost improves gradient boosting for complex categorical variables efficiently.

problem Efficiently handling complex categorical variables with known structure.
method Two methods to overcome computational obstacles in SCDT enumeration for structured categorical variables.
result StructureBoost outperforms existing packages on complex categorical problems.

Develops a new method for decision trees using categorical variable structure.

problem Lack of structure in treating categorical variables as predictors.
method Introduces a mathematical framework to represent categorical structure and generalizes decision trees to utilize this structure.
result Improves prediction accuracy on weather data using the new method.

The Bivariate Dynamic Contagion Processes (BDCP) are a broad class of bivariate point processes characterized by the intensities as a general class of piecewise deterministic Markov processes. The BDCP describes a rich dynamic structure where the system is under the influence of both external and internal factors model…

2014-05-22abs ↗pdf ↗

We present a technique for clustering categorical data by generating many dissimilarity matrices and averaging over them. We begin by demonstrating our technique on low dimensional categorical data and comparing it to several other techniques that have been proposed. Then we give conditions under which our method shoul…

2015-06-26abs ↗pdf ↗