We examine the efficiency of the Asymmetric Power ARCH (APARCH) model in the case where the residuals follow the standardized Pearson type IV distribution. The model is tested with a variety of loss functions and the efficiency is examined via application of several statistical tests and risk measures. The results indi…
A new QHR model extends HR model with a quadratic variance function.
problem Modeling volatility with greater flexibility and stationarity.
method Introducing a quadratic variance function to the HR model, maintaining Markovian property.
result Stationary distribution of the QHR model is Pearson type IV.
Combines cost-sensitive and Neyman-Pearson paradigms for better binary classification.
problem Asymmetric binary classification problems with unequal error severities.
method Develops TUBE-CS algorithm to bridge cost-sensitive and Neyman-Pearson paradigms.
result High-probability control of population type I error.
Adapts Neyman-Pearson classification for both source and target distribution shifts.
problem Minimizing errors while controlling both Type-I and Type-II errors under distribution shifts.
method Derives an adaptive procedure that guarantees improved error rates and adapts to uninformative sources.
result Automatic adaptation to uninformative sources avoids negative transfer.
The market events of 2007-2009 have reinvigorated the search for realistic return models that capture greater likelihoods of extreme movements. In this paper we model the medium-term log-return dynamics in a market with both fundamental and technical traders. This is based on a Poisson trade arrival model with variable…
Most existing binary classification methods target on the optimization of the overall classification risk and may fail to serve some real-world applications such as cancer diagnosis, where users are more concerned with the risk of misclassifying one specific class than the other. Neyman-Pearson (NP) paradigm was introd…
New algorithm controls type I error in NP classification under label noise.
problem Label noise affects NP classification methods, reducing power.
method Proposes a label-noise-adjusted Neyman-Pearson algorithm.
result Improves power while controlling type I error under desired level.
Motivated by problems of anomaly detection, this paper implements the Neyman-Pearson paradigm to deal with asymmetric errors in binary classification with a convex loss. Given a finite collection of classifiers, we combine them and obtain a new classifier that satisfies simultaneously the two following properties with …
Neyman-Pearson testing improves goodness of fit in detecting new physics.
problem Detecting small anomalies in data distributions.
method Employing Neyman-Pearson strategy with a rich parametrized family of models.
result Neyman-Pearson testing is more sensitive to small departures and unbiased towards specific anomalies.
Paper solves the minimal generating set problem for singular Reidemeister moves.
problem Determine minimal generating sets of oriented singular Reidemeister moves.
method Introduced new invariant for singular links to detect type IV moves and provide obstructions.
result Proved exactly 96 distinct inclusion-minimal generating sets for singular moves.
New bounds for Neyman-Pearson region using f-divergences.
problem Bounding the Neyman-Pearson region for hypothesis testing.
method Establishing novel lower and upper bounds using f-divergences. result Best possible lower bound for the Neyman-Pearson boundary using hockey-stick f-divergences. The gain-loss asymmetry, observed in the inverse statistics of stock indices is present for logarithmic return levels that are over 2%, and it is the result of the non-Pearson type auto-correlations in the index. These non-Pearson type correlations can be viewed also as functionally dependent daily volatilities, ext…
New method corrects bias in density ratio estimation for missing data.
problem Missing data bias in density ratio estimation.
method Adapted KLIEP method (M-KLIEP) for MNAR data.
result M-KLIEP restores consistency and minimax optimality.
The article generalizes Pearson correlation to Riemannian manifolds.
problem Analyzing statistical models on non-linear manifolds.
method Reconstitutes Pearson correlation properties and derives a nonlinear generalization.
result Developed the Riemann-Pearson Correlation for manifold analysis.
Value-at-Risk (VaR) and Conditional Value-at-Risk (CVaR) are popular risk measures from academic, industrial and regulatory perspectives. The problem of minimizing CVaR is theoretically known to be of Neyman-Pearson type binary solution. We add a constraint on expected return to investigate the Mean-CVaR portfolio sele…
Develops algorithms for multi-class Neyman-Pearson classification with cost sensitivity.
problem Asymmetric misclassification costs in multi-class classification problems.
method Establishes connection with cost-sensitive learning, proposes two algorithms, extends NP oracle properties.
result Proposes algorithms with theoretical guarantees for multi-class Neyman-Pearson classification.
Motivated by optimal investment problems in mathematical finance, we consider a variational problem of Neyman-Pearson type for law-invariant robust utility functionals and convex risk measures. Explicit solutions are found for quantile-based coherent risk measures and related utility functionals. Typically, these solut…
This paper addresses the challenges in classifying textual data obtained from open online platforms, which are vulnerable to distortion. Most existing classification methods minimize the overall classification error and may yield an undesirably large type I error (relevant textual messages are classified as irrelevant)…
Financial markets analyzed by reducing correlation matrix complexity.
problem Understanding complex financial market correlations.
method Coarse graining Pearson correlation matrices into Guhr matrices by market sectors.
result Significant reduction in the number of relevant variables.
We first study holomorphic isometries from the Poincaré disk into the product of the unit disk and the complex unit n-ball for n≥2. On the other hand, we observe that there exists a holomorphic isometry from the product of the unit disk and the complex unit n-ball into any irreducible bounded symmetric domain …
The paper extends Pearson correlation to multi-variables, useful for noise measurement and feature selection.
problem The standard Pearson correlation coefficient is limited to two variables and doesn't meet the needs for multi-variable analysis.
method The authors use random matrix theory to extend Pearson's correlation coefficient to an arbitrary number of variables.
result The extended correlation coefficient is useful for gauging noise and selecting features, particularly in classification.
Model predicts epileptic seizures with high accuracy using EEG signals.
problem Predicting epileptic seizures with high accuracy for diagnosis and treatment.
method Pearson's product-moment correlation coefficient with a linear classifier on generalized Gaussian modeling.
result 100% effectiveness for sensitivity and specificity greater than 83%.
BGM-IV uses AI to estimate causal effects in complex data.
problem Estimating causal effects in high-dimensional, nonlinear settings with endogeneity.
method Structured latent generative modeling for posterior inference in a causally structured latent space.
result BGM-IV outperforms existing methods in high-dimensional covariate regimes.
Paper discovers valid IVs from data without domain knowledge.
problem Inferring causal effects from observational data with latent confounders.
method Data-driven algorithm based on partial ancestral graphs (PAGs).
result Discovering valid IVs leads to accurate causal effect estimation.
The paper analyzes skewness and kurtosis measures for skew-elliptical distributions.
problem Examining skewness and kurtosis measures for skew-elliptical distributions.
method Deriving exact expressions for skewness and kurtosis measures for skew-elliptical distributions, constructing test statistics, and comparing measures through simulations and real data analysis.
result Exact expressions and test statistics for skewness and kurtosis measures for various skew-elliptical distributions.
For time series comparisons, it has often been observed that z-score normalized Euclidean distances far outperform the unnormalized variant. In this paper we show that a z-score normalized, squared Euclidean Distance is, in fact, equal to a distance based on Pearson Correlation. This has profound impact on many distanc…
A new method learns IV representation from data to estimate causal effects.
problem Inferring causal effects from observational data with latent confounders.
method Disentangled representation learning using Variational AutoEncoder (VAE).
result The proposed method outperforms existing IV-based estimators and VAE-based estimators.
The Neyman-Pearson (NP) paradigm in binary classification seeks classifiers that achieve a minimal type II error while enforcing the prioritized type I error controlled under some user-specified level α. This paradigm serves naturally in applications such as severe disease diagnosis and spam detection, where people h…
DML-IV improves IV regression for learning decision policies by reducing bias.
problem Spurious correlations in offline datasets caused by hidden confounders.
method Double/debiased machine learning (DML) framework to reduce bias in two-stage IV regression.
result DML-IV outperforms state-of-the-art methods and learns high-performing policies.
Ivy combines weak IV candidates to estimate causal effects robustly.
problem Estimating causal effects from observational data using weak or invalid IV candidates.
method Ivy synthesizes multiple weak IV candidates into a robust summary.
result Ivy produces more reliable causal effect estimates compared to allele scores.
Flow IV uses IVs to infer counterfactuals in complex models.
problem Identifying causal effects and counterfactual reasoning in nonseparable outcome models.
method Utilizes instrumental variables and normalizing flows to estimate and infer counterfactual outcomes.
result Identifies a method to make causal inferences from observed data in nonseparable models.
Minimizes indecisions in selective classification to control misclassification rates.
problem Controlling misclassification rates in high-risk scenarios.
method Using indecisions to control misclassification rates, even below Bayes optimal.
result Control of misclassification rates to any user-specified level, even below Bayes optimal.
New RDPC dissimilarity measure improves time series clustering.
problem Improving time series clustering methods for diverse data.
method Combining weighted Pearson correlation with largest element-wise differences.
result RDPC outperforms existing methods in complex datasets.
Entropy measures in their various incarnations play an important role in the study of stochastic time series providing important insights into both the correlative and the causative structure of the stochastic relationships between the individual components of a system. Recent applications of entropic techniques and th…
Characterizes distribution-free rates in unbalanced classification problems.
problem Minimizing error under two different distributions in unbalanced settings.
method Characterizes minimax rates over all pairs of distributions using a geometric condition.
result Identifies a dichotomy between hard and easy classes based on a three-points-separation condition.
This study uses local Gaussian correlation to analyze stock return tails, revealing more sensitive network properties.
problem Misleading results from Pearson correlation in financial networks.
method Local Gaussian correlation coefficient for capturing nonlinear dependence and heavy-tailed distributions.
result Local Gaussian correlation network among negative tails is more sensitive to stock market risks.
Method selects valid IVs from a large set using clustering and test of overidentifying restrictions.
problem Selecting valid instrumental variables from a large set of candidates.
method Agglomerative hierarchical clustering combined with a test of overidentifying restrictions.
result Achieves oracle properties when the largest group of IVs is valid.
The Pearson distance between a pair of random variables X,Y with correlation ρxy, namely, 1-ρxy, has gained widespread use, particularly for clustering, in areas such as gene expression analysis, brain imaging and cyber security. In all these applications it is implicitly assumed/required that the distance …
Develops framework for estimating and improving DTRs with time-varying IV in the presence of unmeasured confounding.
problem Estimating DTRs from observational data with unmeasured confounding.
method Time-varying instrumental variable (IV) framework for estimating and improving DTRs.
result IV-optimal and IV-improved DTRs perform better than DTRs assuming no unmeasured confounding.
We propose a direction of arrival (DOA) estimation method that combines sound-intensity vector (IV)-based DOA estimation and DNN-based denoising and dereverberation. Since the accuracy of IV-based DOA estimation degrades due to environmental noise and reverberation, two DNNs are used to remove such effects from the obs…
Novel quasi-Bayesian method for IV regression using machine learning models.
problem Uncertainty quantification in IV regression with machine learning models.
method Quasi-Bayesian procedure based on kernelized IV models and dual formulation.
result Established minimax optimal contraction rates and scalable inference algorithm.
New RL algorithms correct bias in dynamic data analysis.
problem Dynamic data generation and analysis create endogeneity issues.
method Instrument variable (IV)-based reinforcement learning (RL) algorithms.
result Established theoretical properties of IV-RL algorithms.
Novel method prices call options using Pearson diffusion processes.
problem Pricing European call options with skewness and kurtosis.
method Modeling asset returns with Pearson diffusion processes.
result Proposed method outperforms Black-Scholes and Heston models.
The paper tackles Neyman-Pearson classification control issues.
problem Neyman-Pearson classification's control constraint is hard to satisfy in finite samples.
method Developed refined learning procedures under two accuracy control strategies.
result Proposed methods achieve desired control levels in finite samples.
Study compares GARCH, EWMA, and IV models for GBP/USD and EUR/GBP currency pairs.
problem Predicting 20-day variation in GBP/USD and EUR/GBP currency pairs.
method Applied GARCH, EWMA, and IV models to GBP/USD and EUR/GBP pairs data.
result GARCH models outperform other models in predicting volatility for EUR/GBP, while GARCH with rolling window for GBP/USD.
Let M be an irreducible Riemannian symmetric space. The index i(M) of M is the minimal codimension of a (non-trivial) totally geodesic submanifold of M. The purpose of this note is to determine the index i(M) for all irreducible Riemannian symmetric spaces M of type (II) and (IV).
USP test improves on Pearson's chi-squared and G-test for independence.
problem Deficiencies in Pearson's chi-squared and G-test for independence. method USP test based on U-statistic estimator of population dependence measure. result USP test controls size, handles small cell counts, and detects minimal violations of independence.
This article describes the following results which relate to each other; i) convergence of high dimensional contact structure to codimension one foliation with Reeb component, ii) relation between Nil-type and Sol-type contact submanifolds of S^5, iii) definition of convex Thurston-Bennequin inequality, and iv) general…