New framework for cyclic quantum causal models with graph separation property.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Develops a new framework for causal models on cyclic graphs, solving unique solvability issues.
Improves decentralized learning by teleporting active nodes for better convergence.
The importance of nodes in a network constantly fluctuates based on changes in the network structure as well as changes in external interest. We propose an evolving teleportation adaptation of the PageRank method to capture how changes in external interest influence the importance of a node. This framework seamlessly g…
Introduces NCDawareRank, a new ranking framework for networks.
Here we prove the existence of a new type of the world-sheet string singularities - the cusps that are stable during the finite time. These singularities make the emission of the captured massive quantum particle possible in the frames of the author's model suggested earlier. In aggregate, we have a new mechanism of qu…
Paper improves Lasso for S&P500 index tracking with post-selection inference.
New method corrects selection bias in post-selective inference for Group LASSO.
PS-DME evaluates model performance and reliability after data-dependent selection.
We develop a general approach to valid inference after model selection. At the core of our framework is a result that characterizes the distribution of a post-selection estimator conditioned on the selection event. We specialize the approach to model selection by the lasso to form valid confidence intervals for the sel…
Develops methods to adjust prediction set coverage based on post-selection analysis.
New framework for valid hypothesis testing in complex data settings.
Proposes HSIC-Lasso for selective inference in non-linear data.
While statistics and machine learning offers numerous methods for ensuring generalization, these methods often fail in the presence of adaptivity---the common practice in which the choice of analysis depends on previous interactions with the same dataset. A recent line of work has introduced powerful, general purpose a…
The paper discusses methods for interval estimation of coefficients in penalized regression models for insurance data.
"Which Generative Adversarial Networks (GANs) generates the most plausible images?" has been a frequently asked question among researchers. To address this problem, we first propose an \emph{incomplete} U-statistics estimate of maximum mean discrepancy to measure the distribution discrepancy betwee…
Paper simplifies data carving inference with a parametric distribution.
We propose a novel kernel based post selection inference (PSI) algorithm, which can not only handle non-linearity in data but also structured output such as multi-dimensional and multi-label outputs. Specifically, we develop a PSI algorithm for independence measures, and propose the Hilbert-Schmidt Independence Criteri…
New algorithms for sampling in constrained domains without learning rates.
A method to split a data point into two parts that individually cannot reconstruct the whole, but together can.
New samplers minimize KL divergence for constrained and non-Euclidean geometries.
Measuring divergence between two distributions is essential in machine learning and statistics and has various applications including binary classification, change point detection, and two-sample test. Furthermore, in the era of big data, designing divergence measure that is interpretable and can handle high-dimensiona…
Study examines inference methods after variable selection in Cox models.
Proposes MinPEN framework for estimating relationships in multivariate models.
A decentralized online quantum cash system, called qBitcoin, is given. We design the system which has great benefits of quantization in the following sense. Firstly, quantum teleportation technology is used for coin transaction, which prevents from the owner of the coin keeping the original coin data even after sending…
We investigate learning of the differential geometric structure of a data manifold embedded in a high-dimensional Euclidean space. We first analyze kernel-based algorithms and show that under the usual regularizations, non-probabilistic methods cannot recover the differential geometric structure, but instead find mostl…
ParPIC clusters directed graphs using random walks and diffusion operators.
New method splits unknown covariance Gaussians into independent parts.
New method for estimating high-dimensional binary time series coefficients.
Finding statistically significant high-order interaction features in predictive modeling is important but challenging task. The difficulty lies in the fact that, for a recent applications with high-dimensional covariates, the number of possible high-order interaction features would be extremely large. Identifying stati…
Quantum game theory, whatever opinions may be held due to its abstract physical formalism, have already found various applications even outside the orthodox physics domain. In this paper we introduce the concept of a quantum auction, its advantages and drawbacks. Then we describe the models that have already been put f…
New method clusters directed and undirected graphs without losing directional information.
Due to the increasing availability of high-dimensional empirical applications in many research disciplines, valid simultaneous inference becomes more and more important. For instance, high-dimensional settings might arise in economic studies due to very rich data sets with many potential covariates or in the analysis o…
We address the problem of non-parametric multiple model comparison: given candidate models, decide whether each candidate is as good as the best one(s) or worse than it. We propose two statistical tests, each controlling a different notion of decision errors. The first test, building on the post selection inference…
We propose a statistical inference framework for the component-wise functional gradient descent algorithm (CFGD) under normality assumption for model errors, also known as -Boosting. The CFGD is one of the most versatile tools to analyze data, because it scales well to high-dimensional data sets, allows for a very…
Estimates true Sharpe ratio of selected assets with various methods.
Study on estimating causal effects with limited data and multiple environments.
Paper proposes a new method combining random forests and Lasso selection.
Ever since the proof of asymptotic normality of maximum likelihood estimator by Cramer (1946), it has been understood that a basic technique of the Taylor series expansion suffices for asymptotics of -estimators with smooth/differentiable loss function. Although the Taylor series expansion is a purely deterministic …
Machine learning can help us in solving problems in the context big data analysis and classification, as well as in playing complex games such as Go. But can it also be used to find novel protocols and algorithms for applications such as large-scale quantum communication? Here we show that machine learning can be used …
Refining one's hypotheses in the light of data is a common scientific practice; however, the dependency on the data introduces selection bias and can lead to specious statistical analysis. An approach for addressing this is via conditioning on the selection procedure to account for how we have used the data to generate…
Selective inference for group lasso estimators across various distributions and covariates.
We study the convergence of the predictive surface of regression trees and forests. To support our analysis we introduce a notion of adaptive concentration for regression trees. This approach breaks tree training into a model selection phase in which we pick the tree splits, followed by a model fitting phase where we f…
In this paper, we provide efficient estimators and honest confidence bands for a variety of treatment effects including local average (LATE) and local quantile treatment effects (LQTE) in data-rich environments. We can handle very many control variables, endogenous receipt of treatment, heterogeneous treatment effects,…
DebiNet uses over-parameterized neural networks to improve linear model performance and debiasing.
Reweighted ALPS improves sampling from multimodal distributions using warm start points.
DL/FBF improves GPSR solutions by selecting compact, generalising expressions.
Proposes DR-ME test for interpretable distributional treatment effects.