A new method for modeling insurance claim frequencies using random proportions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper compares different models for time-to-event analysis.
Study compares Cox model and RSF for predicting patient survival, finding RSF superior in certain scenarios.
A new large-scale tabular benchmark for Learning from Label Proportions.
Proposes a proportional masking strategy for better tabular data imputation.
In this note, we study the utility maximization problem on the terminal wealth under proportional transaction costs and bounded random endowment. In particular, we restrict ourselves to the numéraire-based model and work with utility functions only supporting R+. Under the assumption of existence of consistent price sy…
Study examines how insurance affects households prone to proportional losses, especially those near poverty.
New ridge regression bounds for high-dimensional data without proportional growth.
This work efficiently learns linear threshold functions from label proportions using Gaussian distributions.
New pivoting strategy improves trace norm contraction in low-rank approximation.
The paper analyzes data augmentation for precision matrix estimation in high dimensions.
Study of deep linear neural networks with proportional width and depth.
We investigate the effect of the proportional hazards assumption on prognostic and predictive models of the survival time of patients suffering from amyotrophic lateral sclerosis (ALS). We theoretically compare the underlying model formulations of several variants of survival forests and implementations thereof, includ…
We propose and study kernel conjugate gradient methods (KCGM) with random projections for least-squares regression over a separable Hilbert space. Considering two types of random projections generated by randomized sketches and Nyström subsampling, we prove optimal statistical results with respect to variants of norms …
This paper revisits the classic iterative proportional scaling (IPS) from a modern optimization perspective. In contrast to the criticisms made in the literature, we show that based on a coordinate descent characterization, IPS can be slightly modified to deliver coefficient estimates, and from a majorization-minimizat…
In this paper we study the problem of maximizing expected utility from the terminal wealth with proportional transaction costs and random endowment. In the context of the existence of consistent price systems, we consider the duality between the primal utility maximization problem and the dual one, which is set up on t…
The paper tackles resource allocation for arms with unknown and random rewards, achieving optimal regret bounds.
The Cannon-Thurston map's measures become singular with respect to sphere measures.
We present an optimal investment theorem for a currency exchange model with random and possibly discontinuous proportional transaction costs. The investor's preferences are represented by a multivariate utility function, allowing for simultaneous consumption of any prescribed selection of the currencies at a given term…
ETM models improve efficiency in semi-supervised logistic regression.
The Cannon-Thurston map's pushed measures on the circle are singular with respect to sphere measures.
We consider Bernoulli random variables, which are independent conditional on a common random factor determining their probability distribution. We show that certain expected functionals of the proportion of variables in a given state converge at rate as . Based on these results, we …
In this paper, we consider a numéraire-based utility maximization problem under constant proportional transaction costs and random endowment. Assuming that the agent cannot short sell assets and is endowed with a strictly positive contingent claim, a primal optimizer of this utility maximization problem exists. Moreove…
In this paper, we study the optimal control problem for a company whose surplus process evolves as an upward jump diffusion with random return on investment. Three types of practical optimization problems faced by a company that can control its liquid reserves by paying dividends and injecting capital. In the first pro…
Study Gaussian approximation for deep neural networks with random weights.
Kernel ridge regression (KRR) is a standard method for performing non-parametric regression over reproducing kernel Hilbert spaces. Given samples, the time and space complexity of computing the KRR estimate scale as and respectively, and so is prohibitive in many cases. We prop…
We consider a discrete time financial market with proportional transaction costs under model uncertainty, and study a numéraire-based semi-static utility maximization problem with an exponential utility preference. The randomization techniques recently developed in \cite{BDT17} allow us to transform the original proble…
New findings on kernel regression in the quadratic regime, improving understanding of machine learning models.
This paper studies the utility maximization on the terminal wealth with random endowments and proportional transaction costs. To deal with unbounded random payoffs from some illiquid claims, we propose to work with the acceptable portfolios defined via the consistent price system (CPS) such that the liquidation value p…
In dynamic topic modeling, the proportional contribution of a topic to a document depends on the temporal dynamics of that topic's overall prevalence in the corpus. We extend the Dynamic Topic Model of Blei and Lafferty (2006) by explicitly modeling document level topic proportions with covariates and dynamic structure…
Random Reshuffling outperforms Stochastic Gradient Descent in smooth convex optimization.
Random forest is widely exploited as an ensemble learning method. In many practical applications, however, there is still a significant challenge to learn from imbalanced data. To alleviate this limitation, we propose a deep dynamic boosted forest (DDBF), a novel ensemble algorithm that incorporates the notion of hard …
Random column sampling is not guaranteed to yield data sketches that preserve the underlying structures of the data and may not sample sufficiently from less-populated data clusters. Also, adaptive sampling can often provide accurate low rank approximations, yet may fall short of producing descriptive data sketches, es…
We show that real and imaginary parts of equivariant spherical harmonics on have almost surely a single nodal component. Moreover, if the degree of the spherical harmonic is and the equivariance degree is , then the expected genus is proportional to . Hence if $\fra…
We investigate financial market correlations using random matrix theory and principal component analysis. We use random matrix theory to demonstrate that correlation matrices of asset price changes contain structure that is incompatible with uncorrelated random price changes. We then identify the principal components o…
K-nearest neighbor (kNN) search has wide applications in many areas, including data mining, machine learning, statistics and many applied domains. Inspired by the success of ensemble methods and the flexibility of tree-based methodology, we propose random projection forests (rpForests), for kNN search. rpForests finds …
An investor with constant absolute risk aversion trades a risky asset with general Itô-dynamics, in the presence of small proportional transaction costs. In this setting, we formally derive a leading-order optimal trading policy and the associated welfare, expressed in terms of the local dynamics of the frictionless op…
Using public data (Forbes Global 2000) we show that the asset sizes for the largest global firms follow a Pareto distribution in an intermediate range, that is ``interrupted'' by a sharp cut-off in its upper tail, where it is totally dominated by financial firms. This flattening of the distribution contrasts with a lar…
fcHMRF-LIS controls FDR in neuroimaging data, improving power and scalability.
Paper proposes a method to estimate true positive proportion without knowing it.
We analyze the complexity of Gibbs samplers for inference in crossed random effect models used in modern analysis of variance. We demonstrate that for certain designs the plain vanilla Gibbs sampler is not scalable, in the sense that its complexity is worse than proportional to the number of parameters and data. We thu…
We introduce a semi-parametric Bayesian model for survival analysis. The model is centred on a parametric baseline hazard, and uses a Gaussian process to model variations away from it nonparametrically, as well as dependence on covariates. As opposed to many other methods in survival analysis, our framework does not im…
Kernel methods are powerful and flexible approach to solve many problems in machine learning. Due to the pairwise evaluations in kernel methods, the complexity of kernel computation grows as the data size increases; thus the applicability of kernel methods is limited for large scale datasets. Random Fourier Features (R…
Efficiently estimates linear models robust to corrupted data.
We demonstrate by mathematical analysis and systematic computer simulations that redistribution can lead to sustainable growth in a society. The human capital dynamics of each agent is described by a stochastic multiplicative process which, in the long run, leads to the destruction of individual human capital and the e…
The family of admissible positions in a transaction costs model is a random closed set, which is convex in case of proportional transaction costs. However, the convexity fails, e.g. in case of fixed transaction costs or when only a finite number of transfers are possible. The paper presents an approach to measure risks…
Paper improves deep learning for instance-level classification from label proportions.
We explain theoretically a curious empirical phenomenon: "Approximating a matrix by deterministically selecting a subset of its columns with the corresponding largest leverage scores results in a good low-rank matrix surrogate". To obtain provable guarantees, previous work requires randomized sampling of the columns wi…