Study separates interventions on a causal Bayesian network using aggregate observations.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Score-based methods fail with isolated components and incorrect mixing proportions.
Aioli unifies language model data mixing methods and improves performance.
Estimates watermarked content proportions in mixed-source texts.
Training on mixed distributions improves test performance even when components are unrelated.
Proposes a proportional masking strategy for better tabular data imputation.
American options in a multi-asset market model with proportional transaction costs are studied in the case when the holder of an option is able to exercise it gradually at a so-called mixed (randomised) stopping time. The introduction of gradual exercise leads to tighter bounds on the option price when compared to the …
Positive-Unlabeled (PU) learning is an analog to supervised binary classification for the case when only the positive sample is clean, while the negative sample is contaminated with latent instances of positive class and hence can be considered as an unlabeled mixture. The objectives are to classify the unlabeled sampl…
Game (Israeli) options in a multi-asset market model with proportional transaction costs are studied in the case when the buyer is allowed to exercise the option and the seller has the right to cancel the option gradually at a mixed (or randomised) stopping time, rather than instantly at an ordinary stopping time. Allo…
Study compares clustering methods for mixed-type data.
In mix-game which is an extension of minority game, there are two groups of agents; group1 plays the majority game, but the group2 plays the minority game. This paper studies the change of the average winnings of agents and volatilities vs. the change of mixture of agents in mix-game model. It finds that the correlatio…
Tissue heterogeneity is a major confounding factor in studying individual populations that cannot be resolved directly by global profiling. Experimental solutions to mitigate tissue heterogeneity are expensive, time consuming, inapplicable to existing data, and may alter the original gene expression patterns. Here we a…
Mix-IRLS solves imbalanced mixed linear regression problems efficiently.
The problem of developing binary classifiers from positive and unlabeled data is often encountered in machine learning. A common requirement in this setting is to approximate posterior probabilities of positive and negative classes for a previously unseen data point. This problem can be decomposed into two steps: (i) t…
Estimates proportions of LLM-generated text in mixed documents.
Modern machine learning techniques can be used to construct powerful models for difficult collider physics problems. In many applications, however, these models are trained on imperfect simulations due to a lack of truth-level information in the data, which risks the model learning artifacts of the simulation. In this …
Spectral methods improve signal recovery in mixed GLMs with precise asymptotics.
The paper proposes methods to estimate positive examples and learn classifiers from mixed data.
We apply conformal flows of metrics restricted to the orthogonal distribution of a foliation to study the question: Which foliations admit a metric such that the leaves are totally geodesic and the mixed scalar curvature is positive? Our evolution operator includes the integrability tensor of , and for the case …
We introduce and study the flow of metrics on a foliated Riemannian manifold , whose velocity along the orthogonal distribution is proportional to the mixed scalar curvature, $\Sc_{\,\rm mix}$. The flow is used to examine the question: When a foliation admits a metric with a given property of $\Sc_{\,\rm mix}$ (…
The paper tackles high-dimensional mixed linear regression with unknown parameters and proposes methods for estimation, confidence intervals, and hypothesis testing.
Improved density estimation for mixed discrete-continuous data.
This paper studies the correlations of the average winnings of agents and the volatilities of systems based on mix-game model which is an extension of minority game (MG). In mix-game, there are two groups of agents; group1 plays the majority game, but the group2 plays the minority game. The results show that the correl…
A new LDA model with covariates for mixed-membership clusters.
A new method for distilling predictions from a teacher model to a student model without original training data.
Due to diverse nature of data acquisition and modern applications, many contemporary problems involve high dimensional datum $\x \in \R^\d$ whose entries often lie in a union of subspaces and the goal is to find out which entries of $\x$ match with a particular subspace $\sU$, classically called \emph {matched subspace…
Identifying components and estimating mixing weights in unlabeled finite mixtures under marginal independence.
We introduce a novel multivariate random process producing Bernoulli outputs per dimension, that can possibly formalize binary interactions in various graphical structures and can be used to model opinion dynamics, epidemics, financial and biological time series data, etc. We call this a Bernoulli Autoregressive Proces…
We introduce a mixture model for censored durations (C-mix), and develop maximum likelihood inference for the joint estimation of the time distributions and latent regression parameters of the model. We consider a high-dimensional setting, with datasets containing a large number of biomedical covariates. We therefore p…
The study prevents model collapse in overparameterized linear regression by mixing real and synthetic labels.
In this paper the problem of optimal derivative design, profit maximization and risk minimization under adverse selection when multiple agencies compete for the business of a continuum of heterogenous agents is studied. The presence of ties in the agents' best-response correspondences yields discontinuous payoff functi…
The paper studies stability of generative models trained on mixed data.
American options are studied in a general discrete market in the presence of proportional transaction costs, modelled as bid-ask spreads. Pricing algorithms and constructions of hedging strategies, stopping times and martingale representations are presented for short (seller's) and long (buyer's) positions in an Americ…
We propose a novel neural sequence prediction method based on \textit{error-correcting output codes} that avoids exact softmax normalization and allows for a tradeoff between speed and performance. Instead of minimizing measures between the predicted probability distribution and true distribution, we use error-correcti…
This paper computes exact posterior distributions of mixture weights in hierarchical Bayesian models.
New sampling methods for constrained and composite distributions.
We argue that the existing regret matchings for Nash equilibrium approximation conduct "jumpy" strategy updating when the probabilities of future plays are set to be proportional to positive regret measures. We propose a geometrical regret matching which features "smooth" strategy updating. Our approach is simple, intu…
Optimal SD improves ridge regression performance strictly and precisely.
AdaPT-GMM improves multiple testing power with covariates.
Paper tackles model collapse in recursive generative models using a weighted training scheme.
Unified framework for stability and generalization of Push-Sum in decentralized learning over directed graphs.
Study long-only minimum variance portfolio in one-factor market with arbitrary sign betas.
Paper proposes a method to estimate true positive proportion without knowing it.
In life-cycle economics the Samuelson paradigm (Samuelson, 1969) states that the optimal investment is in constant proportions out of lifetime wealth composed of current savings and the present value of future income. It is well known that in the presence of credit constraints this paradigm no longer applies. Instead, …
Study uniform rates for estimating Gaussian mixtures without separation assumption.
We consider an optimal control problem of a property insurance company with proportional reinsurance strategy. The insurance business brings in catastrophe risk, such as earthquake and flood. The catastrophe risk could be partly reduced by reinsurance. The management of the company controls the reinsurance rate and div…
Improved KSD test for better detection of differences in distributions.
Paper improves deep learning for instance-level classification from label proportions.