Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

6.2%12.4%18.6%24.7% · Feb 202619922001200920172026
48 results for non-random sampling

Proposes MGPLL for PL learning with non-random noise.

problem Partial label learning with non-random label noise.
method Bi-directional mapping framework, conditional noise label generation, multi-class predictor, adversarial learning.
result Demonstrates state-of-the-art performance in partial label learning.

The paper proposes a method to assess surrogate heterogeneity in non-randomized data.

problem Lack of methods to evaluate surrogate heterogeneity in non-randomized data.
method Proposes a framework using meta-learners to assess surrogate heterogeneity in real-world data.
result Identifies individuals for whom the surrogate is a valid replacement of the primary outcome.

This paper studies node embeddings of networks, revealing their geometric properties.

problem Understanding the geometric properties of node embeddings in random networks.
method Characterization of ergodic limits, generalization, and convex relaxations of random walk node embedding objectives.
result The optimal node embedding Grammians have rank 1 for a nuclear norm relaxation of the non-randomized objective.

For the pedestrian observer, financial markets look completely random with erratic and uncontrollable behavior. To a large extend, this is correct. At first approximation the difference between real price changes and the random walk model is too small to be detected using traditional time series analysis. However, we s…

2011-08-16abs ↗pdf ↗

Study relaxes identification assumptions for natural direct effects in non-randomized settings.

problem Identifying causal direct effects under unmeasured confounding.
method Developed relaxed conditions for identifying natural direct effects in non-randomized settings.
result Identified natural direct effect under unmeasured confounding conditions.

Theoretical framework explains why few epochs are enough for LLM fine-tuning.

problem Understanding why few epochs are sufficient for LLM fine-tuning.
method Combining early stopping theory with attention-based Neural Tangent Kernel (NTK) for LLMs.
result Formalizes convergence rate of attention-based fine-tuning with respect to sample size.

We present a general framework, the coupled compound Poisson factorization (CCPF), to capture the missing-data mechanism in extremely sparse data sets by coupling a hierarchical Poisson factorization with an arbitrary data-generating model. We derive a stochastic variational inference algorithm for the resulting model …

2017-01-09abs ↗pdf ↗

Independent component analysis (ICA) is a method for recovering statistically independent signals from observations of unknown linear combinations of the sources. Some of the most accurate ICA decomposition methods require searching for the inverse transformation which minimizes different approximations of the Mutual I…

2016-09-22abs ↗pdf ↗

Consider a random smooth Gaussian field G(x):FRG(x):F\to\mathbb{R}, where FF is a compact in Rd\mathbb{R}^d. We derive a formula for average area of a surface generated by the equation G(x)=0G(x)=0 and give some applications. As an auxiliary result we obtain an integral expression for area of a surface induced by zeros of a \e…

2011-02-17abs ↗pdf ↗

Generative AutoEncoders require a chosen probability distribution in latent space, usually multivariate Gaussian. The original Variational AutoEncoder (VAE) uses randomness in encoder - causing problematic distortion, and overlaps in latent space for distinct inputs. It turned out unnecessary: we can instead use determ…

2018-11-12abs ↗pdf ↗

Many real datasets contain values missing not at random (MNAR). In this scenario, investigators often perform list-wise deletion, or delete samples with any missing values, before applying causal discovery algorithms. List-wise deletion is a sound and general strategy when paired with algorithms such as FCI and RFCI, b…

2017-05-25abs ↗pdf ↗

Unified framework for causal inference under sample selection.

problem Causal inference under sample selection with treatment and outcome non-randomness.
method ForestRiesz estimator, Riesz representation framework.
result ForestRiesz estimator yields more stable treatment effect estimates than conventional double machine learning approaches.

Transfer learning improves causal model estimates in small samples.

problem Challenges in estimating individual treatment effects (ITE) from small datasets.
method Treatment Agnostic Representation Networks (TARNet) with transfer learning (TL-TARNet).
result Transfer learning reduces ITE error and bias in small samples.

Paper introduces a neural network training algorithm for noisy data that achieves optimal parameters and replicates real-world behaviors.

problem Theoretical gap between universal approximation theorems and practical machine learning with noisy data.
method Randomized training algorithm for neural networks trained on noisy data samples.
result Trained neural networks achieve optimal parameters and exhibit real-world behaviors like sub-linear complexity and interpolation.

Stock markets are complex systems exhibiting collective phenomena and particular features such as synchronization, fluctuations distributed as power-laws, non-random structures and similarity to neural networks. Such specific properties suggest that markets operate at a very special point. Financial markets are believe…

2013-10-09abs ↗pdf ↗

The paper deals with distribution of singular values of product of random matrices arising in the analysis of deep neural networks. The matrices resemble the product analogs of the sample covariance matrices, however, an important difference is that the population covariance matrices, which are assumed to be non-random…

2020-01-17abs ↗pdf ↗

Bayesian model for cost-effectiveness analysis with subgroup discovery.

problem Statistical challenges in cost-effectiveness analysis, especially with non-random treatment assignment and censored data.
method Developed a nonparametric Bayesian model using Dirichlet and Gamma processes to estimate cost-survival distributions and identify cost-effectiveness subgroups.
result Identified and estimated policy-relevant causal CEA estimands using a Bayesian nonparametric g-computation procedure.

Generative framework improves causal estimation from observational data.

problem Estimating individualized treatment effects from non-randomized data.
method Importance-Weighted Diffusion Distillation (IWDD) combining diffusion models and IPW.
result IWDD achieves state-of-the-art prediction performance and significantly improves causal estimation.

Optimal reinsurance contracts for multiple dependent risks are derived without specific dependency assumptions.

problem Finding optimal reinsurance contracts for multiple dependent risks without assuming their dependency structure.
method Assumes maximal expected utility criterion and independent negotiation of reinsurance for each risk. Derives optimality conditions and shows that under mild assumptions, optimal contracts are classical (non-randomized) type.
result Optimal reinsurance contracts exist and can be classical (non-randomized) type under mild assumptions.

A new method prices time-to-event cash flows using survival analysis.

problem Pricing insurance investment portfolios with time-to-event cash flows.
method Discrete-time survival analysis framework, hazard rate estimators, asymptotic multivariate normality.
result Pricing model yields estimates closer to actual cash flows than non-random models.

This paper considers the ideal gas-like model of trading markets, where each individual is identified as a gas molecule that interacts with others trading in elastic or money-conservative collisions. Traditionally this model introduces different rules of random selection and exchange between pair agents. Real economic …

2009-06-10abs ↗pdf ↗

Study designs for estimating treatment effects in adaptive experiments.

problem Estimating treatment effects under adaptive treatment assignment.
method Propose and analyze IPW and AIPW estimators, establish CLTs under design stability.
result Central limit theorems for IPW and AIPW estimators under design stability.

Financial markets are a typical example of complex systems where interactions between constituents lead to many remarkable features. Here, we show that a pairwise maximum entropy model (or auto-logistic model) is able to describe switches between ordered (strongly correlated) and disordered market states. In this frame…

2012-10-31abs ↗pdf ↗

Proposes SSC for estimating counterfactual survival trajectories from observational data.

problem Challenges in estimating causal effects on time-to-event outcomes from observational data.
method Synthetic Survival Control (SSC) framework for estimating counterfactual hazard trajectories in panel data settings.
result SSC estimates counterfactual hazard trajectories as a weighted combination of observed trajectories from other units.

The paper tackles matrix completion in ultra-sparse sampling, improving imputation accuracy.

problem Matrix completion in ultra-sparse sampling, where each row has only a few entries.
method Estimate row span of matrix or averaged second-moment matrix, normalize and impute missing entries.
result Gradient descent method normalizes and imputes missing entries, achieving low variance and unbiased estimation.

Tests whether a treatment's effect is fully mediated by observed outcomes and identifies causal mechanisms.

problem Understanding how a treatment affects an outcome through intermediate variables.
method Proposes a test to evaluate full mediation and causal mechanism identification, extending to non-randomly assigned treatments.
result A conditionally random treatment is conditionally independent of the outcome given mediators and covariates if full mediation and causal mechanism identification hold.

The price impact for a single trade is estimated by the immediate response on an event time scale, i.e., the immediate change of midpoint prices before and after a trade. We work out the price impacts across a correlated financial market. We quantify the asymmetries of the distributions and of the market structures of …

2017-10-22abs ↗pdf ↗

New algorithm recovers model coefficients and supports from noisy data.

problem Simultaneous estimation and support recovery in linear models with Gaussian noise.
method Projection-based algorithm for STG regularized minimization problem, proving convergence and support recovery guarantees.
result New algorithm outperforms existing methods in support recovery for various data setups.

We empirically analyze the price and liquidity responses to trade signs, traded volumes and signed traded volumes. Utilizing the singular value decomposition, we explore the interconnections of price responses and of liquidity responses across the whole market. The statistical characteristics of their singular vectors …

2017-11-21abs ↗pdf ↗

Given an initial (resp., terminal) probability measure μμ (resp., νν) on Rd\mathbb{R}^d, we characterize those optimal stopping times ττ that maximize or minimize the functional EB0Bτα\mathbb{E} |B_0 - B_τ|^α, α>0α> 0, where (Bt)t(B_t)_t is Brownian motion with initial law B0μB_0\sim μ and with final distribution --once stop…

2017-11-08abs ↗pdf ↗

We analyze cascades of defaults in an interbank loan market. The novel feature of this study is that the network structure and the size distribution of banks are derived from empirical data. We find that the ability of a defaulted institution to start a cascade depends on an interplay of shock size and connectivity. Fu…

2013-10-06abs ↗pdf ↗

A new method uses randomized trials to estimate the strength of unobserved confounding.

problem Unobserved confounding compromises causal conclusions from non-randomized studies.
method Designs a statistical test to detect unobserved confounding strength and estimates a lower bound.
result Estimates an asymptotically valid lower bound on unobserved confounding strength.

As the success of deep learning reaches more grounds, one would like to also envision the potential limits of deep learning. This paper gives a first set of results proving that certain deep learning algorithms fail at learning certain efficiently learnable functions. The results put forward a notion of cross-predictab…

2018-12-16abs ↗pdf ↗