Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

2795588361,115 · Jun 202019922001200920172026
48 results for non-randomized data

The paper proposes a method to assess surrogate heterogeneity in non-randomized data.

problem Lack of methods to evaluate surrogate heterogeneity in non-randomized data.
method Proposes a framework using meta-learners to assess surrogate heterogeneity in real-world data.
result Identifies individuals for whom the surrogate is a valid replacement of the primary outcome.

Proposes MGPLL for PL learning with non-random noise.

problem Partial label learning with non-random label noise.
method Bi-directional mapping framework, conditional noise label generation, multi-class predictor, adversarial learning.
result Demonstrates state-of-the-art performance in partial label learning.

We present a general framework, the coupled compound Poisson factorization (CCPF), to capture the missing-data mechanism in extremely sparse data sets by coupling a hierarchical Poisson factorization with an arbitrary data-generating model. We derive a stochastic variational inference algorithm for the resulting model …

2017-01-09abs ↗pdf ↗

This paper studies node embeddings of networks, revealing their geometric properties.

problem Understanding the geometric properties of node embeddings in random networks.
method Characterization of ergodic limits, generalization, and convex relaxations of random walk node embedding objectives.
result The optimal node embedding Grammians have rank 1 for a nuclear norm relaxation of the non-randomized objective.

Study relaxes identification assumptions for natural direct effects in non-randomized settings.

problem Identifying causal direct effects under unmeasured confounding.
method Developed relaxed conditions for identifying natural direct effects in non-randomized settings.
result Identified natural direct effect under unmeasured confounding conditions.

For the pedestrian observer, financial markets look completely random with erratic and uncontrollable behavior. To a large extend, this is correct. At first approximation the difference between real price changes and the random walk model is too small to be detected using traditional time series analysis. However, we s…

2011-08-16abs ↗pdf ↗

Consider a random smooth Gaussian field G(x):FRG(x):F\to\mathbb{R}, where FF is a compact in Rd\mathbb{R}^d. We derive a formula for average area of a surface generated by the equation G(x)=0G(x)=0 and give some applications. As an auxiliary result we obtain an integral expression for area of a surface induced by zeros of a \e…

2011-02-17abs ↗pdf ↗

Stock markets are complex systems exhibiting collective phenomena and particular features such as synchronization, fluctuations distributed as power-laws, non-random structures and similarity to neural networks. Such specific properties suggest that markets operate at a very special point. Financial markets are believe…

2013-10-09abs ↗pdf ↗

Theoretical framework explains why few epochs are enough for LLM fine-tuning.

problem Understanding why few epochs are sufficient for LLM fine-tuning.
method Combining early stopping theory with attention-based Neural Tangent Kernel (NTK) for LLMs.
result Formalizes convergence rate of attention-based fine-tuning with respect to sample size.

New algorithm recovers model coefficients and supports from noisy data.

problem Simultaneous estimation and support recovery in linear models with Gaussian noise.
method Projection-based algorithm for STG regularized minimization problem, proving convergence and support recovery guarantees.
result New algorithm outperforms existing methods in support recovery for various data setups.

A new method prices time-to-event cash flows using survival analysis.

problem Pricing insurance investment portfolios with time-to-event cash flows.
method Discrete-time survival analysis framework, hazard rate estimators, asymptotic multivariate normality.
result Pricing model yields estimates closer to actual cash flows than non-random models.

We analyze cascades of defaults in an interbank loan market. The novel feature of this study is that the network structure and the size distribution of banks are derived from empirical data. We find that the ability of a defaulted institution to start a cascade depends on an interplay of shock size and connectivity. Fu…

2013-10-06abs ↗pdf ↗

Study neural networks by mapping correlations, revealing essential statistics.

problem Understanding information processing in trained neural networks.
method Characterize neural network as distribution transformations, focusing on correlation functions.
result Higher-order correlations are crucial for internal layers, while input layer captures more.

A new method uses randomized trials to estimate the strength of unobserved confounding.

problem Unobserved confounding compromises causal conclusions from non-randomized studies.
method Designs a statistical test to detect unobserved confounding strength and estimates a lower bound.
result Estimates an asymptotically valid lower bound on unobserved confounding strength.

The paper deals with distribution of singular values of product of random matrices arising in the analysis of deep neural networks. The matrices resemble the product analogs of the sample covariance matrices, however, an important difference is that the population covariance matrices, which are assumed to be non-random…

2020-01-17abs ↗pdf ↗

Optimal reinsurance contracts for multiple dependent risks are derived without specific dependency assumptions.

problem Finding optimal reinsurance contracts for multiple dependent risks without assuming their dependency structure.
method Assumes maximal expected utility criterion and independent negotiation of reinsurance for each risk. Derives optimality conditions and shows that under mild assumptions, optimal contracts are classical (non-randomized) type.
result Optimal reinsurance contracts exist and can be classical (non-randomized) type under mild assumptions.

This paper considers the ideal gas-like model of trading markets, where each individual is identified as a gas molecule that interacts with others trading in elastic or money-conservative collisions. Traditionally this model introduces different rules of random selection and exchange between pair agents. Real economic …

2009-06-10abs ↗pdf ↗

Generative framework improves causal estimation from observational data.

problem Estimating individualized treatment effects from non-randomized data.
method Importance-Weighted Diffusion Distillation (IWDD) combining diffusion models and IPW.
result IWDD achieves state-of-the-art prediction performance and significantly improves causal estimation.

Financial markets are a typical example of complex systems where interactions between constituents lead to many remarkable features. Here, we show that a pairwise maximum entropy model (or auto-logistic model) is able to describe switches between ordered (strongly correlated) and disordered market states. In this frame…

2012-10-31abs ↗pdf ↗

Proposes a new model for testing causal structural priors and synthesizing data.

problem Testing and synthesizing causal structural priors using nonparametric knowledge and neural networks.
method Causal Structural Hypothesis Testing (C-SHT) and Causal Structural Variational Hypothesis Testing (C-SVHT) using deep neural networks.
result Demonstrates out-of-distribution generalization error as a proxy for causal structural prior hypothesis testing.

The price impact for a single trade is estimated by the immediate response on an event time scale, i.e., the immediate change of midpoint prices before and after a trade. We work out the price impacts across a correlated financial market. We quantify the asymmetries of the distributions and of the market structures of …

2017-10-22abs ↗pdf ↗

We empirically analyze the price and liquidity responses to trade signs, traded volumes and signed traded volumes. Utilizing the singular value decomposition, we explore the interconnections of price responses and of liquidity responses across the whole market. The statistical characteristics of their singular vectors …

2017-11-21abs ↗pdf ↗

Given an initial (resp., terminal) probability measure μμ (resp., νν) on Rd\mathbb{R}^d, we characterize those optimal stopping times ττ that maximize or minimize the functional EB0Bτα\mathbb{E} |B_0 - B_τ|^α, α>0α> 0, where (Bt)t(B_t)_t is Brownian motion with initial law B0μB_0\sim μ and with final distribution --once stop…

2017-11-08abs ↗pdf ↗

Generative AutoEncoders require a chosen probability distribution in latent space, usually multivariate Gaussian. The original Variational AutoEncoder (VAE) uses randomness in encoder - causing problematic distortion, and overlaps in latent space for distinct inputs. It turned out unnecessary: we can instead use determ…

2018-11-12abs ↗pdf ↗

Accelerated optimization methods improve robustness and privacy in estimation.

problem Improving robustness and privacy in estimation methods.
method Accelerated gradient methods based on Frank-Wolfe and projected gradient descent, with tailored learning rates and Nesterov's momentum.
result Reduction in iteration complexity, leading to stronger statistical guarantees.

Paper introduces a neural network training algorithm for noisy data that achieves optimal parameters and replicates real-world behaviors.

problem Theoretical gap between universal approximation theorems and practical machine learning with noisy data.
method Randomized training algorithm for neural networks trained on noisy data samples.
result Trained neural networks achieve optimal parameters and exhibit real-world behaviors like sub-linear complexity and interpolation.

DSVGD improves federated learning with fewer communication rounds.

problem Federated learning scalability and trustworthiness.
method Distributed Stein Variational Gradient Descent (DSVGD) for non-parametric Bayesian inference.
result DSVGD achieves comparable accuracy and scalability to other methods, with well-calibrated predictions.

Study on random matrices in deep neural networks with IID entries.

problem Distribution of singular values in product of random matrices for deep neural networks.
method Random matrix theory with a streamlined approach for non-Gaussian data.
result Generalization of macroscopic universality property to non-Gaussian data.

Many real datasets contain values missing not at random (MNAR). In this scenario, investigators often perform list-wise deletion, or delete samples with any missing values, before applying causal discovery algorithms. List-wise deletion is a sound and general strategy when paired with algorithms such as FCI and RFCI, b…

2017-05-25abs ↗pdf ↗

The paper analyzes trade execution strategies for large traders in a stochastic market environment.

problem Analyzing trade execution strategies in a stochastic market with price impact.
method Formulated a Markov game model and used backward induction method of dynamic programming.
result Explicit closed-form execution strategy at Markov perfect equilibrium.

Proposes a method to combine datasets with missing values using Gaussian process latent variables.

problem Combining datasets with missing values under non-Missing at Random (NMAR) missingness.
method Gaussian process latent variable model for non-MAR missing data.
result Valid estimates are obtained using the proposed method, while existing methods provide severely biased estimates.

Proposes SSC for estimating counterfactual survival trajectories from observational data.

problem Challenges in estimating causal effects on time-to-event outcomes from observational data.
method Synthetic Survival Control (SSC) framework for estimating counterfactual hazard trajectories in panel data settings.
result SSC estimates counterfactual hazard trajectories as a weighted combination of observed trajectories from other units.

New method improves causal inference by estimating complex treatment effects with active learning.

problem Traditional causal inference frameworks ignore interference and assume independent treatment effects.
method Active Learning in Causal Inference with Interference (ACI) using Gaussian process and genetic algorithms.
result ACI achieves accurate effects estimation with reduced data requirements in complex interference scenarios.