Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

133266398531 · Jun 202019922001200920172026
48 results for two-phase sampling

Improved statistical inference for expensive data using machine learning predictions.

problem Statistical inference under adaptive two-phase multiwave sampling with expensive measurements.
method Multiwave Predict-Then-Debias estimator combining proxy information and expensive measurements.
result Valid estimators and confidence intervals for M-estimation under adaptive sampling.

Common models for two-phase lipid bilayer membranes are based on an energy that consists of an elastic term for each lipid phase and a line energy at interfaces. Although such an energy controls only the length of interfaces, the membrane surface is usually assumed to be at least C1C^1 across phase boundaries. We consi…

2016-03-01abs ↗pdf ↗

New method improves statistical inference using machine learning-imputed data.

problem Improving statistical inference with imputed data from machine learning.
method Two-phase sampling approach for Z-estimation with ML-imputed outcomes.
result Guaranteed efficiency matching or exceeding classical inference, regardless of prediction quality.

Study shows thresholding scheme converges for mean curvature flow of convex sets.

problem Analyzing convergence of thresholding scheme for mean curvature flow.
method Time discretization using Merriman, Bence and Osher's scheme, focusing on two-phase mean convex settings.
result Time-integrated energy of approximation converges to limit's energy in minimizing movements interpretation.

Estimates vaccine effectiveness and immune correlates in TND studies with missing data.

problem Confounding and missing data in TND studies of vaccine effectiveness and immune correlates.
method Targeted maximum likelihood estimation using a semiparametric logistic regression model.
result Valid causal inference of vaccine effectiveness and immune correlates in TND studies with missing exposure data.

We examine the out-of-equilibrium phase reported by Plerou {\it et. al.} in Nature, {\bf 421}, 130 (2003) using the data of the New York stock market (NYSE) between the years 2001 --2002. We find that the observed two phase phenomenon is an artifact of the definition of the control parameter coupled with the nature of …

2005-02-15abs ↗pdf ↗

MetaNOR learns common nonlocal kernels for efficient metamaterial modeling.

problem Efficiently modeling wave propagation in new metamaterials.
method Meta-learns a common nonlocal kernel from existing tasks and transfers this knowledge to new tasks with minimal data.
result Substantial improvements in sampling efficiency for new metamaterials.

For better classification generative models are used to initialize the model and model features before training a classifier. Typically it is needed to solve separate unsupervised and supervised learning problems. Generative restricted Boltzmann machines and deep belief networks are widely used for unsupervised learnin…

2018-04-25abs ↗pdf ↗

PPG separates policy and value function training phases for better reinforcement learning efficiency.

problem Challenges in traditional reinforcement learning methods for policy and value function optimization.
method Integrates Phasic Policy Gradient framework that splits policy and value function training into distinct phases.
result Significantly improves sample efficiency on Procgen Benchmark compared to PPO.

Enhances neural networks' robustness against adversarial samples without sacrificing clean sample generalization.

problem Limited generalization and time complexity of adversarial training.
method Feature Pyramid Decoder (FPD) framework that integrates denoising and image restoration modules into CNNs and constrains the Lipschitz constant.
result FPD-enhanced CNNs achieve sufficient robustness against general adversarial samples on various datasets.

Reliable training of generative adversarial networks (GANs) typically require massive datasets in order to model complicated distributions. However, in several applications, training samples obey invariances that are \textit{a priori} known; for example, in complex physics simulations, the training data obey universal …

2019-06-04abs ↗pdf ↗

The two phase behavior in financial markets actually means the bifurcation phenomenon, which represents the change of the conditional probability from an unimodal to a bimodal distribution. In this paper, the bifurcation phenomenon in Hang-Seng index is carefully investigated. It is observed that the bifurcation phenom…

2007-12-30abs ↗pdf ↗

Detects anomalies in product health metrics at eBay for better alerts.

problem Detecting anomalies in unsupervised product health metrics at eBay.
method Developed a Moving Metric Detector (MMD) for anomaly detection and a point-wise ranking model for alert retrieval.
result Improves alert precision and avoids alert spamming in eBay production.

We consider the problem of efficiently computing the maximum likelihood estimator in Generalized Linear Models (GLMs) when the number of observations is much larger than the number of coefficients (np1n \gg p \gg 1). In this regime, optimization algorithms can immensely benefit from approximate second order information.…

2015-11-28abs ↗pdf ↗

Matrix completion, i.e., the exact and provable recovery of a low-rank matrix from a small subset of its elements, is currently only known to be possible if the matrix satisfies a restrictive structural constraint---known as {\em incoherence}---on its row and column spaces. In these cases, the subset of elements is sam…

2013-06-12abs ↗pdf ↗

Often the challenge associated with tasks like fraud and spam detection[1] is the lack of all likely patterns needed to train suitable supervised learning models. In order to overcome this limitation, such tasks are attempted as outlier or anomaly detection tasks. We also hypothesize that out- liers have behavioral pat…

2017-12-12abs ↗pdf ↗

Proposes a method to estimate sparse Gaussian graphical models with hidden clustering structure.

problem Modeling statistical relationships between variables with sparsity and clustering.
method Two-phase algorithm using sGS-ADMM for initial point and pALM for solution.
result Demonstrates good performance and efficiency of the proposed model and algorithm on synthetic and real data.

Noisy labels are very common in real-world training data, which lead to poor generalization on test data because of overfitting to the noisy labels. In this paper, we claim that such overfitting can be avoided by "early stopping" training a deep neural network before the noisy labels are severely memorized. Then, we re…

2019-11-19abs ↗pdf ↗

Discrimination-aware classification is receiving an increasing attention in data science fields. The pre-process methods for constructing a discrimination-free classifier first remove discrimination from the training data, and then learn the classifier from the cleaned data. However, they lack a theoretical guarantee f…

2017-02-28abs ↗pdf ↗

Proposes MD-LiNA for multi-domain latent factor causal discovery.

problem Discovering causal structures among latent factors from multi-domain data.
method Multi-Domain Linear Non-Gaussian Acyclic Models (MD-LiNA) with an integrated two-phase algorithm.
result Locally consistent estimators of causal structure among shared latent factors.

DP-SGD can update fewer coordinates while maintaining privacy.

problem How to update fewer coordinates in DP-SGD without losing optimization signal.
method TP-TopK (Two-Phase TopK DP-SGD), a two-phase method for coordinate-sparse private training.
result Private training can update fewer coordinates without losing optimization signal, scaling noise with active dimension \(k\) instead of full dimension \(d\).

A new method infers graph structure and parameters using a single generative flow network.

problem Bayesian Network structure and parameter inference from data.
method Single GFlowNet with two-phase sampling: DAG generation followed by parameter assignment.
result Accurate approximation of joint posterior distribution over graph structure and parameters.

This paper proposes a new AED framework for multi-metric experiments with fixed budget.

problem Statistical power challenges in testing multiple metrics simultaneously.
method Two-phase structure: adaptive exploration followed by validation. SHRVar algorithm with relative-variance-based sampling.
result Achieves provable error probability that decreases exponentially.

We discuss a class of (local and non-local) theories of gravity that share same properties: i) they admit the Einstein spacetime with arbitrary cosmological constant as a solution; ii) the on-shell action of such a theory vanishes and iii) any (cosmological or black hole) horizon in the Einstein spacetime with a positi…

2012-03-13abs ↗pdf ↗

Paper proposes a method to learn linear regression models using multiple pre-trained models.

problem Learning a linear regression model with limited target data.
method Representation transfer learning method using multiple pre-trained models.
result The method achieves better sample complexity compared to baseline methods.

SGD efficiently learns the XOR function with near-optimal sample complexity.

problem Learning the XOR function with a 2-layer neural network.
method Minibatch SGD on a 2-layer neural network with ReLU activations, focusing on signal-finding and signal-heavy phases.
result Achieves population error o(1)o(1) with dextpolylog(d)d \: ext{polylog}(d) samples.

This paper proposes a new global optimization algorithm using deep learning.

problem Developing efficient algorithms for global optimization of non-convex functions.
method Two-phase approach: minimization phase with model-driven deep learning, escaping phase with reinforcement learning.
result The proposed algorithm significantly outperforms classical optimization methods and handles ill-posed functions.