Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

129258387516 · Jun 202019922001200920182026
48 results for Statistical design

Paper develops efficient incomplete U-statistics for degenerate cases.

problem High computational cost and non-standard asymptotic behavior in degenerate U-statistics.
method Characterizes dependence structure using hypergraph theory and combinatorial designs, bypassing traditional Hoeffding decomposition.
result Derives a Berry-Esseen bound for incomplete U-statistics of deterministic designs, enabling Gaussian limiting distributions in degenerate cases.

This paper develops a Bayesian optimal design of experiments for estimating the statistical expectation of a black-box function.

problem Estimating the statistical expectation of a black-box function in complex systems.
method Sequentially querying the black-box function at specific designs selected by an infill-sampling criterion, maximizing expected information gain.
result Derivation of a semi-analytic mathematical formula for expected information gain about the statistical expectation of a physical response.

Introduces a new geometric method for optimal experimental design.

problem Restrictive invariance properties of traditional OED approaches based on probability densities.
method Mutual transport dependence (MTD) using optimal transport theory.
result Demonstrates high-quality designs and flexibility compared to standard methods.

Modern ML methods show unexpected behaviors that contradict classical statistics.

problem Modern machine learning methods exhibit behaviors at odds with classical statistical intuitions.
method Comparison between fixed and random design settings in ML and statistics.
result Moving from fixed to random designs reveals new insights into bias-variance tradeoffs and overfitting.

Designs new functionals for ranking joint probability distributions based on correlations.

problem Ranking joint probability distributions based on their correlations.
method Using first principles from inference, a set of functionals are designed with the Principle of Constant Correlations (PCC) guiding the construction.
result The nn-partite information (NPI) uniquely determines whether inferential transformations preserve, destroy, or create correlations.

Optimal design for multinomial logit models improves assortment selection efficiency.

problem Optimal experimental design for multinomial logit models with feedback.
method Two complementary approaches: MILP reformulation and lifted design.
result Achieves statistical efficiency and scalability for MNL bandits.

The study examines how experimental design choices affect machine learning model performance.

problem Lack of guidelines on choosing experimental designs and machine learning models.
method 12 experimental designs, 7 families of predictive models, 7 test functions, 8 noise settings.
result Guidelines for practical applications of DOE and ML are provided.

Electric vehicles (EVs) have been gaining popularity due to their environmental friendliness and efficiency. EV charging station networks are scalable solutions for supporting increasing numbers of EVs within modern electric grid constraints, yet few tools exist to aid the physical configuration design of new networks.…

2018-04-02abs ↗pdf ↗

The paper optimizes interpolation schedules in generative models to improve sampling accuracy.

problem Improving sampling accuracy in generative models with fewer resources.
method Minimizing the averaged squared Lipschitzness of the drift field, using transfer formulas.
result Designed schedules yield more accurate fine-scale statistics at fixed integrator budget.

In this paper, the optimal mean-reverting portfolio (MRP) design problem is considered, which plays an important role for the statistical arbitrage (a.k.a. pairs trading) strategy in financial markets. The target of the optimal MRP design is to construct a portfolio from the underlying assets that can exhibit a satisfa…

2018-03-08abs ↗pdf ↗

Study phase transitions in identifying infected individuals using group testing.

problem Identifying a set of k infected individuals from a population using pooled tests.
method Two random assignment designs (constant-column and Bernoulli) and polynomial-time inference procedures.
result Sharp phase transitions in statistical and computational limits for detection and recovery problems.

New auction design uses statistical learning to reduce costs and improve fairness.

problem Designing efficient multi-item auctions with reduced implementation costs and fairness.
method Nonparametric density estimation for credible intervals, two new strategies.
result Strategies consistently outperform alternative methods in revenue maximization and cost reduction.

This study improves audit sampling by using sequential procedures with statistical guarantees.

problem Improving audit efficiency and reliability with statistical methods.
method Formulated as a sequential testing problem, defining null and alternative hypotheses, stopping and decision rules, and exact boundary conditions.
result Exact design yields ex ante control of decision error probabilities, and simulation-based implementation approximates this design.

Estimates extreme probabilities using fewer simulations than Monte Carlo.

problem Estimating tail probabilities of complex systems efficiently.
method Builds a statistical surrogate with few evaluations and sequentially improves the estimate.
result Improves estimation of extreme probabilities with fewer simulations.

A quantum circuit designed for efficient statistical model preparation and training.

problem Challenges in preparing and learning statistical models on quantum processors.
method Utilizes the maximum entropy principle to design a statistics-informed parameterized quantum circuit (SI-PQC).
result Improves trainability and interpretability for learning quantum states and classical model parameters.

The paper develops efficient sampling strategies for BRDF data manifolds.

problem Efficiently sampling and measuring BRDF data from high-dimensional manifolds.
method Statistical design of experiments and generalized proactive learning.
result Established more efficient sampling and measurement strategies for BRDF data manifolds.

A new framework learns system design using neural features in function space.

problem Learning system design with neural feature extractors.
method Introduces feature geometry in function space, nesting technique for optimal feature approximation.
result Optimal features found from data samples using off-the-shelf architectures and optimizers.

We propose an optimum mechanism for providing monetary incentives to the data sources of a statistical estimator such as linear regression, so that high quality data is provided at low cost, in the sense that the sum of payments and estimation error is minimized. The mechanism applies to a broad range of estimators, in…

2014-08-11abs ↗pdf ↗

A method to select validation data from a dataset using statistical criteria.

problem Selecting a validation basis from a full dataset for machine learning model validation.
method Adopting a 'design of experiments' point of view and using statistical criteria, particularly Maximum Mean Discrepancy criteria.
result The 'support points' concept is particularly relevant for selecting validation data.

Framework for efficient statistical estimation with privacy guarantees.

problem Statistical estimation problems with differential privacy constraints.
method High-dimensional Propose-Test-Release (HPTR) framework combining exponential mechanism, robust statistics, and resilience.
result Near-optimal utility guarantees and tight local sensitivity bounds for various statistical problems.

New algorithm converts data into sub-gaussian designs efficiently.

problem Efficiently converting large datasets into sub-gaussian random designs for robust performance.
method Algorithmic Gaussianization through sketching and averaging, using LESS embeddings.
result Efficient data sketches nearly indistinguishable from sub-gaussian designs.

This paper proposes a new AED framework for multi-metric experiments with fixed budget.

problem Statistical power challenges in testing multiple metrics simultaneously.
method Two-phase structure: adaptive exploration followed by validation. SHRVar algorithm with relative-variance-based sampling.
result Achieves provable error probability that decreases exponentially.

Value functions struggle to represent transition dynamics, impacting statistical efficiency.

problem Limited representational power of value functions in capturing transition dynamics.
method Case studies of various reinforcement learning problems to explore the limitations of value-based methods.
result Value-based methods can be as efficient as model-based ones in some cases but severely underperform in others due to information loss.

Optimizes experimental design using synthetic controls for better outcomes.

problem Estimating average treatment effects in studies with pre-treatment data.
method Mixed-integer programming for selecting treated and control units and weights.
result Improves mean squared error and statistical power compared to simple alternatives.

Bayesian surrogate models reduce uncertainty in high-dimensional design optimisation problems.

problem Uncertainty in high-dimensional inputs for complex computational models.
method Variational Bayesian inference for constructing statistical surrogates with Gaussian process priors and KL divergence for approximation.
result The RDVGP surrogate provides accurate and versatile approximations for robust structural optimisation.

How should statistical procedures be designed so as to be scalable computationally to the massive datasets that are increasingly the norm? When coupled with the requirement that an answer to an inferential question be delivered within a certain time budget, this question has significant repercussions for the field of s…

2013-09-30abs ↗pdf ↗

Efficiently estimates private least squares with linear error growth.

problem Private estimation of ordinary least squares with bounded residuals and leverage.
method Scaled noise added to a stable nonprivate estimator of the regression vector.
result Near-optimal accuracy guarantee with linear error growth in dimension.

Novel neural architecture improves Bayesian experimental design efficiency.

problem Intractable evaluation of expected information gain (EIG) in Bayesian optimal experimental design.
method Develops a neural architecture that optimizes a single variational model for estimating EIG across many designs, using a lower bound for computational efficiency.
result Significantly improves accuracy in Bayesian experimental design with better sample efficiency.

New algorithms improve experimental design efficiency and approximation quality.

problem Finding optimal subset of vectors for expensive measurements.
method Bayesian experimental design using determinantal point processes.
result Developed efficient algorithms for optimal design under multiple criteria.

We bridge statistical and worst-case approaches to experimental design for linear regression.

problem Designing efficient experiments for linear regression models with arbitrary responses.
method Propose a new experimental design framework for arbitrary response distributions, combining statistical and worst-case approaches.
result Develop efficient randomized design procedures achieving strong variance bounds for unbiased estimators using few responses.

New method uses approximate KLD for intractable likelihood models.

problem Designing experiments for models with intractable likelihoods.
method Derive a lower bound of KLD utility, express it in terms of entropies, and evaluate efficiently.
result Demonstrated the performance of the proposed method through numerical examples.

We simplify information measure computation using learned features.

problem Computing information measures from raw data is computationally expensive.
method Developed a separable design for computing information measures from learned feature representations.
result A variety of information measures can be computed efficiently through learned feature representations.