Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

4998147196 · Jun 202019922001200920182026
48 results for ultra-high dimensions

Paper models ultra-high dimensional data using Gaussian and vine copulas.

problem Modeling ultra-high dimensional non-Gaussian data.
method Divide-and-conquer approach: Gaussian methods for subsets, vine copulas for reconciled model.
result Feasibility and estimation in thousands of dimensions demonstrated.

A new method scales sparse machine learning to ultra-high dimensional problems.

problem Sparse and interpretable machine learning in ultra-high dimensional data.
method Two-phase approach: backbone set determination followed by reduced problem solving.
result The backbone set contains truly relevant features with high probability.

DeepFS uses deep neural networks to select significant features in ultra high-dimensional data.

problem Challenges in traditional feature selection methods for high-dimensional, low-sample-size data.
method Two-step nonparametric approach combining deep neural networks and feature screening.
result DeepFS effectively identifies significant features with high precision for ultra high-dimensional data.

Study predicts price predictability in ultra-high frequency financial data using entropy tests.

problem Tackles predictability of ultra-high frequency financial data.
method Develops statistical tests based on Shannon entropy and Kullback-Leibler divergence to analyze predictability.
result Degree of randomness increases with aggregation level in transaction time.

Study uses Hawkes and diffusion models to analyze stock price dynamics.

problem Analyzing volatility and price dynamics in ultra-high-frequency stock data.
method Combined symmetric Hawkes and diffusion models with maximum likelihood estimation.
result Model provides accurate volatility estimation and dynamics of parameters.

A streaming algorithm estimates quadratic covariation from financial data efficiently.

problem Estimating quadratic covariation from ultra-high-frequency financial data with limited memory.
method Formulated multi-scale, realized kernel, pre-averaging, and modulated realized covariance estimators with fixed bandwidth.
result Fixed bandwidth estimators require higher bandwidth for positive semidefiniteness.

Study uses deep learning to detect BCCs in high-res histopathological images.

problem Detecting BCCs in high-resolution, weakly labeled histopathological images.
method Attention-based deep learning models to process ultra-high resolution images with weak labels.
result Attention-based models achieve almost perfect classification performance (AUC of 0.99).

AutoCompress automatically prunes DNNs to ultra-high compression rates without accuracy loss.

problem Efficiently compressing deep neural networks to reduce storage and computation requirements.
method Automatic hyperparameter determination, ADMM-based structured weight pruning, purification step, heuristic search.
result Achieves ultra-high pruning rates on weights and FLOPs, up to 33x in pruning rate.

A new feature selection method using random forest and Kolmogorov filter.

problem Ultra-high dimensional data feature selection.
method Fused Kolmogorov filter with random forest based recursive feature elimination.
result Selection and L2L_2 consistency under weak conditions.

We present a large-scale study of commonality in liquidity and resilience across assets in an ultra high-frequency (millisecond-timestamped) Limit Order Book (LOB) dataset from a pan-European electronic equity trading facility. We first show that extant work in quantifying liquidity commonality through the degree of ex…

2014-06-20abs ↗pdf ↗

FL-Sailer enables federated learning for scATAC-seq data, reducing dimensionality and noise.

problem Privacy-preserving federated learning for ultra-high dimensional, sparse, and heterogeneous scATAC-seq data.
method FL-Sailer integrates adaptive leverage score sampling and an invariant VAE architecture.
result FL-Sailer converges to an approximate solution with bounded error, surpassing centralized methods.

This paper optimizes deep learning systems for high performance and low energy consumption.

problem Achieving ultra-high energy efficiency and performance for deep neural networks.
method Developed an algorithm-hardware co-optimization framework that reduces computational and storage complexity.
result Achieved at least 152X speedup and 71X energy efficiency gain compared to IBM TrueNorth processor.

This study examines how financial tick data becomes more random with time aggregation.

problem Investigating the randomness of financial tick data over time.
method Applied statistical randomness tests from NIST and TestU01 batteries to ultra-high frequency financial data.
result Financial tick data becomes increasingly random as the aggregation level of transaction time increases.

By studying all the trades and best bids/asks of ultra high frequency snapshots recorded from the order books of a basket of 10 futures assets, we bring qualitative empirical evidence that the impact of a single trade depends on the intertrade time lags. We find that when the trading rate becomes faster, the return var…

2010-10-20abs ↗pdf ↗

Using ultra-high-frequency data extracted from the order flows of 23 stocks traded on the Shenzhen Stock Exchange, we study the empirical regularities of order placement in the opening call auction, cool period and continuous auction. The distributions of relative logarithmic prices against reference prices in the thre…

2007-12-06abs ↗pdf ↗

Study uses multi-kernel Hawkes models to analyze high-frequency price dynamics.

problem Understanding responsive speeds of market participants in high-frequency trading.
method Multi-kernel Hawkes models with conditional Hessian analysis for optimization.
result Existence of multi-kernels (UHF, VHF, HF) in high-frequency price dynamics.

GIDS reduces high-dimensional response and predictor spaces, improving interpretability and computational efficiency.

problem Challenges in modeling interactions among high-dimensional multimodal data.
method Graph Independence Dual Screening (GIDS) framework that reduces both response and predictor dimensions.
result GIDS reduces feature space to 9,000 CpGs and 2,000 transcripts, revealing coordinated regulatory mechanisms.

Paper forecasts financial trading durations using a new point process model.

problem Forecasting limit order book durations in high-frequency financial data.
method Self-exciting flexible residual point process incorporating empirical distributional features.
result The model achieves strong predictive performance compared to alternative approaches.

VOLARE provides standardized realized volatility measures from financial data.

problem Lack of standardized realized volatility measures from ultra-high-frequency data.
method Asset-specific pipeline for cleaning and sampling data, providing a wide range of realized estimators.
result Comprehensive set of realized estimators for equities, exchange rates, and futures.

Paper proposes unsupervised learning for optimizing deep neural networks in time-sensitive applications.

problem Inaccurate optimization solutions from supervised learning for real-time applications.
method Proposes unsupervised deep learning to learn latent function with optimization constraints as supervision.
result DNN can optimize variable problems in ultra-reliable communications without supervision labels.

The article studies a combined L1L_1 and concave regularization method for high-dimensional models.

problem Tackles variable selection and prediction in high-dimensional settings.
method Uses combined L1L_1 and concave penalties to optimize model sparsity and prediction risk.
result Global optimum of the method achieves oracle prediction risk and false sign rate bounds.

A method to identify important features without solving the full problem.

problem Identifying important features in high-dimensional data.
method Persistent reduction using extreme ray identification on a polyhedral cone.
result A subset of features can be guaranteed to have zero coefficients in all optimal solutions.

Through the analysis of a dataset of ultra high frequency order book updates, we introduce a model which accommodates the empirical properties of the full order book together with the stylized facts of lower frequency financial data. To do so, we split the time interval of interest into periods in which a well chosen r…

2013-12-02abs ↗pdf ↗

A new feature screening method using projection correlation and knockoffs controls FDR in high-dimensional data.

problem Feature selection in ultra-high dimensional datasets with heavy-tailed errors and multivariate responses.
method Projection correlation for dependence measurement, knockoffs for FDR control, two-step approach.
result The method controls FDR and ensures sure screening under weak assumptions.