Paper models ultra-high dimensional data using Gaussian and vine copulas.
problem Modeling ultra-high dimensional non-Gaussian data.
method Divide-and-conquer approach: Gaussian methods for subsets, vine copulas for reconciled model.
result Feasibility and estimation in thousands of dimensions demonstrated.
A new method scales sparse machine learning to ultra-high dimensional problems.
problem Sparse and interpretable machine learning in ultra-high dimensional data.
method Two-phase approach: backbone set determination followed by reduced problem solving.
result The backbone set contains truly relevant features with high probability.
DeepFS uses deep neural networks to select significant features in ultra high-dimensional data.
problem Challenges in traditional feature selection methods for high-dimensional, low-sample-size data.
method Two-step nonparametric approach combining deep neural networks and feature screening.
result DeepFS effectively identifies significant features with high precision for ultra high-dimensional data.
In data sets with many more features than observations, independent screening based on all univariate regression models leads to a computationally convenient variable selection method. Recent efforts have shown that in the case of generalized linear models, independent screening may suffice to capture all relevant feat…
New algorithm reduces data access and complexity for big data problems.
problem Efficiently solving problems with large numbers of samples in ultra-high dimensional space.
method Accelerated Variance Reduced Block Coordinate Descent (AVRBCD)
result Achieves an accelerated convergence rate of O(1/k^2) with low per-iteration complexity.
Enhanced Dantzig selector reduces recovery error in high dimensions.
problem Improving recovery accuracy in ultra-high dimensional settings.
method Constrained Dantzig selector with sequential linear programming.
result Achieves convergence rates within a logarithmic factor of the sample size of oracle rates.
Study predicts price predictability in ultra-high frequency financial data using entropy tests.
problem Tackles predictability of ultra-high frequency financial data.
method Develops statistical tests based on Shannon entropy and Kullback-Leibler divergence to analyze predictability.
result Degree of randomness increases with aggregation level in transaction time.
Study uses Hawkes and diffusion models to analyze stock price dynamics.
problem Analyzing volatility and price dynamics in ultra-high-frequency stock data.
method Combined symmetric Hawkes and diffusion models with maximum likelihood estimation.
result Model provides accurate volatility estimation and dynamics of parameters.
A variable screening procedure via correlation learning was proposed Fan and Lv (2008) to reduce dimensionality in sparse ultra-high dimensional models. Even when the true model is linear, the marginal regression can be highly nonlinear. To address this issue, we further extend the correlation learning to marginal nonp…
A streaming algorithm estimates quadratic covariation from financial data efficiently.
problem Estimating quadratic covariation from ultra-high-frequency financial data with limited memory.
method Formulated multi-scale, realized kernel, pre-averaging, and modulated realized covariance estimators with fixed bandwidth.
result Fixed bandwidth estimators require higher bandwidth for positive semidefiniteness.
Study uses deep learning to detect BCCs in high-res histopathological images.
problem Detecting BCCs in high-resolution, weakly labeled histopathological images.
method Attention-based deep learning models to process ultra-high resolution images with weak labels.
result Attention-based models achieve almost perfect classification performance (AUC of 0.99).
A new method for identifying interactions in high dimensions.
problem Challenges in identifying interactions with many covariates.
method Interaction pursuit (IP) procedure: feature screening and selection.
result The method screens interactions separately from main effects, improving effectiveness.
AutoCompress automatically prunes DNNs to ultra-high compression rates without accuracy loss.
problem Efficiently compressing deep neural networks to reduce storage and computation requirements.
method Automatic hyperparameter determination, ADMM-based structured weight pruning, purification step, heuristic search.
result Achieves ultra-high pruning rates on weights and FLOPs, up to 33x in pruning rate.
We propose an algorithm, semismooth Newton coordinate descent (SNCD), for the elastic-net penalized Huber loss regression and quantile regression in high dimensional settings. Unlike existing coordinate descent type algorithms, the SNCD updates each regression coefficient and its corresponding subgradient simultaneousl…
We propose a novel application of the Simultaneous Orthogonal Matching Pursuit (S-OMP) procedure for sparsistant variable selection in ultra-high dimensional multi-task regression problems. Screening of variables, as introduced in \cite{fan08sis}, is an efficient and highly scalable way to remove many irrelevant variab…
A new feature selection method using random forest and Kolmogorov filter.
problem Ultra-high dimensional data feature selection.
method Fused Kolmogorov filter with random forest based recursive feature elimination.
result Selection and L2 consistency under weak conditions. We present a large-scale study of commonality in liquidity and resilience across assets in an ultra high-frequency (millisecond-timestamped) Limit Order Book (LOB) dataset from a pan-European electronic equity trading facility. We first show that extant work in quantifying liquidity commonality through the degree of ex…
A detailed analysis of correlation between stock returns at high frequency is compared with simple models of random walks. We focus in particular on the dependence of correlations on time scales - the so-called Epps effect. This provides a characterization of stochastic models of stock price returns which is appropriat…
Paper uses TCN with attention to predict UHF stock price changes.
problem Predicting discrete dynamic distribution of UHF stock price changes.
method Classified price changes, used TCN with attention mechanism.
result TCN and TCN (attention) models outperform GARCH and LSTM models.
New method selects key features from millions of biological data points.
problem Scalability issue in feature selection for ultra-high dimensional biological data.
method Scaled up HSIC Lasso to handle millions of features.
result Achieves high accuracy with only 20 out of one million features.
Efficiently solves Elastic Net in high dimensions with Newton method.
problem Feature selection in high-dimensional data with non-negligible collinearity.
method Semi-smooth Newton Augmented Lagrangian Method.
result Significantly reduces computational cost compared to competitors.
Social and economic systems are complex adaptive systems, in which heterogenous agents interact and evolve in a self-organized manner, and macroscopic laws emerge from microscopic properties. To understand the behaviors of complex systems, computational experiments based on physical and mathematical models provide a us…
FL-Sailer enables federated learning for scATAC-seq data, reducing dimensionality and noise.
problem Privacy-preserving federated learning for ultra-high dimensional, sparse, and heterogeneous scATAC-seq data.
method FL-Sailer integrates adaptive leverage score sampling and an invariant VAE architecture.
result FL-Sailer converges to an approximate solution with bounded error, surpassing centralized methods.
RaSE screens variables via random subspaces, identifying joint effects.
problem Missing joint effects of predictors in ultra-high dimensional data.
method Random Subspace Ensemble (RaSE) framework combining subspace evaluation criteria.
result RaSE identifies signals with no marginal effect or high-order interactions.
At the ultra high frequency level, the notion of price of an asset is very ambiguous. Indeed, many different prices can be defined (last traded price, best bid price, mid price,...). Thus, in practice, market participants face the problem of choosing a price when implementing their strategies. In this work, we propose …
This paper optimizes deep learning systems for high performance and low energy consumption.
problem Achieving ultra-high energy efficiency and performance for deep neural networks.
method Developed an algorithm-hardware co-optimization framework that reduces computational and storage complexity.
result Achieved at least 152X speedup and 71X energy efficiency gain compared to IBM TrueNorth processor.
This study examines how financial tick data becomes more random with time aggregation.
problem Investigating the randomness of financial tick data over time.
method Applied statistical randomness tests from NIST and TestU01 batteries to ultra-high frequency financial data.
result Financial tick data becomes increasingly random as the aggregation level of transaction time increases.
We study the distributions of event-time returns and clock-time returns at different microscopic timescales using ultra-high-frequency data extracted from the limit-order books of 23 stocks traded in the Chinese stock market in 2003. We find that the returns at the one-trade timescale obey the inverse cubic law. For la…
New method for model selection in high-dimensional misspecified models.
problem Model misspecification and high dimensionality in big data.
method Exploits Bayesian principles in misspecified models, introduces HGBIC_p criterion.
result Established consistency of HGBIC_p in ultra-high dimensions under mild conditions.
Empirical Bayes method improves high-dimensional classification accuracy.
problem High-dimensional classification with sparse mean differences.
method Dirichlet process mixture model and variational Bayes algorithm.
result Effective estimation of mean difference leads to reduced misclassification.
By studying all the trades and best bids/asks of ultra high frequency snapshots recorded from the order books of a basket of 10 futures assets, we bring qualitative empirical evidence that the impact of a single trade depends on the intertrade time lags. We find that when the trading rate becomes faster, the return var…
Proposes method for optimal treatment decisions in big data.
problem Handling ultra-high dimensional data and misspecified models.
method Two-step estimation procedure with penalized regressions.
result Asymptotic properties of the proposed estimators are investigated.
Using ultra-high-frequency data extracted from the order flows of 23 stocks traded on the Shenzhen Stock Exchange, we study the empirical regularities of order placement in the opening call auction, cool period and continuous auction. The distributions of relative logarithmic prices against reference prices in the thre…
We propose a mathematical procedure for finding informed traders in ultra-high frequency trading. We wrote it as Vector ARMA and found condition of its stationarity. For the price exposure complied with ARMA(1,2) we proved that underlying asset price difference can be derived as ARMA(1,1) process. For validation of the…
Study uses multi-kernel Hawkes models to analyze high-frequency price dynamics.
problem Understanding responsive speeds of market participants in high-frequency trading.
method Multi-kernel Hawkes models with conditional Hessian analysis for optimization.
result Existence of multi-kernels (UHF, VHF, HF) in high-frequency price dynamics.
GIDS reduces high-dimensional response and predictor spaces, improving interpretability and computational efficiency.
problem Challenges in modeling interactions among high-dimensional multimodal data.
method Graph Independence Dual Screening (GIDS) framework that reduces both response and predictor dimensions.
result GIDS reduces feature space to 9,000 CpGs and 2,000 transcripts, revealing coordinated regulatory mechanisms.
Study compares Fourier estimators to mitigate asynchrony effects in finance.
problem Impact of asynchrony on instantaneous financial estimates.
method Comparison of Malliavin-Mancino and Cuchiero-Teichmann estimators.
result Malliavin-Mancino estimator produces more stable estimates under asynchrony.
Alpha-norm regularization simplifies marketing demand forecasting.
problem Ultra high-dimensional problems in demand estimation and forecasting.
method Nonconvex alpha-norm objective with coordinate descent and proximal operators.
result Alpha-norm regularization provides accurate out-of-sample estimates for promotion effects.
Paper forecasts financial trading durations using a new point process model.
problem Forecasting limit order book durations in high-frequency financial data.
method Self-exciting flexible residual point process incorporating empirical distributional features.
result The model achieves strong predictive performance compared to alternative approaches.
VOLARE provides standardized realized volatility measures from financial data.
problem Lack of standardized realized volatility measures from ultra-high-frequency data.
method Asset-specific pipeline for cleaning and sampling data, providing a wide range of realized estimators.
result Comprehensive set of realized estimators for equities, exchange rates, and futures.
Paper proposes unsupervised learning for optimizing deep neural networks in time-sensitive applications.
problem Inaccurate optimization solutions from supervised learning for real-time applications.
method Proposes unsupervised deep learning to learn latent function with optimization constraints as supervision.
result DNN can optimize variable problems in ultra-reliable communications without supervision labels.
We have analyzed the statistical probabilities of limit-order book (LOB) shape through building the book using the ultra-high-frequency data from 23 liquid stocks traded on the Shenzhen Stock Exchange in 2003. We find that the averaged LOB shape has a maximum away from the same best price for both buy and sell LOBs. Th…
New method for spot volatility estimation with reduced microstructure noise.
problem Estimating spot volatility from noisy high-frequency data.
method Pre-averaging/kernel estimator to handle microstructure noise.
result Optimal bandwidth selection and kernel functions for minimal variance.
Variable selection is a challenging issue in statistical applications when the number of predictors p far exceeds the number of observations n. In this ultra-high dimensional setting, the sure independence screening (SIS) procedure was introduced to significantly reduce the dimensionality by preserving the true mod…
The article studies a combined L1 and concave regularization method for high-dimensional models.
problem Tackles variable selection and prediction in high-dimensional settings.
method Uses combined L1 and concave penalties to optimize model sparsity and prediction risk. result Global optimum of the method achieves oracle prediction risk and false sign rate bounds.
A method to identify important features without solving the full problem.
problem Identifying important features in high-dimensional data.
method Persistent reduction using extreme ray identification on a polyhedral cone.
result A subset of features can be guaranteed to have zero coefficients in all optimal solutions.
Through the analysis of a dataset of ultra high frequency order book updates, we introduce a model which accommodates the empirical properties of the full order book together with the stylized facts of lower frequency financial data. To do so, we split the time interval of interest into periods in which a well chosen r…
A new feature screening method using projection correlation and knockoffs controls FDR in high-dimensional data.
problem Feature selection in ultra-high dimensional datasets with heavy-tailed errors and multivariate responses.
method Projection correlation for dependence measurement, knockoffs for FDR control, two-step approach.
result The method controls FDR and ensures sure screening under weak assumptions.