Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

8162331 · Jul 201919922001200920172026
48 results for Big T-Rex

Big T-Rex solves FDR-controlled sparse regression on laptops with millions of variables.

problem Scalable FDR-controlled variable selection for high-dimensional data.
method Early terminated random experiments with memory-mapping and permutation-based dummy generation.
result Solves FDR-controlled Lasso problems with 5 million variables on a laptop in 30 minutes.

T-Rex selector selects variables fast and controls FDR in high-dimensional data.

problem Variable selection in high-dimensional data with FDR control.
method Fused solutions of early terminated random experiments.
result FDR control at target level with high variable selection power.

Improved FDR control for sparse financial index tracking.

problem Maintaining FDR control in high-dimensional financial data with strong variable dependencies.
method Expanding T-Rex framework to handle overlapping groups of correlated variables with nearest neighbors penalization.
result Accurately tracks the S&P 500 index using only a small number of stocks.

New method reduces memory usage for high-dimensional variable selection.

problem Scalability issues in high-dimensional variable selection, especially in genomics.
method Adaptive sampling of null features to eliminate dummy matrix materialization.
result Reduces memory and runtime by several orders of magnitude while preserving FDR control.

T-Rex uses EM to fit robust factor models in noisy data.

problem Robustly fitting factor models in high-dimensional data with heavy tails and outliers.
method Expectation-Maximization (EM) algorithm based on Tyler's M-estimator for elliptical distributions.
result Demonstrates robustness in direction-of-arrival estimation and subspace recovery.

Reinforcement learning has exceeded human-level performance in game playing AI with deep learning methods according to the experiments from DeepMind on Go and Atari games. Deep learning solves high dimension input problems which stop the development of reinforcement for many years. This study uses both two techniques t…

2019-09-10abs ↗pdf ↗

Enhances FDR control in variable selection using neural networks.

problem Balancing rigorous error control with statistical power in high-dimensional variable selection.
method Learning-augmented T-Rex Selector framework with a neural network trained on synthetic datasets.
result Achieves superior detection of true variables compared to existing approaches.

In this short note, we formulate three problems relating to nonnegative scalar curvature (NNSC) fill-ins. Loosely speaking, the first two problems focus on: When are (n1)(n-1)-dimensional Bartnik data (Σin1,γi,Hi)\big(Σ_i ^{n-1}, γ_i, H_i\big), i=1,2i=1,2, NNSC-cobordant? (i.e., there is an nn-dimensional compact Riemannian manifold…

2020-01-16abs ↗pdf ↗

New algorithm finds approximate stationary points faster under differential privacy constraints.

problem Finding approximate stationary points of smooth and Lipschitz functions under differential privacy constraints.
method Developed an efficient algorithm that improves convergence rates to stationary points.
result Achieved faster rates of convergence to stationary points in both finite-sum and stochastic settings.

Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…

2017-12-12abs ↗pdf ↗

In this paper we consider the large genus asymptotics for two classes of Siegel-Veech constants associated with an arbitrary connected stratum H(α)\mathcal{H} (α) of Abelian differentials. The first is the saddle connection Siegel-Veech constant cscmi,mj(H(α))c_{\text{sc}}^{m_i, m_j} \big( \mathcal{H} (α) \big) counting saddle conne…

2018-10-11abs ↗pdf ↗

Given (X,ω)(X,ω) compact Kähler manifold and ψM+PSH(X,ω)ψ\in\mathcal{M}^{+}\subset PSH(X,ω) a model type envelope with non-zero mass, i.e. a fixed potential determing some singularities such that X(ω+ddcψ)n>0\int_{X}(ω+dd^{c}ψ)^{n}>0, we prove that the ψψ-relative finite energy class E1(X,ω,ψ)\mathcal{E}^{1}(X,ω,ψ) becomes a complete metric space…

2019-09-09abs ↗pdf ↗

New DP algorithms achieve near-optimal regret bounds for online learning problems.

problem Online learning problems with zero-loss solutions and differential privacy constraints.
method Developed new Differentially Private algorithms with near-optimal regret bounds.
result Achieved near-optimal regret bounds for various online prediction and convex optimization problems.

The following problem is addressed: A 33-manifold MM is endowed with a triple Ω=(Ω1,Ω2,Ω3)Ω= \big(Ω^1,Ω^2,Ω^3\big) of closed 22-forms. One wants to construct a coframing ω=(ω1,ω2,ω3)ω= \big(ω^1,ω^2,ω^3\big) of MM such that, first, dωi=Ωi{\rm d}ω^i = Ω^i for i=1,2,3i=1,2,3, and, second, the Riemannian metric $g=\big(ω^1\big)^2+\big(ω^2\big)^2+\…

2019-08-02abs ↗pdf ↗

This note displays an interesting phenomenon for percentiles of independent but non-identical random variables. Let X1,,XnX_1,\cdots,X_n be independent random variables obeying non-identical continuous distributions and X(1)X(n)X^{(1)}\geq \cdots\geq X^{(n)} be the corresponding order statistics. For any p(0,1)p\in(0,1), we investig…

2018-08-24abs ↗pdf ↗

New insights into spectral statistics of sample covariance matrix for stable linear systems.

problem Estimating high-dimensional stable state transition matrices from noisy data.
method Combining spectral theorem for non-Hermitian operators, concentration of measure, and perturbation theory.
result The spectral radius of the sample covariance matrix exhibits phase transitions in high dimensions.

Uniform volume estimate for Kähler metrics in big cohomology classes.

problem Estimating volume for singular Kähler metrics in big cohomology classes.
method Generalized mixed energy estimate for functions in complex Sobolev space to big cohomology classes.
result Uniform non-collapsing volume estimate for local Kähler metrics.

Study convexity of Mabuchi functional in big cohomology classes.

problem Convexity of Mabuchi functional in big cohomology classes.
method Defined an invariant related to transcendental Fujita approximations and established convexity under vanishing of this invariant.
result Established almost convexity along weak geodesics in big cohomology classes.

Mobile big data contains vast statistical features in various dimensions, including spatial, temporal, and the underlying social domain. Understanding and exploiting the features of mobile data from a social network perspective will be extremely beneficial to wireless networks, from planning, operation, and maintenance…

2016-09-30abs ↗pdf ↗

Currently, the world is witnessing a mounting avalanche of data due to the increasing number of mobile network subscribers, Internet websites, and online services. This trend is continuing to develop in a quick and diverse manner in the form of big data. Big data analytics can process large amounts of raw data and extr…

2018-01-19abs ↗pdf ↗

We consider model-free reinforcement learning for infinite-horizon discounted Markov Decision Processes (MDPs) with a continuous state space and unknown transition kernel, when only a single sample path under an arbitrary policy of the system is available. We consider the Nearest Neighbor Q-Learning (NNQL) algorithm to…

2018-02-12abs ↗pdf ↗

We study the local equivalence problem for real-analytic (Cω\mathcal{C}^ω) hypersurfaces M5C3M^5 \subset \mathbb{C}^3 which, in coordinates (z1,z2,w)C3(z_1, z_2, w) \in \mathbb{C}^3 with w=u+ivw = u+i\, v, are rigid: \[ u \,=\, F\big(z_1,z_2,\overline{z}_1,\overline{z}_2\big), \] with FF independent of vv. Specifically, we study th…

2019-04-04abs ↗pdf ↗

Big Data bring new opportunities to modern society and challenges to data scientists. On one hand, Big Data hold great promises for discovering subtle population patterns and heterogeneities that are not possible with small-scale data. On the other hand, the massive sample size and high dimensionality of Big Data intro…

2013-08-07abs ↗pdf ↗

The paper tackles machine unlearning by designing efficient algorithms for adaptive query classes.

problem Designing efficient unlearning algorithms for machine learning models.
method Formalizes the problem and gives efficient unlearning algorithms for linear and prefix-sum query classes.
result Improved guarantees for stochastic convex optimization with reduced unlearning query complexity.

Once first answers in any dimension to the Green-Griffiths and Kobayashi conjectures for generic algebraic hypersurfaces Xn1Pn(C)\mathbb{X}^{n-1} \subset \mathbb{P}^n(\mathbb{C}) have been reached, the principal goal is to decrease (to improve) the degree bounds, knowing that the `celestial' horizon lies near $d \geqslant 2n…

2019-01-13abs ↗pdf ↗

In this paper, we consider the problem of sequentially optimizing a black-box function ff based on noisy samples and bandit feedback. We assume that ff is smooth in the sense of having a bounded norm in some reproducing kernel Hilbert space (RKHS), yielding a commonly-considered non-Bayesian form of Gaussian process …

2017-05-31abs ↗pdf ↗

Explosive growth in data and availability of cheap computing resources have sparked increasing interest in Big learning, an emerging subfield that studies scalable machine learning algorithms, systems, and applications with Big Data. Bayesian methods represent one important class of statistic methods for machine learni…

2014-11-24abs ↗pdf ↗

Data preprocessing techniques are devoted to correct or alleviate errors in data. Discretization and feature selection are two of the most extended data preprocessing techniques. Although we can find many proposals for static Big Data preprocessing, there is little research devoted to the continuous Big Data problem. A…

2018-10-14abs ↗pdf ↗

Characterizes and analyzes the large scale geometry of big mapping class groups of surfaces.

problem Analyzing the large scale geometry of big mapping class groups of surfaces with a unique maximal end.
method Building on previous work, the paper characterizes and analyzes the large scale geometry of big mapping class groups of surfaces with a unique maximal end.
result Proves that any locally CB big mapping class group is CB generated and gives an explicit criterion for determining which big mapping class groups are CB generated.

Proves existence of Kähler-Einstein metrics in big cohomology classes.

problem Existence of Kähler-Einstein metrics in big cohomology classes.
method Using a divisorial stability condition and Fujita-Odaka type delta invariants, building up from scratch the theory of pluripotential theory.
result Uniform Yau-Tian-Donaldson existence theorem for Kähler-Einstein metrics in the big cohomology class setting.

Big mapping class groups of infinite type surfaces have infinite asymptotic dimension.

problem Understanding asymptotic dimension of big mapping class groups of infinite type surfaces.
method Analyzing big mapping class groups with coarsely bounded generating sets and essential shifts.
result Big mapping class groups of infinite type surfaces have infinite asymptotic dimension.

Paper defines new stability and metrics for complex spaces.

problem Stability and metrics for complex spaces with big cohomology classes.
method Introduces slope stability and Hermitian-Einstein metrics for big cohomology classes.
result Kobayashi Hitchin correspondence and Bogomolov Gieseker inequality proved.

New obstruction prevents certain spacetimes with both big bang and big crunch.

problem Preventing spacetimes with both big bang and big crunch.
method Analyzing initial data sets subject to dominant energy condition and enlargeability obstruction.
result Pairs of spacetimes with both big bang and big crunch are not connected in certain cases.