Big T-Rex solves FDR-controlled sparse regression on laptops with millions of variables.
problem Scalable FDR-controlled variable selection for high-dimensional data.
method Early terminated random experiments with memory-mapping and permutation-based dummy generation.
result Solves FDR-controlled Lasso problems with 5 million variables on a laptop in 30 minutes.
IEN speeds up T-Rex+GVS for fast, efficient GWAS.
problem Efficiently selecting groups of genetic variants in large-scale genomics studies.
method Informed Elastic Net (IEN) as a faster base selector for T-Rex+GVS.
result IEN reduces computation time while maintaining high TPR and FDR control.
T-Rex selector selects variables fast and controls FDR in high-dimensional data.
problem Variable selection in high-dimensional data with FDR control.
method Fused solutions of early terminated random experiments.
result FDR control at target level with high variable selection power.
A critical flaw of existing inverse reinforcement learning (IRL) methods is their inability to significantly outperform the demonstrator. This is because IRL typically seeks a reward function that makes the demonstrator appear near-optimal, rather than inferring the underlying intentions of the demonstrator that may ha…
Improved FDR control for sparse financial index tracking.
problem Maintaining FDR control in high-dimensional financial data with strong variable dependencies.
method Expanding T-Rex framework to handle overlapping groups of correlated variables with nearest neighbors penalization.
result Accurately tracks the S&P 500 index using only a small number of stocks.
New method reduces memory usage for high-dimensional variable selection.
problem Scalability issues in high-dimensional variable selection, especially in genomics.
method Adaptive sampling of null features to eliminate dummy matrix materialization.
result Reduces memory and runtime by several orders of magnitude while preserving FDR control.
Novel framework controls FDR in high-dimensional, dependent data.
problem FDR control failure in high-dimensional, dependent data.
method Dependency-aware T-Rex selector integrating hierarchical graphical models and martingale theory.
result First to control FDR in high-dimensional, dependent data.
Sparse PCA selects variables with FDR control for improved performance.
problem Sparse PCA selects irrelevant variables when maximizing explained variance.
method Proposes FDR-controlled selection using T-Rex selector.
result Significant performance improvement over traditional sparse PCA.
T-Rex uses EM to fit robust factor models in noisy data.
problem Robustly fitting factor models in high-dimensional data with heavy tails and outliers.
method Expectation-Maximization (EM) algorithm based on Tyler's M-estimator for elliptical distributions.
result Demonstrates robustness in direction-of-arrival estimation and subspace recovery.
Reinforcement learning has exceeded human-level performance in game playing AI with deep learning methods according to the experiments from DeepMind on Go and Atari games. Deep learning solves high dimension input problems which stop the development of reinforcement for many years. This study uses both two techniques t…
Enhances FDR control in variable selection using neural networks.
problem Balancing rigorous error control with statistical power in high-dimensional variable selection.
method Learning-augmented T-Rex Selector framework with a neural network trained on synthetic datasets.
result Achieves superior detection of true variables compared to existing approaches.
New problems on NNSC fill-ins for Bartnik data in high dimensions.
problem Conditions for (n−1)-dimensional Bartnik data to be NNSC-cobordant. method Formulating three problems related to nonnegative scalar curvature fill-ins.
result Conditions for (n−1)-dimensional Bartnik data to be NNSC-cobordant. New algorithm finds approximate stationary points faster under differential privacy constraints.
problem Finding approximate stationary points of smooth and Lipschitz functions under differential privacy constraints.
method Developed an efficient algorithm that improves convergence rates to stationary points.
result Achieved faster rates of convergence to stationary points in both finite-sum and stochastic settings.
New algorithm for private non-convex optimization with optimal rates.
problem Private optimization of non-convex functions under KL condition.
method Variance-reduced gradient descent and proximal point method.
result Achieves nearly optimal rates for excess empirical risk.
In Part III of this study, we apply the price dynamical model with big buyers and big sellers developed in Part I of this paper to the daily closing prices of the top 20 banking and real estate stocks listed in the Hong Kong Stock Exchange. The basic idea is to estimate the strength parameters of the big buyers and the…
Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…
In this paper we consider the large genus asymptotics for two classes of Siegel-Veech constants associated with an arbitrary connected stratum H(α) of Abelian differentials. The first is the saddle connection Siegel-Veech constant cscmi,mj(H(α)) counting saddle conne…
Establishes Kobayashi-Hitchin correspondence for nef and big classes.
problem Analyzing stability and positivity in algebraic geometry.
method Introducing adapted currents and metrics to establish correspondence.
result Equality cases of Bogomolov-Gieseker and Miyaoka-Yau inequalities.
Given (X,ω) compact Kähler manifold and ψ∈M+⊂PSH(X,ω) a model type envelope with non-zero mass, i.e. a fixed potential determing some singularities such that ∫X(ω+ddcψ)n>0, we prove that the ψ−relative finite energy class E1(X,ω,ψ) becomes a complete metric space…
This article provides the role of big idea statisticians in future of Big Data Science. We describe the `United Statistical Algorithms' framework for comprehensive unification of traditional and novel statistical methods for modeling Small Data and Big Data, especially mixed data (discrete, continuous).
Geometric quantization extended to big line bundles.
problem Quantization of line bundles with large curvature.
method Proving asymptotic isometry and submultiplicative norms equivalence, showing Mabuchi geodesic rays.
result Bounded submultiplicative filtrations on big line bundles lead to Mabuchi geodesic rays.
New DP algorithms achieve near-optimal regret bounds for online learning problems.
problem Online learning problems with zero-loss solutions and differential privacy constraints.
method Developed new Differentially Private algorithms with near-optimal regret bounds.
result Achieved near-optimal regret bounds for various online prediction and convex optimization problems.
The following problem is addressed: A 3-manifold M is endowed with a triple Ω=(Ω1,Ω2,Ω3) of closed 2-forms. One wants to construct a coframing ω=(ω1,ω2,ω3) of M such that, first, dωi=Ωi for i=1,2,3, and, second, the Riemannian metric $g=\big(ω^1\big)^2+\big(ω^2\big)^2+\…
This note displays an interesting phenomenon for percentiles of independent but non-identical random variables. Let X1,⋯,Xn be independent random variables obeying non-identical continuous distributions and X(1)≥⋯≥X(n) be the corresponding order statistics. For any p∈(0,1), we investig…
New insights into spectral statistics of sample covariance matrix for stable linear systems.
problem Estimating high-dimensional stable state transition matrices from noisy data.
method Combining spectral theorem for non-Hermitian operators, concentration of measure, and perturbation theory.
result The spectral radius of the sample covariance matrix exhibits phase transitions in high dimensions.
Refines spinorial Sobolev inequality on sphere, proving stability and new properties of Killing spinors.
problem Stability of spinorial Sobolev inequalities on the unit sphere.
method Establishes a refined spinorial Sobolev inequality and proves stability.
result Stability inequality and new properties of Killing spinors.
Study assesses 'big data' in materials science, highlighting challenges.
problem Understanding what constitutes 'big data' in materials science.
method Selected examples of machine learning models, data quality, and infrastructure requirements.
result Big data presents unique challenges in materials science.
Uniform volume estimate for Kähler metrics in big cohomology classes.
problem Estimating volume for singular Kähler metrics in big cohomology classes.
method Generalized mixed energy estimate for functions in complex Sobolev space to big cohomology classes.
result Uniform non-collapsing volume estimate for local Kähler metrics.
Study convexity of Mabuchi functional in big cohomology classes.
problem Convexity of Mabuchi functional in big cohomology classes.
method Defined an invariant related to transcendental Fujita approximations and established convexity under vanishing of this invariant.
result Established almost convexity along weak geodesics in big cohomology classes.
Mobile big data contains vast statistical features in various dimensions, including spatial, temporal, and the underlying social domain. Understanding and exploiting the features of mobile data from a social network perspective will be extremely beneficial to wireless networks, from planning, operation, and maintenance…
Currently, the world is witnessing a mounting avalanche of data due to the increasing number of mobile network subscribers, Internet websites, and online services. This trend is continuing to develop in a quick and diverse manner in the form of big data. Big data analytics can process large amounts of raw data and extr…
Big data transforms accounting and auditing, enhancing insights but posing challenges.
problem Challenges in data privacy and security with increased data sources.
method Utilizing AI and machine learning for efficient data analysis and anomaly detection.
result Enhanced analytics tools and continuous learning are key to overcoming challenges.
Unified view on big bang singularities from initial data.
problem Understanding quiescent big bang singularities from initial data.
method Unified geometric perspective on various results.
result Solutions of Oude Groeniger et al. induce data on the singularity.
We consider model-free reinforcement learning for infinite-horizon discounted Markov Decision Processes (MDPs) with a continuous state space and unknown transition kernel, when only a single sample path under an arbitrary policy of the system is available. We consider the Nearest Neighbor Q-Learning (NNQL) algorithm to…
We study the local equivalence problem for real-analytic (Cω) hypersurfaces M5⊂C3 which, in coordinates (z1,z2,w)∈C3 with w=u+iv, are rigid: \[ u \,=\, F\big(z_1,z_2,\overline{z}_1,\overline{z}_2\big), \] with F independent of v. Specifically, we study th…
Big Data bring new opportunities to modern society and challenges to data scientists. On one hand, Big Data hold great promises for discovering subtle population patterns and heterogeneities that are not possible with small-scale data. On the other hand, the massive sample size and high dimensionality of Big Data intro…
The paper tackles machine unlearning by designing efficient algorithms for adaptive query classes.
problem Designing efficient unlearning algorithms for machine learning models.
method Formalizes the problem and gives efficient unlearning algorithms for linear and prefix-sum query classes.
result Improved guarantees for stochastic convex optimization with reduced unlearning query complexity.
Once first answers in any dimension to the Green-Griffiths and Kobayashi conjectures for generic algebraic hypersurfaces Xn−1⊂Pn(C) have been reached, the principal goal is to decrease (to improve) the degree bounds, knowing that the `celestial' horizon lies near $d \geqslant 2n…
In this paper, we consider the problem of sequentially optimizing a black-box function f based on noisy samples and bandit feedback. We assume that f is smooth in the sense of having a bounded norm in some reproducing kernel Hilbert space (RKHS), yielding a commonly-considered non-Bayesian form of Gaussian process …
In this note we show that many subgroups of mapping class groups of infinite-type surfaces without boundary have trivial centers, including all normal subgroups. Using similar techniques, we show that every nontrivial normal subgroup of a big mapping class group contains a nonabelian free group. In contrast, we show th…
Explosive growth in data and availability of cheap computing resources have sparked increasing interest in Big learning, an emerging subfield that studies scalable machine learning algorithms, systems, and applications with Big Data. Bayesian methods represent one important class of statistic methods for machine learni…
Data preprocessing techniques are devoted to correct or alleviate errors in data. Discretization and feature selection are two of the most extended data preprocessing techniques. Although we can find many proposals for static Big Data preprocessing, there is little research devoted to the continuous Big Data problem. A…
Characterizes and analyzes the large scale geometry of big mapping class groups of surfaces.
problem Analyzing the large scale geometry of big mapping class groups of surfaces with a unique maximal end.
method Building on previous work, the paper characterizes and analyzes the large scale geometry of big mapping class groups of surfaces with a unique maximal end.
result Proves that any locally CB big mapping class group is CB generated and gives an explicit criterion for determining which big mapping class groups are CB generated.
Proves existence of Kähler-Einstein metrics in big cohomology classes.
problem Existence of Kähler-Einstein metrics in big cohomology classes.
method Using a divisorial stability condition and Fujita-Odaka type delta invariants, building up from scratch the theory of pluripotential theory.
result Uniform Yau-Tian-Donaldson existence theorem for Kähler-Einstein metrics in the big cohomology class setting.
Big data analytics improves healthcare through early detection and quality life.
problem Limited access to healthcare data hinders evidence-based decision-making.
method Analysis of healthcare data using various tools and techniques.
result Big data analytics enhances healthcare quality and patient outcomes.
Big mapping class groups of infinite type surfaces have infinite asymptotic dimension.
problem Understanding asymptotic dimension of big mapping class groups of infinite type surfaces.
method Analyzing big mapping class groups with coarsely bounded generating sets and essential shifts.
result Big mapping class groups of infinite type surfaces have infinite asymptotic dimension.
Paper defines new stability and metrics for complex spaces.
problem Stability and metrics for complex spaces with big cohomology classes.
method Introduces slope stability and Hermitian-Einstein metrics for big cohomology classes.
result Kobayashi Hitchin correspondence and Bogomolov Gieseker inequality proved.
New obstruction prevents certain spacetimes with both big bang and big crunch.
problem Preventing spacetimes with both big bang and big crunch.
method Analyzing initial data sets subject to dominant energy condition and enlargeability obstruction.
result Pairs of spacetimes with both big bang and big crunch are not connected in certain cases.