Big T-Rex solves FDR-controlled sparse regression on laptops with millions of variables.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
LLMs can simulate human investment attitudes based on personality traits.
Big data analytics improves healthcare through early detection and quality life.
Study reveals LLM personas have two distinct components: frame-robust aggregated traits and frame-dependent geometric features.
Involutions generate mapping class groups of infinite surfaces.
Big Data is one of the major challenges of statistical science and has numerous consequences from algorithmic and theoretical viewpoints. Big Data always involve massive data but they also often include online data and data heterogeneity. Recently some statistical methods have been adapted to process Big Data, like lin…
Delivering useful hydrological forecasts is critical for urban and agricultural water management, hydropower generation, flood protection and management, drought mitigation and alleviation, and river basin planning and management, among others. In this work, we present and appraise a new simple and flexible methodology…
Study identifies personality traits from dance movements in music.
We propose a new heavy-tailed distribution --- Gaussian-Chain (GC) distribution, which is inspirited by the hierarchical structures prevailing in social organizations. We determine the mean, variance and kurtosis of the Gaussian-Chain distribution to show its heavy-tailed property, and compute the tail distribution tab…
This paper surveys enterprise financial risk analysis from Big Data and LLMs perspectives.
As a non-parametric Bayesian model which produces informative predictive distribution, Gaussian process (GP) has been widely used in various fields, like regression, classification and optimization. The cubic complexity of standard GP however leads to poor scalability, which poses challenges in the era of big data. Hen…
We propose a tri-modal architecture to predict Big Five personality trait scores from video clips with different channels for audio, text, and video data. For each channel, stacked Convolutional Neural Networks are employed. The channels are fused both on decision-level and by concatenating their respective fully conne…
In this short note, we formulate three problems relating to nonnegative scalar curvature (NNSC) fill-ins. Loosely speaking, the first two problems focus on: When are -dimensional Bartnik data , , NNSC-cobordant? (i.e., there is an -dimensional compact Riemannian manifold…
New algorithm finds approximate stationary points faster under differential privacy constraints.
New algorithm for private non-convex optimization with optimal rates.
In Part III of this study, we apply the price dynamical model with big buyers and big sellers developed in Part I of this paper to the daily closing prices of the top 20 banking and real estate stocks listed in the Hong Kong Stock Exchange. The basic idea is to estimate the strength parameters of the big buyers and the…
Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…
In this paper we consider the large genus asymptotics for two classes of Siegel-Veech constants associated with an arbitrary connected stratum of Abelian differentials. The first is the saddle connection Siegel-Veech constant counting saddle conne…
Establishes Kobayashi-Hitchin correspondence for nef and big classes.
Given compact Kähler manifold and a model type envelope with non-zero mass, i.e. a fixed potential determing some singularities such that , we prove that the relative finite energy class becomes a complete metric space…
This article provides the role of big idea statisticians in future of Big Data Science. We describe the `United Statistical Algorithms' framework for comprehensive unification of traditional and novel statistical methods for modeling Small Data and Big Data, especially mixed data (discrete, continuous).
Geometric quantization extended to big line bundles.
New DP algorithms achieve near-optimal regret bounds for online learning problems.
The following problem is addressed: A -manifold is endowed with a triple of closed -forms. One wants to construct a coframing of such that, first, for , and, second, the Riemannian metric $g=\big(ω^1\big)^2+\big(ω^2\big)^2+\…
This note displays an interesting phenomenon for percentiles of independent but non-identical random variables. Let be independent random variables obeying non-identical continuous distributions and be the corresponding order statistics. For any , we investig…
New insights into spectral statistics of sample covariance matrix for stable linear systems.
We use bank-level balance sheet data from 2005 to 2010 to study interactions within the banking system of five emerging countries: Argentina, Brazil, Mexico, South Africa, and Taiwan. For each country we construct a financial network based on the leverage ratio dependence between each pair of banks, and find results th…
Refines spinorial Sobolev inequality on sphere, proving stability and new properties of Killing spinors.
Study assesses 'big data' in materials science, highlighting challenges.
Over the last few years, traffic data has been exploding and the transportation discipline has entered the era of big data. It brings out new opportunities for doing data-driven analysis, but it also challenges traditional analytic methods. This paper proposes a new Divide and Combine based approach to do K means clust…
Uniform volume estimate for Kähler metrics in big cohomology classes.
Study convexity of Mabuchi functional in big cohomology classes.
Mobile big data contains vast statistical features in various dimensions, including spatial, temporal, and the underlying social domain. Understanding and exploiting the features of mobile data from a social network perspective will be extremely beneficial to wireless networks, from planning, operation, and maintenance…
Currently, the world is witnessing a mounting avalanche of data due to the increasing number of mobile network subscribers, Internet websites, and online services. This trend is continuing to develop in a quick and diverse manner in the form of big data. Big data analytics can process large amounts of raw data and extr…
Big data transforms accounting and auditing, enhancing insights but posing challenges.
Unified view on big bang singularities from initial data.
We consider model-free reinforcement learning for infinite-horizon discounted Markov Decision Processes (MDPs) with a continuous state space and unknown transition kernel, when only a single sample path under an arbitrary policy of the system is available. We consider the Nearest Neighbor Q-Learning (NNQL) algorithm to…
We study the local equivalence problem for real-analytic () hypersurfaces which, in coordinates with , are rigid: \[ u \,=\, F\big(z_1,z_2,\overline{z}_1,\overline{z}_2\big), \] with independent of . Specifically, we study th…
Big Data bring new opportunities to modern society and challenges to data scientists. On one hand, Big Data hold great promises for discovering subtle population patterns and heterogeneities that are not possible with small-scale data. On the other hand, the massive sample size and high dimensionality of Big Data intro…
The paper tackles machine unlearning by designing efficient algorithms for adaptive query classes.
Once first answers in any dimension to the Green-Griffiths and Kobayashi conjectures for generic algebraic hypersurfaces have been reached, the principal goal is to decrease (to improve) the degree bounds, knowing that the `celestial' horizon lies near $d \geqslant 2n…
In this paper, we consider the problem of sequentially optimizing a black-box function based on noisy samples and bandit feedback. We assume that is smooth in the sense of having a bounded norm in some reproducing kernel Hilbert space (RKHS), yielding a commonly-considered non-Bayesian form of Gaussian process …
In this note we show that many subgroups of mapping class groups of infinite-type surfaces without boundary have trivial centers, including all normal subgroups. Using similar techniques, we show that every nontrivial normal subgroup of a big mapping class group contains a nonabelian free group. In contrast, we show th…
Explosive growth in data and availability of cheap computing resources have sparked increasing interest in Big learning, an emerging subfield that studies scalable machine learning algorithms, systems, and applications with Big Data. Bayesian methods represent one important class of statistic methods for machine learni…
Data preprocessing techniques are devoted to correct or alleviate errors in data. Discretization and feature selection are two of the most extended data preprocessing techniques. Although we can find many proposals for static Big Data preprocessing, there is little research devoted to the continuous Big Data problem. A…
Characterizes and analyzes the large scale geometry of big mapping class groups of surfaces.
Proves existence of Kähler-Einstein metrics in big cohomology classes.
Big mapping class groups of infinite type surfaces have infinite asymptotic dimension.