Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In Part III of this study, we apply the price dynamical model with big buyers and big sellers developed in Part I of this paper to the daily closing prices of the top 20 banking and real estate stocks listed in the Hong Kong Stock Exchange. The basic idea is to estimate the strength parameters of the big buyers and the…
Study assesses 'big data' in materials science, highlighting challenges.
This article provides the role of big idea statisticians in future of Big Data Science. We describe the `United Statistical Algorithms' framework for comprehensive unification of traditional and novel statistical methods for modeling Small Data and Big Data, especially mixed data (discrete, continuous).
New algorithm finds approximate stationary points faster under differential privacy constraints.
The paper proposes a new model using financial big data to improve portfolio risk analysis.
Extends K-stability theory to projective klt pairs with a big anticanonical class.
The paper tackles machine unlearning by designing efficient algorithms for adaptive query classes.
Given compact Kähler manifold and a model type envelope with non-zero mass, i.e. a fixed potential determing some singularities such that , we prove that the relative finite energy class becomes a complete metric space…
Two algorithms improve fitting autoregressive models for big data.
In this short note, we formulate three problems relating to nonnegative scalar curvature (NNSC) fill-ins. Loosely speaking, the first two problems focus on: When are -dimensional Bartnik data , , NNSC-cobordant? (i.e., there is an -dimensional compact Riemannian manifold…
This paper uses information theory to improve risk modeling in big data.
Explosive growth in data and availability of cheap computing resources have sparked increasing interest in Big learning, an emerging subfield that studies scalable machine learning algorithms, systems, and applications with Big Data. Bayesian methods represent one important class of statistic methods for machine learni…
We consider model-free reinforcement learning for infinite-horizon discounted Markov Decision Processes (MDPs) with a continuous state space and unknown transition kernel, when only a single sample path under an arbitrary policy of the system is available. We consider the Nearest Neighbor Q-Learning (NNQL) algorithm to…
Improves EM algorithm for better local optima in mixture models.
Localized big bang singularities found without background solutions.
New algorithm for private non-convex optimization with optimal rates.
The paper improves smoothed analysis for online problems with adaptive adversaries.
We study the local equivalence problem for real-analytic () hypersurfaces which, in coordinates with , are rigid: \[ u \,=\, F\big(z_1,z_2,\overline{z}_1,\overline{z}_2\big), \] with independent of . Specifically, we study th…
The Era of Big Data has forced researchers to explore new distributed solutions for building fuzzy classifiers, which often introduce approximation errors or make strong assumptions to reduce computational and memory requirements. As a result, Big Data classifiers might be expected to be inferior to those designed for …
In this paper we consider the large genus asymptotics for two classes of Siegel-Veech constants associated with an arbitrary connected stratum of Abelian differentials. The first is the saddle connection Siegel-Veech constant counting saddle conne…
This paper investigates to identify the requirement and the development of machine learning-based mobile big data analysis through discussing the insights of challenges in the mobile big data (MBD). Furthermore, it reviews the state-of-the-art applications of data analysis in the area of MBD. Firstly, we introduce the …
This study designs a financial risk control platform using big data and machine learning.
Establishes Kobayashi-Hitchin correspondence for nef and big classes.
Geometric quantization extended to big line bundles.
New DP algorithms achieve near-optimal regret bounds for online learning problems.
The following problem is addressed: A -manifold is endowed with a triple of closed -forms. One wants to construct a coframing of such that, first, for , and, second, the Riemannian metric $g=\big(ω^1\big)^2+\big(ω^2\big)^2+\…
We show that the Big Bang singularity of the Friedmann-Lemaitre-Robertson-Walker model does not raise major problems to General Relativity. We prove a theorem showing that the Einstein equation can be written in a non-singular form, which allows the extension of the spacetime before the Big Bang. The physical interpret…
This note displays an interesting phenomenon for percentiles of independent but non-identical random variables. Let be independent random variables obeying non-identical continuous distributions and be the corresponding order statistics. For any , we investig…
New insights into spectral statistics of sample covariance matrix for stable linear systems.
Refines spinorial Sobolev inequality on sphere, proving stability and new properties of Killing spinors.
New conic quadratic formulations improve outlier detection in regression models.
In real world industrial applications of topic modeling, the ability to capture gigantic conceptual space by learning an ultra-high dimensional topical representation, i.e., the so-called "big model", is becoming the next desideratum after enthusiasms on "big data", especially for fine-grained downstream tasks such as …
Improved algorithm reduces stochastic gradient complexity for large-scale learning problems.
Uniform volume estimate for Kähler metrics in big cohomology classes.
MiMuon optimizer improves generalization for large models by reducing generalization error.
New algorithm learns halfspaces with adversarial noise efficiently.
Study convexity of Mabuchi functional in big cohomology classes.
Mobile big data contains vast statistical features in various dimensions, including spatial, temporal, and the underlying social domain. Understanding and exploiting the features of mobile data from a social network perspective will be extremely beneficial to wireless networks, from planning, operation, and maintenance…
Currently, the world is witnessing a mounting avalanche of data due to the increasing number of mobile network subscribers, Internet websites, and online services. This trend is continuing to develop in a quick and diverse manner in the form of big data. Big data analytics can process large amounts of raw data and extr…
To solve the big topic modeling problem, we need to reduce both time and space complexities of batch latent Dirichlet allocation (LDA) algorithms. Although parallel LDA algorithms on the multi-processor architecture have low time and space complexities, their communication costs among processors often scale linearly wi…
The paper proves an equilibrium in a limited stock market participation model with power utilities.
Einstein's equation, in its standard form, breaks down at the Big Bang singularity. A new version, equivalent to Einstein's whenever the latter is defined, but applicable in wider situations, is proposed. The new equation remains smooth at the Big Bang singularity of the Friedmann-Lemaitre-Robertson-Walker model. It is…
Big data transforms accounting and auditing, enhancing insights but posing challenges.
New method improves likelihood-free parameter estimation in complex models.
Study uses ML to analyze financial behavior in big data.
We study the statistical performance of semidefinite programming (SDP) relaxations for clustering under random graph models. Under the Synchronization model, Censored Block Model and Stochastic Block Model, we show that SDP achieves an error rate of the form \[ \exp\Big[-\big(1-o(1)\big)\bar{n} I^* \Bi…
Unified view on big bang singularities from initial data.