Explosive growth in data and availability of cheap computing resources have sparked increasing interest in Big learning, an emerging subfield that studies scalable machine learning algorithms, systems, and applications with Big Data. Bayesian methods represent one important class of statistic methods for machine learni…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Big data transforms accounting and auditing, enhancing insights but posing challenges.
Study assesses 'big data' in materials science, highlighting challenges.
New DP algorithms achieve near-optimal regret bounds for online learning problems.
We consider model-free reinforcement learning for infinite-horizon discounted Markov Decision Processes (MDPs) with a continuous state space and unknown transition kernel, when only a single sample path under an arbitrary policy of the system is available. We consider the Nearest Neighbor Q-Learning (NNQL) algorithm to…
New algorithm learns halfspaces with adversarial noise efficiently.
Improved algorithm reduces stochastic gradient complexity for large-scale learning problems.
This study designs a financial risk control platform using big data and machine learning.
Improves EM algorithm for better local optima in mixture models.
New algorithm for private non-convex optimization with optimal rates.
Efficiently learns sparse halfspaces with noisy labels.
The paper tackles machine unlearning by designing efficient algorithms for adaptive query classes.
This paper investigates to identify the requirement and the development of machine learning-based mobile big data analysis through discussing the insights of challenges in the mobile big data (MBD). Furthermore, it reviews the state-of-the-art applications of data analysis in the area of MBD. Firstly, we introduce the …
New method improves efficiency analysis with big data.
We propose the online machine learning for big data analysis with heterogeneity. We performed an experiment to compare the accuracy of each iteration between batch one and online one. It is possible to converge quickly with the same accuracy as the batch one.
Paper optimizes a big data and ML risk monitoring system for financial markets.
Big data comes in various ways, types, shapes, forms and sizes. Indeed, almost all areas of science, technology, medicine, public health, economics, business, linguistics and social science are bombarded by ever increasing flows of data begging to analyzed efficiently and effectively. In this paper, we propose a rough …
In this short note, we formulate three problems relating to nonnegative scalar curvature (NNSC) fill-ins. Loosely speaking, the first two problems focus on: When are -dimensional Bartnik data , , NNSC-cobordant? (i.e., there is an -dimensional compact Riemannian manifold…
Study uses ML to analyze financial behavior in big data.
Big models pretrain and fine-tune for semi-supervised learning on ImageNet.
New algorithm finds approximate stationary points faster under differential privacy constraints.
The Era of Big Data has forced researchers to explore new distributed solutions for building fuzzy classifiers, which often introduce approximation errors or make strong assumptions to reduce computational and memory requirements. As a result, Big Data classifiers might be expected to be inferior to those designed for …
With the spreading prevalence of Big Data, many advances have recently been made in this field. Frameworks such as Apache Hadoop and Apache Spark have gained a lot of traction over the past decades and have become massively popular, especially in industries. It is becoming increasingly evident that effective big data a…
UCBVI-γ algorithm minimizes regret in discounted MDPs.
The paper improves smoothed analysis for online problems with adaptive adversaries.
Automatically adjusts model size for continual Gaussian processes.
Meta-learning helps use small data from many tasks to compensate for lack of big data.
We introduce a means of automating machine learning (ML) for big data tasks, by performing scalable stochastic Bayesian optimisation of ML algorithm parameters and hyper-parameters. More often than not, the critical tuning of ML algorithm parameters has relied on domain expertise from experts, along with laborious hand…
In Part III of this study, we apply the price dynamical model with big buyers and big sellers developed in Part I of this paper to the daily closing prices of the top 20 banking and real estate stocks listed in the Hong Kong Stock Exchange. The basic idea is to estimate the strength parameters of the big buyers and the…
Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…
In this paper we consider the large genus asymptotics for two classes of Siegel-Veech constants associated with an arbitrary connected stratum of Abelian differentials. The first is the saddle connection Siegel-Veech constant counting saddle conne…
Establishes Kobayashi-Hitchin correspondence for nef and big classes.
Given compact Kähler manifold and a model type envelope with non-zero mass, i.e. a fixed potential determing some singularities such that , we prove that the relative finite energy class becomes a complete metric space…
The need for new methods to deal with big data is a common theme in most scientific fields, although its definition tends to vary with the context. Statistical ideas are an essential part of this, and as a partial response, a thematic program on statistical inference, learning, and models in big data was held in 2015 i…
This article provides the role of big idea statisticians in future of Big Data Science. We describe the `United Statistical Algorithms' framework for comprehensive unification of traditional and novel statistical methods for modeling Small Data and Big Data, especially mixed data (discrete, continuous).
Geometric quantization extended to big line bundles.
Efficiently learns halfspaces with malicious noise, near-optimal label complexity.
This paper deals with bandit online learning problems involving feedback of unknown delay that can emerge in multi-armed bandit (MAB) and bandit convex optimization (BCO) settings. MAB and BCO require only values of the objective function involved that become available through feedback, and are used to estimate the gra…
This paper surveys AUC maximization for big data and AI.
Finding efficient and provable methods to solve non-convex optimization problems is an outstanding challenge in machine learning and optimization theory. A popular approach used to tackle non-convex problems is to use convex relaxation techniques to find a convex surrogate for the problem. Unfortunately, convex relaxat…
Big Data is one of the major challenges of statistical science and has numerous consequences from algorithmic and theoretical viewpoints. Big Data always involve massive data but they also often include online data and data heterogeneity. Recently some statistical methods have been adapted to process Big Data, like lin…
The following problem is addressed: A -manifold is endowed with a triple of closed -forms. One wants to construct a coframing of such that, first, for , and, second, the Riemannian metric $g=\big(ω^1\big)^2+\big(ω^2\big)^2+\…
This note displays an interesting phenomenon for percentiles of independent but non-identical random variables. Let be independent random variables obeying non-identical continuous distributions and be the corresponding order statistics. For any , we investig…
New insights into spectral statistics of sample covariance matrix for stable linear systems.
Refines spinorial Sobolev inequality on sphere, proving stability and new properties of Killing spinors.
Interpretability has always been a major concern for fuzzy rule-based classifiers. The usage of human-readable models allows them to explain the reasoning behind their predictions and decisions. However, when it comes to Big Data classification problems, fuzzy rule-based classifiers have not been able to maintain the g…
Uniform volume estimate for Kähler metrics in big cohomology classes.
Paper introduces slow kill for efficient large-scale variable screening.