Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

3877741,1611,548 · Jun 202019922001200920172026
48 results for big learning

Explosive growth in data and availability of cheap computing resources have sparked increasing interest in Big learning, an emerging subfield that studies scalable machine learning algorithms, systems, and applications with Big Data. Bayesian methods represent one important class of statistic methods for machine learni…

2014-11-24abs ↗pdf ↗

New DP algorithms achieve near-optimal regret bounds for online learning problems.

problem Online learning problems with zero-loss solutions and differential privacy constraints.
method Developed new Differentially Private algorithms with near-optimal regret bounds.
result Achieved near-optimal regret bounds for various online prediction and convex optimization problems.

We consider model-free reinforcement learning for infinite-horizon discounted Markov Decision Processes (MDPs) with a continuous state space and unknown transition kernel, when only a single sample path under an arbitrary policy of the system is available. We consider the Nearest Neighbor Q-Learning (NNQL) algorithm to…

2018-02-12abs ↗pdf ↗

New algorithm learns halfspaces with adversarial noise efficiently.

problem Learning halfspaces in the presence of adversarial noise.
method Polynomial-time Perceptron-like online active learning algorithm.
result Near-optimal label and sample complexity with isotropic log-concave marginal distribution.

Improved algorithm reduces stochastic gradient complexity for large-scale learning problems.

problem High stochastic gradient complexity for large-scale learning problems.
method Hybrid Stochastic-Deterministic Minibatch Proximal Gradient (HSDMPG) algorithm.
result Achieves nearly optimal generalization in less than a single pass over data.

This study designs a financial risk control platform using big data and machine learning.

problem Traditional risk management models are inadequate for modern financial complexities.
method Big data mining, real-time streaming data processing, statistical analysis, and precise customer behavior mining.
result The platform effectively identifies and responds to potential risks in real-time.

The paper tackles machine unlearning by designing efficient algorithms for adaptive query classes.

problem Designing efficient unlearning algorithms for machine learning models.
method Formalizes the problem and gives efficient unlearning algorithms for linear and prefix-sum query classes.
result Improved guarantees for stochastic convex optimization with reduced unlearning query complexity.

This paper investigates to identify the requirement and the development of machine learning-based mobile big data analysis through discussing the insights of challenges in the mobile big data (MBD). Furthermore, it reviews the state-of-the-art applications of data analysis in the area of MBD. Firstly, we introduce the …

2018-08-02abs ↗pdf ↗

We propose the online machine learning for big data analysis with heterogeneity. We performed an experiment to compare the accuracy of each iteration between batch one and online one. It is possible to converge quickly with the same accuracy as the batch one.

2019-06-15abs ↗pdf ↗

Paper optimizes a big data and ML risk monitoring system for financial markets.

problem Traditional risk monitoring methods are inadequate for modern financial markets due to data complexity and volume.
method Four-layer architecture integrating big data and advanced ML algorithms (LSTM, RF, GB).
result Significantly enhances efficiency and accuracy in risk management, especially in market crash risk detection.

In this short note, we formulate three problems relating to nonnegative scalar curvature (NNSC) fill-ins. Loosely speaking, the first two problems focus on: When are (n1)(n-1)-dimensional Bartnik data (Σin1,γi,Hi)\big(Σ_i ^{n-1}, γ_i, H_i\big), i=1,2i=1,2, NNSC-cobordant? (i.e., there is an nn-dimensional compact Riemannian manifold…

2020-01-16abs ↗pdf ↗

Big models pretrain and fine-tune for semi-supervised learning on ImageNet.

problem Learning from few labeled examples with a large amount of unlabeled data.
method Unsupervised pretraining of a big ResNet model followed by supervised fine-tuning and distillation.
result 73.9% ImageNet top-1 accuracy with just 1% of labels (\le13 labeled images per class).

New algorithm finds approximate stationary points faster under differential privacy constraints.

problem Finding approximate stationary points of smooth and Lipschitz functions under differential privacy constraints.
method Developed an efficient algorithm that improves convergence rates to stationary points.
result Achieved faster rates of convergence to stationary points in both finite-sum and stochastic settings.

With the spreading prevalence of Big Data, many advances have recently been made in this field. Frameworks such as Apache Hadoop and Apache Spark have gained a lot of traction over the past decades and have become massively popular, especially in industries. It is becoming increasingly evident that effective big data a…

2017-11-25abs ↗pdf ↗

The paper improves smoothed analysis for online problems with adaptive adversaries.

problem Online prediction, discrepancy minimization, and online optimization with adaptive adversaries.
method General technique to prove smoothed guarantees against adaptive adversaries, reducing to simpler oblivious adversaries.
result Strong smoothed guarantees for three online problems, matching or improving previous results.

Meta-learning helps use small data from many tasks to compensate for lack of big data.

problem How to leverage small labeled data from many tasks to improve learning when big labeled data is scarce.
method Introduced a novel spectral approach to efficiently utilize small data tasks with the help of medium data tasks.
result The total number of examples necessary with only small data tasks scales similarly as when big data tasks are available.

Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…

2017-12-12abs ↗pdf ↗

In this paper we consider the large genus asymptotics for two classes of Siegel-Veech constants associated with an arbitrary connected stratum H(α)\mathcal{H} (α) of Abelian differentials. The first is the saddle connection Siegel-Veech constant cscmi,mj(H(α))c_{\text{sc}}^{m_i, m_j} \big( \mathcal{H} (α) \big) counting saddle conne…

2018-10-11abs ↗pdf ↗

Given (X,ω)(X,ω) compact Kähler manifold and ψM+PSH(X,ω)ψ\in\mathcal{M}^{+}\subset PSH(X,ω) a model type envelope with non-zero mass, i.e. a fixed potential determing some singularities such that X(ω+ddcψ)n>0\int_{X}(ω+dd^{c}ψ)^{n}>0, we prove that the ψψ-relative finite energy class E1(X,ω,ψ)\mathcal{E}^{1}(X,ω,ψ) becomes a complete metric space…

2019-09-09abs ↗pdf ↗

The need for new methods to deal with big data is a common theme in most scientific fields, although its definition tends to vary with the context. Statistical ideas are an essential part of this, and as a partial response, a thematic program on statistical inference, learning, and models in big data was held in 2015 i…

2015-09-09abs ↗pdf ↗

Efficiently learns halfspaces with malicious noise, near-optimal label complexity.

problem Learning ss-sparse halfspaces under malicious label noise.
method Active learning algorithm with instance reweighting and empirical risk minimization.
result Near-optimal label complexity of O(slog4d/ε)O(s \log^4 d / ε) and noise tolerance Ω(ε)Ω(ε).

This paper deals with bandit online learning problems involving feedback of unknown delay that can emerge in multi-armed bandit (MAB) and bandit convex optimization (BCO) settings. MAB and BCO require only values of the objective function involved that become available through feedback, and are used to estimate the gra…

2018-07-09abs ↗pdf ↗

Big Data is one of the major challenges of statistical science and has numerous consequences from algorithmic and theoretical viewpoints. Big Data always involve massive data but they also often include online data and data heterogeneity. Recently some statistical methods have been adapted to process Big Data, like lin…

2015-11-26abs ↗pdf ↗

The following problem is addressed: A 33-manifold MM is endowed with a triple Ω=(Ω1,Ω2,Ω3)Ω= \big(Ω^1,Ω^2,Ω^3\big) of closed 22-forms. One wants to construct a coframing ω=(ω1,ω2,ω3)ω= \big(ω^1,ω^2,ω^3\big) of MM such that, first, dωi=Ωi{\rm d}ω^i = Ω^i for i=1,2,3i=1,2,3, and, second, the Riemannian metric $g=\big(ω^1\big)^2+\big(ω^2\big)^2+\…

2019-08-02abs ↗pdf ↗

This note displays an interesting phenomenon for percentiles of independent but non-identical random variables. Let X1,,XnX_1,\cdots,X_n be independent random variables obeying non-identical continuous distributions and X(1)X(n)X^{(1)}\geq \cdots\geq X^{(n)} be the corresponding order statistics. For any p(0,1)p\in(0,1), we investig…

2018-08-24abs ↗pdf ↗

New insights into spectral statistics of sample covariance matrix for stable linear systems.

problem Estimating high-dimensional stable state transition matrices from noisy data.
method Combining spectral theorem for non-Hermitian operators, concentration of measure, and perturbation theory.
result The spectral radius of the sample covariance matrix exhibits phase transitions in high dimensions.

Uniform volume estimate for Kähler metrics in big cohomology classes.

problem Estimating volume for singular Kähler metrics in big cohomology classes.
method Generalized mixed energy estimate for functions in complex Sobolev space to big cohomology classes.
result Uniform non-collapsing volume estimate for local Kähler metrics.