Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

2955908841,179 · Jun 202019922001200920172026
48 results for Big Data analytics

Currently, the world is witnessing a mounting avalanche of data due to the increasing number of mobile network subscribers, Internet websites, and online services. This trend is continuing to develop in a quick and diverse manner in the form of big data. Big data analytics can process large amounts of raw data and extr…

2018-01-19abs ↗pdf ↗

Big data from phone calls improves credit scoring models and profits.

problem Improving credit scoring models to enhance financial inclusion.
method Combining call-detail records and traditional data to build scorecards using social network analytics.
result Combining call-detail records with traditional data significantly increases model performance and profit.

Study shows big winner stocks significantly impact passive and active investment strategies.

problem Impact of big winner stocks on passive and active investment strategies.
method Numerical and analytical techniques applied to historical stock price data.
result Concentrated portfolios underperform equally weighted indexes due to missing big winner stocks.

Loss Data Analytics is an interactive, online, freely available text. The idea behind the name Loss Data Analytics is to integrate classical loss data models from applied probability with modern analytic tools. In particular, we seek to recognize that big data (including social media and usage based insurance) are here…

2018-08-20abs ↗pdf ↗

Machine learning guides clinicians in predictive modeling using big data.

problem Insufficient understanding of machine learning among clinicians hinders its adoption.
method Provides a series of guides on machine learning principles, resampling, model evaluation, and coding.
result Clinicians need methodological rigor and clarity to use machine learning effectively.

Big data trend has enforced the data-centric systems to have continuous fast data streams. In recent years, real-time analytics on stream data has formed into a new research field, which aims to answer queries about what-is-happening-now with a negligible delay. The real challenge with real-time stream data processing …

2016-12-27abs ↗pdf ↗

Tensor completion is a problem of filling the missing or unobserved entries of partially observed tensors. Due to the multidimensional character of tensors in describing complex datasets, tensor completion algorithms and their applications have received wide attention and achievement in areas like data mining, computer…

2017-11-28abs ↗pdf ↗

The world is witnessing an unprecedented growth of cyber-physical systems (CPS), which are foreseen to revolutionize our world {via} creating new services and applications in a variety of sectors such as environmental monitoring, mobile-health systems, intelligent transportation systems and so on. The {information and …

2018-10-29abs ↗pdf ↗

This study analyzes and predicts airline delays using machine learning models.

problem Improving the accuracy of predicting airline flight delays.
method The study combines airline and weather datasets, using various machine learning models (Logistic Regression, Naive Bayes, K-NN, Decision Tree, Random Forest) to predict flight delays.
result The Random Forest model achieved an accuracy of 82% in predicting flight delays of 15 minutes or more.

The following problem is addressed: A 33-manifold MM is endowed with a triple Ω=(Ω1,Ω2,Ω3)Ω= \big(Ω^1,Ω^2,Ω^3\big) of closed 22-forms. One wants to construct a coframing ω=(ω1,ω2,ω3)ω= \big(ω^1,ω^2,ω^3\big) of MM such that, first, dωi=Ωi{\rm d}ω^i = Ω^i for i=1,2,3i=1,2,3, and, second, the Riemannian metric $g=\big(ω^1\big)^2+\big(ω^2\big)^2+\…

2019-08-02abs ↗pdf ↗

Proposes a Big Data framework for SC forecasting, including data preprocessing and machine learning.

problem Improving SC forecasting accuracy and efficiency.
method Data collection, preprocessing, machine learning model training, hyperparameter tuning, performance evaluation.
result Optimized SC forecasting models enhance workforce, inventory, and overall SC performance.

With the advent of Web 2.0, various types of data are being produced every day. This has led to the revolution of big data. Huge amount of structured and unstructured data are produced in financial markets. Processing these data could help an investor to make an informed investment decision. In this paper, a framework …

2018-11-17abs ↗pdf ↗

This paper surveys enterprise financial risk analysis from Big Data and LLMs perspectives.

problem Predicting future financial risk of enterprises.
method Systematic literature review of enterprise financial risk analysis approaches from Big Data and LLMs perspectives.
result Offers a holistic synthesis of research methods and key insights.

Three great theorems of Thurston read: Haken manifolds are hyperbolic; big ramified coverings are hyperbolic; big surgeries are hyperbolic. Recent developments indicate that the later two theorems are essentially a corollary of the first, that is there are much more Haken manifolds than expected by Thurston. In fact Fr…

1999-03-11abs ↗pdf ↗

Predicting the next position of movable objects has been a problem for at least the last three decades, referred to as trajectory prediction. In our days, the vast amounts of data being continuously produced add the big data dimension to the trajectory prediction problem, which we are trying to tackle by creating a λ-A…

2019-09-29abs ↗pdf ↗

We study the local equivalence problem for real-analytic (Cω\mathcal{C}^ω) hypersurfaces M5C3M^5 \subset \mathbb{C}^3 which, in coordinates (z1,z2,w)C3(z_1, z_2, w) \in \mathbb{C}^3 with w=u+ivw = u+i\, v, are rigid: \[ u \,=\, F\big(z_1,z_2,\overline{z}_1,\overline{z}_2\big), \] with FF independent of vv. Specifically, we study th…

2019-04-04abs ↗pdf ↗

Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…

2017-12-12abs ↗pdf ↗

Proves initial data on big bang singularities for Einstein-nonlinear scalar field equations lead to unique solutions.

problem Initial data on big bang singularities for Einstein equations.
method Geometric formulation of initial data, proving existence and uniqueness of solutions.
result Initial data on the singularity for the Einstein-nonlinear scalar field equations in 4 spacetime dimensions lead to a unique development of the data.

Let XX be a compact Kähler unibranch complex analytic space of pure dimension. Fix a big class αα with smooth representative θθ and a model potential φφ with positive mass. We define and the study non-pluripolar products of quasi-plurisubharmonic functions on XX. We study the spaces Ep(X,θ;[φ])\mathcal{E}^p(X,θ;[φ]) of fin…

2019-07-16abs ↗pdf ↗

Mobile big data contains vast statistical features in various dimensions, including spatial, temporal, and the underlying social domain. Understanding and exploiting the features of mobile data from a social network perspective will be extremely beneficial to wireless networks, from planning, operation, and maintenance…

2016-09-30abs ↗pdf ↗

New method improves likelihood-free parameter estimation in complex models.

problem Estimating parameters in simulation-based models with unknown likelihood.
method Nested multi-time-scale stochastic approximation (NMTS) method.
result Eliminates bias and accelerates convergence in likelihood-free inference.

As an emerging research direction, online streaming feature selection deals with sequentially added dimensions in a feature space while the number of data instances is fixed. Online streaming feature selection provides a new, complementary algorithmic methodology to enrich online feature selection, especially targets t…

2016-03-02abs ↗pdf ↗

Data preprocessing techniques are devoted to correct or alleviate errors in data. Discretization and feature selection are two of the most extended data preprocessing techniques. Although we can find many proposals for static Big Data preprocessing, there is little research devoted to the continuous Big Data problem. A…

2018-10-14abs ↗pdf ↗