Currently, the world is witnessing a mounting avalanche of data due to the increasing number of mobile network subscribers, Internet websites, and online services. This trend is continuing to develop in a quick and diverse manner in the form of big data. Big data analytics can process large amounts of raw data and extr…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Big data analytics improves healthcare through early detection and quality life.
Study uses ML to analyze financial behavior in big data.
Efficiently augments triplet data for better data analytics.
Big data transforms accounting and auditing, enhancing insights but posing challenges.
Future buildings will offer new convenience, comfort, and efficiency possibilities to their residents. Changes will occur to the way people live as technology involves into people's lives and information processing is fully integrated into their daily living activities and objects. The future expectation of smart build…
Big data from phone calls improves credit scoring models and profits.
Study shows big winner stocks significantly impact passive and active investment strategies.
Loss Data Analytics is an interactive, online, freely available text. The idea behind the name Loss Data Analytics is to integrate classical loss data models from applied probability with modern analytic tools. In particular, we seek to recognize that big data (including social media and usage based insurance) are here…
Recent years have witnessed amazing outcomes from "Big Models" trained by "Big Data". Most popular algorithms for model training are iterative. Due to the surging volumes of data, we can usually afford to process only a fraction of the training data in each iteration. Typically, the data are either uniformly sampled or…
Machine learning guides clinicians in predictive modeling using big data.
The paper extends Nakamaye's theorem to non-closed forms on complex manifolds.
Big data trend has enforced the data-centric systems to have continuous fast data streams. In recent years, real-time analytics on stream data has formed into a new research field, which aims to answer queries about what-is-happening-now with a negligible delay. The real challenge with real-time stream data processing …
Statistical analysis (SA) is a complex process to deduce population properties from analysis of data. It usually takes a well-trained analyst to successfully perform SA, and it becomes extremely challenging to apply SA to big data applications. We propose to use deep neural networks to automate the SA process. In parti…
In Divide & Recombine (D&R), big data are divided into subsets, each analytic method is applied to subsets, and the outputs are recombined. This enables deep analysis and practical computational performance. An innovate D\&R procedure is proposed to compute likelihood functions of data-model (DM) parameters for big dat…
Tensor completion is a problem of filling the missing or unobserved entries of partially observed tensors. Due to the multidimensional character of tensors in describing complex datasets, tensor completion algorithms and their applications have received wide attention and achievement in areas like data mining, computer…
New method predicts customer churn using mixed-penalty logistic regression.
The world is witnessing an unprecedented growth of cyber-physical systems (CPS), which are foreseen to revolutionize our world {via} creating new services and applications in a variety of sectors such as environmental monitoring, mobile-health systems, intelligent transportation systems and so on. The {information and …
This study analyzes and predicts airline delays using machine learning models.
In recent years, analyzing task-based fMRI (tfMRI) data has become an essential tool for understanding brain function and networks. However, due to the sheer size of tfMRI data, its intrinsic complex structure, and lack of ground truth of underlying neural activities, modeling tfMRI data is hard and challenging. Previo…
Paper presents a data preprocessing method for PHM models.
Interactive model analysis, the process of understanding, diagnosing, and refining a machine learning model with the help of interactive visualization, is very important for users to efficiently solve real-world artificial intelligence and data mining problems. Dramatic advances in big data analytics has led to a wide …
Mobile networks possess information about the users as well as the network. Such information is useful for making the network end-to-end visible and intelligent. Big data analytics can efficiently analyze user and network information, unearth meaningful insights with the help of machine learning tools. Utilizing big da…
The following problem is addressed: A -manifold is endowed with a triple of closed -forms. One wants to construct a coframing of such that, first, for , and, second, the Riemannian metric $g=\big(ω^1\big)^2+\big(ω^2\big)^2+\…
Proposes a Big Data framework for SC forecasting, including data preprocessing and machine learning.
Machine learning knot invariants with physics applications.
With the advent of Web 2.0, various types of data are being produced every day. This has led to the revolution of big data. Huge amount of structured and unstructured data are produced in financial markets. Processing these data could help an investor to make an informed investment decision. In this paper, a framework …
The tremendous growth of positioning technologies and GPS enabled devices has produced huge volumes of tracking data during the recent years. This source of information constitutes a rich input for data analytics processes, either offline (e.g. cluster analysis, hot motion discovery) or online (e.g. short-term forecast…
The purpose of this paper is to establish a Nadel vanishing theorem for big line bundles with multiplier ideal sheaves of singular metrics admitting an analytic Zariski decomposition (such as, metrics with minimal singularities and Siu's metrics). For this purpose, we apply the theory of harmonic integrals and generali…
Marketing analytics is a diverse field, with both academic researchers and practitioners coming from a range of backgrounds including marketing, expert systems, statistics, and operations research. This paper provides an integrative review at the boundary of these areas. The aim is to give researchers in the intelligen…
This paper surveys enterprise financial risk analysis from Big Data and LLMs perspectives.
Three great theorems of Thurston read: Haken manifolds are hyperbolic; big ramified coverings are hyperbolic; big surgeries are hyperbolic. Recent developments indicate that the later two theorems are essentially a corollary of the first, that is there are much more Haken manifolds than expected by Thurston. In fact Fr…
Predicting the next position of movable objects has been a problem for at least the last three decades, referred to as trajectory prediction. In our days, the vast amounts of data being continuously produced add the big data dimension to the trajectory prediction problem, which we are trying to tackle by creating a λ-A…
We study the local equivalence problem for real-analytic () hypersurfaces which, in coordinates with , are rigid: \[ u \,=\, F\big(z_1,z_2,\overline{z}_1,\overline{z}_2\big), \] with independent of . Specifically, we study th…
Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…
Study shows volumes of complex classes can be represented by convex bodies.
Proves initial data on big bang singularities for Einstein-nonlinear scalar field equations lead to unique solutions.
This article provides the role of big idea statisticians in future of Big Data Science. We describe the `United Statistical Algorithms' framework for comprehensive unification of traditional and novel statistical methods for modeling Small Data and Big Data, especially mixed data (discrete, continuous).
Let be a compact Kähler unibranch complex analytic space of pure dimension. Fix a big class with smooth representative and a model potential with positive mass. We define and the study non-pluripolar products of quasi-plurisubharmonic functions on . We study the spaces of fin…
New problems on NNSC fill-ins for Bartnik data in high dimensions.
It is proved by Kawamata that the canonical bundle of a projective manifold is semi-ample if it is big and nef. We give an analytic proof using the Ricci flow, degeneration of Riemannian manifolds and -theory. Combined with our earlier results, we construct unique singular Kahler-Einstein metrics with a global Rie…
Deep learning faces adoption challenges in business analytics.
Study assesses 'big data' in materials science, highlighting challenges.
Mobile big data contains vast statistical features in various dimensions, including spatial, temporal, and the underlying social domain. Understanding and exploiting the features of mobile data from a social network perspective will be extremely beneficial to wireless networks, from planning, operation, and maintenance…
New method improves likelihood-free parameter estimation in complex models.
Nowadays processing of Big Security Data, such as log messages, is commonly used for intrusion detection purposed. Its heterogeneous nature, as well as combination of numerical and categorical attributes does not allow to apply the existing data mining methods directly on the data without feature preprocessing. Therefo…
As an emerging research direction, online streaming feature selection deals with sequentially added dimensions in a feature space while the number of data instances is fixed. Online streaming feature selection provides a new, complementary algorithmic methodology to enrich online feature selection, especially targets t…
Data preprocessing techniques are devoted to correct or alleviate errors in data. Discretization and feature selection are two of the most extended data preprocessing techniques. Although we can find many proposals for static Big Data preprocessing, there is little research devoted to the continuous Big Data problem. A…