Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

3416811,0221,362 · Jun 202019922001200920182026
48 results for big data analysis

This study designs a financial risk control platform using big data and machine learning.

problem Traditional risk management models are inadequate for modern financial complexities.
method Big data mining, real-time streaming data processing, statistical analysis, and precise customer behavior mining.
result The platform effectively identifies and responds to potential risks in real-time.

The paper proposes a new model using financial big data to improve portfolio risk analysis.

problem Addressing potential information loss in portfolio risk measurement.
method Uses financial big data to incorporate out-of-target-portfolio information and overcomes the curse of dimensionality.
result The use of financial big data improves small portfolio risk analysis.

Astronomy needs efficient machine learning and image analysis for big data.

problem Efficient machine learning and image analysis for big astronomical data.
method Exemplary results, challenges, and methodological advancements in machine learning and image analysis.
result Astronomy pushes the boundaries of data analysis in machine learning.

Big Data bring new opportunities to modern society and challenges to data scientists. On one hand, Big Data hold great promises for discovering subtle population patterns and heterogeneities that are not possible with small-scale data. On the other hand, the massive sample size and high dimensionality of Big Data intro…

2013-08-07abs ↗pdf ↗

LSAR efficiently estimates AR models for big time series data.

problem Efficiently analyzing large-scale time series data with high accuracy.
method Developed a fast algorithm to estimate leverage scores and an efficient LSAR algorithm for fitting AR models.
result LSAR algorithm finds maximum likelihood estimates with high probability and improved worst-case running time.

Neuroscience faces new challenges in data analysis as datasets grow richer.

problem How to analyze large, complex neuroscientific datasets effectively.
method Development of non-parametric, generative models combining frequentist and Bayesian approaches.
result New statistical methods will be essential for extracting meaningful insights from neuroscientific data.

This paper surveys enterprise financial risk analysis from Big Data and LLMs perspectives.

problem Predicting future financial risk of enterprises.
method Systematic literature review of enterprise financial risk analysis approaches from Big Data and LLMs perspectives.
result Offers a holistic synthesis of research methods and key insights.

This paper analyzes interactive model analysis for machine learning.

problem Understanding, diagnosing, and refining machine learning models.
method Classification of relevant work into understanding, diagnosis, and refinement categories.
result Exploration of future research opportunities in interactive model analysis.

New method for tensor completion using nonconvex dual total variation.

problem Tensor completion from partial measurements with exponential-family noise.
method Proposed dual-TV (DTV) regularizers for tensor completion under exponential-family noise.
result Theoretical upper bounds on recovery error for tensor completion.

Develops an efficient online robust PCA method for big data.

problem Efficiency and robustness in processing big data with changing subspaces.
method Online moving window robust principal component analysis (OMWRPCA) with change point detection.
result Successfully tracks both slowly and abruptly changing subspaces and detects change points.

The paper proposes a simple algorithm for finding low-dimensional manifolds in big data.

problem Finding low-dimensional manifolds in large datasets.
method A novel algorithm for identifying low-dimensional manifolds in high-dimensional data.
result The algorithm effectively finds low-dimensional manifolds in data, outperforming classical methods like PCA and Isomap.

Parallelizes Bayesian MCMC for big data, improving efficiency and speed.

problem Efficiently analyzing large Bayesian hierarchical models with big data.
method Two-stage approach: first stage estimates group-specific parameters in parallel, second stage uses stage 1 posteriors as proposals.
result Agrees with full data analysis but with increased efficiency and reduced computation times.

CCA helps find hidden connections in complex biomedical data.

problem Analyzing large, multi-variable datasets in biology and medicine.
method Canonical correlation analysis (CCA) for exploring relationships between two sets of variables.
result CCA uncovers essential hidden associations between diverse data types.

Framework optimizes portfolios using big data from financial markets.

problem Optimizing investment decisions with structured and unstructured financial data.
method 5-stage methodology including DEA, text mining, clustering, ranking, and heuristics for portfolio optimization.
result Helps investors select, weight, and manage assets for informed investment decisions.

Divide data into subsets, analyze each, and recombine results for likelihood function computation.

problem Computing likelihood functions for large and complex data.
method Divide & Recombine (D&R) procedure to estimate density parameters of likelihood model (LM) from MCMC draws.
result The method successfully computes likelihood functions for logistic regression data model.

Proposes a Big Data framework for SC forecasting, including data preprocessing and machine learning.

problem Improving SC forecasting accuracy and efficiency.
method Data collection, preprocessing, machine learning model training, hyperparameter tuning, performance evaluation.
result Optimized SC forecasting models enhance workforce, inventory, and overall SC performance.

The paper improves smoothed analysis for online problems with adaptive adversaries.

problem Online prediction, discrepancy minimization, and online optimization with adaptive adversaries.
method General technique to prove smoothed guarantees against adaptive adversaries, reducing to simpler oblivious adversaries.
result Strong smoothed guarantees for three online problems, matching or improving previous results.

DeepDeath models predict underlying causes of death using big data.

problem Predicting health trajectories in large populations.
method Design and apply two model classes: Hadoop-based ensemble of random forests and DeepDeath (RNN with LSTMs).
result DeepDeath model outperforms N-gram models, learning temporal data aspects.

Unsupervised learning identifies phases and transitions in complex systems.

problem Discovering hidden patterns and phases in large datasets.
method Raw spin configurations were analyzed using principal component analysis and clustering.
result Unsupervised learning successfully identifies physical concepts like order parameter and structure factor.

Bayesian SVARs improve model construction and policy analysis in big data.

problem Manual selection of variables in SVAR models limits their applicability in big data.
method Develops a Bayesian methodology for constructing information sets and retaining the largest system.
result Output increases with housing production over household credit in SVAR models.