Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

2885768641,152 · Jun 202019922001200920172026
48 results for financial big data

This study designs a financial risk control platform using big data and machine learning.

problem Traditional risk management models are inadequate for modern financial complexities.
method Big data mining, real-time streaming data processing, statistical analysis, and precise customer behavior mining.
result The platform effectively identifies and responds to potential risks in real-time.

The paper proposes a new model using financial big data to improve portfolio risk analysis.

problem Addressing potential information loss in portfolio risk measurement.
method Uses financial big data to incorporate out-of-target-portfolio information and overcomes the curse of dimensionality.
result The use of financial big data improves small portfolio risk analysis.

Paper optimizes a big data and ML risk monitoring system for financial markets.

problem Traditional risk monitoring methods are inadequate for modern financial markets due to data complexity and volume.
method Four-layer architecture integrating big data and advanced ML algorithms (LSTM, RF, GB).
result Significantly enhances efficiency and accuracy in risk management, especially in market crash risk detection.

This paper surveys enterprise financial risk analysis from Big Data and LLMs perspectives.

problem Predicting future financial risk of enterprises.
method Systematic literature review of enterprise financial risk analysis approaches from Big Data and LLMs perspectives.
result Offers a holistic synthesis of research methods and key insights.

Neural networks model financial data with Lévy processes.

problem Forecasting chaotic financial time series with big jumps.
method Lévy-induced stochastic differential equation network approximated by neural networks.
result The method improves prediction accuracy using non-Gaussian Lévy processes.

Big data from phone calls improves credit scoring models and profits.

problem Improving credit scoring models to enhance financial inclusion.
method Combining call-detail records and traditional data to build scorecards using social network analytics.
result Combining call-detail records with traditional data significantly increases model performance and profit.

A new method combines federated learning and logistic regression for better credit scoring.

problem Improving credit scoring models while protecting data privacy.
method Projected gradient-based vertical federated learning (FL-LRBC) for logistic regression.
result Significant improvement in AUC and KS statistics due to data enrichment.

An analysis of the stylized facts in financial time series is carried out. We find that, instead of the heavy tails in asset return distributions, the slow decay behaviour in autocorrelation functions of absolute returns is actually directly related to the degree of clustering of large fluctuations within the financial…

2010-02-01abs ↗pdf ↗

Paper proposes a method to evaluate SME credit risk using meta paths.

problem Evaluate credit risk of small and medium-sized enterprises with limited data.
method Exploits the representative power of information networks and meta paths to infer SME financial status.
result Meta path feature effectively identifies SMEs with credit risks.

A method uses Wasserstein clustering to simplify financial data analysis.

problem Processing and analyzing granular financial data with missing values and identifying clusters.
method Variant of Lloyd's algorithm applied to probability distributions, using Wasserstein barycenters.
result Demonstrated usefulness in financial regulation context.

Agent-based modeling is a powerful simulation technique to understand the collective behavior and microscopic interaction in complex financial systems. Recently, the concept for determining the key parameters of the agent-based models from empirical data instead of setting them artificially was suggested. We first revi…

2017-03-04abs ↗pdf ↗

Study shows big winner stocks significantly impact passive and active investment strategies.

problem Impact of big winner stocks on passive and active investment strategies.
method Numerical and analytical techniques applied to historical stock price data.
result Concentrated portfolios underperform equally weighted indexes due to missing big winner stocks.

AI-driven framework improves enterprise financial audits and risk identification.

problem Manual auditing is inefficient and limited by data complexity and evolving fraud tactics.
method Machine learning algorithms (SVM, RF, KNN) applied to a dataset of audit project counts, violations, and fraud instances.
result Random Forest achieves best performance with F1-score of 0.9012, identifying fraud and compliance anomalies.

FinRL-Meta creates diverse market environments for DRL in finance.

problem Inaccurate financial data and diverse market environments challenge DRL in finance.
method Open-source data processing tools, hundreds of market environments, and multiprocessing.
result FinRL-Meta improves DRL accuracy and speed in financial simulations.

With the advent of Web 2.0, various types of data are being produced every day. This has led to the revolution of big data. Huge amount of structured and unstructured data are produced in financial markets. Processing these data could help an investor to make an informed investment decision. In this paper, a framework …

2018-11-17abs ↗pdf ↗

Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…

2017-12-12abs ↗pdf ↗

GRTR framework uses graph regularization to improve financial forecasting.

problem High computational costs and economic domain knowledge loss in tensor models.
method Graph-Regularized Tensor Regression (GRTR) framework incorporating economic domain knowledge.
result Improved performance in multi-way financial forecasting with reduced computational costs.

MegazordNet combines stats and ML for better financial time series forecasting.

problem Forecasting financial time series is challenging due to its chaotic nature.
method MegazordNet integrates statistical features with a deep learning model.
result MegazordNet outperforms single statistical and machine learning methods in S&P 500 stock price prediction.

The task of predicting future stock values has always been one that is heavily desired albeit very difficult. This difficulty arises from stocks with non-stationary behavior, and without any explicit form. Hence, predictions are best made through analysis of financial stock data. To handle big data sets, current conven…

2019-04-17abs ↗pdf ↗

We study the most famous example of a large financial market: the Arbitrage Pricing Model, where investors can trade in a one-period setting with countably many assets admitting a factor structure. We consider the problem of maximising expected utility in this setting. Besides establishing the existence of optimizers u…

2019-07-12abs ↗pdf ↗

Unified approach for clustering financial multiplex networks.

problem Lack of methods to capture interconnections between assets over time.
method Tensor-based unified local and global clustering coefficients for multiplex networks.
result Unified clustering coefficients effectively describe dependencies between assets over time.

Currently, the world is witnessing a mounting avalanche of data due to the increasing number of mobile network subscribers, Internet websites, and online services. This trend is continuing to develop in a quick and diverse manner in the form of big data. Big data analytics can process large amounts of raw data and extr…

2018-01-19abs ↗pdf ↗

Mobile big data contains vast statistical features in various dimensions, including spatial, temporal, and the underlying social domain. Understanding and exploiting the features of mobile data from a social network perspective will be extremely beneficial to wireless networks, from planning, operation, and maintenance…

2016-09-30abs ↗pdf ↗

Machine learning and blockchain are two of the most noticeable technologies in recent years. The first one is the foundation of artificial intelligence and big data, and the second one has significantly disrupted the financial industry. Both technologies are data-driven, and thus there are rapidly growing interests in …

2019-09-12abs ↗pdf ↗

The study examines order flow in financial markets using fractional Lévy stable motion.

problem Challenges in selecting the best models for financial time series data.
method Investigates order disbalance time series from the perspective of fractional Lévy stable motion.
result Orders exhibit stable anti-correlation for 18 randomly selected stocks.

Kelly criterion, that maximizes the expectation value of the logarithm of wealth for bookmaker bets, gives an advantage over different class of strategies. We use projective symmetries for a explanation of this fact. Kelly's approach allows for an interesting financial interpretation of the Boltzmann/Shannon entropy. A…

2006-07-18abs ↗pdf ↗

In this short note, we formulate three problems relating to nonnegative scalar curvature (NNSC) fill-ins. Loosely speaking, the first two problems focus on: When are (n1)(n-1)-dimensional Bartnik data (Σin1,γi,Hi)\big(Σ_i ^{n-1}, γ_i, H_i\big), i=1,2i=1,2, NNSC-cobordant? (i.e., there is an nn-dimensional compact Riemannian manifold…

2020-01-16abs ↗pdf ↗

Data preprocessing techniques are devoted to correct or alleviate errors in data. Discretization and feature selection are two of the most extended data preprocessing techniques. Although we can find many proposals for static Big Data preprocessing, there is little research devoted to the continuous Big Data problem. A…

2018-10-14abs ↗pdf ↗

Big Data bring new opportunities to modern society and challenges to data scientists. On one hand, Big Data hold great promises for discovering subtle population patterns and heterogeneities that are not possible with small-scale data. On the other hand, the massive sample size and high dimensionality of Big Data intro…

2013-08-07abs ↗pdf ↗

Predict stock trends using news sentiment and technical indicators in Spark.

problem Predicting the stock market trend is challenging due to multiple influencing factors.
method Created a machine learning classification problem with features from technical indicators and news sentiment scores.
result Random Forest model achieved 63.58% test accuracy in Spark.

It had been believed in the conventional practice that the risk of a bank going bankrupt is lessened in a straightforward manner by transferring the risk of loan defaults. But the failure of American International Group in 2008 posed a more complex aspect of financial contagion. This study presents an extension of the …

2014-09-25abs ↗pdf ↗