Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

2845688521,136 · Jun 202019922001200920182026
48 results for industrial big data

Proposes a methodology to improve data science ROI by addressing key business questions.

problem Companies often fail to maximize data science value, focusing on basic analysis.
method Categorizes and answers 'The Big Three' questions using data science methods.
result Shows how to apply the methodology to real business use cases.

Robinhood users react strongly to overnight price changes and big losers, trading quickly after extreme losses.

problem Understanding trading behavior of Robinhood users, especially in high-frequency trading scenarios.
method Analyzed intraday and overnight price changes, focusing on big losers and gainers.
result Robinhood users react more to overnight price changes and big losers, trading quickly after extreme losses.

In real world industrial applications of topic modeling, the ability to capture gigantic conceptual space by learning an ultra-high dimensional topical representation, i.e., the so-called "big model", is becoming the next desideratum after enthusiasms on "big data", especially for fine-grained downstream tasks such as …

2014-11-10abs ↗pdf ↗

This research tackles unsupervised topic extraction in noisy social media data.

problem Capturing customer insights from social media data is challenging due to noise and heterogeneity.
method The research presents three nonparametric approaches based on the Variational Autoencoder framework: Embedded Dirichlet Process, Embedded Hierarchical Dirichlet Process, and time-aware Dynamic Embedded Dirichlet Process.
result The models achieve equal to better performance than state-of-the-art methods in topic extraction from noisy social media data.

What is a systematic way to efficiently apply a wide spectrum of advanced ML programs to industrial scale problems, using Big Models (up to 100s of billions of parameters) on Big Data (up to terabytes or petabytes)? Modern parallelization strategies employ fine-grained operations and scheduling beyond the classic bulk-…

2013-12-30abs ↗pdf ↗

Study tackles imbalanced data in car insurance claims prediction.

problem Predicting rare events (claims) in car insurance with imbalanced data.
method Various machine learning techniques (logistic-regression, decision tree, random forest, xgBoost, feed-forward network) applied to imbalanced dataset.
result Comparison of machine learning algorithms' performance in claim occurrence prediction.

Paper proposes a method to evaluate SME credit risk using meta paths.

problem Evaluate credit risk of small and medium-sized enterprises with limited data.
method Exploits the representative power of information networks and meta paths to infer SME financial status.
result Meta path feature effectively identifies SMEs with credit risks.

The paper proposes an AI and IIoT framework for improved maintenance.

problem Current maintenance practices need improvement with AI and IIoT.
method Review of reliability modeling, introduction of Intelligent Maintenance framework, and novel probabilistic deep learning approach.
result Demonstrated novel probabilistic deep learning reliability modelling in Turbofan Engine Degradation Dataset.

DSCOVR improves distributed optimization for big data with less communication and synchronization.

problem Efficiently optimizing large linear models with convex loss functions over distributed systems.
method Randomized primal-dual block coordinate algorithms with doubly stochastic coordinate optimization and variance reduction.
result DSCOVR algorithms require less overall computation and communication compared to other first-order distributed algorithms.

Method estimates causal effects from incremental data, overcoming missing data challenges.

problem Estimating causal effects from non-stationary, incrementally available observational data.
method Continual Causal Effect Representation Learning
result Method achieves continual causal effect estimation without compromising original data.

NetDP predicts loan defaults using network data, addressing cold-start issues.

problem Cold-start problem in default prediction for new users.
method Combines unsupervised and supervised network representations, using parameter-server for scalability.
result Effectiveness in cold-start problem, especially for new users.

The rise of Big Data has led to new demands for Machine Learning (ML) systems to learn complex models with millions to billions of parameters, that promise adequate capacity to digest massive datasets and offer powerful predictive analytics thereupon. In order to run ML algorithms at such scales, on a distributed clust…

2015-12-31abs ↗pdf ↗

AI enhances bank credit risk management through deep learning and data analysis.

problem Inaccurate credit decisions and potential risks in bank credit risk management.
method Innovative application of AI technology, including deep learning and big data analysis.
result AI provides more accurate and comprehensive credit decision support, reducing risks and losses.

LightGCNet simplifies AI for soft sensors, reducing complexity and training time.

problem Complex and resource-intensive deep learning models for soft sensors.
method LightGCNet uses compact angle constraints and node pool strategy for efficient learning.
result LightGCNet achieves small network size, fast learning, and good generalization.

This research improves debt collection strategies using advanced machine learning.

problem Accurate estimation of propensity to pay and cashflow for optimal debt collection.
method Developed a machine learning framework with pre-processing and model selection.
result The proposed model outperforms current industry strategies.

Paper analyzes electricity price and demand data to detect cyber-attacks using time series methods.

problem Detecting cyber-attacks in electricity price and demand data.
method Time series analysis, including moving average, moving standard deviation, and augmented Dickey-Fuller test.
result Identified anomalies in the data using time-series stationary criteria.

InfDetect detects e-commerce insurance fraud using graph analysis.

problem Detecting fraudulent claims in e-commerce insurance with multiple parties involved.
method Developed a large-scale fraud detection system InfDetect using graph-based approaches.
result InfDetect successfully detected thousands of fraudulent claims and saved money daily.

Sabrina integrates financial data and domain knowledge for better visualization.

problem Scattered financial data across various sources makes it hard for analysts to understand the economy.
method Sabrina uses a pipeline to fuse firm-specific and macroeconomic data, visualizing it in a unified interface.
result Sabrina aids financial analysts in their analysis process, as shown in a user study.

Big Data classifiers perform similarly to Small Data classifiers, suggesting scalability tradeoffs.

problem Comparing Big Data classifiers to Small Data classifiers for performance and scalability.
method Empirical study comparing Big Data classifiers to Small Data classifiers.
result Big Data classifiers are slightly inferior but catching up with Small Data classifiers.

Big Data bring new opportunities to modern society and challenges to data scientists. On one hand, Big Data hold great promises for discovering subtle population patterns and heterogeneities that are not possible with small-scale data. On the other hand, the massive sample size and high dimensionality of Big Data intro…

2013-08-07abs ↗pdf ↗