We offer open-source code for a fundamental industry classification.
problem Creating accurate industry classifications from public data.
method Developed open-source code with modules for data handling.
result Improved trading signals through various industry classifications.
We give complete algorithms and source code for constructing (multilevel) statistical industry classifications, including methods for fixing the number of clusters at each level (and the number of levels). Under the hood there are clustering algorithms (e.g., k-means). However, what should we cluster? Correlations? Ret…
Develops MIS, a probabilistic model for multi-industry classification.
problem GICS's limitation of assigning each firm to exactly one industry, especially for diversified firms.
method Topic modeling to probabilistically assign firms to multiple industries based on business descriptions.
result Demonstrates MIS's ability to flexibly assign firms to multiple industries with relevance probabilities.
Machine learning creates non-factor covariance matrices for risk models.
problem Creating robust risk models for financial portfolios.
method Developed an explicit algorithm and source code for machine learning risk models.
result Machine learning models outperform traditional risk models in empirical backtests.
Analyzes Indian chemical industry post-Covid.
problem Global uncertainty impacts chemical industry performance.
method Fundamental analysis of key players and trends.
result Various geopolitical and macroeconomic trends shape industry performance.
Improves industry classification for diversified companies.
problem Traditional industry classification struggles with multi-sector conglomerates.
method Bayesian Non-Parametrics, Markov Updating, and hierarchical modeling.
result MIS-2 provides a measurable improvement over GICS in predicting future correlations.
Novel financial time-series data representation improves industry sector classification.
problem Classifying industries using historical stock returns time-series data.
method Proposed a novel representation based on stock returns embeddings for time-series data, overcoming representational challenges of conventional approaches.
result Substantial performance improvements over baselines using conventional representations.
Neural model learns company embeddings from data and news.
problem Subjective industry classification schemes in finance.
method Multimodal neural model training company embeddings.
result Objective company representations capture nuanced relationships.
AI agent predicts industry and product/service codes for companies.
problem Manual curation of company data is expensive and prone to errors.
method Hierarchical multi-class industry code classifier with multi-label product/service code classifier.
result High accuracy (92-96%) achieved with limited labeled data.
Dataset for rare event prediction in industrial multivariate time series.
problem Building a model to predict rare events in a multivariate time series dataset.
method The dataset is used to build a classification model for early prediction of rare events.
result The dataset can be used for various types of models including classification and exploration.
Enhances insurance loss models using InsurTech data and machine learning.
problem Traditional insurance loss models lack predictive accuracy due to limited data sources.
method Combining proprietary claims data with InsurTech data and applying machine learning techniques.
result Improved predictive accuracy of the loss model through machine learning.
Deep learning classifies over 94% of crystallization images accurately.
problem Classifying macromolecular crystallization outcomes from various experiments.
method Deep convolutional neural networks trained on a large annotated dataset.
result More than 94% of test images correctly labeled, regardless of origin.
Study finds fundamental analysis useful for predicting stock prices in China's transitional economy.
problem Investment predictability in China's transitional economy.
method Examined 3 industries (media, power, steel) with 3 types of correlation on 25 financial determinants of 60 Chinese companies over 4 years.
result Fundamental analysis can predict stock prices in China's transitional economy, contradicting the Efficient Market Hypothesis.
Fault detection in industrial plants is a hot research area as more and more sensor data are being collected throughout the industrial process. Automatic data-driven approaches are widely needed and seen as a promising area of investment. This paper proposes an effective machine learning algorithm to predict industrial…
Study compares classification techniques to predict customer churn in banking.
problem Predicting customer churn in banking industry.
method Comparison of six supervised classification techniques (ANN and random forest) on 10000 European bank customers.
result ANN structure with five nodes in a single hidden layer is the best performing classifier.
This research develops a dynamic risk management system for industrial companies.
problem Risk assessment and management in industrial enterprises.
method Qualitative and quantitative analysis, systematic risk classification, dynamic system development.
result Effective risk management strategies formed through dynamic risk management system and risk assessment methods.
Deep learning outperforms classical methods in categorizing BIM images.
problem Classifying building designs from BIM models.
method Used classical machine learning (HOG + SVM) and deep learning models (pre-trained and custom-designed networks).
result Deep learning models achieve significantly higher accuracy (above 89%) compared to classical methods (57%).
Efficient neural network ensembles improve image classification reliability and uncertainty quantification.
problem Uncertainty in neural network predictions for industrial image classification.
method Investigated efficient neural network ensembles (snapshot, batch, multi-input multi-output) for image classification reliability and uncertainty quantification.
result Batch ensemble is a cost-effective and competitive alternative to deep ensembles, offering savings in training and test time.
New AI method improves anomaly detection across different IIoT sensors.
problem Poor performance of anomaly detection models when applied to different machines.
method Robust AI method using pre-processing and multiple models on different pumps.
result Models perform well across different environments and types of pumps.
Wind turbine status classification models are improved for cross-site applicability.
problem Standard models don't generalize across different wind turbines.
method Data normalization and convolutional neural networks with feature-space extension.
result Trained models successfully classified wind turbines from different sites.
Prototype learns automotive industry ontology from unstructured data.
problem Automatic learning of domain-specific ontologies from unstructured text data.
method Two-stage classification system: first classifier for concepts and irrelevant collocates, second classifier for concept types.
result Prototype validated with automotive industry complaint and repair data.
LLMs outperform strong tabular baselines on industrial car retrofit prediction.
problem Industrial retrofit planning on structured operational data
method Embedding features, direct prompted classification, and ML+LLM stacking
result LLMs outperform strong tabular baselines on industrial car retrofit prediction
Hierarchical AI multi-agent framework optimizes equity portfolios in China's A-share market.
problem Optimizing equity portfolios in China's A-share market using AI and multi-agent systems.
method A hierarchical multi-agent design integrating macro, firm-level, and reinforcement learning approaches.
result Consistently outperforms benchmarks and state-of-the-art systems on risk-adjusted returns and drawdown control.
New dataset for industrial machine sounds to aid maintenance.
problem Lack of public datasets for industrial machine sounds.
method Recorded normal and anomalous sounds of industrial machines.
result Assists in automated facility maintenance development.
Proposes a model to update industrial data predictions based on temporal changes.
problem Improving prediction accuracy in industrial data analytics by addressing changing conditions over time.
method Integrates similarity and loss functions to estimate and update prediction models adaptively.
result The data renewal model enhances prediction accuracy by identifying and updating model changes.
New risk theory for 'Pay-for-Performance' models.
problem How to price and hedge operational and financial risks in new business models.
method Developed a new risk theory and calculation method for 'Pay-for-Performance' models.
result Presented a model for determining risk premiums including both financial and operational risks.
Rocket algorithm classifies time-series data efficiently using random projections and natural sparsity.
problem Time-series classification challenges in diverse fields.
method Random convolutional kernels, non-linear transformation, compressed sensing framework.
result Rocket algorithm preserves discriminative patterns in time-series data and expresses inherent sparsity.
This paper explores how NLP enhances insurance data analysis.
problem Traditional insurance data limitations and need for alternative data.
method Application of NLP techniques to transform and analyze unstructured text data.
result NLP techniques improve insurance data analysis and risk assessment.
We give a simple explicit algorithm for building multi-factor risk models. It dramatically reduces the number of or altogether eliminates the risk factors for which the factor covariance matrix needs to be computed. This is achieved via a nested "Russian-doll" embedding: the factor covariance matrix itself is modeled v…
The study designs inherently interpretable machine learning models for high-risk sectors.
problem The need for transparent and explainable machine learning models in regulated industries.
method Qualitative template based on feature effects and model architecture constraints for assessing inherent interpretability.
result Demonstrates the design and evaluation of an interpretable ReLU DNN model for predicting credit default.
This paper reviews early time series classification methods.
problem Minimizing class prediction delay in time-sensitive applications.
method Divided into four categories: prefix based, shapelet based, model based, and miscellaneous approaches.
result Demonstrates reasonable performance in various applications.
Large language models learn company embeddings from SEC filings.
problem Lack of a rigorous definition of company similarity.
method Pre-trained and finetuned large language models (LLMs) to learn embeddings from SEC filings.
result LLMs can reproduce GICS classifications and indicate similar financial performance.
We give a complete algorithm and source code for constructing what we refer to as heterotic risk models (for equities), which combine: i) granularity of an industry classification; ii) diagonality of the principal component factor covariance matrix for any sub-cluster of stocks; and iii) dramatic reduction of the facto…
Hybrid model combines PCA and RNN for better aerospace stock price prediction.
problem Challenges in predicting stock prices of aerospace companies due to market uncertainty and complexity.
method Combination of Principal Component Analysis (PCA) and Recurrent Neural Networks (RNN).
result PCA improves both accuracy and efficiency of stock price prediction.
Survey of deep causal models for industrial applications.
problem Estimating causal effects using deep learning.
method Deep causal models map covariates to a representation space and use objective functions for unbiased counterfactual data estimation.
result Comprehensive overview of deep causal models with industry applications.
Study shows how China's stock market reflects economic demand changes during COVID-19.
problem Understanding how stock market volatility is influenced by economic demand changes.
method Divided industries into demand-oriented groups and analyzed spillover networks.
result Spillover effects from demand-oriented sectors to consumption-oriented sectors increased during the outbreak.
Proposes a deep learning churn prediction system for telecom using TL and meta-classification.
problem Churn prediction challenges in telecom due to large data, high dimensions, and imbalanced data.
method Transfer Learning (TL) and Ensemble-based Meta-Classification. Two stages: TL on Deep CNNs, then GP-AdaBoost meta-classifier.
result TL-DeepE system achieved 75.4% and 68.2% prediction accuracy on Orange and Cell2cell datasets, respectively.
Improved logistic models for better interpretability in regulated industries.
problem Limited interpretability in machine learning models for regulated industries.
method Blend logistic regression with machine learning techniques for enhanced interpretability.
result Solid performance with minimal analyst effort and enhanced interpretability.
Hotel booking chatbot handles daily tens of thousands of searches.
problem Improving hotel booking through conversational AI.
method Frame-based dialogue management system with ML models for intent, entity recognition, and info retrieval.
result Chatbot deployed on a commercial scale, handling hotel searches daily.
We discuss a general dynamic replication approach to counterparty credit risk modeling. This leads to a fundamental jump-process backward stochastic differential equation (BSDE) for the credit risk adjusted portfolio value. We then reduce the fundamental BSDE to a continuous BSDE. Depending on the close out value conve…
UBMF tackles fault diagnosis in imbalanced industrial data with enhanced accuracy and adaptability.
problem Fault diagnosis challenges in imbalanced industrial data.
method Integrates four key modules: data perturbation, cross-task feature extraction, uncertainty-based filtering, and Bayesian meta-knowledge integration.
result Achieves an average improvement of 42.22% across ten diagnostic tasks.
Paper predicts bearing degradation stages for pharmaceutical industry maintenance.
problem Predicting when to maintain specific parts of production machines.
method AutoEncoder-based k-means segmentation of high-frequency vibration data.
result Framework generates reliable predictions for bearing degradation stages.
System automates discovery and classification of training videos for career progression.
problem Difficulties in planning and navigating career paths due to changing job requirements and emerging sectors.
method Extracted educational videos, built a machine learning classifier, and optimized probability thresholds.
result Significant improvements in model performance by incorporating video attributes.
New method classifies nonlinear time series using deep CNNs and bispectra.
problem Classifying nonlinear time series data effectively.
method Combines HOSA with deep CNNs.
result Effective classification of nonlinear time series data.
Deep learning improves defect classification in real-time surface inspection.
problem Real-time defect classification in manufacturing industry using limited datasets.
method Convolutional Neural Networks (CNNs) designed for speed and accuracy, neural data augmentation for class imbalance.
result 98.0% accuracy in binary defect classification with 22,000 labeled images.
Homotopy classification for certain 4-manifolds with dihedral fundamental groups.
problem Classifying the homotopy types of specific 4-manifolds with dihedral fundamental groups.
method Using quadratic 2-type and combining with results from Hambleton-Kreck and Bauer.
result Homotopy types of finite oriented Poincaré 4-complexes are determined by their quadratic 2-type when fundamental group is dihedral.
New classification for 4-manifolds with specific fundamental groups.
problem Classifying 4-manifolds with fundamental group Z/ptimesZ. method Using quadratic 2-type and intersection form, along with generalizations to infinite groups.
result Homotopy type determined by quadratic 2-type and Z/p-valued intersection form. This paper evaluates AutoML tools for ML tasks.
problem Efficiency in machine learning for ML engineers.
method Evaluation and comparison of AutoML tools on various datasets.
result Performance and advantages/disadvantages of AutoML tools.