Analyzes 6M Python notebooks and 2M enterprise DS pipelines to guide investments in data science.
problem Challenges in following the rapidly evolving landscape of data science technologies and applications.
method Downloaded and analyzed over 6M Python notebooks and 2M enterprise DS pipelines, performing statistical and comparative analyses.
result Identifies actionable conclusions for system builders and technology bets for practitioners based on current trends.
Study reveals similarities in knowledge flows between pharmaceutical and AI industries.
problem Understanding the dynamics of drug pipelines in global pharmaceutical industry.
method Multilayer network analysis of drug pipeline, global supply chain, and ownership data.
result Proven similarities in knowledge flows between pharmaceutical and AI industries.
In recent years, the interest in Big Data sources has been steadily growing within the Official Statistic community. The Italian National Institute of Statistics (Istat) is currently carrying out several Big Data pilot studies. One of these studies, the ICT Big Data pilot, aims at exploiting massive amounts of textual …
This paper surveys enterprise financial risk analysis from Big Data and LLMs perspectives.
problem Predicting future financial risk of enterprises.
method Systematic literature review of enterprise financial risk analysis approaches from Big Data and LLMs perspectives.
result Offers a holistic synthesis of research methods and key insights.
This study identifies financial risk paths in digital-transformed enterprises.
problem Identifying financial risks in digital-transformed enterprises.
method DEMATEL-ISM-MICMAC method.
result Political and economic environment affects enterprise's financial structure.
Paper proposes a method to evaluate SME credit risk using meta paths.
problem Evaluate credit risk of small and medium-sized enterprises with limited data.
method Exploits the representative power of information networks and meta paths to infer SME financial status.
result Meta path feature effectively identifies SMEs with credit risks.
Paper assesses financial potential for enterprise development.
problem Determining financial potential for enterprise development.
method Stages of financial potential assessment based on literature analysis.
result Proposes a mechanism for managing enterprise financial potential.
Next generation of embedded Information and Communication Technology (ICT) systems are collaborative systems able to perform autonomous tasks. The remarkable expansion of the embedded ICT market, together with the rise and breakthroughs of Artificial Intelligence (AI), have put the focus on the Edge as it stands as one…
Study combines intra-risk and contagion risk for SME bankruptcy prediction.
problem Predicting bankruptcy risk of SMEs considering both intra-risk and contagion risk.
method Proposes a novel model using Graph Neural Networks to combine intra-risk and contagion risk.
result Model outperforms state-of-the-art methods in bankruptcy prediction.
Graph neural networks improve SME credit risk assessment.
problem Improving credit risk assessment for small and medium enterprises (SMEs).
method Graph neural networks were used to model the relationships between financial indicators of enterprises, creating a graph structure and embedding representations for credit risk prediction.
result The proposed model accurately predicts enterprise credit levels, demonstrating robustness and effectiveness.
We present a macroeconomic agent-based model that combines several mechanisms operating at the same timescale, while remaining mathematically tractable. It comprises enterprises and workers who compete in a job market and a commodity goods market. The model is stock-flow consistent; a bank lends money charging interest…
For researching the association between coal enterprise management and return in financial market, this paper applies the method of time difference relevance and PageRank method to seek the leader-index of a stock set containing 21 coal enterprises in A-share market and score those stocks. Based on the return in 2011, …
Margin trading and short selling boost green tech innovation in China.
problem Encouraging green technology innovation in Chinese companies.
method Quasi-experimental research using panel data of Chinese listed companies, double difference model.
result Margin trading and short selling increase green tech innovation significantly.
Study reveals holiday effect on China's time-honored brands, especially alcoholic beverages.
problem Understanding holiday impact on China's time-honored brands.
method Event study using listed companies of China's time-honored brands from 2012-2021.
result Time-honored brand stocks show significant post-holiday effect during Chinese New Year, alcoholic beverages more sensitive.
SHAP Distance assesses semantic fidelity of synthetic tabular data.
problem Semantic fidelity of synthetic tabular data is not well evaluated.
method SHAP Distance, defined as cosine distance between global SHAP attribution vectors.
result SHAP Distance detects semantic discrepancies overlooked by standard measures.
A new method to protect enterprise data privacy in AI models.
problem Enterprise data leakage risks in AI models.
method ABack, a training-free mechanism using Hidden State Model.
result Improves privacy utility by up to 15% over strong baselines.
The paper shows supply chain features improve cyber risk prediction.
problem Predicting cyber risk from supply chain attributes.
method Machine learning, external supply chain features, AUC improvement.
result Supply chain network features improve AUC by 2.3%.
This research develops a dynamic risk management system for industrial companies.
problem Risk assessment and management in industrial enterprises.
method Qualitative and quantitative analysis, systematic risk classification, dynamic system development.
result Effective risk management strategies formed through dynamic risk management system and risk assessment methods.
Chebyshev polynomials analyze Czech enterprises' stock dynamics.
problem Analyzing stock dynamics of enterprises not following normal distribution.
method Chebyshev polynomial decomposition of stock time series.
result Allows effective analysis of stock dynamics without variance and correlation.
AI-driven framework improves enterprise financial audits and risk identification.
problem Manual auditing is inefficient and limited by data complexity and evolving fraud tactics.
method Machine learning algorithms (SVM, RF, KNN) applied to a dataset of audit project counts, violations, and fraud instances.
result Random Forest achieves best performance with F1-score of 0.9012, identifying fraud and compliance anomalies.
To assure cyber security of an enterprise, typically SIEM (Security Information and Event Management) system is in place to normalize security event from different preventive technologies and flag alerts. Analysts in the security operation center (SOC) investigate the alerts to decide if it is truly malicious or not. H…
Being one of the most important factors of economic growth of the country, innovations became one of the key vectors in Russian economic policy. In this field technology parks are one of the most effective instruments which can provide growth of innovative activity in sectors, regions and economies. In this paper, we m…
Generative AI agents improve ERP systems by automating complex financial tasks.
problem Static, rule-based workflows limit adaptability and intelligence in ERP systems.
method Introducing Generative Business Process AI Agents (GBPAs) that integrate generative AI with business process modeling and multi-agent orchestration.
result GBPAs achieve up to 40% reduction in processing time and 94% drop in error rate.
Networked-guarantee loans may cause the systemic risk related concern of the government and banks in China. The prediction of default of enterprise loans is a typical extremely imbalanced prediction problem, and the networked-guarantee make this problem more difficult to solve. Since the guaranteed loan is a debt oblig…
AI improves MSME credit scoring using bank statement data.
problem Lack of access to financing for MSMEs due to traditional credit scoring methods.
method Developed a cash flow-based pipeline using bank statement data for machine learning credit scoring.
result Bank statement features significantly improve credit scoring models, achieving AUROC of 0.806.
The basic financial purpose of an enterprise is maximization of its value. Trade credit management should also contribute to realization of this fundamental aim. Many of the current asset management models that are found in financial management literature assume book profit maximization as the basic financial purpose. …
FinanceBench benchmarks LLMs on financial QA, revealing limitations.
problem Evaluating LLMs' performance on financial question answering.
method Developed a comprehensive test suite (FinanceBench) with 10,231 questions, tested 16 models, and manually reviewed answers.
result Existing LLMs have significant limitations for financial QA, especially GPT-4-Turbo.
Financial institutions face new model risks with AI, requiring enhanced model risk management.
problem New model risks from Generative AI applications in financial institutions.
method Enhanced model risk framework with additional testing and controls.
result Financial institutions need to enhance their model risk management for Generative AI applications.
It was not until the beginning of the 1990s that the effects of information and communication technology on economic growth as well as on the profitability of enterprises raised the interest of researchers. After giving a general description on the relationship between a more intense use of ICT devices and dynamic econ…
AVATAR uses a surrogate model to quickly evaluate ML pipelines, saving time and resources.
problem Time-consuming evaluation of ML pipelines limits exploration of complex models.
method AVATAR employs a surrogate model to assess pipeline validity without execution.
result AVATAR accelerates ML pipeline evaluation, improving efficiency in complex scenarios.
This study proposes a new model for predicting financial distress in SMEs using machine learning.
problem Challenges in predicting financial distress for SMEs due to ambiguity and limited data.
method Feature selection algorithm based on element credits and data source collection. Incorporates financial statements, governance qualities, and market data with a Relevant Vector Machine.
result The proposed model improves financial distress prediction efficiency with fewer characteristic factors.
Digital transformation boosts corporate financial asset allocation, especially short-term.
problem Understanding how digital transformation affects corporate financial decisions.
method Fixed-effects models and staggered DID design using A-share listed companies data.
result Digital transformation significantly promotes corporate financial asset allocation, more pronounced in short-term.
Paper presents a new method for better financial market forecasting.
problem Traditional investment strategies fail to capture market nuances and risks.
method Combines deep learning, factor integration, and correlated stock analysis.
result Enhanced diversification and performance capture in financial markets.
SmallML predicts customer churn for SMEs with small data, improving accuracy by 24.2 points.
problem AI exclusion of SMEs due to data scale mismatch.
method Bayesian transfer learning with hierarchical pooling and conformal prediction.
result 96.7% AUC on 100 obs SMEs, 24.2 point improvement over logistic regression.
Hyperparameter tuning of multi-stage pipelines introduces a significant computational burden. Motivated by the observation that work can be reused across pipelines if the intermediate computations are the same, we propose a pipeline-aware approach to hyperparameter tuning. Our approach optimizes both the design and exe…
Classical Machine Learning (ML) pipelines often comprise of multiple ML models where models, within a pipeline, are trained in isolation. Conversely, when training neural network models, layers composing the neural models are simultaneously trained using backpropagation. We argue that the isolated training scheme of ML…
Study proposes a statistical testing framework for evaluating clustering pipelines.
problem Quantifying the statistical reliability of clustering results from data analysis pipelines.
method Selective inference-based statistical testing framework for clustering pipelines.
result The proposed test controls the type I error rate and is effective in validating clustering results.
Investigates fairness in pipeline models where individuals may drop out.
problem Fairness in pipeline models where individuals may drop out and subsequent stages depend on remaining individuals.
method Rigorous framework for evaluating fairness guarantees, showing that naïve auditing is insufficient and dependence must exist between stages.
result Fairness in pipelines can be arbitrary, even with just two stages, and requires dependence between stages.
Most of the existing solutions to enterprise threat management are preventive approaches prescribing means to prevent policy violations with varying degrees of success. In this paper we consider the complementary scenario where a number of security violations have already occurred, or security threats, or vulnerabiliti…
Data science relies on pipelines that are organized in the form of interdependent computational steps. Each step consists of various candidate algorithms that maybe used for performing a particular function. Each algorithm consists of several hyperparameters. Algorithms and hyperparameters must be optimized as a whole …
Automates supervised learning pipeline design with matrix and tensor factorization.
problem Designing effective supervised learning pipelines with many choices.
method Uses matrix and tensor factorization to model pipeline search space and develops greedy experiment design protocols.
result Demonstrates the effectiveness of the approach on real-world classification problems.
New framework uses MABs for better WLAN performance.
problem Maximizing enterprise WLAN performance with complex AP and station interactions.
method Empowered APs and stations with agents using Thompson sampling to learn optimal channel and AP selection.
result Adaptive framework using MABs outperforms static configurations in various scenarios.
Data science relies on pipelines that are organized in the form of interdependent computational steps. Each step consists of various candidate algorithms that maybe used for performing a particular function. Each algorithm consists of several hyperparameters. Algorithms and hyperparameters must be optimized as a whole …
New model optimizes oil product distribution via pipelines.
problem Optimizing oil product distribution via pipelines.
method Discrete-time mixed integer linear programming model.
result Significant reductions in pipeline operational cost.
Pipeline decomposes portfolio optimization problems into smaller, solvable subproblems.
problem Large-scale portfolio optimization with constraints.
method Decomposition pipeline with preprocessing, clustering, and risk rebalancing.
result Pipeline reduces problem size by 80% and computation time.
Paper proposes a statistical test for feature selection pipelines using selective inference.
problem Assessing the significance of feature selection pipelines in data analysis.
method Selective inference technique applied to feature selection pipelines composed of various algorithms.
result The proposed statistical test controls false positive feature selection probabilities.
New approach identifies and explains errors in machine learning pipelines.
problem Challenges in identifying and explaining errors in complex machine learning pipelines.
method Uses iteration and provenance to automatically infer root causes of failures.
result Significantly improves precision and recall compared to state-of-the-art methods.
Pipelined Backpropagation trains large models without batches efficiently.
problem Training large models efficiently on hardware with limited batch sizes.
method Fine-grained Pipelined Backpropagation with Spike Compensation and Linear Weight Prediction.
result Fine-grained Pipelined Backpropagation with a batch size of one matches the accuracy of SGD for multiple networks.