MOPO-LSI offers a user guide for sustainable investments.
problem Sustainable investment optimization challenges.
method Open-source library for multi-objective portfolio optimization.
result User-friendly guide for MOPO-LSI version 1.0.
OMBA learns product and user representations for better online market basket analysis.
problem Limited ability to uncover rarely occurring and temporal associations in MBA.
method Jointly learns product and user representations, captures temporal dynamics, scalable online method.
result OMBA outperforms state-of-the-art methods by 21% on real-world datasets.
USTAR combines multiple social media modalities to model user-guided activities.
problem Lack of comprehensive spatiotemporal activity models using all social media modalities.
method Online learning method embedding locations, time, text, and users into a single space, incorporating NGTSM and GTSM records, and using collaborative filtering.
result USTAR significantly improves region and keyword retrieval compared to state-of-the-art methods.
A guide to using low-pass graph filters for network data.
problem Understanding and processing graph data with low-pass filters.
method Definition and application of low-pass graph filters to graph data.
result Low-pass filters effectively retain lower frequency graph data contents.
GUIDE-VAE generates user-guided data with improved realism and performance.
problem Generating data points for multi-user datasets while considering user information.
method Conditional generative model that integrates user embeddings and a pattern dictionary-based covariance composition.
result GUIDE-VAE outperforms conventional VAEs in multi-user settings, especially under data imbalance.
RuleKit aids in creating interpretable models for various data types.
problem Creating interpretable models for different data types.
method Sequential covering induction algorithm for classification, regression, and survival problems.
result Facilitates verification of hypotheses about data dependencies.
GuideR learns rules guided by user preferences for classification, regression, and survival analysis.
problem Lack of user preferences in rule learning algorithms.
method Guided sequential covering approach.
result User preferences improve rule quality in classification, regression, and survival analysis.
ColaBO accelerates optimization with user beliefs.
problem Optimizing expensive functions with limited prior knowledge.
method General Bayesian framework for incorporating user beliefs.
result Significant optimization acceleration with accurate prior information.
A new decision tree algorithm handles missing values efficiently.
problem Data sets with missing values.
method Proposes a decision tree algorithm that allows user-guided data partitioning and handles missing values.
result The algorithm produces more accurate and interpretable results than common procedures without imputations.
This guide clarifies techniques for assessing and comparing model calibration and performance.
problem Assessing and comparing the calibration and performance of predictive models in insurance and actuarial practice.
method Clarifies statistical techniques for assessing model calibration and comparing models, emphasizing the importance of specifying the prediction target functional and choosing the appropriate scoring function.
result Provides guidance for the practical choice of scoring functions and illustrates results with real data case studies.
Incremental method for graph Laplacian eigenpairs improves clustering efficiency.
problem Determining the number of clusters in spectral clustering.
method Incremental computation of graph Laplacian eigenpairs.
result Efficiently computes the K-th smallest eigenpair. Seglearn segments time series data for machine learning tasks.
problem Handling multivariate sequence and contextual data for classification, regression, and forecasting.
method Sliding window segmentation approach within a scikit-learn compatible pipeline.
result Efficient learning of time series data for various machine learning tasks.
A system uses randomisation to explore data guided by user knowledge and interests.
problem Creating an efficient exploratory data analysis system aware of user background knowledge and interests.
method Model user background knowledge with tiles, use constrained randomisation for efficient implementation, and apply linear projection pursuit to find informative views.
result The method is robust under noise, fast for interactive use, and gives understandable results for real-world data.
DoubleML is a Python library for causal inference using machine learning.
problem Estimating causal parameters in complex models with machine learning.
method Double machine learning framework for valid statistical inference.
result High flexibility and easy extension for various model specifications.
ODTLearn learns optimal decision trees for predictive and prescriptive tasks.
problem Learning optimal decision trees for high-stakes predictive and prescriptive tasks.
method Mixed-integer optimization framework and object-oriented design.
result Implementation of optimal decision trees for various tasks.
The smallest eigenvalues and the associated eigenvectors (i.e., eigenpairs) of a graph Laplacian matrix have been widely used for spectral clustering and community detection. However, in real-life applications the number of clusters or communities (say, K) is generally unknown a-priori. Consequently, the majority of …
Survey on random features for kernel approximation, focusing on algorithms, theory, and practical applications.
problem Efficiently approximating kernel methods for large-scale problems.
method Random features techniques to speed up kernel methods.
result Need for a high number of random features for good approximation quality.
Efficient private matrix analysis algorithms for recent variants.
problem Private analysis of recent matrix updates.
method Identifying sufficient conditions on positive semidefinite matrices.
result First efficient differentially private algorithms for various matrix analysis tasks.
Lecture notes on analysis tools for X-ray tomography.
problem Understanding X-ray tomography using mathematical analysis.
method Overview of analysis tools and ideas, minimal assumptions.
result Broad overview of analysis tools for X-ray tomography.
The paper explores machine learning in mobile big data analysis.
problem Challenges in mobile big data analysis.
method Discussion and review of existing methods.
result Identification of main challenges and future directions.
This paper introduces compositional data analysis for financial ratios, improving industry-level analysis.
problem Statistical issues with standard financial ratios at industry level.
method Compositional data analysis techniques for financial ratios.
result Improved analysis of financial ratios using compositional data methods.
In this dissertation, the main goal is visualisation of financial time series. We expect that visualisation of financial time series will be a useful auxiliary for technical analysis. Firstly, we review the technical analysis methods and test our trading rules, which are built by the essential concepts of technical ana…
Paper combines geometry and time-series analysis for spatiotemporal data.
problem Multivariate time-series data from multiple sensors.
method Combines manifold learning, Riemannian geometry, and spectral analysis.
result Proposes Riemannian multi-resolution analysis (RMRA) for dynamic mode extraction.
In this paper the exact linear relation between the leading eigenvectors of the modularity matrix and the singular vectors of an uncentered data matrix is developed. Based on this analysis the concept of a modularity component is defined, and its properties are developed. It is shown that modularity component analysis …
Paper uses dynamic analysis to detect malware with PHMMs.
problem Malware detection using static and dynamic analysis techniques.
method Hidden Markov Models (HMMs) and Profile Hidden Markov Models (PHMMs) trained on API call sequences.
result PHMMs outperform HMMs in malware detection.
This paper simplifies complex geometry for shape analysis.
problem Understanding interactions between differential geometry and functional analysis.
method Provides an overview of infinite-dimensional Riemannian manifolds and metrics.
result Roadmap for beginners in computational anatomy and shape analysis.
Interactive DR framework for comparing datasets.
problem Limited flexibility in existing DR methods for comparative analysis.
method Unified linear comparative analysis (ULCA) with interactive optimization and visualization.
result ULCA and optimization algorithm improve comparative analysis efficiency and flexibility.
This paper reviews R packages for automating data analysis tasks.
problem Time-consuming Exploratory Data Analysis in large, noisy data sets.
method Systematic review of 12 R packages for autoEDA.
result Identifies automated tasks and areas for future development.
Improves transparency of deep neural networks through feature and consistency analysis.
problem Black-box nature of deep learning inference limits transparency for safety-critical systems.
method Structural and linguistic feature analysis, consistency analysis.
result 75% of human workers found input data and results consistent, 70% found inference and results consistent.
Proposes a method to optimize class mean preservation in kernel-based feature spaces.
problem Optimizing the selection of kernel subspace for better performance.
method Component analysis method for kernel-based dimensionality reduction that optimally preserves class mean distances.
result Discriminant analysis version of the proposed method provides insights into feature space properties.
Proposes a multivariate regression model for better analysis of multiple datasets.
problem Insufficient performance of single-dataset analysis in integrative studies.
method Sparse estimation for variable and group selection, alternating direction method of multipliers algorithm.
result Demonstrated improved performance through simulations and real data analysis.
Combines topological and geometric approaches to data analysis.
problem Understanding when and how geometric objects intersect.
method Connects topological and geometric concepts of curvature.
result Reconceptualizes curvature and links it to hyperconvexity.
New method uses topological data analysis to study stock market crashes.
problem Characterizing and predicting stock market crashes.
method Topological data analysis, persistence landscape, dynamic time series analysis.
result Demonstrates effectiveness of new method for Flash Crash characterization and prediction.
PARSEC compresses text for sentiment analysis with minimal loss in accuracy.
problem Compressing text data for sentiment analysis without losing accuracy.
method Uses Parts-of-Speech tags to compress text intelligently.
result Accurate compression is possible with minimal loss in sentiment classification accuracy.
A primer on navigating causal analysis problems.
problem Estimating causal effects from observational data.
method Four schools of thought for causal analysis.
result Conceptual map for causal analysis.
Novel sensitivity analysis for neural networks improves model interpretability.
problem Improving interpretability of complex probabilistic models.
method Bayesian neural networks with latent variables and sensitivity analysis.
result Increases interpretability of black-box probabilistic models.
Analyzes stock trends and e-commerce user behavior using Twitter data.
problem Understanding the relationship between stock prices, stock news, and e-commerce user behavior.
method Cross-domain analysis using Hadoop, Hive, and Tableau on three datasets.
result Identified correlations between stock sentiment, stock trends, and e-commerce user behavior.
Genetic programming optimizes Gaussian kernels for better sentiment analysis.
problem Improving accuracy of sentiment analysis in text.
method Genetic Programming applied to evolve more effective Gaussian kernels.
result The evolved kernels outperform traditional Gaussian Processes in sentiment analysis.
This work improves group data analysis using modified tensor decompositions.
problem Improving group data analysis models for better signal modeling.
method Introduces a new generalization of block tensor decomposition for group data analysis.
result Demonstrates improved performance in multilabel classification and clustering tasks.
Study analyzes Disney stock market performance using machine learning.
problem Forecasting stock market performance of Disney.
method Exploratory data analysis, feature engineering, model selection (linear regression).
result Linear regression model performed best.
This study analyzes data science vocabulary changes over 13 years.
problem Understanding evolution of data science terms over time.
method Exploratory Data Analysis, Latent Semantic Analysis, Latent Dirichlet Analysis, N-grams Analysis.
result Identified new vocabulary and its incorporation into scientific literature.
Analyzes changes in cryptocurrency market structure.
problem Understanding shifts in cryptocurrency market dynamics.
method Structural change analysis techniques.
result Identifies key structural changes in the market.
FinSphere improves stock analysis quality with AI and expert-curated data.
problem Lack of objective evaluation metrics and depth in stock analysis by FinLLMs.
method Developed AnalyScore, curated Stocksis dataset, and FinSphere AI agent.
result FinSphere outperforms general and domain-specific LLMs in generating high-quality stock analysis reports.
Random forest identifies key features for diagnosing machine faults.
problem Identifying specific factors causing machine faults.
method Gaussian mixture model clustering, spectrum analysis, random forest classification.
result Identified significant features for different machine states.
This study examines whether PCA can effectively identify nitrogen pollution sources in rivers.
problem Identifying pollution sources in rivers for effective environmental management.
method Principal Component Analysis and its modifications, along with Independent Component Analysis and Factor Analysis, are applied to nitrogen pollution source identification.
result PCA and related techniques can be powerful tools for uncovering nitrogen pollution sources in rivers.
Solves optimal stopping problem for financial technical analysis.
problem Optimal stopping problem for technical analysis models.
method Wide-class dynamics modeling support/resistance lines.
result Solution to optimal stopping problem for technical analysis.
We present a unifying framework which reduces the construction of probabilistic component analysis techniques to a mere selection of the latent neighbourhood, thus providing an elegant and principled framework for creating novel component analysis models as well as constructing probabilistic equivalents of deterministi…
Paper introduces probabilistic methods to approximate archetypal analysis, reducing complexity.
problem Inherent computational complexity of archetypal analysis limits its practical applicability.
method Two preprocessing techniques: dimensionality reduction and representation cardinality reduction, using probabilistic geometry.
result The method effectively reduces scaling and provides near-optimal solutions for prediction errors.