Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

114228342456 · Jun 202019922001200920182026
48 results for user-guided analysis

OMBA learns product and user representations for better online market basket analysis.

problem Limited ability to uncover rarely occurring and temporal associations in MBA.
method Jointly learns product and user representations, captures temporal dynamics, scalable online method.
result OMBA outperforms state-of-the-art methods by 21% on real-world datasets.

USTAR combines multiple social media modalities to model user-guided activities.

problem Lack of comprehensive spatiotemporal activity models using all social media modalities.
method Online learning method embedding locations, time, text, and users into a single space, incorporating NGTSM and GTSM records, and using collaborative filtering.
result USTAR significantly improves region and keyword retrieval compared to state-of-the-art methods.

GUIDE-VAE generates user-guided data with improved realism and performance.

problem Generating data points for multi-user datasets while considering user information.
method Conditional generative model that integrates user embeddings and a pattern dictionary-based covariance composition.
result GUIDE-VAE outperforms conventional VAEs in multi-user settings, especially under data imbalance.

GuideR learns rules guided by user preferences for classification, regression, and survival analysis.

problem Lack of user preferences in rule learning algorithms.
method Guided sequential covering approach.
result User preferences improve rule quality in classification, regression, and survival analysis.

This guide clarifies techniques for assessing and comparing model calibration and performance.

problem Assessing and comparing the calibration and performance of predictive models in insurance and actuarial practice.
method Clarifies statistical techniques for assessing model calibration and comparing models, emphasizing the importance of specifying the prediction target functional and choosing the appropriate scoring function.
result Provides guidance for the practical choice of scoring functions and illustrates results with real data case studies.

Seglearn segments time series data for machine learning tasks.

problem Handling multivariate sequence and contextual data for classification, regression, and forecasting.
method Sliding window segmentation approach within a scikit-learn compatible pipeline.
result Efficient learning of time series data for various machine learning tasks.

A system uses randomisation to explore data guided by user knowledge and interests.

problem Creating an efficient exploratory data analysis system aware of user background knowledge and interests.
method Model user background knowledge with tiles, use constrained randomisation for efficient implementation, and apply linear projection pursuit to find informative views.
result The method is robust under noise, fast for interactive use, and gives understandable results for real-world data.

ODTLearn learns optimal decision trees for predictive and prescriptive tasks.

problem Learning optimal decision trees for high-stakes predictive and prescriptive tasks.
method Mixed-integer optimization framework and object-oriented design.
result Implementation of optimal decision trees for various tasks.

The smallest eigenvalues and the associated eigenvectors (i.e., eigenpairs) of a graph Laplacian matrix have been widely used for spectral clustering and community detection. However, in real-life applications the number of clusters or communities (say, KK) is generally unknown a-priori. Consequently, the majority of …

2015-12-23abs ↗pdf ↗

Survey on random features for kernel approximation, focusing on algorithms, theory, and practical applications.

problem Efficiently approximating kernel methods for large-scale problems.
method Random features techniques to speed up kernel methods.
result Need for a high number of random features for good approximation quality.

This paper introduces compositional data analysis for financial ratios, improving industry-level analysis.

problem Statistical issues with standard financial ratios at industry level.
method Compositional data analysis techniques for financial ratios.
result Improved analysis of financial ratios using compositional data methods.

Paper combines geometry and time-series analysis for spatiotemporal data.

problem Multivariate time-series data from multiple sensors.
method Combines manifold learning, Riemannian geometry, and spectral analysis.
result Proposes Riemannian multi-resolution analysis (RMRA) for dynamic mode extraction.

In this paper the exact linear relation between the leading eigenvectors of the modularity matrix and the singular vectors of an uncentered data matrix is developed. Based on this analysis the concept of a modularity component is defined, and its properties are developed. It is shown that modularity component analysis …

2015-10-19abs ↗pdf ↗

This paper simplifies complex geometry for shape analysis.

problem Understanding interactions between differential geometry and functional analysis.
method Provides an overview of infinite-dimensional Riemannian manifolds and metrics.
result Roadmap for beginners in computational anatomy and shape analysis.

Interactive DR framework for comparing datasets.

problem Limited flexibility in existing DR methods for comparative analysis.
method Unified linear comparative analysis (ULCA) with interactive optimization and visualization.
result ULCA and optimization algorithm improve comparative analysis efficiency and flexibility.

Improves transparency of deep neural networks through feature and consistency analysis.

problem Black-box nature of deep learning inference limits transparency for safety-critical systems.
method Structural and linguistic feature analysis, consistency analysis.
result 75% of human workers found input data and results consistent, 70% found inference and results consistent.

Proposes a method to optimize class mean preservation in kernel-based feature spaces.

problem Optimizing the selection of kernel subspace for better performance.
method Component analysis method for kernel-based dimensionality reduction that optimally preserves class mean distances.
result Discriminant analysis version of the proposed method provides insights into feature space properties.

Proposes a multivariate regression model for better analysis of multiple datasets.

problem Insufficient performance of single-dataset analysis in integrative studies.
method Sparse estimation for variable and group selection, alternating direction method of multipliers algorithm.
result Demonstrated improved performance through simulations and real data analysis.

This work improves group data analysis using modified tensor decompositions.

problem Improving group data analysis models for better signal modeling.
method Introduces a new generalization of block tensor decomposition for group data analysis.
result Demonstrates improved performance in multilabel classification and clustering tasks.

This study analyzes data science vocabulary changes over 13 years.

problem Understanding evolution of data science terms over time.
method Exploratory Data Analysis, Latent Semantic Analysis, Latent Dirichlet Analysis, N-grams Analysis.
result Identified new vocabulary and its incorporation into scientific literature.

FinSphere improves stock analysis quality with AI and expert-curated data.

problem Lack of objective evaluation metrics and depth in stock analysis by FinLLMs.
method Developed AnalyScore, curated Stocksis dataset, and FinSphere AI agent.
result FinSphere outperforms general and domain-specific LLMs in generating high-quality stock analysis reports.

This study examines whether PCA can effectively identify nitrogen pollution sources in rivers.

problem Identifying pollution sources in rivers for effective environmental management.
method Principal Component Analysis and its modifications, along with Independent Component Analysis and Factor Analysis, are applied to nitrogen pollution source identification.
result PCA and related techniques can be powerful tools for uncovering nitrogen pollution sources in rivers.

We present a unifying framework which reduces the construction of probabilistic component analysis techniques to a mere selection of the latent neighbourhood, thus providing an elegant and principled framework for creating novel component analysis models as well as constructing probabilistic equivalents of deterministi…

2013-03-13abs ↗pdf ↗

Paper introduces probabilistic methods to approximate archetypal analysis, reducing complexity.

problem Inherent computational complexity of archetypal analysis limits its practical applicability.
method Two preprocessing techniques: dimensionality reduction and representation cardinality reduction, using probabilistic geometry.
result The method effectively reduces scaling and provides near-optimal solutions for prediction errors.