Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

114228342456 · Jun 202019922001200920182026
48 results for large surveys

Survey of large language models in financial prediction and trading.

problem Improving predictability and robustness of financial predictions and trading decisions.
method Task-centered taxonomy, review of empirical evidence, design patterns, benchmarks, and challenges analysis.
result Improved predictability and robustness of financial predictions and trading decisions through large language models.

The rapid development of computing power and efficient Markov Chain Monte Carlo (MCMC) simulation algorithms have revolutionized Bayesian statistics, making it a highly practical inference method in applied work. However, MCMC algorithms tend to be computationally demanding, and are particularly slow for large datasets…

2018-07-23abs ↗pdf ↗

An efficient LDP protocol for QMLE with improved practicality and theoretical guarantees.

problem Difficult implementation of existing LDP QMLE for large-scale surveys.
method Developed an alternative LDP protocol without long waiting time, high communication cost, and derivative boundedness assumptions.
result Sufficient conditions for consistency and asymptotic normality of the protocol.

Survey on automating geometry problem solving with large models.

problem Automating geometric problem solving with spatial understanding and logical reasoning.
method Synthesizes GPS advancements through benchmark construction, parsing, and reasoning paradigms.
result Unified analytical paradigm and emerging opportunities identified.

Study shows non-systematic bias in customer satisfaction surveys limits data value.

problem Non-systematic bias in customer satisfaction surveys limits data value.
method Used real customer satisfaction survey data of a large retail bank to show the irreducible error and suggest thoughtful survey design methods.
result A thoughtful survey design can reduce non-systematic error in customer satisfaction surveys.

Deep learning models outperform MICE in large survey imputation but with hyperparameter tuning.

problem Comparing deep learning and MICE for missing data imputation in large surveys.
method Extensive simulation studies comparing four machine learning-based MI methods: MICE with classification trees, MICE with random forests, generative adversarial imputation networks, and multiple imputation using denoising autoencoders.
result MICE with classification trees consistently outperforms deep learning methods in terms of bias, mean squared error, and coverage.

This paper surveys enterprise financial risk analysis from Big Data and LLMs perspectives.

problem Predicting future financial risk of enterprises.
method Systematic literature review of enterprise financial risk analysis approaches from Big Data and LLMs perspectives.
result Offers a holistic synthesis of research methods and key insights.

Reduce survey questions to scale market research without annoying customers.

problem Performing market research by surveying customers with many questions is inefficient and annoying.
method Used Bayesian networks to model and reduce the number of questions asked to customers.
result Demonstrated the effectiveness of the approach using an example of broadband customer segmentation.

Content based image retrieval, a technique which uses visual contents of image to search images from large scale image databases according to users' interests. This paper provides a comprehensive survey on recent technology used in the area of content based face image retrieval. Nowadays digital devices and photo shari…

2014-02-20abs ↗pdf ↗

Digital personas improve survey results for stable attributes but fail for subjective responses.

problem When can digital personas reliably approximate human survey findings?
method Using LISS panel, constructed personas from background variables and survey histories, tested against held-out post-cutoff answers.
result Digital personas improve alignment with human response distributions for stable attributes but fail for subjective responses.

Method completes mixed matrix from complex surveys with heterogeneous missingness.

problem Recovering a mixed dataframe matrix from complex survey sampling with different missingness patterns.
method Two-stage procedure: logistic regression for missingness modeling, and weighted log-likelihood maximization with low-rank constraint.
result The proposed method achieves sublinear convergence and shows superior performance compared to existing methods.

This paper surveys fairness notions in ML and recommends the most suitable one for real-world scenarios.

problem Ensuring ML systems do not discriminate against specific individuals or sub-populations.
method Identifying fairness-related characteristics of real-world scenarios and analyzing the behavior of fairness notions.
result A decision diagram to recommend the most suitable fairness notion for specific setups.

Survey of LLMs in finance tasks, highlighting progress and challenges.

problem Transforming financial practices with advanced LLMs.
method Exploration of various financial tasks, categorization, and analysis of methodologies.
result Unlocking novel opportunities for financial applications with LLMs.

Survey of group actions on hyperbolic spaces, focusing on mapping class groups and Out(F_n).

problem Understanding the large scale geometry of mapping class groups and Out(F_n) using their actions on hyperbolic spaces.
method Analysis of hyperbolic groups and construction of projection complexes.
result Significant understanding of Out(F_n) lags behind mapping class groups.

Survey of privacy-preserving distributed deep learning methods.

problem Protecting confidential patterns in data during distributed deep learning.
method Comparison of federated learning, split learning, large batch SGD, and privacy-preserving techniques.
result Trade-offs between computational resources, data leakage, and communication efficiency.

Survey on random features for kernel approximation, focusing on algorithms, theory, and practical applications.

problem Efficiently approximating kernel methods for large-scale problems.
method Random features techniques to speed up kernel methods.
result Need for a high number of random features for good approximation quality.

Survey examines distillation methods for large language models.

problem Efficiently compress large language models while preserving their capabilities.
method Knowledge Distillation and Dataset Distillation techniques.
result Integrating KD and DD can produce more effective and scalable compression strategies.

This paper is a survey on the {\em Zimmer program}. In it's broadest form, this program seeks an understanding of actions of large groups on compact manifolds. The goals of this survey are (1)(1) to put in context the original questions and conjectures of Zimmer and Gromov that motivated the program, (2)(2) to indicate t…

2008-09-28abs ↗pdf ↗

Survey of knowledge distillation for resource-limited devices.

problem Deploying large deep learning models on resource-limited devices.
method Knowledge distillation using a smaller model trained with information from a larger model.
result A new metric (distillation metric) for comparing different knowledge distillation algorithms.

Survey on LLMs for time series analytics across various domains.

problem Cross-modality gap between LLMs and time series data.
method Taxonomy of approaches, cross-modality strategies, and experiments on multimodal datasets.
result Effective combinations of textual data and cross-modality strategies enhance time series analytics.

Ensemble CNNs improve mode classification in smartphone travel surveys.

problem Classifying transportation modes from smartphone travel survey data.
method Developed an ensemble of CNN models with different architectures and hyper-parameters, combined using average voting, majority voting, optimal weights, and a Random Forest meta-learner.
result The ensemble method with Random Forest as meta-learner achieved 91.8% accuracy, surpassing other methods.

Survey of minimal generating sets for nonorientable mapping class groups.

problem Challenges in generating minimal sets for nonorientable surfaces.
method Detailed analysis of various generating sets, including torsions, involutions, and commutators.
result For large genus, both Mod(Ng)\mathrm{Mod}(N_{g}) and Tg\mathcal{T}_{g} are generated by two elements.