Deep learning classifies galaxies from Dark Energy Survey in just 8 minutes.
problem Classifying galaxies from large-scale surveys.
method Transfer learning from pre-trained neural networks.
result Achieved state-of-the-art accuracy of 99.6% in galaxy classification.
Survey examines LLMs in financial trading.
problem Using LLMs to outperform professional traders in finance.
method Comprehensive review of current research on LLMs in financial trading.
result LLMs can potentially outperform professional traders in backtesting.
Survey on using large models to train smaller datasets in NLP.
problem Lack of large datasets and computing resources for NLP tasks.
method Analysis of recent transfer learning approaches in NLP.
result Increased demand for transfer learning in NLP due to large models.
Survey of large language models in financial prediction and trading.
problem Improving predictability and robustness of financial predictions and trading decisions.
method Task-centered taxonomy, review of empirical evidence, design patterns, benchmarks, and challenges analysis.
result Improved predictability and robustness of financial predictions and trading decisions through large language models.
The rapid development of computing power and efficient Markov Chain Monte Carlo (MCMC) simulation algorithms have revolutionized Bayesian statistics, making it a highly practical inference method in applied work. However, MCMC algorithms tend to be computationally demanding, and are particularly slow for large datasets…
An efficient LDP protocol for QMLE with improved practicality and theoretical guarantees.
problem Difficult implementation of existing LDP QMLE for large-scale surveys.
method Developed an alternative LDP protocol without long waiting time, high communication cost, and derivative boundedness assumptions.
result Sufficient conditions for consistency and asymptotic normality of the protocol.
Survey on automating geometry problem solving with large models.
problem Automating geometric problem solving with spatial understanding and logical reasoning.
method Synthesizes GPS advancements through benchmark construction, parsing, and reasoning paradigms.
result Unified analytical paradigm and emerging opportunities identified.
Survey on metric SYZ conjecture and non-archimedean geometry.
problem Existence of special Lagrangian fibrations on Calabi-Yau manifolds.
method Pluripotential theory and non-archimedean geometry.
result Subtleties and open questions in the conjectural picture.
Study shows non-systematic bias in customer satisfaction surveys limits data value.
problem Non-systematic bias in customer satisfaction surveys limits data value.
method Used real customer satisfaction survey data of a large retail bank to show the irreducible error and suggest thoughtful survey design methods.
result A thoughtful survey design can reduce non-systematic error in customer satisfaction surveys.
Deep learning models outperform MICE in large survey imputation but with hyperparameter tuning.
problem Comparing deep learning and MICE for missing data imputation in large surveys.
method Extensive simulation studies comparing four machine learning-based MI methods: MICE with classification trees, MICE with random forests, generative adversarial imputation networks, and multiple imputation using denoising autoencoders.
result MICE with classification trees consistently outperforms deep learning methods in terms of bias, mean squared error, and coverage.
Survey of LLMs in finance tasks, including adoption and performance.
problem Utilizing large language models in financial tasks.
method Review of current approaches, decision framework for adoption.
result Synthesizes state-of-the-art for LLMs in finance.
Automated classification of astronomical light curves for LSST.
problem Handling massive astronomical data from LSST.
method Gradient boosting of decision trees, feature extraction and selection, augmentation.
result Achieved one of the top results in the PLAsTiCC challenge.
Enhances price sentiment index using survey comments.
problem Improving accuracy in price sentiment analysis.
method Classified comments from Economy Watchers Survey using LLMs.
result Higher correlation with existing indices.
Survey on spectral gaps of random hyperbolic surfaces.
problem Understanding spectral gaps of random hyperbolic surfaces.
method Brief survey on geometry and spectra, discussion of results by Hide-Magee, Anantharaman-Monk, and Hide-Macera-Thomas.
result Near optimal spectral gaps for random surfaces.
Survey finds LLMs match human economic expectations closely.
problem Understanding human economic expectations and their deviations.
method Survey of LLM's expectations based on news articles.
result LLM's expectations closely match existing surveys and exhibit deviations.
This paper surveys enterprise financial risk analysis from Big Data and LLMs perspectives.
problem Predicting future financial risk of enterprises.
method Systematic literature review of enterprise financial risk analysis approaches from Big Data and LLMs perspectives.
result Offers a holistic synthesis of research methods and key insights.
This paper surveys large-scale machine learning methods for efficient data analysis.
problem Efficiently processing large-scale data with machine learning models.
method Divided into three categories: model simplification, optimization approximation, and computation parallelism.
result Blueprint for future developments in large-scale machine learning.
Reduce survey questions to scale market research without annoying customers.
problem Performing market research by surveying customers with many questions is inefficient and annoying.
method Used Bayesian networks to model and reduce the number of questions asked to customers.
result Demonstrated the effectiveness of the approach using an example of broadband customer segmentation.
Content based image retrieval, a technique which uses visual contents of image to search images from large scale image databases according to users' interests. This paper provides a comprehensive survey on recent technology used in the area of content based face image retrieval. Nowadays digital devices and photo shari…
Digital personas improve survey results for stable attributes but fail for subjective responses.
problem When can digital personas reliably approximate human survey findings?
method Using LISS panel, constructed personas from background variables and survey histories, tested against held-out post-cutoff answers.
result Digital personas improve alignment with human response distributions for stable attributes but fail for subjective responses.
Method completes mixed matrix from complex surveys with heterogeneous missingness.
problem Recovering a mixed dataframe matrix from complex survey sampling with different missingness patterns.
method Two-stage procedure: logistic regression for missingness modeling, and weighted log-likelihood maximization with low-rank constraint.
result The proposed method achieves sublinear convergence and shows superior performance compared to existing methods.
We build a deep reinforcement learning (RL) agent that can predict the likelihood of an individual testing positive for malaria by asking questions about their household. The RL agent learns to determine which survey question to ask next and when to stop to make a prediction about their likelihood of malaria based on t…
Survey on-device ML challenges and future directions.
problem Training machine learning models on-device with limited resources.
method Reformulated as resource constrained learning, comparing techniques from various AI areas.
result Identification of open challenges and future research directions.
Survey of knowledge distillation for model compression.
problem Deploying large deep models on resource-limited devices.
method Knowledge distillation from large teacher models to small student models.
result Effective model compression and acceleration.
New GAN model deblends galaxy images with high accuracy and speed.
problem Deblending blended galaxy images in dense regions of the universe.
method Branched generative adversarial network (GAN) to produce images of deblended galaxies.
result High peak signal-to-noise ratio and structural similarity scores compared to ground truth images.
This paper surveys fairness notions in ML and recommends the most suitable one for real-world scenarios.
problem Ensuring ML systems do not discriminate against specific individuals or sub-populations.
method Identifying fairness-related characteristics of real-world scenarios and analyzing the behavior of fairness notions.
result A decision diagram to recommend the most suitable fairness notion for specific setups.
New method detects careless responding in long surveys.
problem Careless responding in long surveys threatens internal validity.
method Detects a changepoint in combined measurements of carelessness.
result Highly accurate in identifying carelessness onset.
Survey of LLMs in finance tasks, highlighting progress and challenges.
problem Transforming financial practices with advanced LLMs.
method Exploration of various financial tasks, categorization, and analysis of methodologies.
result Unlocking novel opportunities for financial applications with LLMs.
Survey of group actions on hyperbolic spaces, focusing on mapping class groups and Out(F_n).
problem Understanding the large scale geometry of mapping class groups and Out(F_n) using their actions on hyperbolic spaces.
method Analysis of hyperbolic groups and construction of projection complexes.
result Significant understanding of Out(F_n) lags behind mapping class groups.
Survey of alignment techniques for large language models.
problem Ensuring large language models align with human values.
method Analysis of diverse alignment methods and training paradigms.
result Preference-based methods offer more flexibility for nuanced alignment.
Survey of financial LLMs in finance.
problem Limited research on Financial LLMs (FinLLMs).
method Chronological overview of PLMs, comparison of techniques, performance evaluations, and advanced tasks.
result Compilation of accessible datasets and benchmarks for AI research in finance.
LLMs improve financial analysis by processing large data sets.
problem Traditional financial analysis methods struggle with large data volumes.
method Integrating LLMs for enhanced data processing and analysis.
result LLMs offer new capabilities for real-time financial decision-making.
Study uses machine learning to analyze patient survey data for Lyme disease.
problem Understanding patient responses to treatment and disease progression.
method Applied various machine learning techniques to a patient registry.
result Identified key features that predict patient responses to antibiotic treatment.
Survey of privacy-preserving distributed deep learning methods.
problem Protecting confidential patterns in data during distributed deep learning.
method Comparison of federated learning, split learning, large batch SGD, and privacy-preserving techniques.
result Trade-offs between computational resources, data leakage, and communication efficiency.
Survey on random features for kernel approximation, focusing on algorithms, theory, and practical applications.
problem Efficiently approximating kernel methods for large-scale problems.
method Random features techniques to speed up kernel methods.
result Need for a high number of random features for good approximation quality.
Survey on random walks on mapping class groups and their properties.
problem Understanding random walks on mapping class groups.
method Analyzing actions on Teichmüller spaces and curve complexes.
result Laws of large numbers and central limit theorems for random walks.
Survey examines distillation methods for large language models.
problem Efficiently compress large language models while preserving their capabilities.
method Knowledge Distillation and Dataset Distillation techniques.
result Integrating KD and DD can produce more effective and scalable compression strategies.
Survey on DNNs for speech processing, focusing on limited data challenges.
problem Challenges in training DNNs for speech tasks with limited data.
method Overview of techniques for few-shot speech processing.
result Promising few-shot techniques for speech processing.
This paper is a survey on the {\em Zimmer program}. In it's broadest form, this program seeks an understanding of actions of large groups on compact manifolds. The goals of this survey are (1) to put in context the original questions and conjectures of Zimmer and Gromov that motivated the program, (2) to indicate t…
Survey of knowledge distillation for resource-limited devices.
problem Deploying large deep learning models on resource-limited devices.
method Knowledge distillation using a smaller model trained with information from a larger model.
result A new metric (distillation metric) for comparing different knowledge distillation algorithms.
CAESAR source finder improves automated source extraction for ASKAP surveys.
problem Automated source extraction challenges in ASKAP surveys.
method Extended CAESAR source finder for compact and extended sources.
result Improved algorithm performances and scalability for future ASKAP surveys.
Survey on LLMs for time series analytics across various domains.
problem Cross-modality gap between LLMs and time series data.
method Taxonomy of approaches, cross-modality strategies, and experiments on multimodal datasets.
result Effective combinations of textual data and cross-modality strategies enhance time series analytics.
Ensemble CNNs improve mode classification in smartphone travel surveys.
problem Classifying transportation modes from smartphone travel survey data.
method Developed an ensemble of CNN models with different architectures and hyper-parameters, combined using average voting, majority voting, optimal weights, and a Random Forest meta-learner.
result The ensemble method with Random Forest as meta-learner achieved 91.8% accuracy, surpassing other methods.
Survey connects and systematizes transfer learning research.
problem Reduce dependence on target domain data for target learners.
method Systematic review of 40+ transfer learning approaches.
result Importance of choosing appropriate transfer learning models.
Survey on privacy issues in deep learning and proposed solutions.
problem Privacy concerns in deep learning models due to sensitive data.
method Review of existing privacy techniques and gaps in research.
result Identification of test-time inference privacy as a research gap.
Survey of minimal generating sets for nonorientable mapping class groups.
problem Challenges in generating minimal sets for nonorientable surfaces.
method Detailed analysis of various generating sets, including torsions, involutions, and commutators.
result For large genus, both Mod(Ng) and Tg are generated by two elements. LLMs can predict CFO responses to economic surveys
problem Measuring business sentiment
method Prompting an LLM to role-play as a CFO
result LLM reproduces individual human responses
Survey on data collection challenges in machine learning.
problem Data scarcity and need for labeled data in machine learning.
method Comprehensive study of data acquisition, labeling, and improvement techniques.
result Identification of research challenges in data collection.