Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

25.0%50.0%75.0%100.0% · Dec 199219922001200920182026
48 results for publicly available accounting information

Study improves detection of accounting fraud using machine learning.

problem Global concern of accounting fraud threatening financial stability.
method Machine learning methods to differentiate between fraud and non-fraud companies.
result Out-of-sample results suggest great potential in detecting falsified financial statements.

Study shows publicly available news impacts financial markets.

problem Impact of publicly available news on financial markets.
method Extracted news from Common Crawl, identified relevant companies, used sentiment analysis and information theory.
result Publicly available news has significant impact on financial markets.

Study analyzes financial intermediation costs in decentralized lending protocols.

problem Understanding the cost of financial intermediation in decentralized lending protocols.
method Analysis of publicly available data on rates, supply, borrow activity, and accounts.
result Ex-post margins are 1% and lower for stablecoin markets.

A boosting algorithm PB-MVBoost learns weights for view-specific classifiers and views.

problem Improving multiview learning by balancing accuracy and diversity.
method Iteratively learns weights for view-specific classifiers and views using a PAC-Bayes multiview C-Bound.
result Efficiency demonstrated on three publicly available datasets.

This study detects fake and automated accounts on Instagram.

problem Fake engagement on Instagram leads to financial loss and wrong audience targeting.
method Two datasets were created and machine learning algorithms like Naive Bayes, Logistic Regression, Support Vector Machines, Neural Networks, and cost-sensitive genetic algorithm were applied.
result 86% accuracy for automated accounts and 96% for fake accounts were achieved.

Health care is one of the most exciting frontiers in data mining and machine learning. Successful adoption of electronic health records (EHRs) created an explosion in digital clinical data available for analysis, but progress in machine learning for healthcare research has been difficult to measure because of the absen…

2017-03-22abs ↗pdf ↗

In this paper we examine inefficiencies and information disparity in the Japanese stock market. By carefully analysing information publicly available on the internet, an `outsider' to conventional statistical arbitrage strategies--which are based on market microstructure, company releases, or analyst reports--can never…

2010-03-03abs ↗pdf ↗

This research examines relationship between staging of Venture Capital (VC) investments and social feedback visible in publicly available data on the Web. We address the question of Venture Capital investment sensitivity to performance and prospects of new venture, given as likelihood of obtaining future financing, ava…

2012-12-30abs ↗pdf ↗

A new method uses RF's out-of-bag errors for multiple imputation.

problem Missing data in biomedical studies and lack of prediction uncertainty.
method Constructs conditional distributions from the empirical distribution of out-of-bag prediction errors.
result Valid multiple imputation results achieved without parametric assumptions.

Evaluates transfer learning methods in dynamic data availability scenarios.

problem Real-world data availability varies over time, leading to unrealistic TL method evaluations.
method Proposes a data manipulation framework to simulate varying data availability and domain transformations.
result Demonstrates the usefulness of the framework on proprietary and publicly available datasets.

Enhanced word embeddings boost multiclass text classification accuracy.

problem Improving multiclass text classification accuracy using pre-trained embeddings.
method Proposed word-class embeddings (WCEs) to enhance pre-trained word embeddings.
result WCEs significantly improve multiclass text classification accuracy.

RAGIC predicts stock intervals with risk considerations, improving prediction accuracy and coverage.

problem Limited success in predicting stock market outcomes due to stochastic nature and risk oversight.
method RAGIC uses a GAN with a risk module and temporal module to generate risk-sensitive stock intervals.
result RAGIC achieves a consistent 95% coverage with narrow interval widths, balancing accuracy and risk.

Marich extracts high-fidelity models from public data with minimal queries.

problem Creating an accurate replica of a target ML model using few queries.
method Sequentially selects informative queries to maximize entropy and reduce model mismatch.
result Extracted models achieve 60-95% of target model's accuracy with 1,000-8,500 queries.

Study shows context-specific models improve swipe gesture authentication for smartphone users.

problem Improving swipe gesture-based continuous authentication for smartphones.
method Conducted experiments on HMOG dataset with 100 subjects, analyzing authentication error in different scenarios.
result Context-specific models are needed for different smartphone usage and human activity scenarios.

Deep learning models (aka Deep Neural Networks) have revolutionized many fields including computer vision, natural language processing, speech recognition, and is being increasingly used in clinical healthcare applications. However, few works exist which have benchmarked the performance of the deep learning models with…

2017-10-23abs ↗pdf ↗

New feature selection methods improve uplift modeling accuracy.

problem Overfitting and poor interpretability in feature selection for uplift models.
method Explicitly designed feature selection methods inspired by statistics and information theory.
result Proposed methods outperform traditional feature selection methods in uplift modeling.

Model learns to discover and disambiguate entities and relations in text streams.

problem Learning to follow and resolve mentions in a continuous text stream.
method End-to-end trainable memory network for online, one-shot learning.
result Improves disambiguation and discovery skills with minimal supervision.

We consider the problem of discrete-time signal denoising, focusing on a specific family of non-linear convolution-type estimators. Each such estimator is associated with a time-invariant filter which is obtained adaptively, by solving a certain convex optimization problem. Adaptive convolution-type estimators were dem…

2018-03-29abs ↗pdf ↗

Mixture of multi-task GPs for clustering and prediction of functional data.

problem Handling multi-task learning, clustering, and prediction for functional data.
method A mixture of multi-task Gaussian processes with a variational EM algorithm for hyper-parameter optimization.
result Enhanced predictive performance for group-structured data.

VTrackIt creates a synthetic dataset with infrastructure and vehicle info for AVs.

problem Lack of infrastructure and pooled vehicle info in existing AV datasets.
method Developed VTrackIt, a synthetic dataset with intelligent infrastructure and pooled vehicle info, and introduced InfraGAN for trajectory predictions.
result VTrackIt reduces high-risk edge cases in AV trajectory predictions.

Unsupervised learning filters tweets for emergency services during crises.

problem Challenges in filtering relevant information from social web data during disasters.
method Multi-task domain adversarial attention network for unsupervised domain adaptation.
result The multi-task model outperforms single task models in filtering relevant tweets.

In this study, we propose a new statical approach for high-dimensionality reduction of heterogenous data that limits the curse of dimensionality and deals with missing values. To handle these latter, we propose to use the Random Forest imputation's method. The main purpose here is to extract useful information and so r…

2017-07-02abs ↗pdf ↗

Study tackles hate speech against journalists on social media.

problem Hate speech against journalists on social media remains prevalent despite efforts.
method Defined journalist-specific hate speech, annotated tweets, trained deep learning models, and proposed an ensemble model.
result Proposed ensemble model outperforms individual models in detecting journalist-targeted hate speech.

A new screening method for high-dimensional data reduces computational cost.

problem Challenges in variable selection for ultrahigh-dimensional linear regression.
method Ordering absolute sample ridge partial correlations to screen variables.
result The method provides sure screening property without strong assumptions.

I introduce Forecastable Component Analysis (ForeCA), a novel dimension reduction technique for temporally dependent signals. Based on a new forecastability measure, ForeCA finds an optimal transformation to separate a multivariate time series into a forecastable and an orthogonal white noise space. I present a converg…

2012-05-21abs ↗pdf ↗

MIMIC-Extract transforms EHR data for reproducible healthcare machine learning.

problem Lack of accessible, standardized healthcare data for machine learning.
method Open-source pipeline for converting raw EHR data into usable dataframes.
result Demonstrates utility through benchmark tasks and baseline results.

Mindful active learning improves activity recognition using wearable sensors.

problem Activity recognition using wearable sensors with human cognitive and physical limitations.
method Introduces mindful active learning, a framework that considers human memory and query budget.
result EMMA framework achieves higher accuracy than traditional methods, especially with limited query budgets and weak human memory.

Deep learning models combine text and time-series data for better taxi demand forecasts.

problem Accurate taxi demand forecasting in event areas.
method Two deep learning architectures using word embeddings, convolutional layers, and attention mechanisms.
result The models significantly reduce forecast error by fusing text and time-series data.

Detecting PE malware files is now commonly approached using statistical and machine learning models. While these models commonly use features extracted from the structure of PE files, we propose that icons from these files can also help better predict malware. We propose an innovative machine learning approach to extra…

2017-12-10abs ↗pdf ↗

Bayesian variational inference improves medical image segmentation confidence.

problem Improving interpretability and confidence in deep learning models for medical image segmentation.
method Encoder-decoder architecture based on variational inference for segmenting brain tumor images.
result The model segments brain tumors with both aleatoric and epistemic uncertainty.

New method detects information leakage using approximate Bayes predictor.

problem Unintentional exposure of sensitive information via observable data.
method Statistical learning theory and information theory framework, approximating Bayes predictor's log-loss and accuracy.
result MI can be accurately estimated to detect ILs, outperforming state-of-the-art baselines.

The paper introduces a new method to select high-quality clustering solutions in k-means.

problem Selecting the optimal number of clusters in k-means clustering.
method The paper introduces a new method to estimate the degrees of freedom in k-means clustering, which is used for model selection.
result The proposed method for selecting high-quality clustering solutions is competitive and reliable.

Enhanced travel time prediction using deep neural networks and road network information.

problem Improving travel time estimation using deep learning models.
method Proposes incorporating road network information into deep learning models for travel time prediction.
result Improved travel time prediction, especially with limited training data.

Funnelling improves cross-lingual text classification accuracy.

problem Classifying documents in multiple languages more accurately than individual language classifiers.
method A two-tier classification system using posterior probabilities from language-dependent classifiers.
result Funnelling significantly outperforms state-of-the-art baselines in multilingual text classification.

Proposes a novel imputation network for clinical time series data.

problem Missing value imputation in clinical time series data with sparsity, irregularity, and high-dimensionality.
method Variational-recurrent imputation network that considers correlated features, temporal dynamics, and uncertainty.
result The proposed method outperformed state-of-the-art methods on real-world EHR datasets.