Study improves detection of accounting fraud using machine learning.
problem Global concern of accounting fraud threatening financial stability.
method Machine learning methods to differentiate between fraud and non-fraud companies.
result Out-of-sample results suggest great potential in detecting falsified financial statements.
Study shows publicly available news impacts financial markets.
problem Impact of publicly available news on financial markets.
method Extracted news from Common Crawl, identified relevant companies, used sentiment analysis and information theory.
result Publicly available news has significant impact on financial markets.
A number of important applied problems in engineering, finance and medicine can be formulated as a problem of anomaly detection. A classical approach to the problem is to describe a normal state using a one-class support vector machine. Then to detect anomalies we quantify a distance from a new observation to the const…
Study analyzes financial intermediation costs in decentralized lending protocols.
problem Understanding the cost of financial intermediation in decentralized lending protocols.
method Analysis of publicly available data on rates, supply, borrow activity, and accounts.
result Ex-post margins are 1% and lower for stablecoin markets.
A boosting algorithm PB-MVBoost learns weights for view-specific classifiers and views.
problem Improving multiview learning by balancing accuracy and diversity.
method Iteratively learns weights for view-specific classifiers and views using a PAC-Bayes multiview C-Bound.
result Efficiency demonstrated on three publicly available datasets.
This study detects fake and automated accounts on Instagram.
problem Fake engagement on Instagram leads to financial loss and wrong audience targeting.
method Two datasets were created and machine learning algorithms like Naive Bayes, Logistic Regression, Support Vector Machines, Neural Networks, and cost-sensitive genetic algorithm were applied.
result 86% accuracy for automated accounts and 96% for fake accounts were achieved.
Health care is one of the most exciting frontiers in data mining and machine learning. Successful adoption of electronic health records (EHRs) created an explosion in digital clinical data available for analysis, but progress in machine learning for healthcare research has been difficult to measure because of the absen…
In this paper we examine inefficiencies and information disparity in the Japanese stock market. By carefully analysing information publicly available on the internet, an `outsider' to conventional statistical arbitrage strategies--which are based on market microstructure, company releases, or analyst reports--can never…
This research examines relationship between staging of Venture Capital (VC) investments and social feedback visible in publicly available data on the Web. We address the question of Venture Capital investment sensitivity to performance and prospects of new venture, given as likelihood of obtaining future financing, ava…
A new method uses RF's out-of-bag errors for multiple imputation.
problem Missing data in biomedical studies and lack of prediction uncertainty.
method Constructs conditional distributions from the empirical distribution of out-of-bag prediction errors.
result Valid multiple imputation results achieved without parametric assumptions.
Evaluates transfer learning methods in dynamic data availability scenarios.
problem Real-world data availability varies over time, leading to unrealistic TL method evaluations.
method Proposes a data manipulation framework to simulate varying data availability and domain transformations.
result Demonstrates the usefulness of the framework on proprietary and publicly available datasets.
Enhanced word embeddings boost multiclass text classification accuracy.
problem Improving multiclass text classification accuracy using pre-trained embeddings.
method Proposed word-class embeddings (WCEs) to enhance pre-trained word embeddings.
result WCEs significantly improve multiclass text classification accuracy.
RAGIC predicts stock intervals with risk considerations, improving prediction accuracy and coverage.
problem Limited success in predicting stock market outcomes due to stochastic nature and risk oversight.
method RAGIC uses a GAN with a risk module and temporal module to generate risk-sensitive stock intervals.
result RAGIC achieves a consistent 95% coverage with narrow interval widths, balancing accuracy and risk.
Marich extracts high-fidelity models from public data with minimal queries.
problem Creating an accurate replica of a target ML model using few queries.
method Sequentially selects informative queries to maximize entropy and reduce model mismatch.
result Extracted models achieve 60-95% of target model's accuracy with 1,000-8,500 queries.
Green bond leaks impact equity markets, altering investor reactions.
problem Green bond leaks affect equity market reactions.
method Identified 259 instances of pre-announcement leaks in 2,036 green bond headlines.
result News leaks significantly alter equity trading dynamics and investor reactions.
Study shows context-specific models improve swipe gesture authentication for smartphone users.
problem Improving swipe gesture-based continuous authentication for smartphones.
method Conducted experiments on HMOG dataset with 100 subjects, analyzing authentication error in different scenarios.
result Context-specific models are needed for different smartphone usage and human activity scenarios.
Deep learning models (aka Deep Neural Networks) have revolutionized many fields including computer vision, natural language processing, speech recognition, and is being increasingly used in clinical healthcare applications. However, few works exist which have benchmarked the performance of the deep learning models with…
New feature selection methods improve uplift modeling accuracy.
problem Overfitting and poor interpretability in feature selection for uplift models.
method Explicitly designed feature selection methods inspired by statistics and information theory.
result Proposed methods outperform traditional feature selection methods in uplift modeling.
CatBoost boosts performance on datasets with categorical features.
problem Handling categorical features in gradient boosting.
method Gradient boosting library with GPU and CPU implementations.
result Outperforms existing implementations on popular datasets.
Model learns to discover and disambiguate entities and relations in text streams.
problem Learning to follow and resolve mentions in a continuous text stream.
method End-to-end trainable memory network for online, one-shot learning.
result Improves disambiguation and discovery skills with minimal supervision.
We consider the problem of discrete-time signal denoising, focusing on a specific family of non-linear convolution-type estimators. Each such estimator is associated with a time-invariant filter which is obtained adaptively, by solving a certain convex optimization problem. Adaptive convolution-type estimators were dem…
Mixture of multi-task GPs for clustering and prediction of functional data.
problem Handling multi-task learning, clustering, and prediction for functional data.
method A mixture of multi-task Gaussian processes with a variational EM algorithm for hyper-parameter optimization.
result Enhanced predictive performance for group-structured data.
VTrackIt creates a synthetic dataset with infrastructure and vehicle info for AVs.
problem Lack of infrastructure and pooled vehicle info in existing AV datasets.
method Developed VTrackIt, a synthetic dataset with intelligent infrastructure and pooled vehicle info, and introduced InfraGAN for trajectory predictions.
result VTrackIt reduces high-risk edge cases in AV trajectory predictions.
Unsupervised learning filters tweets for emergency services during crises.
problem Challenges in filtering relevant information from social web data during disasters.
method Multi-task domain adversarial attention network for unsupervised domain adaptation.
result The multi-task model outperforms single task models in filtering relevant tweets.
In this study, we propose a new statical approach for high-dimensionality reduction of heterogenous data that limits the curse of dimensionality and deals with missing values. To handle these latter, we propose to use the Random Forest imputation's method. The main purpose here is to extract useful information and so r…
Study tackles hate speech against journalists on social media.
problem Hate speech against journalists on social media remains prevalent despite efforts.
method Defined journalist-specific hate speech, annotated tweets, trained deep learning models, and proposed an ensemble model.
result Proposed ensemble model outperforms individual models in detecting journalist-targeted hate speech.
A new screening method for high-dimensional data reduces computational cost.
problem Challenges in variable selection for ultrahigh-dimensional linear regression.
method Ordering absolute sample ridge partial correlations to screen variables.
result The method provides sure screening property without strong assumptions.
Extracts StarCraft II tournament data for AI and ML studies.
problem Lack of accessible esports data for scientific use.
method Gathered and processed StarCraft II tournament replays using an API parser library.
result The largest publicly available StarCraft II esports dataset.
New algorithm clusters data and learns kernels without relaxing constraints.
problem Learning kernels or distance metrics from pairwise constraints without losing generalization.
method Joint clustering and kernel learning without relaxing constraints.
result Outperforms existing approaches on diverse datasets.
I introduce Forecastable Component Analysis (ForeCA), a novel dimension reduction technique for temporally dependent signals. Based on a new forecastability measure, ForeCA finds an optimal transformation to separate a multivariate time series into a forecastable and an orthogonal white noise space. I present a converg…
High frequency trading has led to widespread efforts to reduce information propagation delays between physically distant exchanges. Using relativistically correct millisecond-resolution tick data, we document a 3-millisecond decrease in one-way communication time between the Chicago and New York areas that has occurred…
MIMIC-Extract transforms EHR data for reproducible healthcare machine learning.
problem Lack of accessible, standardized healthcare data for machine learning.
method Open-source pipeline for converting raw EHR data into usable dataframes.
result Demonstrates utility through benchmark tasks and baseline results.
Mindful active learning improves activity recognition using wearable sensors.
problem Activity recognition using wearable sensors with human cognitive and physical limitations.
method Introduces mindful active learning, a framework that considers human memory and query budget.
result EMMA framework achieves higher accuracy than traditional methods, especially with limited query budgets and weak human memory.
Deep learning models combine text and time-series data for better taxi demand forecasts.
problem Accurate taxi demand forecasting in event areas.
method Two deep learning architectures using word embeddings, convolutional layers, and attention mechanisms.
result The models significantly reduce forecast error by fusing text and time-series data.
Accurate real-time tracking of influenza outbreaks helps public health officials make timely and meaningful decisions that could save lives. We propose an influenza tracking model, ARGO (AutoRegression with GOogle search data), that uses publicly available online search data. In addition to having a rigorous statistica…
Develops workflow for synthetic insurance datasets.
problem Lack of realistic publicly available insurance datasets.
method Uses CTGAN neural network architecture to generate tabular data.
result Synthesized datasets evaluated positively in multiple aspects.
Detecting PE malware files is now commonly approached using statistical and machine learning models. While these models commonly use features extracted from the structure of PE files, we propose that icons from these files can also help better predict malware. We propose an innovative machine learning approach to extra…
Nowadays, hyperspectral image classification widely copes with spatial information to improve accuracy. One of the most popular way to integrate such information is to extract hierarchical features from a multiscale segmentation. In the classification context, the extracted features are commonly concatenated into a lon…
Bayesian variational inference improves medical image segmentation confidence.
problem Improving interpretability and confidence in deep learning models for medical image segmentation.
method Encoder-decoder architecture based on variational inference for segmenting brain tumor images.
result The model segments brain tumors with both aleatoric and epistemic uncertainty.
TAPER learns unified patient EHR representations for healthcare tasks.
problem Irregular and multimodal data in electronic health records.
method Transformer networks and BERT for embedding structured and unstructured data.
result TAPER model outperforms on mortality, readmission, and length of stay tasks.
New method detects information leakage using approximate Bayes predictor.
problem Unintentional exposure of sensitive information via observable data.
method Statistical learning theory and information theory framework, approximating Bayes predictor's log-loss and accuracy.
result MI can be accurately estimated to detect ILs, outperforming state-of-the-art baselines.
The paper introduces a new method to select high-quality clustering solutions in k-means.
problem Selecting the optimal number of clusters in k-means clustering.
method The paper introduces a new method to estimate the degrees of freedom in k-means clustering, which is used for model selection.
result The proposed method for selecting high-quality clustering solutions is competitive and reliable.
Proposed SMO algorithm for OC-SVM+ significantly outperforms non-sequential algorithms.
problem One-class SVM with privileged information
method Sequential Minimal Optimization (SMO) algorithm
result Finite-time convergence established
Study improves engagement prediction in educational videos.
problem Addressing cold-start problem in educational recommenders.
method Introduced VLE dataset with content and engagement features, conducted experiments.
result VLE dataset leads to better engagement prediction models.
Enhanced travel time prediction using deep neural networks and road network information.
problem Improving travel time estimation using deep learning models.
method Proposes incorporating road network information into deep learning models for travel time prediction.
result Improved travel time prediction, especially with limited training data.
Funnelling improves cross-lingual text classification accuracy.
problem Classifying documents in multiple languages more accurately than individual language classifiers.
method A two-tier classification system using posterior probabilities from language-dependent classifiers.
result Funnelling significantly outperforms state-of-the-art baselines in multilingual text classification.
Deep learning predicts fit for fashion e-commerce.
problem Predicting correct fit for customer satisfaction and cost reduction.
method Deep learning content-collaborative approach using customer and article embeddings.
result Significant improvement over state-of-the-art methods.
Proposes a novel imputation network for clinical time series data.
problem Missing value imputation in clinical time series data with sparsity, irregularity, and high-dimensionality.
method Variational-recurrent imputation network that considers correlated features, temporal dynamics, and uncertainty.
result The proposed method outperformed state-of-the-art methods on real-world EHR datasets.