Stock markets show unusual overnight and intraday returns.
problem Unusual patterns of overnight and intraday returns in stock markets.
method Analyzed features of the returns to deduce the cause.
result The only plausible explanation for these returns is that they are due to market manipulation.
Silence on suspicious stock market patterns persists despite lack of plausible explanations.
problem Suspicious patterns in stock market returns not explained or warned about.
method Analysis of correspondence and market data over five years.
result People aware of suspicious patterns chose not to alert the public.
EnsemFDet detects fraud by solving subproblems on small graphs, scaling up e-commerce fraud detection.
problem Detecting and preventing fraud in e-commerce, especially in large-scale bipartite graphs.
method EnsemFDet uses an ensemble approach to decompose the problem into smaller subproblems, solving them in parallel.
result EnsemFDet is up to 100x faster than state-of-the-art methods while maintaining high accuracy.
We study trade-based manipulation of stock prices from the perspective of complex trading networks constructed by using detailed information of trades. A stock trading network consists of nodes and directed links, where every trader is a node and a link is formed from one trader to the other if the former sells shares …
Cincer cleans both new and past data by identifying and relabeling suspicious and counter-examples.
problem Sequential learning under label noise, especially in applications with human supervision.
method Cincer uses example-based explanations to identify and relabel suspicious and counter-examples, leveraging Fisher information matrix approximation.
result Cincer achieves better data and models by clarifying the model's suspicions, especially with FIM approximation.
Detects organized fraudsters in insurance claims with high precision.
problem Fraudulent insurance claims lead to heavy financial losses.
method Developed a novel data-driven procedure using graph learning algorithms.
result Achieves more than 80% precision in fraud detection.
RevTrack identifies suspicious subgraphs on blockchain for AML.
problem Detecting money laundering in cryptocurrency transactions.
method Graph-based machine learning, tracking initial senders and final receivers.
result RevClassify outperforms state-of-the-art subgraph classification techniques in cost and accuracy.
Framework detects suspicious money laundering flows in large transaction graphs.
problem Detecting money laundering in large, complex transaction networks.
method Adapted framework for domain-specific constraints, including weighting method for edge significance.
result Framework outperforms state-of-the-art solutions in efficiency and effectiveness for large datasets.
This paper reviews statistical and machine learning methods for anti-money laundering.
problem Lack of scientific literature on statistical and machine learning methods for anti-money laundering.
method Client risk profiling and suspicious behavior flagging.
result Client risk profiling involves diagnostics, while suspicious behavior flagging involves non-disclosed features and hand-crafted risk indices.
Framework detects and ranks suspicious market manipulation using temporal convolutions and expert assessment.
problem Detecting and deterring rogue agents in financial markets.
method Weakly supervised learning, expert assessment, similarity search.
result Promising preliminary results in detecting and ranking suspicious market manipulation.
DeepQuarantine detects and quarantines suspicious emails.
problem High-quality spam detection and prevention.
method Convolutional Neural Networks on MIME headers for deep feature extraction.
result DQ enhances spam detection with high precision.
AIMM-X monitors markets for suspicious behavior using transparent scoring.
problem Detecting market manipulation from benign mechanisms.
method Combines microstructure signals and public attention signals for anomaly detection.
result Transparent scoring allows tracing and understanding flagged windows.
New approach makes adversarial examples less suspicious without changing perceptual salience.
problem Robustness of deep neural networks to unsuspicious adversarial examples.
method Splitting images into foreground and background, allowing larger perturbations in background while maintaining low cognitive salience.
result Dual-perturbation attacks are effective against classifiers robust to conventional attacks and adversarial training yields more robust classifiers.
Network analysis helps prevent money laundering by identifying risky clients and suspicious clusters.
problem Preventing money laundering using social network analysis.
method Real-world data analysis, network metrics, predictive models, visual analysis.
result Risk profiles can be predicted using social network metrics.
Detects malicious accounts in permissionless blockchains using graph properties and ML.
problem Identifying and classifying malicious accounts in permissionless blockchains.
method Temporal graph properties, ML algorithms (ExtraTreesClassifier, K-Means), cosine similarity, behavior change analysis.
result ExtraTreesClassifier performs best in detecting malicious accounts on Ethereum blockchain.
Defense against adversarial attacks by manipulating feature thickness.
problem Vulnerability of machine learning models to adversarial attacks.
method Feature Manipulation (FM)-Defense using a combo-variational autoencoder.
result Detection and purification of adversarial examples with high accuracy.
Paper improves DNS typo-squatting detection with ensemble model.
problem Detecting and preventing DNS typo-squatting attacks.
method Ensemble-based feature selection and bagging classification model.
result The proposed framework achieves high accuracy and precision in identifying typo-squatting domains.
Machine learning detects survey validity from user behavior.
problem Detecting valid responses in web surveys.
method Uses mouse activity and machine learning models (LSTM, HMM).
result Predicts survey validity without analyzing specific answers.
One of the challenges of using machine learning techniques with medical data is the frequent dearth of source image data on which to train. A representative example is automated lung cancer diagnosis, where nodule images need to be classified as suspicious or benign. In this work we propose an automatic synthetic lung …
We introduce a comprehensive and statistical framework in a model free setting for a complete treatment of localized data corruptions due to severe noise sources, e.g., an occluder in the case of a visual recording. Within this framework, we propose i) a novel algorithm to efficiently separate, i.e., detect and localiz…
Anomaly detection scores from VAE gradients improve tumor detection.
problem Improving anomaly detection in medical imaging.
method Using Variational Autoencoders to approximate anomaly ratings.
result Variance Autoencoder gradient-based ratings outperform other methods in tumor detection.
We present a data mining approach for profiling bank clients in order to support the process of detection of anti-money laundering operations. We first present the overall system architecture, and then focus on the relevant component for this paper. We detail the experiments performed on real world data from a financia…
Oddnet detects anomalies in dynamic networks using time series methods.
problem Detecting anomalies in temporal networks (e.g., transport, social networks).
method Feature-based network anomaly detection using time series methods.
result Demonstrated effectiveness on synthetic and real-world datasets.
Many machine learning systems rely on data collected in the wild from untrusted sources, exposing the learning algorithms to data poisoning. Attackers can inject malicious data in the training dataset to subvert the learning process, compromising the performance of the algorithm producing errors in a targeted or an ind…
GCAN detects fake news on social media with explanations.
problem Detecting fake news on social media with explanations.
method Graph-aware Co-Attention Networks (GCAN).
result GCAN significantly outperforms state-of-the-art methods in accuracy.
Method detects insider trading using trading data and dimensionality reduction.
problem Identifying insider trading in large datasets.
method Unsupervised machine learning, principal component analysis, autoencoders.
result Identifies suspicious trading behavior based on reconstruction errors.
Paper proposes a new topology for AML analysis using Poincaré embeddings.
problem Complex money laundering schemes and regulatory constraints hinder AML analysis and information sharing.
method Proposes a new topology for AML analysis using Poincaré embeddings.
result Demonstrates improved AML analysis and information sharing through Poincaré embeddings.
An assumption-free automatic check of medical images for potentially overseen anomalies would be a valuable assistance for a radiologist. Deep learning and especially Variational Auto-Encoders (VAEs) have shown great potential in the unsupervised learning of data distributions. In principle, this allows for such a chec…
Model detects market anomalies using a Hawkes process with hidden Markov chain.
problem Detecting high-frequency market manipulation in cryptocurrency trades.
method Developed a Markov-modulated Hawkes process with piecewise constant excitation kernels.
result Demonstrated the model's effectiveness in detecting suspicious trading activities.
Novel framework monitors cardiac image segmentation models in real-time.
problem Ensuring continuous high model performance and segmentation results in clinics.
method Formulated as anomaly detection, the framework derives surrogate quality measures for segmentation.
result Demonstrated accurate, fast, and scalable quality control monitoring.
Applying deep learning methods to mammography assessment has remained a challenging topic. Dense noise with sparse expressions, mega-pixel raw data resolution, lack of diverse examples have all been factors affecting performance. The lack of pixel-level ground truths have especially limited segmentation methods in push…
Trimming helps in conformal prediction when it separates anomaly scores.
problem Effectiveness of trimming in conformal prediction under contamination.
method Analyse fixed-threshold trimming as a replacement of the contaminated calibration law with a retained law.
result Trimming helps when it separates anomaly scores, reducing clean-target coverage to a one-dimensional score-CDF transfer problem.
Anomaly detection aims to distinguish observations that are rare and different from the majority. While most existing algorithms assume that instances are i.i.d., in many practical scenarios, links describing instance-to-instance dependencies and interactions are available. Such systems are called attributed networks. …
Uncertainty estimation in deep neural networks is essential for designing reliable and robust AI systems. Applications such as video surveillance for identifying suspicious activities are designed with deep neural networks (DNNs), but DNNs do not provide uncertainty estimates. Capturing reliable uncertainty estimates i…
The CAPM's market returns are endogenously determined, affecting all assets' expected returns.
problem The standard CAPM's market return assumption is not endogenously consistent.
method Demonstrates the impact of endogenously determined market returns on asset returns and the range of feasible market returns.
result Expected returns are influenced by all assets' risks, and market returns are limited by asset distribution.
Paper detects social media influencers affecting financial markets.
problem Impact of social media influencers on financial markets.
method Developed an early warning system for detecting suspicious social network activity.
result Discrepancy in meme and non-meme stocks' reactions to social networks.
We investigate the two components of the total daily return (close-to-close), the overnight return (close-to-open) and the daytime return (open-to-close), as well as the corresponding volatilities of the 2215 NYSE stocks from 1988 to 2007. The tail distribution of the volatility, the long-term memory in the sequence, a…
Method predicts rarity of image features to support research integrity investigations.
problem Difficulty in determining if image reuse is by chance or intentional.
method Statistical estimation of ORB features' chance occurrence across PubMed Open Access Subset dataset.
result The method produces decreasingly smaller p-values for more complex imagery, supporting null hypothesis.
The study finds significant power-law cross correlations in Bitcoin's return-volatility dynamics.
problem Investigating asymmetry in Bitcoin's return-volatility relationships.
method Analysis of daily and high-frequency Bitcoin data to identify cross correlations.
result Power-law cross correlations between returns and future volatilities are observed, indicating long-range dependencies.
We simulate a series of daily returns from intraday price movements initiated by microstructure elements. Significant evidence is found that daily returns and daily return volatility exhibit first order autocorrelation, but trading volume and daily return volatility are not correlated, while intraday volatility is. We …
Diversification return is an incremental return earned by a rebalanced portfolio of assets. The diversification return of a rebalanced portfolio is often incorrectly ascribed to a reduction in variance. We argue that the underlying source of the diversification return is the rebalancing, which forces the investor to se…
The paper evaluates machine learning cyber defenses using log data against adversarial attacks.
problem Evaluating the robustness of machine learning cyber defenses against adversarial attacks.
method Developed a testing framework using deep reinforcement learning and adversarial natural language processing.
result Higher dropout levels increase robustness, with 90% dropout probability showing the highest robustness.
The paper links labor income risk to stock returns using industry portfolio returns.
problem Understanding the impact of sectoral shifts on stock returns.
method Using cross-industry dispersion (CID) as a proxy for unemployment risk, the paper examines the relationship between stock returns and the sensitivity of returns to CID innovations.
result Stocks with high sensitivity to CID have lower expected returns, suggesting they are more exposed to sectoral shifts and unemployment risk.
The paper uses PCA and HMM to forecast stock returns outperforming buy-and-hold.
problem Predicting stock returns accurately.
method Applied PCA to covariance matrix of S&P 500 stocks, used HMM on principal components, and forecasted stock returns.
result The model outperforms buy-and-hold strategy in terms of annualized Sharpe ratio.
The paper explores how market-based returns depend on past trade values.
problem Improving accuracy in forecasting market-based average and volatility of returns.
method Derives the dependence of market-based volatility and higher statistical moments of returns on statistical moments and correlations of current and past trade values.
result Market-based statistical moments can be approximated by a finite number of moments, improving forecast reliability.
A new method models financial returns by separating sign and magnitude, improving forecasting accuracy.
problem Capturing nonlinear predictability in financial return dynamics.
method Decomposes returns into sign and magnitude components, using a joint distribution model.
result Significantly outperforms traditional linear models in forecasting U.S. stock market returns.
We present a simple microstructure model of financial returns that combines (i) the well-known ARFIMA process applied to tick-by-tick returns, (ii) the bid-ask bounce effect, (iii) the fat tail structure of the distribution of returns and (iv) the non-Poissonian statistics of inter-trade intervals. This model allows us…
Regression Trees analyze stock returns, revealing market excess return as the most informative factor.
problem Understanding informational content of three factors in stock returns.
method Joint regression tree analysis of daily stock return data for 5 major US corporations.
result The market excess return factor is always the most informative in all cases (solo and joint).