We introduce Microsoft Machine Learning for Apache Spark (MMLSpark), an ecosystem of enhancements that expand the Apache Spark distributed computing library to tackle problems in Deep Learning, Micro-Service Orchestration, Gradient Boosting, Model Interpretability, and other areas of modern computation. Furthermore, we…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We report on an open-source implementation for distributed function minimization on top of Apache Spark by using gradient and quasi-Newton methods. We show-case it with an application to Optimal Transport and some scalability tests on classification and regression problems.
With the spreading prevalence of Big Data, many advances have recently been made in this field. Frameworks such as Apache Hadoop and Apache Spark have gained a lot of traction over the past decades and have become massively popular, especially in industries. It is becoming increasingly evident that effective big data a…
Apache Spark is a popular open-source platform for large-scale data processing that is well-suited for iterative machine learning tasks. In this paper we present MLlib, Spark's open-source distributed machine learning library. MLlib provides efficient functionality for a wide range of learning settings and includes sev…
JAMPI improves matrix multiplication in Spark, boosting performance by up to 24%.
PyODDS automates outlier detection for new data sources.
BreachRadar detects points-of-compromise in bank transactions to prevent fraud.
Training deep networks is expensive and time-consuming with the training period increasing with data size and growth in model parameters. In this paper, we provide a framework for distributed training of deep networks over a cluster of CPUs in Apache Spark. The framework implements both Data Parallelism and Model Paral…
In the era of big data, practical applications in various domains continually generate large-scale time-series data. Among them, some data show significant or potential periodicity characteristics, such as meteorological and financial data. It is critical to efficiently identify the potential periodic patterns from mas…
CFS (Correlation-Based Feature Selection) is an FS algorithm that has been successfully applied to classification problems in many domains. We describe Distributed CFS (DiCFS) as a completely redesigned, scalable, parallel and distributed version of the CFS algorithm, capable of dealing with the large volumes of data t…
In this paper we present a new algorithm for computing a low rank approximation of the product by taking only a single pass of the two matrices and . The straightforward way to do this is to (a) first sketch and individually, and then (b) find the top components using PCA on the sketch. Our algori…
The AMIDST Toolbox is a software for scalable probabilistic machine learning with a spe- cial focus on (massive) streaming data. The toolbox supports a flexible modeling language based on probabilistic graphical models with latent variables and temporal dependencies. The specified models can be learnt from large data s…
We consider the problem of learning a high-dimensional but low-rank matrix from a large-scale dataset distributed over several machines, where low-rankness is enforced by a convex trace norm constraint. We propose DFW-Trace, a distributed Frank-Wolfe algorithm which leverages the low-rank structure of its updates to ac…
In this paper, a neural network-based stock price prediction and trading system using technical analysis indicators is presented. The model developed first converts the financial time series data into a series of buy-sell-hold trigger signals using the most commonly preferred technical analysis indicators. Then, a Mult…
Over the past decade, several approaches have been introduced for short-term traffic prediction. However, providing fine-grained traffic prediction for large-scale transportation networks where numerous detectors are geographically deployed to collect traffic data is still an open issue. To address this issue, in this …
With large volumes of health care data comes the research area of computational phenotyping, making use of techniques such as machine learning to describe illnesses and other clinical concepts from the data itself. The "traditional" approach of using supervised learning relies on a domain expert, and has two main limit…
Supervised learning algorithms are nowadays successfully scaling up to datasets that are very large in volume, leveraging the potential of in-memory cluster-computing Big Data frameworks. Still, massive datasets with a number of large-domain categorical features are a difficult challenge for any classifier. Most off-th…
Good atlases are defined for effective orbifolds, and a spark complex is constructed on each good atlas. It is proved that this process is 2-functorial with compatible systems playing as morphisms between good atlases, and that the spark character 2-functor factors through this 2-functor.
A d-bar-analogue of differential characters for complex manifolds is introduced and studied using a new theory of homological spark complexes. Many essentially different spark complexes are shown to have isomorphic groups of spark classes. This has many consequences: It leads to an analytic representation of O*-gerbes …
It is crucial to provide compatible treatment schemes for a disease according to various symptoms at different stages. However, most classification methods might be ineffective in accurately classifying a disease that holds the characteristics of multiple treatment stages, various symptoms, and multi-pathogenesis. More…
Protein interactions constitute the fundamental building block of almost every life activity. Identifying protein communities from Protein-Protein Interaction (PPI) networks is essential to understand the principles of cellular organization and explore the causes of various diseases. It is critical to integrate multipl…
We present GluonCV and GluonNLP, the deep learning toolkits for computer vision and natural language processing based on Apache MXNet (incubating). These toolkits provide state-of-the-art pre-trained models, training scripts, and training logs, to facilitate rapid prototyping and promote reproducible research. We also …
We study the Harvey-Lawson spark characters of level p on complex manifolds. Presenting Deligne cohomology classes by sparks of level , we give an explicit analytic product formula for Deligne cohomology. We also define refined Chern classes in Deligne cohomology for holomorphic vector bundles over complex manifolds…
We give a new description of the ring structure on the differential characters of a smooth manifold via the smooth hyperspark complex. We show the explicit product formula, and as an application, calculate the product for differential characters of the unit circle. Applying the presentation of spark classes by smooth h…
We introduce a new homological machine for the study of secondary geometric invariants. The objects, called spark complexes, occur in many areas of mathematics. The theory is applied here to establish the equivalence of a large family of spark complexes which appear naturally in geometry, topology and physics. These co…
Road accidents are an important issue of our modern societies, responsible for millions of deaths and injuries every year in the world. In Quebec only, in 2018, road accidents are responsible for 359 deaths and 33 thousands of injuries. In this paper, we show how one can leverage open datasets of a city like Montreal, …
In recent years, analyzing task-based fMRI (tfMRI) data has become an essential tool for understanding brain function and networks. However, due to the sheer size of tfMRI data, its intrinsic complex structure, and lack of ground truth of underlying neural activities, modeling tfMRI data is hard and challenging. Previo…
Predict stock trends using news sentiment and technical indicators in Spark.
Training deep networks is a time-consuming process, with networks for object recognition often requiring multiple days to train. For this reason, leveraging the resources of a cluster to speed up training is an important area of work. However, widely-popular batch-processing computational frameworks like MapReduce and …
The rise of Online Social Networks (OSNs) has caused an insurmountable amount of interest from advertisers and researchers seeking to monopolize on its features. Researchers aim to develop strategies for determining how information is propagated among users within an OSN that is captured by diffusion or influence model…
The theory of differential characters is developed completely from a de Rham - Federer viewpoint. Characters are defined as equivalence classes of special currents, called sparks, which appear naturally in the theory of singular connections. There are many different spaces of currents which yield the character groups. …
We introduce a notion of K-semistability for Sasakian manifolds. This extends to the irregular case the orbifold K-semistability of Ross-Thomas. Our main result is that a Sasakian manifold with constant scalar curvature is necessarily K-semistable. As an application, we show how one can recover the volume minimization …
Many machine learning models, such as logistic regression~(LR) and support vector machine~(SVM), can be formulated as composite optimization problems. Recently, many distributed stochastic optimization~(DSO) methods have been proposed to solve the large-scale composite optimization problems, which have shown better per…
Spark Transformer achieves high sparsity in FFN and attention without sacrificing model quality.
Data preprocessing techniques are devoted to correct or alleviate errors in data. Discretization and feature selection are two of the most extended data preprocessing techniques. Although we can find many proposals for static Big Data preprocessing, there is little research devoted to the continuous Big Data problem. A…
In the same way that a contact manifold determines and is determined by a symplectic cone, a Sasaki manifold determines and is determined by a suitable Kahler cone. Kahler-Sasaki geometry is the geometry of these cones. This paper presents a symplectic action-angle coordinates approach to toric Kahler geometry and how …
We introduce GraSPy, a Python library devoted to statistical inference, machine learning, and visualization of random graphs and graph populations. This package provides flexible and easy-to-use algorithms for analyzing and understanding graphs with a scikit-learn compliant API. GraSPy can be downloaded from Python Pac…
Feature selection (FS) is a key research area in the machine learning and data mining fields, removing irrelevant and redundant features usually helps to reduce the effort required to process a dataset while maintaining or even improving the processing algorithm's accuracy. However, traditional algorithms designed for …
The paper shows that random frames have full spark with high probability.
A natural map from Lawson homology to Deligne cohomology groups for smooth complex projective varieties is constructed by using the Harvey-Lawson spark complexes. We also compare this to Abel-Jacobi type constructions by others.
DADApy analyzes high-dimensional data manifolds in Python.
China Vanke Co. faced a hostile takeover by Baoneng Group, sparking controversy.
We review the state of the art of our understanding of the conformal geometry of the irrational rotation algebra. This was sparked by a paper by Cohen and Connes. We review the more recent progress made by Connes and the second named author and the work of the authors of this review.
TailedTS dataset benchmarks heavy-tailed time series forecasting and periodicity quantification.
This paper describes a distributed MapReduce implementation of the minimum Redundancy Maximum Relevance algorithm, a popular feature selection method in bioinformatics and network inference problems. The proposed approach handles both tall/narrow and wide/short datasets. We further provide an open source implementation…
The Basel II Accords have sparked increased interest in the development of approaches based on internal ratings systems and have initiated the elaboration of models for remote ratings forecasts based on external ones as part of Risk Management and Early Warning Systems. This article evaluates the peculiarities of curre…
By comparing Deligne complex and Aeppli-Bott-Chern complex, we construct a differential cohomology that plays the role of Harvey-Lawson spark group , and a cohomology that plays the role of Deligne cohomology for every …
Topic models such as Latent Dirichlet Allocation (LDA) have been widely used in information retrieval for tasks ranging from smoothing and feedback methods to tools for exploratory search and discovery. However, classical methods for inferring topic models do not scale up to the massive size of today's publicly availab…