Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

2579 · Oct 201819922001200920172026
48 results for Apache Spark

We introduce Microsoft Machine Learning for Apache Spark (MMLSpark), an ecosystem of enhancements that expand the Apache Spark distributed computing library to tackle problems in Deep Learning, Micro-Service Orchestration, Gradient Boosting, Model Interpretability, and other areas of modern computation. Furthermore, we…

2018-10-20abs ↗pdf ↗

We report on an open-source implementation for distributed function minimization on top of Apache Spark by using gradient and quasi-Newton methods. We show-case it with an application to Optimal Transport and some scalability tests on classification and regression problems.

2019-09-17abs ↗pdf ↗

With the spreading prevalence of Big Data, many advances have recently been made in this field. Frameworks such as Apache Hadoop and Apache Spark have gained a lot of traction over the past decades and have become massively popular, especially in industries. It is becoming increasingly evident that effective big data a…

2017-11-25abs ↗pdf ↗

Apache Spark is a popular open-source platform for large-scale data processing that is well-suited for iterative machine learning tasks. In this paper we present MLlib, Spark's open-source distributed machine learning library. MLlib provides efficient functionality for a wide range of learning settings and includes sev…

2015-05-26abs ↗pdf ↗

BreachRadar detects points-of-compromise in bank transactions to prevent fraud.

problem Detecting and preventing bank transaction fraud caused by data breaches.
method A distributed alternating algorithm that assigns probabilities to different locations being compromised.
result BreachRadar achieves over 90% precision and recall in detecting compromised cards.

CFS (Correlation-Based Feature Selection) is an FS algorithm that has been successfully applied to classification problems in many domains. We describe Distributed CFS (DiCFS) as a completely redesigned, scalable, parallel and distributed version of the CFS algorithm, capable of dealing with the large volumes of data t…

2019-01-31abs ↗pdf ↗

In this paper we present a new algorithm for computing a low rank approximation of the product ATBA^TB by taking only a single pass of the two matrices AA and BB. The straightforward way to do this is to (a) first sketch AA and BB individually, and then (b) find the top components using PCA on the sketch. Our algori…

2016-10-21abs ↗pdf ↗

The AMIDST Toolbox is a software for scalable probabilistic machine learning with a spe- cial focus on (massive) streaming data. The toolbox supports a flexible modeling language based on probabilistic graphical models with latent variables and temporal dependencies. The specified models can be learnt from large data s…

2017-04-04abs ↗pdf ↗

With large volumes of health care data comes the research area of computational phenotyping, making use of techniques such as machine learning to describe illnesses and other clinical concepts from the data itself. The "traditional" approach of using supervised learning relies on a domain expert, and has two main limit…

2016-12-26abs ↗pdf ↗

Supervised learning algorithms are nowadays successfully scaling up to datasets that are very large in volume, leveraging the potential of in-memory cluster-computing Big Data frameworks. Still, massive datasets with a number of large-domain categorical features are a difficult challenge for any classifier. Most off-th…

2018-05-10abs ↗pdf ↗

A d-bar-analogue of differential characters for complex manifolds is introduced and studied using a new theory of homological spark complexes. Many essentially different spark complexes are shown to have isomorphic groups of spark classes. This has many consequences: It leads to an analytic representation of O*-gerbes …

2005-12-12abs ↗pdf ↗

We study the Harvey-Lawson spark characters of level p on complex manifolds. Presenting Deligne cohomology classes by sparks of level pp, we give an explicit analytic product formula for Deligne cohomology. We also define refined Chern classes in Deligne cohomology for holomorphic vector bundles over complex manifolds…

2008-08-12abs ↗pdf ↗

We give a new description of the ring structure on the differential characters of a smooth manifold via the smooth hyperspark complex. We show the explicit product formula, and as an application, calculate the product for differential characters of the unit circle. Applying the presentation of spark classes by smooth h…

2008-08-05abs ↗pdf ↗

We introduce a new homological machine for the study of secondary geometric invariants. The objects, called spark complexes, occur in many areas of mathematics. The theory is applied here to establish the equivalence of a large family of spark complexes which appear naturally in geometry, topology and physics. These co…

2003-06-11abs ↗pdf ↗

Predict stock trends using news sentiment and technical indicators in Spark.

problem Predicting the stock market trend is challenging due to multiple influencing factors.
method Created a machine learning classification problem with features from technical indicators and news sentiment scores.
result Random Forest model achieved 63.58% test accuracy in Spark.

Training deep networks is a time-consuming process, with networks for object recognition often requiring multiple days to train. For this reason, leveraging the resources of a cluster to speed up training is an important area of work. However, widely-popular batch-processing computational frameworks like MapReduce and …

2015-11-19abs ↗pdf ↗

We introduce a notion of K-semistability for Sasakian manifolds. This extends to the irregular case the orbifold K-semistability of Ross-Thomas. Our main result is that a Sasakian manifold with constant scalar curvature is necessarily K-semistable. As an application, we show how one can recover the volume minimization …

2012-04-10abs ↗pdf ↗

Many machine learning models, such as logistic regression~(LR) and support vector machine~(SVM), can be formulated as composite optimization problems. Recently, many distributed stochastic optimization~(DSO) methods have been proposed to solve the large-scale composite optimization problems, which have shown better per…

2016-01-30abs ↗pdf ↗

Data preprocessing techniques are devoted to correct or alleviate errors in data. Discretization and feature selection are two of the most extended data preprocessing techniques. Although we can find many proposals for static Big Data preprocessing, there is little research devoted to the continuous Big Data problem. A…

2018-10-14abs ↗pdf ↗

We introduce GraSPy, a Python library devoted to statistical inference, machine learning, and visualization of random graphs and graph populations. This package provides flexible and easy-to-use algorithms for analyzing and understanding graphs with a scikit-learn compliant API. GraSPy can be downloaded from Python Pac…

2019-03-29abs ↗pdf ↗

Feature selection (FS) is a key research area in the machine learning and data mining fields, removing irrelevant and redundant features usually helps to reduce the effort required to process a dataset while maintaining or even improving the processing algorithm's accuracy. However, traditional algorithms designed for …

2018-11-01abs ↗pdf ↗

We review the state of the art of our understanding of the conformal geometry of the irrational rotation algebra. This was sparked by a paper by Cohen and Connes. We review the more recent progress made by Connes and the second named author and the work of the authors of this review.

2018-10-24abs ↗pdf ↗

TailedTS dataset benchmarks heavy-tailed time series forecasting and periodicity quantification.

problem Benchmarking robustness of time series models under heavy-tailed distributions.
method Derived from Wikipedia page views, introduces periodicity quantification and robust loss functions.
result Standard Gaussian models degrade on high-volume page categories, while robust alternatives perform consistently.

This paper describes a distributed MapReduce implementation of the minimum Redundancy Maximum Relevance algorithm, a popular feature selection method in bioinformatics and network inference problems. The proposed approach handles both tall/narrow and wide/short datasets. We further provide an open source implementation…

2017-09-07abs ↗pdf ↗

The Basel II Accords have sparked increased interest in the development of approaches based on internal ratings systems and have initiated the elaboration of models for remote ratings forecasts based on external ones as part of Risk Management and Early Warning Systems. This article evaluates the peculiarities of curre…

2016-07-05abs ↗pdf ↗

By comparing Deligne complex and Aeppli-Bott-Chern complex, we construct a differential cohomology H^(X,,)\widehat{H}^*(X, *, *) that plays the role of Harvey-Lawson spark group H^(X,)\widehat{H}^*(X, *), and a cohomology HABC(X;Z(,))H^*_{ABC}(X; \Z(*, *)) that plays the role of Deligne cohomology HD(X;Z())H^*_{\mathcal{D}}(X; \Z(*)) for every …

2014-11-03abs ↗pdf ↗