JAMPI improves matrix multiplication in Spark, boosting performance by up to 24%.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Good atlases are defined for effective orbifolds, and a spark complex is constructed on each good atlas. It is proved that this process is 2-functorial with compatible systems playing as morphisms between good atlases, and that the spark character 2-functor factors through this 2-functor.
Apache Spark is a popular open-source platform for large-scale data processing that is well-suited for iterative machine learning tasks. In this paper we present MLlib, Spark's open-source distributed machine learning library. MLlib provides efficient functionality for a wide range of learning settings and includes sev…
We introduce Microsoft Machine Learning for Apache Spark (MMLSpark), an ecosystem of enhancements that expand the Apache Spark distributed computing library to tackle problems in Deep Learning, Micro-Service Orchestration, Gradient Boosting, Model Interpretability, and other areas of modern computation. Furthermore, we…
A d-bar-analogue of differential characters for complex manifolds is introduced and studied using a new theory of homological spark complexes. Many essentially different spark complexes are shown to have isomorphic groups of spark classes. This has many consequences: It leads to an analytic representation of O*-gerbes …
With the spreading prevalence of Big Data, many advances have recently been made in this field. Frameworks such as Apache Hadoop and Apache Spark have gained a lot of traction over the past decades and have become massively popular, especially in industries. It is becoming increasingly evident that effective big data a…
We study the Harvey-Lawson spark characters of level p on complex manifolds. Presenting Deligne cohomology classes by sparks of level , we give an explicit analytic product formula for Deligne cohomology. We also define refined Chern classes in Deligne cohomology for holomorphic vector bundles over complex manifolds…
We report on an open-source implementation for distributed function minimization on top of Apache Spark by using gradient and quasi-Newton methods. We show-case it with an application to Optimal Transport and some scalability tests on classification and regression problems.
We give a new description of the ring structure on the differential characters of a smooth manifold via the smooth hyperspark complex. We show the explicit product formula, and as an application, calculate the product for differential characters of the unit circle. Applying the presentation of spark classes by smooth h…
We introduce a new homological machine for the study of secondary geometric invariants. The objects, called spark complexes, occur in many areas of mathematics. The theory is applied here to establish the equivalence of a large family of spark complexes which appear naturally in geometry, topology and physics. These co…
Predict stock trends using news sentiment and technical indicators in Spark.
Training deep networks is a time-consuming process, with networks for object recognition often requiring multiple days to train. For this reason, leveraging the resources of a cluster to speed up training is an important area of work. However, widely-popular batch-processing computational frameworks like MapReduce and …
The theory of differential characters is developed completely from a de Rham - Federer viewpoint. Characters are defined as equivalence classes of special currents, called sparks, which appear naturally in the theory of singular connections. There are many different spaces of currents which yield the character groups. …
We introduce a notion of K-semistability for Sasakian manifolds. This extends to the irregular case the orbifold K-semistability of Ross-Thomas. Our main result is that a Sasakian manifold with constant scalar curvature is necessarily K-semistable. As an application, we show how one can recover the volume minimization …
Many machine learning models, such as logistic regression~(LR) and support vector machine~(SVM), can be formulated as composite optimization problems. Recently, many distributed stochastic optimization~(DSO) methods have been proposed to solve the large-scale composite optimization problems, which have shown better per…
Spark Transformer achieves high sparsity in FFN and attention without sacrificing model quality.
In the same way that a contact manifold determines and is determined by a symplectic cone, a Sasaki manifold determines and is determined by a suitable Kahler cone. Kahler-Sasaki geometry is the geometry of these cones. This paper presents a symplectic action-angle coordinates approach to toric Kahler geometry and how …
Feature selection (FS) is a key research area in the machine learning and data mining fields, removing irrelevant and redundant features usually helps to reduce the effort required to process a dataset while maintaining or even improving the processing algorithm's accuracy. However, traditional algorithms designed for …
The paper shows that random frames have full spark with high probability.
A natural map from Lawson homology to Deligne cohomology groups for smooth complex projective varieties is constructed by using the Harvey-Lawson spark complexes. We also compare this to Abel-Jacobi type constructions by others.
China Vanke Co. faced a hostile takeover by Baoneng Group, sparking controversy.
We review the state of the art of our understanding of the conformal geometry of the irrational rotation algebra. This was sparked by a paper by Cohen and Connes. We review the more recent progress made by Connes and the second named author and the work of the authors of this review.
PyODDS automates outlier detection for new data sources.
CFS (Correlation-Based Feature Selection) is an FS algorithm that has been successfully applied to classification problems in many domains. We describe Distributed CFS (DiCFS) as a completely redesigned, scalable, parallel and distributed version of the CFS algorithm, capable of dealing with the large volumes of data t…
This paper describes a distributed MapReduce implementation of the minimum Redundancy Maximum Relevance algorithm, a popular feature selection method in bioinformatics and network inference problems. The proposed approach handles both tall/narrow and wide/short datasets. We further provide an open source implementation…
The Basel II Accords have sparked increased interest in the development of approaches based on internal ratings systems and have initiated the elaboration of models for remote ratings forecasts based on external ones as part of Risk Management and Early Warning Systems. This article evaluates the peculiarities of curre…
BreachRadar detects points-of-compromise in bank transactions to prevent fraud.
By comparing Deligne complex and Aeppli-Bott-Chern complex, we construct a differential cohomology that plays the role of Harvey-Lawson spark group , and a cohomology that plays the role of Deligne cohomology for every …
Topic models such as Latent Dirichlet Allocation (LDA) have been widely used in information retrieval for tasks ranging from smoothing and feedback methods to tools for exploratory search and discovery. However, classical methods for inferring topic models do not scale up to the massive size of today's publicly availab…
The study of private inference has been sparked by growing concern regarding the analysis of data when it stems from sensitive sources. We present the first method for private Bayesian inference in exponential families that properly accounts for noise introduced by the privacy mechanism. It is efficient because it work…
In this work, we develop a distributed least squares approximation (DLSA) method that is able to solve a large family of regression problems (e.g., linear regression, logistic regression, and Cox's model) on a distributed system. By approximating the local objective function using a local quadratic form, we are able to…
Least Angle Regression is a promising technique for variable selection applications, offering a nice alternative to stepwise regression. It provides an explanation for the similar behavior of LASSO (-penalized regression) and forward stagewise regression, and provides a fast implementation of both. The idea has…
Training deep networks is expensive and time-consuming with the training period increasing with data size and growth in model parameters. In this paper, we provide a framework for distributed training of deep networks over a cluster of CPUs in Apache Spark. The framework implements both Data Parallelism and Model Paral…
We generalize some of the results of Harvey, Lawson and Latschev about transgression formulas. The focus here is on flowing forms via vertical vector fields, especially Morse-Bott-Smale vector fields. We prove a very general transgression formula including also a version covering non-compact situations. Among applicati…
Bounded cohomology of groups was first studied by Gromov in 1982. Since then it has sparked much research in Geometric Group Theory. However, it is notoriously hard to explicitly compute bounded cohomology, even for most basic `non-positively curved' groups. On the other hand, there is a well-known interpretation of or…
Building on an idea laid out by Martelli--Sparks--Yau, we use the Duistermaat-Heckman localization formula and an extension of it to give rational and explicit expressions of the volume, the total transversal scalar curvature and the Einstein--Hilbert functional, seen as functionals on the Sasaki cone (Reeb cone). Stud…
Differentiable ABMs face challenges in inference and optimisation.
We leverage a streaming architecture based on ELK, Spark and Hadoop in order to collect, store, and analyse database connection logs in near real-time. The proposed system investigates outliers using unsupervised learning; widely adopted clustering and classification algorithms for log data, highlighting the subtle var…
liquidSVM is a package written in C++ that provides SVM-type solvers for various classification and regression tasks. Because of a fully integrated hyper-parameter selection, very carefully implemented solvers, multi-threading and GPU support, and several built-in data decomposition strategies it provides unprecedented…
In this paper we show that any good toric contact manifold has well defined cylindrical contact homology and describe how it can be combinatorially computed from the associated moment cone. As an application we compute the cylindrical contact homology of a particularly nice family of examples that appear in the work of…
In the era of big data, practical applications in various domains continually generate large-scale time-series data. Among them, some data show significant or potential periodicity characteristics, such as meteorological and financial data. It is critical to efficiently identify the potential periodic patterns from mas…
Communication remains the most significant bottleneck in the performance of distributed optimization algorithms for large-scale machine learning. In this paper, we propose a communication-efficient framework, CoCoA, that uses local computation in a primal-dual setting to dramatically reduce the amount of necessary comm…
Current machine learning systems operate, almost exclusively, in a statistical, or model-free mode, which entails severe theoretical limits on their power and performance. Such systems cannot reason about interventions and retrospection and, therefore, cannot serve as the basis for strong AI. To achieve human level int…
We survey what is known about minimal surfaces in that are complete, embedded, and have finite total curvature. The only classically known examples of such surfaces were the plane and the catenoid. The discovery by Costa, early in the last decade, of a new example that proved to be embedded sparked a great…
New method simplifies individual claims reserving.
Explosive growth in data and availability of cheap computing resources have sparked increasing interest in Big learning, an emerging subfield that studies scalable machine learning algorithms, systems, and applications with Big Data. Bayesian methods represent one important class of statistic methods for machine learni…
New method explains survival analysis models using median-SHAP.
The matrix-completion problem has attracted a lot of attention, largely as a result of the celebrated Netflix competition. Two popular approaches for solving the problem are nuclear-norm-regularized matrix approximation (Candes and Tao, 2009, Mazumder, Hastie and Tibshirani, 2010), and maximum-margin matrix factorizati…