Exact distributed algorithm trains Random Forest models on very large datasets.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In last years, we hear about halal products in non- Muslim societies including European and American ones. In France, for example, sales of halal products sold in stores during the year 2010, increased 23 % and represented 5.5 billion euros, including 1.1 billion for the fast food, and it has not stopped growing since.…
Pinterest uses Graph Convolutional Networks for web-scale recommendations.
Efficiently selects nearest neighbors for labeling to speed up active learning.
RandNE efficiently embeds billion-scale networks with random projection.
This article focuses on the work of O. Chanel and G. Chichilnisky (2013) on the flaws of expected utility theory while assessing the value of life. Expected utility is a fundamental tool in decision theory. However, it does not fit with the experimental results when it comes to catastrophic outcomes ---see, for example…
RNE tackles scalable recommendation for billion-scale scenarios.
On December 16, Zynga, the well-known social game developing company went public. This event is following other recent IPOs in the world of social networking companies, such as Groupon, Linkedin or Pandora to cite a few. With a valuation close to 7 billion USD at the time when it went public, Zynga has become the bigge…
Efficient kernel methods for large datasets using GPU acceleration.
We present a system that enables rapid model experimentation for tera-scale machine learning with trillions of non-zero features, billions of training examples, and millions of parameters. Our contribution to the literature is a new method (SA L-BFGS) for changing batch L-BFGS to perform in near real-time by using stat…
S2M optimizes mining for diverse data subpopulations.
PyTorch-BigGraph scales graph embeddings to large graphs.
Efficiently trains GPs with billions of inducing inputs using Tensor Train decomposition.
Paper proposes efficient inner product approximation for hybrid sparse and dense vectors.
We propose BlackOut, an approximation algorithm to efficiently train massive recurrent neural network language models (RNNLMs) with million word vocabularies. BlackOut is motivated by using a discriminative loss, and we describe a new sampling strategy which significantly reduces computation while improving stability, …
Large GNNs trained with Graph Parallelism improve atomic simulation accuracy.
Credit risk management in Italy is characterized, in the period June 2008 to June 2012, by frequent (frequency=0.5 cycles per year) and intense (peak amplitude: mean=39.2 billion Euros, s.e.=2.83 billion Euros) quarterly contractions and expansions around the mean (915.4 billion Euros, s.e.=3.59 billion Euros) of the n…
BloombergGPT is a large language model trained on financial data, outperforming existing models on financial tasks.
We present a system and a set of techniques for learning linear predictors with convex losses on terascale datasets, with trillions of features, {The number of features here refers to the number of non-zero entries in the data matrix.} billions of training examples and millions of parameters in an hour using a cluster …
Algorithm finds ribbon disks for alternating knots, resolving sliceness for most prime knots.
Graphlets are induced subgraphs of a large network and are important for understanding and modeling complex networks. Despite their practical importance, graphlets have been severely limited to applications and domains with relatively small graphs. Most previous work has focused on exact algorithms, however, it is ofte…
We study distributed stochastic convex optimization under the delayed gradient model where the server nodes perform parameter updates, while the worker nodes compute stochastic gradients. We discuss, analyze, and experiment with a setup motivated by the behavior of real-world distributed computation networks, where the…
Efficiently trains large GMMs with millions to billions of parameters.
We tackle the problem of inferring node labels in a partially labeled graph where each node in the graph has multiple label types and each label type has a large number of possible labels. Our primary example, and the focus of this paper, is the joint inference of label types such as hometown, current city, and employe…
HessFormer enables distributed Hessian computation for large models.
Bayesian optimization techniques have been successfully applied to robotics, planning, sensor placement, recommendation, advertising, intelligent user interfaces and automatic algorithm configuration. Despite these successes, the approach is restricted to problems of moderate dimension, and several workshops on Bayesia…
This paper compares communication efficiency of split learning and federated learning in various scenarios.
Gaussian processes (GPs) are powerful non-parametric function estimators. However, their applications are largely limited by the expensive computational cost of the inference procedures. Existing stochastic or distributed synchronous variational inferences, although have alleviated this issue by scaling up GPs to milli…
Deep learning shows neural networks can be trained with limited data, revealing a low-dimensional manifold of optimal configurations.
New bounds show large language models can generalize beyond training data.
Stochastic variational inference (SVI), the state-of-the-art algorithm for scaling variational inference to large-datasets, is inherently serial. Moreover, it requires the parameters to fit in the memory of a single processor; this is problematic when the number of parameters is in billions. In this paper, we propose e…
A financial data provider shares insights on managing complexity in processing 18 billion daily notifications.
New models capture complex genetic causes of diseases.
We present a hybrid algorithm for Bayesian topic models that combines the efficiency of sparse Gibbs sampling with the scalability of online stochastic inference. We used our algorithm to analyze a corpus of 1.2 million books (33 billion words) with thousands of topics. Our approach reduces the bias of variational infe…
Horizon is Facebook's open RL platform for large, slow feedback datasets.
This study improves mid-cap equity performance with a data-driven, market-neutral approach.
The capacity of a neural network to absorb information is limited by its number of parameters. Conditional computation, where parts of the network are active on a per-example basis, has been proposed in theory as a way of dramatically increasing model capacity without a proportional increase in computation. In practice…
MACH reduces memory usage for extreme classification by hashing.
On December 16th, 2011, Zynga, the well-known social game developing company went public. This event followed other recent IPOs in the world of social networking companies, such as Groupon or Linkedin among others. With a valuation close to 7 billion USD at the time when it went public, Zynga became one of the biggest …
Infinite Tucker Decomposition (InfTucker) and random function prior models, as nonparametric Bayesian models on infinite exchangeable arrays, are more powerful models than widely-used multilinear factorization methods including Tucker and PARAFAC decomposition, (partly) due to their capability of modeling nonlinear rel…
TRASHFIRE improves model robustness by analyzing training rates and costs.
New bounds on stick number of knots found using random polygon generation.
Bayesian Layers adds uncertainty to neural networks, enabling faster experimentation and scalability.
The paper simplifies influence computations for large-scale machine learning models.
Study quantifies inefficiencies in U.S. equity markets, identifying open/close periods and affected stocks.
Bayesian model learns optimal number of latent dimensions for Boolean data.
We present a novel methodology to determine the fundamental value of firms in the social-networking sector based on two ingredients: (i) revenues and profits are inherently linked to its user basis through a direct channel that has no equivalent in other sectors; (ii) the growth of the number of users can be calibrated…
Unified framework for generating meteorological time series from text.