Sparse codes improve optimal control tasks with correlated inputs.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Investigates neural codes and their embeddings, proving conjectures and introducing new code types.
Graph-Structured Cache improves code completion and variable naming tasks.
SLM models code syntax as trees to generate any programming language code.
Adversarial attacks found to be effective on code models.
Distributed gradient descent (DGD) is an efficient way of implementing gradient descent (GD), especially for large data sets, by dividing the computation tasks into smaller subtasks and assigning to different computing servers (CSs) to be executed in parallel. In standard parallel execution, per-iteration waiting time …
We study the concept of a code (or shift) space for a generalized iterated function system (GIFS in short). We prove that relations between GIFSs and their code spaces are analogous to the case of classical IFSs. As an application, we consider the problem of connectedness of attractors of GIFSs. Many of our results are…
Hybrid approach reduces computation time and decoding complexity.
We give complete algorithms and source code for constructing statistical risk models, including methods for fixing the number of risk factors. One such method is based on eRank (effective rank) and yields results similar to (and further validates) the method set forth in an earlier paper by one of us. We also give a co…
This work proposes a meta-learning approach for better adaptation of source code models.
We provide complete source code for building a fundamental industry classification based on publically available and freely downloadable data. We compare various fundamental industry classifications by running a horserace of short-horizon trading signals (alphas) utilizing open source heterotic risk models (https://ssr…
We propose a new method for learning word representations using hierarchical regularization in sparse coding inspired by the linguistic study of word meanings. We show an efficient learning algorithm based on stochastic proximal methods that is significantly faster than previous approaches, making it possible to perfor…
PredNet fails to fully adhere to predictive coding principles.
We provide complete source code for a front-end GUI and its back-end counterpart for a stock market visualization tool. It is built based on the "functional visualization" concept we discuss, whereby functionality is not sacrificed for fancy graphics. The GUI, among other things, displays a color-coded signal (computed…
We present a neural model for representing snippets of code as continuous distributed vectors ("code embeddings"). The main idea is to represent a code snippet as a single fixed-length , which can be used to predict semantic properties of the snippet. This is performed by decomposing code to a col…
New method improves matrix completion with functional maps.
The paper predicts run times for Gaussian chemistry code.
ICQ improves high-dimensional similarity search without sacrificing precision.
LumièreNet creates lecture videos from audio narration.
Locally learned synaptic failure enables complete Bayesian inference.
A new stochastic solver improves Convolutional Sparse Coding efficiency.
It is time-consuming and error-prone to implement inference procedures for each new probabilistic model. Probabilistic programming addresses this problem by allowing a user to specify the model and having a compiler automatically generate an inference procedure for it. For this approach to be practical, it is important…
Machine learning model predicts DFT total energy to complete basis set limit.
The power of sparse signal modeling with learned over-complete dictionaries has been demonstrated in a variety of applications and fields, from signal processing to statistical inference and machine learning. However, the statistical properties of these models, such as under-fitting or over-fitting given sets of data, …
GMBL uses graph embedding to learn binary codes from multiple views for clustering.
Slow running or straggler tasks can significantly reduce computation speed in distributed computation. Recently, coding-theory-inspired approaches have been applied to mitigate the effect of straggling, through embedding redundancy in certain linear computational steps of the optimization algorithm, thus completing the…
We give complete algorithms and source code for constructing (multilevel) statistical industry classifications, including methods for fixing the number of clusters at each level (and the number of levels). Under the hood there are clustering algorithms (e.g., k-means). However, what should we cluster? Correlations? Ret…
Paper proves method for calculating NML code length works for continuous models.
A new framework improves tensor completion accuracy by considering numerical priors.
This paper simplifies ANS for statisticians, making it easier to use.
A modular framework for knowledge distillation simplifies experiments and reproducibility.
Complete classification of Deligne-Mostow lattice representations into PGL(3,C).
A popular approach within the signal processing and machine learning communities consists in modelling signals as sparse linear combinations of atoms selected from a learned dictionary. While this paradigm has led to numerous empirical successes in various fields ranging from image to audio processing, there have only …
A popular approach within the signal processing and machine learning communities consists in modelling signals as sparse linear combinations of atoms selected from a learned dictionary. While this paradigm has led to numerous empirical successes in various fields ranging from image to audio processing, there have only …
We give a complete algorithm and source code for constructing what we refer to as heterotic risk models (for equities), which combine: i) granularity of an industry classification; ii) diagonality of the principal component factor covariance matrix for any sub-cluster of stocks; and iii) dramatic reduction of the facto…
Two novel methods improve network embedding for completely-imbalanced labels.
We propose a sparse-coding framework for activity recognition in ubiquitous and mobile computing that alleviates two fundamental problems of current supervised learning approaches. (i) It automatically derives a compact, sparse and meaningful feature representation of sensor data that does not rely on prior expert know…
High-throughput machine learning predicts thousands of diagnosis codes with high accuracy.
This paper introduces a method to find complete and interpretable concept-based explanations for deep neural networks.
LMConv improves autoregressive models for image generation and completion.
Sparse representations have proven their efficiency in solving a wide class of inverse problems encountered in signal and image processing. Conversely, enforcing the information to be spread uniformly over representation coefficients exhibits relevant properties in various applications such as digital communications. A…
Hashing techniques have been applied broadly in retrieval tasks due to their low storage requirements and high speed of processing. Many hashing methods based on a single view have been extensively studied for information retrieval. However, the representation capacity of a single view is insufficient and some discrimi…
AI generates a sequence of death causes from hospital records.
The main result is the construction of ergodic transversal measures of full support on the space of all k-surfaces of a compact hyperbolic 3-manifold. This space is a laminated space, each of its leaf being identified with a "complete" k-surface, i.e. a surface of constant (extrinsic) curvature k, where k belongs to ]0…
We analyze MDL for binary classification, quantifying overfitting and underfitting.
Estimate collapsibility of causal effects in CPDAGs via strong d-convex hulls.
Flat plumbing basket surfaces of links were introduced to study the geometry of the complement of the links. These flat plumbing basket surface can be presented by a sequential presentation known as flat plumbing basket code first found by Furihata, Hirasawa and Kobayashi. The minimum number of flat plumbings to obtain…
Method infers depth from sparse points and camera motion.