Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

22456789 · Jun 202019922001200920182026
48 results for Compressive Streaming

Method decomposes streaming data into sparse and low-rank components from compressive measurements.

problem Online decomposing compressive streaming data efficiently.
method Solves nn-1\ell_1 cluster-weighted minimization to decompose sparse and low-rank components.
result Outperforms existing methods for numerical and video data.

Paper introduces a streaming compression method for monitoring pedestrian events on footbridges.

problem Storage and analysis of high-rate sensor data from instrumented infrastructure is computationally challenging.
method Develops a streaming feature-based compression method to preserve key patterns and features of pedestrian events.
result Demonstrates the trade-off between compression and accuracy during and between pedestrian events.

Balancing graph summarization and change detection in streaming data.

problem Balancing compression rate in graph summarization and accuracy in change detection.
method Introducing a probabilistic hierarchical latent variable model and optimizing parameters based on the minimum description length principle to balance the trade-off.
result Guaranteed suppression of Type I error probability (false alarms) in change detection.

Adaptive Quantization Modules enable online continual compression of non-i.i.d data streams.

problem Learning to compress and store a dataset from a non-i.i.d data stream, only observing each sample once.
method Discrete auto-encoders and Adaptive Quantization Modules (AQM) to control compression ability.
result Significant gains on continual learning benchmarks with AQM replacing episodic memory.

Sublinear memory sketch finds nearest neighbors in streaming data.

problem Finding nearest neighbors in large datasets with limited memory.
method Combines LSH, online kernel density estimation, and compressed sensing to achieve sublinear memory.
result Achieves sublinear memory performance on stable queries, reporting nearest neighbors efficiently.

Sketches linear classifiers using Weight-Median Sketch for efficient data stream analysis.

problem Efficiently learning and analyzing data streams with limited memory.
method Introduces Weight-Median Sketch for compressed linear classifier learning over data streams.
result Memory-limited execution of various analyses over streams, including feature selection and mutual information estimation.

We introduce a recursive algorithm for performing compressed sensing on streaming data. The approach consists of a) recursive encoding, where we sample the input stream via overlapping windowing and make use of the previous measurement in obtaining the next one, and b) recursive decoding, where the signal estimate from…

2013-12-17abs ↗pdf ↗

Unified SVD compression fails in practical tasks, highlighting the importance of per layer activation reconstruction.

problem The failure of a unified SVD compression method in practical tasks like perplexity and accuracy.
method Unified optimization problem for SVD based compression methods, focusing on cross-layer coupling.
result Downstream metrics like perplexity and accuracy degrade severely compared to standard per layer SVD LLM.

Signed compression progress on a sealed audit is goodhart-resistant.

problem Intrinsic motivation for agents to improve their world models by compressing experience.
method Rewarding agents for the signed decrease of a fixed sealed-audit loss.
result Cumulative reward telescopes exactly to endpoint audit improvement, preventing infinite reward push while true audit performance stagnates.

Bayesian nonparametric CMS improves frequency estimation for power-law data.

problem Estimating frequencies of low-frequency tokens in power-law data streams.
method Developed a learning-augmented count-min sketch using a normalized inverse Gaussian process prior.
result The approach achieves remarkable performance in estimating low-frequency tokens.

A new data-oblivious sketch for logistic regression reduces data size while maintaining approximation accuracy.

problem Efficiently solving logistic regression in one pass over a data stream.
method Data-oblivious sketching approach that reduces data size to poly(μdlog n) weighted points.
result Sketching reduces data size significantly and provides approximation guarantees.

Compressed Counting (CC) [22] was recently proposed for estimating the ath frequency moments of data streams, where 0 < a <= 2. CC can be used for estimating Shannon entropy, which can be approximated by certain functions of the ath frequency moments as a -> 1. Monitoring Shannon entropy for anomaly detection (e.g., DD…

2012-05-09abs ↗pdf ↗

Faster and accurate JPEG2000 image classification without reconstruction.

problem Efficiently classify j2k-compressed images without reconstructing them.
method Train a deep CNN using DWT coefficients directly from j2k-compressed images, using different augmentation techniques.
result Achieved faster and more accurate classification of j2k images without additional computation.

ADL uses a self-constructing network to handle dynamic data streams.

problem Catastrophic forgetting in deep neural networks.
method ADL employs a flexible network structure with drift detection and pruning mechanisms.
result ADL consistently outperforms other continual learning methods in dynamic environments.

Neural codec for high-fidelity audio compression.

problem Efficiently compress audio while maintaining high quality.
method End-to-end neural network architecture with quantized latent space, single multiscale spectrogram adversary, loss balancer mechanism, and lightweight Transformer compression.
result 40% compression with no loss in quality, faster than real-time.

This work proposes ACTC for adaptive distributed learning under communication constraints.

problem Adaptive distributed learning in networks with communication constraints.
method ACTC (Adapt-Compress-Then-Combine) strategy with diffusion exchange of compressed updates.
result ACTC iterates converge to the optimizer with significant bit savings.

Develops an online Gaussian process method that maintains convergence guarantees without sample complexity issues.

problem The computational intractability of Gaussian processes with streaming data.
method Parsimonious Online Gaussian Processes (POG) that maintains asymptotic consistency with bounded memory.
result POG preserves convergence guarantees to the population posterior with finite memory, even for constant error radius.

CADNN optimizes DNN execution on smartphones for real-time inference.

problem Executing Deep Neural Networks on mobile devices with low latency and high accuracy.
method Advanced model compression and architecture-aware optimization.
result CADNN outperforms state-of-the-art frameworks in DNN execution on mobile devices.

We consider the problem of selecting non-zero entries of a matrix AA in order to produce a sparse sketch of it, BB, that minimizes AB2\|A-B\|_2. For large m×nm \times n matrices, such that nmn \gg m (for example, representing nn observations over mm attributes) we give sampling distributions that exhibit four importa…

2013-11-19abs ↗pdf ↗

DiffSketch combines privacy and communication efficiency in distributed learning.

problem Privacy and communication efficiency in distributed machine learning.
method DiffSketch uses Count Sketch for data stream summarization to achieve both privacy and efficiency.
result DiffSketch provides strong differential privacy guarantees and significant communication compression.

Paper proposes an integrated M&D approach for large multistream data.

problem Inability to progress in monitoring and diagnostics due to high-dimensionality and volume of multistream data.
method Adaptive Principal Component monitoring (APC) and Principal Component Signal Recovery (PCSR).
result The integrated M&D approach enables early detection and streamlined SPC.

Learning parameters from voluminous data can be prohibitive in terms of memory and computational requirements. We propose a "compressive learning" framework where we estimate model parameters from a sketch of the training data. This sketch is a collection of generalized moments of the underlying probability distributio…

2016-06-09abs ↗pdf ↗

A new algorithm predicts periodic time series data efficiently in cloud environments.

problem Efficiently identifying and predicting periodic patterns in large-scale time-series data.
method Proposes a Periodicity-based Parallel Time Series Prediction (PPTSP) algorithm using TSDCA, MTSPPR, and PTSP methods.
result Significant improvements in prediction accuracy and performance compared to existing algorithms.

Sketched SGD reduces communication in distributed SGD by sketching gradients.

problem Limited communication in large-scale distributed training of neural networks.
method Introducing Sketched SGD, an algorithm that communicates sketches of gradients instead of full gradients.
result Sketched SGD reduces communication from O(d)\mathcal{O}(d) or O(W)\mathcal{O}(W) to O(logd)\mathcal{O}(\log d), achieving up to 40x reduction in total communication cost.

A new method classifies multiple correlated data streams simultaneously.

problem Classifying multiple correlated data streams in practical scenarios.
method Double-Coupling Support Vector Machines (DC-SVM) considers both internal and external correlations.
result The proposed method outperforms traditional methods on artificial and real-world data streams.

stream-learn is a Python library for analyzing data streams with various drift types.

problem Analyzing drifting and imbalanced data streams.
method Synthetic data stream generator, evaluation methodologies, and imbalanced binary classification metrics.
result Efficient implementation of classifiers for data stream analysis.

PySAD offers a unified Python framework for efficient streaming anomaly detection.

problem Efficient anomaly detection in streaming data with strict constraints.
method Unified architecture with 17+ streaming algorithms, specialized components, and support for multiple learning paradigms.
result PySAD enables real-time processing with bounded memory and is compatible with other Python frameworks.

Context improves one-class classifiers in dynamic data streams.

problem Improving one-class classification in data streams with limited training data.
method Proposes using context to guide one-class classifier learning in data streams, presenting three frameworks.
result The use of context can improve the performance of streaming one-class classifiers.

Paper improves submodular streaming with better approximation, less memory, and lower complexity.

problem Maximizing submodular functions in streaming with a cardinality constraint.
method Sieve-Streaming++ with one pass, O(k)O(k) memory, and (1/2)(1/2)-approximation; adaptive complexity reduction.
result Achieves (1/2)(1/2)-approximation with O(k)O(k) memory and low adaptive complexity.