Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

82163245326 · Jun 202019922001200920172026
48 results for somatic variant calling

Deep Bayesian neural networks improve somatic variant calling accuracy.

problem Improving accuracy in pinpointing somatic variants from next-gen sequencing data.
method Deep Bayesian Recurrent Neural Networks (RNNs) for somatic variant calling.
result Deep Bayesian RNNs provide more reliable confidence intervals for variant calls.

Deep neural network improves cancer mutation calls with confidence.

problem Improving accuracy and confidence in somatic variant calls from cancer sequencing.
method Deep Bayesian Recurrent Neural Network (RNN) with flexible priors.
result Enhanced confidence in mutation calls without performance degradation.

Modeling correlated mutations in cancer for personalized treatment.

problem Identifying mutations for personalized cancer therapy in heterogeneous profiles.
method Proposed correlated zero-inflated negative binomial process with mixed beta-Bernoulli and variational inference.
result Identified biologically relevant correlations between somatic mutations.

Paper tackles cancer mutation data challenges by creating useful low-dimensional representations.

problem Challenges in analyzing and using cancer mutation data for classification and clustering.
method Flatsomatic: variational autoencoders (VAEs) to create latent representations of somatic profiles.
result VAE embeddings perform better than PCA for clustering and equally well for classification.

Flatsomatic compresses cancer mutation data with VAEs, maintaining predictive power.

problem Compressing somatic mutation profiles in cancer while preserving predictive power.
method Flatsomatic uses a Variational Auto Encoder (VAE) with MLP architecture, optimizing evidence lower bound and beta-VAE for latent space regularization.
result Flatsomatic embeddings maintain predictive power of original data, reducing dimensionality from 8,298 to 64.

In this paper, we introduce and study various kinds of decomposition complexity. First, we give a characterization of residually finite groups having finite decomposition complexity (FDC). Secondly, we introduce equi-variant straight FDC (sFDC), and prove that a group having equi-variant sFDC if and only if its box spa…

2015-09-29abs ↗pdf ↗

Stochastic gradient descent~(SGD) and its variants have become more and more popular in machine learning due to their efficiency and effectiveness. To handle large-scale problems, researchers have recently proposed several parallel SGD methods for multicore systems. However, existing parallel SGD methods cannot achieve…

2015-08-24abs ↗pdf ↗

We present a novel method for extracting cancer signatures by applying statistical risk models (http://ssrn.com/abstract=2732453) from quantitative finance to cancer genome data. Using 1389 whole genome sequenced samples from 14 cancers, we identify an "overall" mode of somatic mutational noise. We give a prescription …

2016-04-29abs ↗pdf ↗

Novel CE-method variants reduce local minima convergence with fewer function evaluations.

problem Local minima and expensive function evaluations in optimization.
method Surrogate model-based CE-method variants to reduce local minima convergence.
result Surrogate model-based approach reduces local minima convergence using fewer function evaluations.

Proposes a neural framework to select subsets efficiently across different models.

problem Lack of generalizability in subset selection methods for unseen architectures.
method Introduces a trainable subset selection framework, SubSelNet, that uses attention-based neural gadgets and subset samplers.
result SubSelNet generalizes across architectures and outperforms existing methods.

Paper presents a new method to detect process differences at the trace level using mutual fingerprints.

problem Low-level granularity in process variant analysis leads to many false differences.
method Develops a mutual fingerprint technique to encode entire process traces for comparison.
result Mutual fingerprint method reveals significant differences not detected by existing techniques.

Many contemporary statistical learning methods assume a Euclidean feature space. This paper presents a method for defining similarity based on hyperspherical geometry and shows that it often improves the performance of support vector machine compared to other competing similarity measures. Specifically, the idea of usi…

2017-02-05abs ↗pdf ↗

Unified framework for efficient Frank-Wolfe optimization of Dominant Set Clustering.

problem Optimizing Dominant Set Clustering with various Frank-Wolfe algorithms.
method Unified framework for pairwise, standard, and away-steps Frank-Wolfe algorithms, with explicit convergence rates.
result Explicit convergence rates for Frank-Wolfe methods in Dominant Set Clustering.

The Johnson filtration of the mapping class group of a compact, oriented surface is the descending series consisting of the kernels of the actions on the nilpotent quotients of the fundamental group of the surface. Each term of the Johnson filtration admits a Johnson homomorphism, whose kernel is the next term in the f…

2017-07-24abs ↗pdf ↗

We introduce a variant of the kk-nearest neighbor classifier in which kk is chosen adaptively for each query, rather than supplied as a parameter. The choice of kk depends on properties of each neighborhood, and therefore may significantly vary between different points. (For example, the algorithm will use larger $k…

2019-05-29abs ↗pdf ↗

A new gradient quantization scheme improves communication efficiency in distributed training.

problem Efficiently compressing gradients for parallel training of large models.
method Proposes a new gradient quantization scheme with theoretical guarantees and empirical performance.
result The new scheme matches and exceeds the performance of existing methods.

In this paper, we briefly review the basic scheme of the pseudoinverse learning (PIL) algorithm and present some discussions on the PIL, as well as its variants. The PIL algorithm, first presented in 1995, is a non-gradient descent and non-iterative learning algorithm for multi-layer neural networks and has several adv…

2018-05-20abs ↗pdf ↗

The goal of this study is to explain and examine the statistical underpinnings of the Bollinger Band methodology. We start off by elucidating the rolling regression time series model and deriving its explicit relationship to Bollinger Bands. Next we illustrate the use of Bollinger Bands in pairs trading and prove the e…

2012-12-20abs ↗pdf ↗

Traditional Recurrent Neural Networks assume vectorized data as inputs. However many data from modern science and technology come in certain structures such as tensorial time series data. To apply the recurrent neural networks for this type of data, a vectorisation process is necessary, while such a vectorisation leads…

2017-08-01abs ↗pdf ↗

A new bandit problem where experiments can be interrupted if results are not promising.

problem Interruptible multi-armed bandit problem with a threshold for cumulative reward.
method Formalized survival regret, identified key components (regret and probability of ruin), derived lower bounds and optimal policies.
result No policy can achieve sublinear survival regret, but optimal policies minimize survival regret in a Pareto sense.

Stochastic particle-optimization sampling (SPOS) is a recently-developed scalable Bayesian sampling framework that unifies stochastic gradient MCMC (SG-MCMC) and Stein variational gradient descent (SVGD) algorithms based on Wasserstein gradient flows. With a rigorous non-asymptotic convergence theory developed recently…

2018-11-20abs ↗pdf ↗

We introduce a groupoid ${\mathbf{ΠMG}}}$, called the fundamental modular groupoid, which is a variant of Penner's mapping class groupoid. We study how it relates to the surface mapping class groups and Thompson's group T\mathsf T. We also introduce larger groupoid ΩMG\mathbf{ΩMG}, which is related to outer automorphis…

2018-07-23abs ↗pdf ↗

New geometric variant of factorization homology for conformally flat manifolds.

problem Defining invariants of conformally flat manifolds.
method Introducing a metric-dependent geometric variant of factorization homology.
result Left Kan extensions of conformally flat dd-disk algebras define invariants of conformally flat manifolds.

A new NMF variant tackles underdetermined problems with sparse and separable assumptions.

problem Underdetermined blind source separation, especially multispectral image unmixing.
method Sparse Separable Nonnegative Matrix Factorization (SSNMF) combining separability and sparsity assumptions. Algorithm based on SNPA and sparse nonnegative least squares.
result In noiseless settings, the algorithm recovers true underlying sources.

This paper reviews recent advances in Bayesian nonparametric techniques for constructing and performing inference in infinite hidden Markov models. We focus on variants of Bayesian nonparametric hidden Markov models that enhance a posteriori state-persistence in particular. This paper also introduces a new Bayesian non…

2014-06-30abs ↗pdf ↗

In unsupervised outlier ensembles, the absence of ground truth makes the combination of base outlier detectors a challenging task. Specifically, existing parallel outlier ensembles lack a reliable way of selecting competent base detectors, affecting accuracy and stability, during model combination. In this paper, we pr…

2018-12-04abs ↗pdf ↗

Analyzing the temporal behavior of nodes in time-varying graphs is useful for many applications such as targeted advertising, community evolution and outlier detection. In this paper, we present a novel approach, STWalk, for learning trajectory representations of nodes in temporal graphs. The proposed framework makes u…

2017-11-11abs ↗pdf ↗