tf_geometric simplifies graph deep learning in TensorFlow.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
TF-GNN simplifies graph neural networks in TensorFlow.
One of the fundamental tasks in understanding genomics is the problem of predicting Transcription Factor Binding Sites (TFBSs). With more than hundreds of Transcription Factors (TFs) as labels, genomic-sequence based TFBS prediction is a challenging multi-label classification task. There are two major biological mechan…
Time-frequency (TF) representations provide powerful and intuitive features for the analysis of time series such as audio. But still, generative modeling of audio in the TF domain is a subtle matter. Consequently, neural audio synthesis widely relies on directly modeling the waveform and previous attempts at unconditio…
We describe TF-Replicator, a framework for distributed machine learning designed for DeepMind researchers and implemented as an abstraction over TensorFlow. TF-Replicator simplifies writing data-parallel and model-parallel research code. The same models can be effortlessly deployed to different cluster architectures (i…
Improved TF-IDF for word relevance in health-care social media documents.
A new term weighting scheme TF-IDFC-RF outperforms others in sentiment analysis.
To survive environmental conditions, cells transcribe their response activities into encoded mRNA sequences in order to produce certain amounts of protein concentrations. The external conditions are mapped into the cell through the activation of special proteins called transcription factors (TFs). Due to the difficult …
Orthogonal matching pursuit (OMP) is a widely used compressive sensing (CS) algorithm for recovering sparse signals in noisy linear regression models. The performance of OMP depends on its stopping criteria (SC). SC for OMP discussed in literature typically assumes knowledge of either the sparsity of the signal to be e…
TF-Coder simplifies tensor manipulation programming in TensorFlow.
We study on a new kind of surface covered by translation and factorable (TF-type) surfaces in the three dimensional Euclidean space. We consider I and III Laplace-Beltrami operator surfaces of a TF-type surface. Then we obtain degrees and classes of algebraic surfaces of the surfaces using eliminate methods on software…
A training-free message passing module improves hypergraph neural networks.
Unified approach to trend-following systems, deriving exact relationships and expected returns.
ARA combines aggregated RAPPOR and Tf-Idf estimation for centralized DP analysis.
We consider a portfolio allocation problem for trend following (TF) strategies on multiple correlated assets. Under simplifying assumptions of a Gaussian market and linear TF strategies, we derive analytical formulas for the mean and variance of the portfolio return. We construct then the optimal portfolio that maximiz…
Paper solves convertible bond valuation using finite elements with penalty method.
TF-MoDISco (Transcription Factor Motif Discovery from Importance Scores) is an algorithm for identifying motifs from basepair-level importance scores computed on genomic sequence data. This technical note focuses on version v0.5.6.5. The implementation is available at https://github.com/kundajelab/tfmodisco/tree/v0.5.6…
We investigate the problem of wealth distribution from the viewpoint of asset exchange. Robust nature of Pareto's law across economies, ideologies and nations suggests that this could be an outcome of trading strategies. However, the simple asset exchange models fail to reproduce this feature. A yardsale(YS) model in w…
TF Boosted Trees (TFBT) is a new open-sourced frame-work for the distributed training of gradient boosted trees. It is based on TensorFlow, and its distinguishing features include a novel architecture, automatic loss differentiation, layer-by-layer boosting that results in smaller ensembles and faster prediction, princ…
Develops probabilistic models for gene regulatory network inference.
Improves document summarization by combining word embeddings and n-grams.
Orthogonal matching pursuit (OMP) and orthogonal least squares (OLS) are widely used for sparse signal reconstruction in under-determined linear regression problems. The performance of these compressed sensing (CS) algorithms depends crucially on the \textit{a priori} knowledge of either the sparsity of the signal ($k_…
The paper shows non-aspherical path components in G2-moduli spaces.
PCGS-TF uses a Transformer to adaptively control expert switching in non-stationary environments.
Paper explores applying TDA to text classification, improving model performance.
Paper introduces ZIPTF and C-ZIPTF for better tensor factorization of zero-inflated count data.
A Pathology report is arguably one of the most important documents in medicine containing interpretive information about the visual findings from the patient's biopsy sample. Each pathology report has a retention period of up to 20 years after the treatment of a patient. Cancer registries process and encode high volume…
A multi-task model tackles citation purpose classification with limited data.
Anchors explain text model decisions by highlighting key words.
While deep learning has shown tremendous success in a wide range of domains, it remains a grand challenge to incorporate physical principles in a systematic manner to the design, training, and inference of such models. In this paper, we aim to predict turbulent flow by learning its highly nonlinear dynamics from spatio…
Anchors explains text classifiers by highlighting key words.
In this paper, we compare various methods to compress a text using a neural model. We find that extracting tokens as latent variables significantly outperforms the state-of-the-art discrete latent variable models such as VQ-VAE. Furthermore, we compare various extractive compression schemes. There are two best-performi…
X-DC improves speech separation by making DNNs more interpretable.
While there has been much recent progress using deep learning techniques to separate speech and music audio signals, these systems typically require large collections of isolated sources during the training process. When extending audio source separation algorithms to more general domains such as environmental monitori…
This study analyzes app reviews to understand students' behavior in the app market.
BERT outperforms traditional machine learning in text classification tasks.
In text mining, information retrieval, and machine learning, text documents are commonly represented through variants of sparse Bag of Words (sBoW) vectors (e.g. TF-IDF). Although simple and intuitive, sBoW style representations suffer from their inherent over-sparsity and fail to capture word-level synonymy and polyse…
In spite of the amazing results obtained by deep learning in many applications, a real intelligent behavior of an agent acting in a complex environment is likely to require some kind of higher-level symbolic inference. Therefore, there is a clear need for the definition of a general and tight integration between low-le…
Punctuated Equilibrium (PE) states that after long periods of evolutionary quiescence, species evolution can take place in short time intervals, where sudden differentiation makes new species emerge and some species extinct. In this paper, we introduce and study the effect of punctuated equilibrium on two different ass…
We consider the two logarithmic strain measures\[ω_{\rm iso}=\|\mathrm{dev}_n\log U\|=\|\mathrm{dev}_n\log \sqrt{F^TF}\|\quad\text{ and }\quad ω_{\rm vol}=|\mathrm{tr}(\log U)|=|\mathrm{tr}(\log\sqrt{F^TF})|\,,\]which are isotropic invariants of the Hencky strain tensor , and show that they can be uniquely char…
Influence functions help study large language model generalization, revealing surprising decay patterns.
In previous papers (arxiv:math/0612370 and arxiv:0909.1342) we defined the C*-algebra and the longitudinal pseudodifferential calculus of any singular foliation (M,F). Here we construct the analytic index of an elliptic operator as a KK-theory element, and prove that the same element can be obtained from an "adiabatic …
OLinear forecasts time series more efficiently by transforming data orthogonally.
Study detects fake news in Brazilian Portuguese using machine learning.
E-commerce websites such as Amazon, Alibaba, Flipkart, and Walmart sell billions of products. Machine learning (ML) algorithms involving products are often used to improve the customer experience and increase revenue, e.g., product similarity, recommendation, and price estimation. The products are required to be repres…
Study the limit of Calabi-Yau metrics with degenerate skeletons.
This paper describes multichannel speech enhancement for improving automatic speech recognition (ASR) in noisy environments. Recently, the minimum variance distortionless response (MVDR) beamforming has widely been used because it works well if the steering vector of speech and the spatial covariance matrix (SCM) of no…
The task of determining item similarity is a crucial one in a recommender system. This constitutes the base upon which the recommender system will work to determine which items are more likely to be enjoyed by a user, resulting in more user engagement. In this paper we tackle the problem of determining song similarity …