tf_geometric simplifies graph deep learning in TensorFlow.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
TF-GNN simplifies graph neural networks in TensorFlow.
One of the fundamental tasks in understanding genomics is the problem of predicting Transcription Factor Binding Sites (TFBSs). With more than hundreds of Transcription Factors (TFs) as labels, genomic-sequence based TFBS prediction is a challenging multi-label classification task. There are two major biological mechan…
Time-frequency (TF) representations provide powerful and intuitive features for the analysis of time series such as audio. But still, generative modeling of audio in the TF domain is a subtle matter. Consequently, neural audio synthesis widely relies on directly modeling the waveform and previous attempts at unconditio…
A new term weighting scheme TF-IDFC-RF outperforms others in sentiment analysis.
We study on a new kind of surface covered by translation and factorable (TF-type) surfaces in the three dimensional Euclidean space. We consider I and III Laplace-Beltrami operator surfaces of a TF-type surface. Then we obtain degrees and classes of algebraic surfaces of the surfaces using eliminate methods on software…
We describe TF-Replicator, a framework for distributed machine learning designed for DeepMind researchers and implemented as an abstraction over TensorFlow. TF-Replicator simplifies writing data-parallel and model-parallel research code. The same models can be effortlessly deployed to different cluster architectures (i…
Keyword extraction has received an increasing attention as an important research topic which can lead to have advancements in diverse applications such as document context categorization, text indexing and document classification. In this paper we propose STF-IDF, a novel semantic method based on TF-IDF, for scoring wo…
We consider the two logarithmic strain measures\[ω_{\rm iso}=\|\mathrm{dev}_n\log U\|=\|\mathrm{dev}_n\log \sqrt{F^TF}\|\quad\text{ and }\quad ω_{\rm vol}=|\mathrm{tr}(\log U)|=|\mathrm{tr}(\log\sqrt{F^TF})|\,,\]which are isotropic invariants of the Hencky strain tensor , and show that they can be uniquely char…
To survive environmental conditions, cells transcribe their response activities into encoded mRNA sequences in order to produce certain amounts of protein concentrations. The external conditions are mapped into the cell through the activation of special proteins called transcription factors (TFs). Due to the difficult …
Orthogonal matching pursuit (OMP) is a widely used compressive sensing (CS) algorithm for recovering sparse signals in noisy linear regression models. The performance of OMP depends on its stopping criteria (SC). SC for OMP discussed in literature typically assumes knowledge of either the sparsity of the signal to be e…
TF-Coder simplifies tensor manipulation programming in TensorFlow.
A training-free message passing module improves hypergraph neural networks.
Paper explores applying TDA to text classification, improving model performance.
We consider a portfolio allocation problem for trend following (TF) strategies on multiple correlated assets. Under simplifying assumptions of a Gaussian market and linear TF strategies, we derive analytical formulas for the mean and variance of the portfolio return. We construct then the optimal portfolio that maximiz…
TF-MoDISco (Transcription Factor Motif Discovery from Importance Scores) is an algorithm for identifying motifs from basepair-level importance scores computed on genomic sequence data. This technical note focuses on version v0.5.6.5. The implementation is available at https://github.com/kundajelab/tfmodisco/tree/v0.5.6…
TF Boosted Trees (TFBT) is a new open-sourced frame-work for the distributed training of gradient boosted trees. It is based on TensorFlow, and its distinguishing features include a novel architecture, automatic loss differentiation, layer-by-layer boosting that results in smaller ensembles and faster prediction, princ…
Paper solves convertible bond valuation using finite elements with penalty method.
The paper shows non-aspherical path components in G2-moduli spaces.
PCGS-TF uses a Transformer to adaptively control expert switching in non-stationary environments.
Improves document summarization by combining word embeddings and n-grams.
Orthogonal matching pursuit (OMP) and orthogonal least squares (OLS) are widely used for sparse signal reconstruction in under-determined linear regression problems. The performance of these compressed sensing (CS) algorithms depends crucially on the \textit{a priori} knowledge of either the sparsity of the signal ($k_…
A Pathology report is arguably one of the most important documents in medicine containing interpretive information about the visual findings from the patient's biopsy sample. Each pathology report has a retention period of up to 20 years after the treatment of a patient. Cancer registries process and encode high volume…
We investigate the problem of wealth distribution from the viewpoint of asset exchange. Robust nature of Pareto's law across economies, ideologies and nations suggests that this could be an outcome of trading strategies. However, the simple asset exchange models fail to reproduce this feature. A yardsale(YS) model in w…
Develops probabilistic models for gene regulatory network inference.
In this paper we study the ergodic theory of the geodesic flow on negatively curved geometrically finite manifolds. We prove that the measure theoretic entropy is upper semicontinuous when there is no loss of mass. In case we are losing mass, the critical exponents of parabolic subgroups of the fundamental group have a…
While ubiquitous, textual sources of information such as company reports, social media posts, etc. are hardly included in prediction algorithms for time series, despite the relevant information they may contain. In this work, openly accessible daily weather reports from France and the United-Kingdom are leveraged to pr…
Paper introduces ZIPTF and C-ZIPTF for better tensor factorization of zero-inflated count data.
Differential privacy(DP) has now become a standard in case of sensitive statistical data analysis. The two main approaches in DP is local and central. Both the approaches have a clear gap in terms of data storing,amount of data to be analyzed, analysis, speed etc. Local wins on the speed. We have tested the state of th…
A multi-task model tackles citation purpose classification with limited data.
While there has been much recent progress using deep learning techniques to separate speech and music audio signals, these systems typically require large collections of isolated sources during the training process. When extending audio source separation algorithms to more general domains such as environmental monitori…
This study analyzes app reviews to understand students' behavior in the app market.
X-DC improves speech separation by making DNNs more interpretable.
In this paper, we compare various methods to compress a text using a neural model. We find that extracting tokens as latent variables significantly outperforms the state-of-the-art discrete latent variable models such as VQ-VAE. Furthermore, we compare various extractive compression schemes. There are two best-performi…
Anchors explain text model decisions by highlighting key words.
In previous papers (arxiv:math/0612370 and arxiv:0909.1342) we defined the C*-algebra and the longitudinal pseudodifferential calculus of any singular foliation (M,F). Here we construct the analytic index of an elliptic operator as a KK-theory element, and prove that the same element can be obtained from an "adiabatic …
In text mining, information retrieval, and machine learning, text documents are commonly represented through variants of sparse Bag of Words (sBoW) vectors (e.g. TF-IDF). Although simple and intuitive, sBoW style representations suffer from their inherent over-sparsity and fail to capture word-level synonymy and polyse…
While deep learning has shown tremendous success in a wide range of domains, it remains a grand challenge to incorporate physical principles in a systematic manner to the design, training, and inference of such models. In this paper, we aim to predict turbulent flow by learning its highly nonlinear dynamics from spatio…
Anchors explains text classifiers by highlighting key words.
Study detects fake news in Brazilian Portuguese using machine learning.
BERT outperforms traditional machine learning in text classification tasks.
E-commerce websites such as Amazon, Alibaba, Flipkart, and Walmart sell billions of products. Machine learning (ML) algorithms involving products are often used to improve the customer experience and increase revenue, e.g., product similarity, recommendation, and price estimation. The products are required to be repres…
In spite of the amazing results obtained by deep learning in many applications, a real intelligent behavior of an agent acting in a complex environment is likely to require some kind of higher-level symbolic inference. Therefore, there is a clear need for the definition of a general and tight integration between low-le…
Study the limit of Calabi-Yau metrics with degenerate skeletons.
This paper describes multichannel speech enhancement for improving automatic speech recognition (ASR) in noisy environments. Recently, the minimum variance distortionless response (MVDR) beamforming has widely been used because it works well if the steering vector of speech and the spatial covariance matrix (SCM) of no…
Punctuated Equilibrium (PE) states that after long periods of evolutionary quiescence, species evolution can take place in short time intervals, where sudden differentiation makes new species emerge and some species extinct. In this paper, we introduce and study the effect of punctuated equilibrium on two different ass…
Sharp estimates for heat flow on nonconvex domains.
OLinear forecasts time series more efficiently by transforming data orthogonally.