Neural networks combining multiple data sources can reverse preferences, affecting decision reliability.
problem Preference reversals in neural networks under pooled data.
method Formalized through Case-Based Decision Theory, analyzed Gram geometry, introduced regularization, and developed auditing methods.
result Pooled refitting can reverse shared preferences, and conditions for preserving preferences are derived.
Paper proposes a method to recover point configurations from noisy distance data.
problem Recovering point configurations from noisy distance data.
method Robust Euclidean Distance Geometry via Dual Basis (RoDEoDB) algorithm.
result Exact recovery guarantees for point configuration and Gram matrix under mild conditions.
Spectral graph sparsification preserves geometry of GNN embeddings.
problem Maintaining geometric properties of graph neural network embeddings during sparsification.
method Proving spectral sparsification preserves squared pairwise distances, class means, and covariance structure in embedding space.
result Spectral sparsification preserves the geometry of learned embeddings in GNNs.
To any compact Riemann surface of genus g one may assign a principally polarized abelian variety of dimension g, the Jacobian of the Riemann surface. The Jacobian is a complex torus, and a Gram matrix of the lattice of a Jacobian is called a period Gram matrix. This paper provides upper and lower bounds for all the ent…
We present NN-grams, a novel, hybrid language model integrating n-grams and neural networks (NN) for speech recognition. The model takes as input both word histories as well as n-gram counts. Thus, it combines the memorization capacity and scalability of an n-gram model with the generalization ability of neural network…
Generalized Gram determinant for 3-manifold invariants.
problem Invariants of 3-manifolds and bilinear forms evaluation.
method Evaluation of a bilinear form in the annulus for non-intersecting connections in the disc.
result Closed formula for the generalized Gram determinant.
Eliciting semantic similarity between concepts in the biomedical domain remains a challenging task. Recent approaches founded on embedding vectors have gained in popularity as they risen to efficiently capture semantic relationships The underlying idea is that two words that have close meaning gather similar contexts. …
We investigate the Gram determinant of the bilinear form based on curves in a planar surface, with a focus on the disk with two holes. We prove that the determinant based on n−1 curves divides the determinant based on n curves. Motivated by the work on Gram determinants based on curves in a disk and curves in an an…
Study Gram determinants in knot theory, focusing on a Möbius band determinant.
problem Closed formula for the Gram determinant of type (Mb)1. method Survey of Gram determinants, focusing on a Möbius band determinant.
result Speculation on closed formula for (Mb)1 Gram determinant. KiloGrams finds top-k large n-grams for malware classification.
problem Lack of efficient methods for large n-grams in malware classification.
method Developed a fast method for finding top-k large n-grams.
result Large n-grams improve malware classification and provide interpretable features.
The paper connects Chebyshev polynomials and Gram determinants on Möbius bands.
problem Exploring the relationship between Chebyshev polynomials and Gram determinants on Möbius bands.
method Analyzing Mersenne numbers and Chebyshev polynomials, proving conjectures, and developing algorithms.
result A factor of the Gram determinant supports a conjecture about its closed formula involving Chebyshev polynomials.
Any-gram kernels are a flexible and efficient way to employ bag-of-n-gram features when learning from textual data. They are also compatible with the use of word embeddings so that word similarities can be accounted for. While the original any-gram kernels are implemented on top of tree kernels, we propose a new approa…
Effective Gram matrix predicts deep network generalization.
problem Understanding and predicting deep network generalization.
method Derived a differential equation governing generalization gap, analyzed with effective Gram matrix.
result Effective Gram matrix accurately predicts test loss during training.
New Gram determinant from Möbius band connects to annulus case.
problem Exploring new Gram determinants in knot theory.
method Skein theoretic approach and bilinear forms.
result Proves important results about new Gram determinant structure.
Deep learning methods exhibit promising performance for predictive modeling in healthcare, but two important challenges remain: -Data insufficiency:Often in healthcare predictive modeling, the sample size is insufficient for deep learning methods to achieve satisfactory results. -Interpretation:The representations lear…
Simple bounds for covariance and Gram matrices across various settings.
problem Capturing the behavior of smaller eigenvalues in covariance and Gram matrices.
method General-purpose theorem converting uniform bounds into relative bounds.
result Sharper control of eigenvalues across the spectrum.
Proposes a faster Transformer decoding method by truncating target-side self-attention windows.
problem Efficiency in Transformer decoding with minimal BLEU score loss.
method N-gram assumption to truncate target-side self-attention windows.
result N-gram masked self-attention model maintains BLEU score for N values from 4 to 8. Paper proposes detecting OOD examples using Gram matrices and in-distribution data.
problem Detecting OOD examples with confidence and without OOD data.
method Characterize activity patterns with Gram matrices and identify anomalies in values.
result High OOD detection rates achieved without OOD data.
Corrected CBOW performs similarly to Skip-gram.
problem CBOW embeddings underperform Skip-gram embeddings in word2vec.
method Fixed a bug in CBOW gradient update to improve performance.
result Corrected CBOW embeddings are competitive with Skip-gram on various tasks.
New method corrects missing data bias in dimension reduction.
problem Missing data complicates high-dimensional data analysis.
method Developed a bias-corrected Gram matrix for heterogeneous missingness.
result Proposed method improves dimension reduction techniques significantly.
Deep kernel processes unify various models using Gram matrices and kernel functions.
problem Unified representation of various deep learning models.
method Defining deep kernel processes with progressively transformed Gram matrices and sampling from inverse Wishart distributions.
result Deep Gaussian processes, BNNs, infinite BNNs, and infinite BNNs with bottlenecks can all be written as deep kernel processes.
Layer normalization with activations prevents Gram matrix rank collapse at initialization.
problem Rank collapse in Gram matrices at initialization slows training in deep networks.
method Proved that layer normalization, with activation layers, biases Gram matrix towards identity matrix at exponential rate.
result Layer normalization with activations biases Gram matrix towards identity matrix at exponential rate with depth at initialization.
We use the Jones-Wenzl idempotents to construct a basis of Temperley-Lieb algebra TL_n. This allows a short calculation for a Gram determinant of Lickorish's bilinear form on the Temperley-Lieb algebra.
New metric learning approach for tree data reduces computation cost.
problem Efficiently computing distances between ordered labeled trees.
method Introduced pq-grams and a differentiable weighted pq-gram distance, combined with LMNN for optimization.
result Significantly reduces computation time for tree classification problems.
In this paper we show that the matrix of chromatic joins and the Gram matrix of the Temperley-Lieb algebra are similar (after rescaling), with the change of basis given by diagonal matrices.
We show that the skip-gram formulation of word2vec trained with negative sampling is equivalent to a weighted logistic PCA. This connection allows us to better understand the objective, compare it to other word embedding methods, and extend it to higher dimensional models.
Here we present a novel approach to statistical analysis of financial time series. The approach is based on n-grams frequency dictionaries derived from the quantized market data. Such dictionaries are studied by evaluating their information capacity using relative entropy. A specific quantization of (originally conti…
Generates low-dimensional node vectors for graphs with privacy while preserving structural preferences.
problem Publishing graph node vectors can leak sensitive individual information.
method SE-PrivGEmb, a skip-gram based technique with a unified noise tolerance mechanism and negative sampling probabilities.
result Our method outperforms existing methods in structural equivalence and link prediction tasks.
Transformers learn rich in-context dependencies efficiently.
problem Understanding how transformers learn long-range dependencies efficiently.
method Approximation and dynamics analysis of induction head mechanisms.
result Abrupt transition from lazy to rich mechanisms during training.
A clustering algorithm uses the left Gram matrix for high dimensional data.
problem Clustering high dimensional data with many features and few objects.
method The algorithm uses the normalized left Gram matrix G = XX'/P to cluster objects based on row means.
result The algorithm provides the most accurate cluster configuration more than twice as often as competitors.
Study one-dimensional topological theories with linear generating functions.
problem Understanding one-dimensional topological theories with defects.
method Construct bases of hom spaces for decorated unoriented one-dimensional cobordisms.
result Gram determinant and linear generating functions constructed.
Recurrent neural network (RNN) language models (LMs) and Long Short Term Memory (LSTM) LMs, a variant of RNN LMs, have been shown to outperform traditional N-gram LMs on speech recognition tasks. However, these models are computationally more expensive than N-gram LMs for decoding, and thus, challenging to integrate in…
Linearized attention fails to converge to NTK limit even at large widths.
problem Understanding the convergence of attention mechanisms to the kernel regime.
method Analyzes linearized attention and its relationship to the NTK limit, considering practical widths and conditions.
result Linearized attention does not converge to its NTK limit at any practical width, revealing a fundamental trade-off.
This paper extends the convergence rate of DEQs with ReLU to any general activation.
problem Proving global convergence rate for DEQs with general activations.
method Developed a novel population Gram matrix and new form of dual activation with Hermite polynomial expansion.
result Gradient descent converges to a globally optimal solution at a linear rate for DEQs with general activations.
Improved neural model predicts gender from tweets.
problem Predicting gender from Twitter text.
method RNN model with attention, LSA-reduced n-gram features.
result Improved model achieves state-of-the-art performance on English tweets.
A new method learns dynamic graph representations from time-varying data.
problem Learning dynamic graph representations from time-varying data.
method Higher-order skip-gram with negative sampling (HOSGNS) for tensor factorization.
result HOSGNS outperforms state-of-the-art methods in downstream tasks.
Proposes SNML for selecting word2vec Skip-gram dimensionality.
problem Selecting optimal dimensionality for word2vec Skip-gram models.
method Information criteria (AIC, BIC, SNML) applied to SG and SG Negative Sampling models.
result SNML outperforms AIC and BIC, selecting closer optimal dimensionality.
New method models portfolios with leptokurtic risk factors using Gram-Charlier expansions.
problem Modeling portfolios with excess kurtosis.
method GC-like expansions of the hyperbolic-secant law to account for leptokurtosis.
result Portfolio distribution with risk factors modeled as GC-like expansions of the HS law.
We consider the space M of ordered quadruples of distinct points in the boundary of complex hyperbolic n-space, chn, up to its holomorphic isometry group PU(n,1). One of the important problems in complex hyperbolic geometry is to construct and describe a moduli space for M. For $n=2…
GRAM enhances deep RL for reliable real-world deployment.
problem Generalizing deep RL across in-distribution and out-of-distribution scenarios.
method Introduces a robust adaptation module and a joint training pipeline.
result GRAM achieves strong generalization performance in simulations and hardware.
Develops a universal Hermitian projective calculus for complex hyperbolic two-space
problem Complex hyperbolic geometry
method Algebraic invariant calculus
result Denominator-cleared identities for various geometric quantities
This work shows dimension regularization can replace skip-gram negative sampling for graph embeddings, improving efficiency and performance.
problem Efficiently enforcing dissimilarity among node embeddings in graph learning.
method Dimension regularization as an alternative to skip-gram negative sampling.
result Dimension regularization is a more efficient approach to enforcing dissimilarity in graph embeddings.
Deep generative models can learn to generate realistic-looking images, but many of the most effective methods are adversarial and involve a saddlepoint optimization, which requires a careful balancing of training between a generator network and a critic network. Maximum mean discrepancy networks (MMD-nets) avoid this i…
A transformer model improves spell correction with hierarchical attention.
problem Improving spell correction accuracy and speed.
method Multi encoder-single decoder transformer architecture with hierarchical attention.
result Significant improvement in CER, WER, and SER error rates.
A new score measures data reliability without ground truth.
problem Assessing reliability of datasets without access to ground truth.
method Define ground-truth-based orderings and propose Gram determinant score.
result Gram determinant score effectively captures data quality across diverse observation processes.
This research improves graph embeddings by optimizing node sampling with centrality weights.
problem Improving the accuracy and efficiency of graph embeddings using Skip-Gram methods.
method Implemented and analyzed four graph embedding techniques with different centrality-weighted sampling distributions.
result Centrality-weighted sampling leads to improved accuracy and faster learning times.
The dominant language models (LMs) such as n-gram and neural network (NN) models represent sentence probabilities in terms of conditionals. In contrast, a new trans-dimensional random field (TRF) LM has been recently introduced to show superior performances, where the whole sentence is modeled as a random field. In thi…
In this paper, we solve a problem posed by Rodica Simion regarding type B Gram determinants. We present this in a fashion influenced by the work of W.B.R.Lickorish on Witten-Reshetikhin-Turaev invariants of 3-manifolds. The roots of the determinant were predicted by Dabkowski and Przytycki, and the complete factorizati…