Binary embeddings speed up graph data retrieval.
problem Efficiently retrieving graphical data.
method Binary valued embeddings modeled as coin flips with varying bias, optimized using continuous optimization techniques.
result Binary embeddings outperform other methods on various datasets.
A fast binary embedding method preserves Euclidean distances in high-dimensional data.
problem Preserving Euclidean distances in high-dimensional datasets.
method Stable noise-shaping quantization of Ax with A a sparse Gaussian random matrix, followed by a linear transformation. result Euclidean distances are approximated by the ℓ1 norm on binary sequences, leading to accurate binary codes. G-Net constructs binary neural networks with high accuracy using randomized binary embeddings.
problem Creating high-accuracy binary neural networks with theoretical guarantees.
method Proposes a novel floating-point G-Net family with randomized binary embeddings and theoretical accuracy guarantees.
result Empirically, G-Net achieves almost 30% higher accuracy on CIFAR-10 compared to prior HDC models.
New methods predict links in hypergraphs with multiple entities.
problem Link prediction in knowledge hypergraphs with non-binary relations.
method Introduce HSimplE and HypE embedding-based methods for hypergraphs.
result Proposed methods outperform baselines in hypergraph prediction.
This paper improves binary embeddings and quantized compressed sensing methods.
problem Distance-preserving binary embeddings and quantization for compressed sensing.
method Quantization of fast Johnson-Lindenstrauss embeddings and bounded orthonormal systems.
result Quantization methods yield reconstruction errors that decay polynomially and exponentially in the number of measurements.
Develops a binarized GNN for more efficient graph embeddings.
problem Real-valued GNN parameters limit efficiency and scalability.
method Integrates binarization into GNN-based graph embedding approaches.
result Binarized graph neural network (BGN) achieves state-of-the-art performance with significant efficiency gains.
Unified methodology for evaluating neural embeddings in link prediction tasks.
problem Evaluating the quality of neural embeddings for link prediction in knowledge graphs.
method Comparing, combining, and extending different methodologies for link prediction on graph-based data.
result Training neural embeddings globally for the entire graph outperforms local training.
Bayesian optimization for high-dimensional combinatorial spaces using embeddings.
problem Optimizing expensive functions over large, complex input spaces.
method Dictionary-based ordinal embeddings for high-dimensional combinatorial structures, using Gaussian process models.
result The proposed method outperforms state-of-the-art BO methods on diverse real-world benchmarks.
PoET-BiN reduces power consumption in neural networks on embedded devices.
problem Power inefficiency in neural network implementations on embedded platforms.
method Look-Up Table based implementation with a modified Decision Tree approach.
result Near state-of-the-art results with up to 6 orders of magnitude energy reduction.
GMBL uses graph embedding to learn binary codes from multiple views for clustering.
problem Lack of complete structure and complementary information from multiple views in single-view hash clustering methods.
method Graph-based Multi-view Binary Learning (GMBL) using Laplacian matrix to preserve data structure and assign weights to views.
result GMBL outperforms previous methods in clustering performance on multiple datasets.
New tests for binary classification regression functions without distribution assumptions.
problem Testing regression functions in binary classification without distributional assumptions.
method Conditional kernel mean embeddings and resampling-based framework.
result Distribution-free hypothesis tests with exact type I error control.
Binary embedding of high-dimensional data requires long codes to preserve the discriminative power of the input space. Traditional binary coding methods often suffer from very high computation and storage costs in such a scenario. To address this problem, we propose Circulant Binary Embedding (CBE) which generates bina…
NodeSig efficiently computes binary node embeddings for scalable graph analysis.
problem Scalability issues in graph representation learning models.
method NodeSig uses random walk diffusion probabilities and stable random projections to compute binary node embeddings efficiently.
result NodeSig achieves a good balance between accuracy and efficiency on node classification and link prediction tasks.
B-CP reduces knowledge graph model size by replacing real-valued embeddings with binary values.
problem Storage inefficiency in vector embeddings for large knowledge graphs.
method Binarized CANDECOMP/PARAFAC (B-CP) decomposition algorithm.
result B-CP reduces model size by more than an order of magnitude while maintaining task performance.
Paper proposes IIQ for compressing embedding vectors.
problem Memory issues in representing large vocabularies.
method Isotropic iterative quantization (IIQ) for binary compression.
result More than 30x compression ratio with comparable performance.
New method debiases word embeddings for multiclass settings like race and religion.
problem Word embeddings in online texts perpetuate human stereotypes, including race and religion.
method Proposes a novel methodology to debias word embeddings in multiclass settings.
result Demonstrates robust multiclass debiasing that maintains NLP task efficacy.
This work learns shared word embeddings for acoustic and phonetic sequences.
problem Mapping variable-length acoustic and phonetic sequences to fixed-dimensional vectors.
method Weak supervision and binary classification task to predict word similarity.
result Best model achieves an F1 score of 0.95 for binary classification.
We use mathematical induction to prove that the horizontal composition in the class of coherently diagonal complexes is indeed a binary operation. That is to say, the embedding of two coherently diagonal complexes in an alternating planar diagram produces a coherently diagonal complex.
Paper explores understanding of neural source code embeddings.
problem Lack of understanding of contents and characteristics of code2vec embeddings.
method Small case study using code2vec embeddings to create binary SVM classifiers and compare performance with handcrafted features.
result Code2vec embeddings perform similarly to handcrafted features and have more evenly distributed information gains.
Efficiently learns quantizable embeddings for fast search.
problem Learning binary hamming code representations for search efficiency.
method Directly learns a quantizable embedding representation and sparse binary hash code end-to-end.
result Achieves state-of-the-art search accuracy and significant speedup.
Cubic predicts stock market indices by fusing stock latent embeddings and converting to binary classification.
problem Challenges in predicting stock market indices due to isolated time series treatment and simple regression.
method Fusion of stock latent embeddings, binary encoding classification, and confidence-guided prediction.
result Cubic outperforms state-of-the-art baselines in stock index prediction tasks.
This paper provides an embedding perspective to consensus clustering.
problem Consensus clustering combines multiple clustering results.
method Transfer categorical partitions to binary coding, spectral embedding, etc.
result Unified two major categories of consensus clustering and connected it to graph embedding.
Improves training of binary neural networks for mobile devices.
problem Training accurate binary neural networks for mobile devices.
method Systematic evaluation of network architectures and hyperparameters.
result Increased accuracy by increasing the number of connections in the network.
Binary neural networks trained from scratch achieve state-of-the-art results.
problem Training accurate binary neural networks from scratch is challenging.
method No prior knowledge and simple training strategy used.
result Achieved state-of-the-art results on standard benchmark datasets.
Study 1-bit compressive sensing with generative models, improving recovery accuracy.
problem Accurately recover sparse vectors from binary measurements with generative models.
method Analyzes noiseless and noisy 1-bit measurements with i.i.d.~Gaussian and Lipschitz continuous generative priors, proving sample complexity bounds and stability properties.
result Proves sample complexity bounds and stability properties for 1-bit compressive sensing with generative models.
QNNs can't distinguish binary signals from their negations, revealing a new symmetry.
problem Understanding the behavior of QNNs in binary pattern classification.
method Presented and analyzed a new form of invariance (negational symmetry) in QNNs.
result QNNs cannot differentiate a quantum binary signal and its negational counterpart in binary classification tasks.
Method finds reference products for a given item.
problem Finding relevant products for a given item.
method Product representation learning and fingerprint-type vector searching.
result The method outperforms peer services in search return rate and precision.
New algorithm trains binary-activation, multi-level RNNs for noise-resilient, ADC-/DAC-free PIM inference.
problem Training noise-resilient, ADC-/DAC-free neural networks.
method Binary activations and multi-level weights for eNVM-based processing-in-memory circuits.
result Higher accuracy and noise resilience for recurrent networks compared to existing methods.
PPC learns binary codes from data similarities and dissimilarities.
problem Creating efficient binary codes from data similarities and dissimilarities.
method PPC learns binary codes by modeling attractive and repulsive forces in a signed graph.
result PPC achieves superior results in nearest-neighbor searches compared to spectral methods.
TuckER predicts missing facts in knowledge graphs using tensor decomposition.
problem Predicting missing facts in knowledge graphs.
method TuckER uses Tucker decomposition of binary tensor representations.
result TuckER outperforms state-of-the-art models in link prediction.
Fruit fly brain network learns word embeddings using sparse binary codes.
problem Learning semantic word representations from text.
method Inspired by mushroom body neural network, sparse binary hash codes.
result Fruit fly network achieves comparable NLP performance with reduced resources.
MeliusNet improves binary neural networks to match MobileNet-v1 accuracy.
problem Achieving high accuracy with binary neural networks on mobile devices.
method Alternating DenseBlocks and ImprovementBlocks to increase feature capacity and quality.
result MeliusNet matches MobileNet-v1 accuracy on ImageNet, improving binary network performance.
Simpler approach to link prediction using complex embeddings.
problem Link prediction in large knowledge bases.
method Latent factorization with complex valued embeddings.
result Arguably simpler approach outperforms state-of-the-art models.
Infinite fractal tree solves shortest connection problem.
problem Finding the shortest connection for a fractal set.
method Constructing an infinite planar self-similar binary tree.
result The tree is the unique solution to the Steiner problem.
We pose causal inference as the problem of learning to classify probability distributions. In particular, we assume access to a collection {(Si,li)}i=1n, where each Si is a sample drawn from the probability distribution of Xi×Yi, and li is a binary label indicating whether "Xi→Yi" or …
Graphs from van der Corput sequence embed into Chamanara surface.
problem Embedding graphs from van der Corput sequence into surfaces.
method Constructed 4-regular graphs from van der Corput sequence and Kronecker sequence, embedded into torus and Chamanara surface. result Graphs from van der Corput sequence embed into Chamanara surface with one edge removal.
Survey on embedding techniques in source code.
problem Applying word embedding techniques to source code.
method Collection and categorization of articles from related work and scholarly searches.
result Word embedding has been successfully applied to various granularities of source code.
Extends graph encoder embedding to weighted graphs and matrices.
problem Classifying vertices in various graph types efficiently.
method Graph encoder embedding applied to weighted graphs, distance matrices, and kernel matrices.
result The method achieves asymptotic normality, enabling optimal classification.
Simplifies multi-label classification with stochastic sketch strategy.
problem Complex training processes in multi-label classification.
method Simple stochastic sketch strategy for multi-label classification.
result Competitive performance without complex training processes.
A new activation function improves credit scoring accuracy for imbalanced datasets.
problem Imbalanced datasets in credit scoring lead to underestimation of misclassification costs.
method Introduces ASIG, an asymmetric adjusted Sigmoid function.
result ASIG-embedded classifier outperforms traditional classifiers across various imbalance ratios.
LLMs outperform strong tabular baselines on industrial car retrofit prediction.
problem Industrial retrofit planning on structured operational data
method Embedding features, direct prompted classification, and ML+LLM stacking
result LLMs outperform strong tabular baselines on industrial car retrofit prediction
Machine learning predicts ECHR judgments on human rights violations.
problem Predicting the outcome of ECHR judgments on human rights violations.
method Auto-sklearn for model selection, N-grams, word embeddings, doc2vec, echr2vec for feature extraction, cross-validation for accuracy assessment.
result Features from echr2vec embedding provided the highest cross-validation accuracy for 5 Articles, overall test accuracy was 68.83%.
The paper tackles estimating vectors from binary comparisons, providing bounds and adaptive strategies.
problem Estimating a vector from binary comparisons of preference.
method Theoretical bounds and adaptive strategies for estimating vectors from noisy and randomized comparisons.
result Stable embedding of the space of target vectors and significant gains from adaptive distribution changes.
We give a simple explicit algorithm for building multi-factor risk models. It dramatically reduces the number of or altogether eliminates the risk factors for which the factor covariance matrix needs to be computed. This is achieved via a nested "Russian-doll" embedding: the factor covariance matrix itself is modeled v…
New algorithm B++&C improves hierarchical clustering on large deep embedding datasets.
problem Scaling up hierarchical clustering to massive datasets of deep embeddings.
method Proposes B++&C algorithm for practical hierarchical clustering, introduces B2SAT&C for theoretical approximation.
result Achieves 5%/20% improvement on MW/CKMM objectives compared to classic methods.
A new method learns efficient kernel embeddings from data.
problem Efficient implementation of kernel methods for large datasets.
method Hybrid Constrained Optimization (CVEM) with ADMM and specific RFF formulations.
result Computationally efficient kernel embeddings significantly improve kernel method performance.
The distance metric plays an important role in nearest neighbor (NN) classification. Usually the Euclidean distance metric is assumed or a Mahalanobis distance metric is optimized to improve the NN performance. In this paper, we study the problem of embedding arbitrary metric spaces into a Euclidean space with the goal…
Study compares different levels of supervision for training graph embeddings in wireless networks.
problem Improving power control in wireless interference networks.
method Training graph neural networks (GNNs) with different levels of supervision (supervised, unsupervised, self-supervised).
result Different levels of supervision impact system-level throughput, convergence, and generalization.