We prove that the binary classifiers of bit strings generated by random wide deep neural networks with ReLU activation function are biased towards simple functions. The simplicity is captured by the following two properties. For any given input bit string, the average Hamming distance of the closest input bit string wi…
Two methods for model adaptation compared; fine-tuning outperforms Best-of-N in realizable settings.
problem Comparing methods for adapting large language models to new tasks.
method Supervised fine-tuning vs. Best-of-N approach.
result Supervised fine-tuning outperforms Best-of-N in realizable settings.
A model is developed to study the effectiveness of innovation and its impact on structure creation and structure change on agent-based societies. The abstract model that is developed is easily adapted to any particular field. In any interacting environment, the agents receive something from the environment (the other a…
Differentially private data structures for estimating distances between strings.
problem Estimating distances between query strings and database strings while ensuring privacy.
method Proposes differentially private data structures for Hamming and edit distances using randomized response technique.
result Efficient data structures that provide accurate distance estimates with strong privacy guarantees.
We study the dynamics of co-evolution of producers and customers described by bit-strings representing individual traits. Individual ''size-like'' properties are controlled by binary encounters which outcome depends upon a recognition process. Depending upon the parameter set-up, mutual selection of producers and custo…
Recently, randomly mapping vectorial data to strings of discrete symbols (i.e., sketches) for fast and space-efficient similarity searches has become popular. Such random mapping is called similarity-preserving hashing and approximates a similarity metric by using the Hamming distance. Although many efficient similarit…
The bits-back argument suggests that latent variable models can be turned into lossless compression schemes. Translating the bits-back argument into efficient and practical lossless compression schemes for general latent variable models, however, is still an open problem. Bits-Back with Asymmetric Numeral Systems (BB-A…
New techniques improve 16-bit training accuracy without 32-bit units.
problem Training deep learning models with only 16-bit floating-point units.
method Studied BFloat16 units and applied stochastic rounding and Kahan summation techniques.
result Up to 7% absolute validation accuracy gain in 16-bit-FPU training.
Graph Weighted Models (GWMs) have recently been proposed as a natural generalization of weighted automata over strings and trees to arbitrary families of labeled graphs (and hypergraphs). A GWM generically associates a labeled graph with a tensor network and computes a value by successive contractions directed by its e…
Paper improves DNN accelerator robustness against bit errors with energy savings.
problem Bit errors in quantized DNN weights reduce energy efficiency.
method Combines robust fixed-point quantization, weight clipping, and random bit error training.
result Significantly improves robustness against random bit errors with high energy savings.
Paper offers robust recovery for 1-bit sensing with partial Gaussian circulant matrices.
problem Accurately recovering vectors from 1-bit measurements using structured matrices.
method Correlation-based optimization with randomly signed partial Gaussian circulant matrices and generative models.
result Recovery guarantees match those for i.i.d. Gaussian matrices but with faster computation.
Identifies all perturbative vacua in bosonic string theory.
problem Identifying all perturbative vacua in bosonic string theory.
method Completely identified perturbative vacua through string fluctuations.
result Derivation of path-integrals up to any order from fluctuations.
Majority bit estimation in noisy random recursive DAGs.
problem Estimating the majority bit in a noisy random recursive DAG.
method Majority rule among nodes, with bit flipping and noisy channel.
result Identification of the threshold for p at which majority rule yields errors. The state-of-the-art hardware platforms for training Deep Neural Networks (DNNs) are moving from traditional single precision (32-bit) computations towards 16 bits of precision -- in large part due to the high energy efficiency and smaller bit storage associated with using reduced-precision representations. However, un…
Generalizes bits back coding for time-series models with latent Markov structures.
problem Efficiently compressing time-series data with latent Markov structures.
method Extends bits back coding to time-series models with latent Markov structures, including HMMs and LGSSMs.
result Effective for small scale models, promising for larger scale settings like video compression.
ADD embeds a 48-bit message into images, achieving high accuracy and speed.
problem Embedding high-fidelity messages into images to detect authenticity and source.
method Two-stage process: linear combination and addition of watermark to image, followed by decoding.
result ADD achieves 100% decoding accuracy for 48-bit watermarking, with minimal performance drop under various distortions.
Topological data analysis classifies encrypted bits with success.
problem Classifying encrypted data with traditional machine learning methods.
method Persistent homology for generating topological features, machine learning pipeline.
result Successfully classifies encrypted data, outperforming classical models.
Bayesian Bits unifies quantization and pruning through gradient optimization.
problem Joint mixed precision quantization and pruning for efficient neural networks.
method Gradient-based optimization with a novel bit width decomposition and learnable stochastic gates.
result Bayesian Bits achieves better accuracy vs. efficiency trade-off compared to static bit width networks.
New protocols show 1-bit mean estimation can be order-optimal without interaction.
problem Can 1-bit mean estimation be optimal without interaction?
method Adaptive and non-adaptive threshold and interval queries, with one adaptive transition.
result Arbitrary non-adaptive quantizers can match the adaptive rate, suggesting interaction is not necessary.
One-bit quantization improves inference speed for Random Features models.
problem Efficient inference on resource-constrained devices.
method Analysis of one-bit quantization in Random Features model.
result Asymptotically, quantizing weights except the last incurs no loss in generalization error.
In recent years a lot of attention has been paid to topological spaces which are a bit more general than smooth manifolds - orbifolds. Orbifolds are intuitively speaking manifolds with some singularities. The formal definition is also modelled on that of manifolds, an orbifold is a topological space which locally is ho…
One-bit feedback suffices for a bandit problem's optimal strategy.
problem Optimal strategy for multi-armed bandit problem with limited feedback.
method Coding and decoding schemes for one-bit feedback to mimic full-reward feedback.
result Regret ratio approaches 1 with one-bit feedback.
Low-bit training framework reduces energy consumption in CNNs.
problem Reducing energy consumption in convolutional neural networks.
method Low-bit training framework using MLS tensor format with dynamic quantization.
result Achieves superior trade-off between accuracy and bit-width.
The goal of standard 1-bit compressive sensing is to accurately recover an unknown sparse vector from binary-valued measurements, each indicating the sign of a linear function of the vector. Motivated by recent advances in compressive sensing with generative models, where a generative modeling assumption replaces the u…
Extends Goldberg's result for string links over surfaces.
problem Generalized string links over surfaces.
method Extension of Goldberg's result.
result Exact sequence for generalized string links over surfaces.
A virtual n-string α is a collection of n oriented smooth generic loops on a surface M. A stabilization of α is a surgery that results in attaching a handle to M along disks avoiding α, and the inverse operation is a destabilization of α. We consider virtual n-strings up to virtual homotopy, i.e., seq…
Cobordism of virtual string links on n strands is a combinatorial generalization of link cobordism. There exists a bijection between virtual string links up to cobordisms and elements of the group Zn(n−1). This paper also shows that virtual string links up to unwelded equivalence are classified by those…
String topology coproduct and Turaev cobracket computed for surfaces.
problem Computing string topology coproduct and cobracket for surfaces.
method Algorithm to compute coproduct of cyclic words in terms of generators of the fundamental group.
result String cobracket is the negative of Turaev cobracket.
Numbers and numerical vectors account for a large portion of data. However, recently the amount of string data generated has increased dramatically. Consequently, classifying string data is a common problem in many fields. The most widely used approach to this problem is to convert strings into numerical vectors using …
Bit threads provide an alternative description of holographic entanglement, replacing the Ryu-Takayanagi minimal surface with bulk curves connecting pairs of boundary points. We use bit threads to prove the monogamy of mutual information (MMI) property of holographic entanglement entropies. This is accomplished using t…
Developing tools for computing string amplitudes with hyperbolic vertices.
problem Computing off-shell string amplitudes with new vertices.
method Constructing local coordinates and investigating limits for hyperbolic three-string vertex.
result Derived conservation laws and performed sample computations.
Defines a generalized string concept for abstract root systems.
problem Generalizing the concept of strings to abstract root systems.
method Introduces a new definition for Φ-strings in abstract root systems. result Defines a new set of elements in Σ based on a given λ and subset Φ of simple roots. New algorithm tackles batched stochastic linear bandits with 1-bit communication constraints.
problem Stochastic linear bandits with 1-bit communication constraints.
method Phased-elimination algorithms based on G-optimal designs and 1-bit mean estimation.
result Achieves near-optimal regret bounds for broad scaling regimes.
Paper develops an efficient mean estimator for 1-bit communication constraints.
problem Mean estimation under 1-bit communication constraints.
method Adaptive mean estimator based on randomized threshold queries.
result Order-optimal sample complexity in various tail regimes.
Reduced precision computation for deep neural networks is one of the key areas addressing the widening compute gap driven by an exponential growth in model size. In recent years, deep learning training has largely migrated to 16-bit precision, with significant gains in performance and energy efficiency. However, attemp…
Paper studies distributed learning with limited communication bits, achieving optimal error exponents.
problem Distributed hypothesis testing with constant communication bits.
method Geometric approach in distribution spaces, encoding empirical distributions to transmission bits.
result Optimal achievable error exponents and coding schemes for various communication constraints.
Enhances hashing for fast retrieval with correlated bits.
problem Fast retrieval and small memory footprint for large-scale information retrieval.
method Employing Boltzmann machine distribution as variational posterior to model correlations among hash code bits.
result Significant performance gains achieved by effectively modeling correlations among hash code bits.
Derives path integrals for perturbative strings on various backgrounds.
problem Calculating path integrals for strings on curved backgrounds.
method Derives path integrals from string geometry theory by considering fluctuations around string backgrounds.
result Derives path integrals of all order perturbative strings on various backgrounds.
Paper proposes a CNN-based method for estimating intra frame bits and quality.
problem Efficient video delivery and bit allocation in video coding.
method Deep learning approach using CNNs trained on original frames and encoded distortions.
result Accurate estimation of intra frame bits and quality for better bit allocation.
Develops Palatini formalism in generalized geometry for string theory.
problem Formulating Palatini variation in generalized geometry.
method Palatini formalism within generalized Riemannian geometry of Courant algebroids.
result Natural emergence of generalized Levi-Civita connection and string effective actions.
String geometry theory uniquely determines classical action with T-symmetry.
problem Non-renormalizability and loop corrections in string theory.
method Distinguishes effects of β and ħ parameters, proving no loop corrections.
result No loop corrections in string geometry theory, avoiding non-renormalizability.
Paper proposes a 1-bit quantization scheme for high-dimensional statistical estimation.
problem High-dimensional statistical estimation with limited data.
method Uniformly dithered 1-bit quantization for sparse covariance matrix estimation, sparse linear regression, and matrix completion.
result Near minimax rates in sub-Gaussian regime and improved rates in heavy-tailed regime.
Emerging resistive random-access memory (ReRAM) has recently been intensively investigated to accelerate the processing of deep neural networks (DNNs). Due to the in-situ computation capability, analog ReRAM crossbars yield significant throughput improvement and energy reduction compared to traditional digital methods.…
Improves matrix multiplication throughput for asymmetric bit-width operands.
problem Matrix multiplications between asymmetric bit-width operands, especially 8- and 4-bit, are not efficiently handled by existing SIMD instructions.
method Proposes a new SIMD matrix multiplication instruction that uses mixed precision on inputs (8- and 4-bit) and accumulates into 16-bit output, improving throughput.
result Offers 2x improvement in throughput compared to existing symmetric-operand-size instructions, with negligible overflow.
A new method for 1-bit matrix completion that is faster and more accurate.
problem Estimating a low-rank matrix from binary observations.
method Majorization-Minimization Gauss-Newton (MMGN) method.
result MMGN outperforms existing methods in accuracy and speed.
String backgrounds yield simplified Hull-Strominger system solutions.
problem Solving the simplified Hull-Strominger system in various geometries.
method Variational argument using string action, gradient Ricci solitons, and symmetry reduction.
result Canonical symmetry and transverse geometry properties derived.
BEGIN network models binary data without parametric assumptions.
problem Conditional independence in non-parametric families of binary data.
method BEGIN network models binary data using sparse linear representations and block factorizations.
result BEGIN network captures conditional independence for arbitrary binary and multinomial variables.
A virtual n-string is a chord diagram with n core circles and a collection of arrows between core circles. We consider virtual n-strings up to virtual homotopy, compositions of flat virtual Reidemeister moves on chord diagrams. Given a virtual 1-string α, Turaev associated a based matrix that encodes invariants…