Deep neural network learns hierarchical language family structure.
problem Language family dependency and external information affect classification accuracy.
method Hierarchical attentive units learn auxiliary tasks for robust internal representation.
result Improved classification accuracy on small and big language corpora.
Word embeddings are a powerful approach for unsupervised analysis of language. Recently, Rudolph et al. (2016) developed exponential family embeddings, which cast word embeddings in a probabilistic framework. Here, we develop dynamic embeddings, building on exponential family embeddings to capture how the meanings of w…
Groups with specific curvature have a regular language of geodesics.
problem Understanding the language of geodesics in non-positively curved triangle groups.
method Proving finitely many cone types and regularity of geodesic languages.
result The language of lexicographically first geodesics is regular and satisfies the fellow traveller property.
Triangulation filters spurious circuits in multilingual models.
problem Unreliable explanations of multilingual models across languages.
method Formalizes reference families and introduces triangulation as a causal acceptance rule.
result Triangulation provides a falsifiable standard for mechanistic claims.
Predicts language model performance from public models without training.
problem Understanding how language model performance varies with scale.
method Observational approach building scaling laws from publicly available models.
result Smooth, sigmoidal behavior of emergent phenomena is predictable from small models.
Improved dropout training by choosing deterministic models from a family of conditional models.
problem Improving regularization-heavy language models.
method Understanding dropout as MAP estimation for a family of conditional models and selecting the best deterministic model.
result Deterministic dropout is the best approximation to the true objective, leading to substantial improvements.
Sloth predicts LLM performance using latent skills across families.
problem Variations in benchmark performance due to differences in training configurations and data processing across model families.
method Sloth uses publicly available benchmark data and assumes LLM performance is driven by latent skills influenced by model size and training tokens. It exploits correlations across benchmarks to provide accurate predictions.
result Sloth predicts LLM performance accurately and offers insights into scaling behaviors for complex tasks.
Families of objects appear in several contexts, like algebraic topology, theory of deformations, theoretical physics, etc. An unified coordinate-free algebraic framework for families of geometrical quantities is presented here, which allows one to work without introducing ad hoc spaces, by using the language of differe…
SCC identifies 21 programming languages with 75% accuracy.
problem Classifying code snippets among 21 programming languages.
method Multinomial Naive Bayes classifier trained on Stack Overflow posts.
result SCC achieves 75% accuracy, significantly higher than PLI.
The study examines how language models learn to represent the world, identifying conditions for ecological veridicality.
problem Understanding when language models learn to represent the world accurately and how this learning process can fail.
method Analyzes the Bayes-optimal next-token cross-entropy decomposition and the role of training ecology in shaping model representations.
result The minimum-complexity zero-excess solution is the quotient partition by training equivalence, and this solution is not preserved in in-context learning or per-task adaptation.
Paper learns GWMs for picture graphs using gradient methods.
problem Learning GWMs for picture graphs.
method Gradient-based methods for regression and classification.
result Gradient-based methods can learn GWMs for picture graphs.
Study uses three sources to evaluate language models fairly.
problem Bias in offline model evaluation due to confounded model choice.
method Combines observational logs, randomized experiments, and simulators.
result Randomized experiment and simulator together recover causal model values.
Neural network framework for language recognition considers sequence information and improves accuracy.
problem Challenging task of automatic language identification in noisy conditions.
method Proposes a neural network framework with bidirectional LSTM and attention modeling for relevance weighting.
result Significant improvements over conventional methods in noisy conditions and multi-speaker speech.
New models avoid lookahead bias by training on past data only.
problem Lookahead bias in language models.
method Chronologically consistent training on data before a knowledge-cutoff date.
result Elimination of lookahead bias in predictions.
Fine-tuning improves information conveyance in language models by reorganizing uncertainty into more informative sequences.
problem Uncertainty reduction in large language models through fine-tuning is not fully understood, especially regarding output length.
method Proposed Canopy Entropy ( C E ⋆ \mathrm{CE}^\star CE ⋆ ) to measure uncertainty in both output length and sequence, capturing total Shannon entropy. result Fine-tuned models exhibit stronger positive correlation between entropy rate and semantic diversity, indicating more informative and semantically meaningful generations.
New language model shows context length impacts generation quality and reasoning ability.
problem Analyzing the impact of context length and reasoning on autoregressive generation.
method Introduced synthetic hierarchical languages, used an exact k-gram ansatz, derived asymptotic predictions, and validated empirically.
result Reasoning models with limited context can generate sequences from the true language, improving exponentially over standard models.
EFA extends self-attention to handle mixed data types and dynamic relevance.
problem Handling high-dimensional, mixed data types with dynamic relevance.
method Probabilistic generative model using self-attention and latent factor model.
result EFA consistently outperforms existing models in complex latent structure capture and reconstruction.
DatedGPT prevents lookahead bias in financial forecasting models.
problem Lookahead bias in large language models trained on internet-scale data.
method Time-aware pretraining with annual data cutoffs and instruction fine-tuning.
result Models' knowledge is effectively bounded by their data cutoff year, improving forecasting validity.
We explore Jaeger's state model for the HOMFLYPT polynomial. We reformulate this model in the language of Gauss diagrams and use it to obtain Gauss diagram formulas for a two-parameter family of Vassiliev invariants coming from the HOMFLYPT polynomial. These formulas are new already for invariants of degree 3.
The study compares feed-forward and attention layers in language models.
problem Understanding the role of feed-forward and attention layers in language models.
method Empirical and theoretical analysis in a synthetic setting.
result Feed-forward layers learn simple distributional associations, while attention layers focus on in-context reasoning.
The study provides a generalization bound for a family of implicit networks.
problem Theoretical understanding of implicit networks' generalization is limited.
method A generalization bound is derived for a family of implicit networks using a covering number argument for Rademacher complexity.
result A theoretical generalization bound is established for implicit networks.
Transformers can learn new tasks from diverse pretraining data but struggle with out-of-domain tasks.
problem Transformer models' ability to learn new tasks in-context is limited by their pretraining data coverage.
method Investigation of transformer models trained on ( x , f ( x ) ) (x, f(x)) ( x , f ( x )) pairs, comparing in-context learning capabilities across different task families. result Transformers can identify and learn within task families in their pretraining data but fail with out-of-domain tasks.
SLED improves factuality in LLMs without external knowledge.
problem Unreliable or factually incorrect outputs from large language models.
method Contrasts final layer logits with early layers' logits, uses approximate gradient to refine outputs.
result Consistently improves factual accuracy over existing methods.
Study uses NLP to analyze emotions and challenges of young people with IDD.
problem Challenges faced by young people with IDD during transition to adulthood.
method Natural language processing, unsupervised machine learning, topic modeling.
result NLP methods can assist psychologists in analyzing emotions and summarizing key topics.
We introduce a method for using deep neural networks to amortize the cost of inference in models from the family induced by universal probabilistic programming languages, establishing a framework that combines the strengths of probabilistic programming and deep learning methods. We call what we do "compilation of infer…
TROLL improves RL for LLMs by replacing clipping with a trust region projection.
problem Clipping in RL for LLMs causes instability and suboptimal performance.
method TROLL uses a discrete differentiable trust region projection to replace clipping, balancing computational cost and effectiveness.
result TROLL consistently outperforms PPO-like clipping in training speed, stability, and final success rates.
Paper standardizes data licensing for AI and ML.
problem Unclear and ambiguous data licensing language in AI and ML.
method Developed a new family of data license language (Montreal Data License) and a web-based tool to generate it.
result Clearer tools and concepts for data use in AI and ML markets.
Unsupervised method improves word vectors by suppressing high variance features.
problem Improving semantic information in word vectors.
method Using conceptors to suppress high variance features in word vectors.
result Post-processed word vectors outperform existing alternatives in lexical evaluation tasks.
KodeXv0.1 improves financial question answering over GPT-4.
problem Lack of specialized financial language models.
method Custom training on financial documents, RAG-aware 4bit LoRA tuning.
result KodeX-8Bv0.1 outperforms GPT-4 by up to 9.24% in financial tasks.
Framework synthesizes programs for simulating complex models and estimating parameters.
problem Parameter estimation for complex models requires manual encoding of fixed model structures.
method Combines LLMs for program synthesis with neural simulation-based inference.
result Identifies plausible model families from open-ended prompts with high accuracy.
Researchers calculate complexity of billiard paths in regular polygons.
problem Calculating the complexity of billiard paths in regular polygons.
method Counting saddle connections on lattice surfaces, focusing on combinatorial length.
result They answered a question about billiard language complexity in regular polygons.
PLUS pre-trains protein sequences with structural info, improving performance.
problem Lack of labeled protein sequences for training models.
method PLUS combines masked language modeling with same-family prediction for pre-training.
result PLUS-RNN outperforms other models in protein biology tasks.
LeDeepChef learns to play multiple cooking games well.
problem Designing a general RL agent for multiple games of the same family.
method Actor-critic framework, action-space pruning, hierarchical RL, specialized module.
result LeDeepChef outperformed competitors on a diverse set of cooking games.
Self-improvement refines language models by verifying their own outputs.
problem Improving language models without external feedback.
method Formalizing self-improvement as sharpening, using the model itself as a verifier.
result RLHF-based self-improvement can outperform SFT-based methods.
Framework ensures alignment between humans and machines in LLMs.
problem Human-machine misalignment in LLMs scoring mechanisms.
method Lightweight calibration framework for blackbox models.
result Provably guarantees alignment between humans and machines.
TFB simplifies Bayesian LLM uncertainty estimation without extra training.
problem Estimating uncertainty in LLM responses remains challenging.
method Training-Free Bayesianization (TFB) that transforms low-rank adapters into Bayesian ones without additional training.
result TFB achieves superior uncertainty estimation and generalization compared to existing methods.
Language models allocate information storage, not collapsing into uniform representations.
problem Incomplete neural collapse in language model representations.
method Analyzing variance and information sharing across 14 models, proving an information floor.
result Within-class variance is allocated information storage, not collapsed into uniform representations.
Adaptive batch size schedules improve language model training efficiency and generalization.
problem Dilemma of choosing batch sizes in large-scale model training.
method General-purpose adaptive batch size schedules compatible with data and model parallelism.
result Adaptive batch size schedules outperform constant batch sizes and heuristic warmup schedules.
GSI improves efficiency of large language model inference.
problem Efficiently guiding test-time alignment in large language models.
method Combines soft best-of- n n n scaling with a reward model and speculative samples. result Achieves higher accuracy and reduced latency compared to standard methods.
The paper proposes methods to control errors in language generation models using textual entailment.
problem The lack of a correctness metric hinders applying principled methods to language generation tasks.
method The paper leverages textual entailment to evaluate correctness and proposes two selective generation algorithms: SGen^Sup and SGen^Semi.
result The proposed algorithms control the false discovery rate with respect to textual entailment and achieve comparable selection efficiency to baselines.
Spectral measurements reveal hidden representation geometry in language model training.
problem Hidden internal representation in language model training is hard to examine.
method Empirical protocol using activation covariance and per-sample gradient SVD spectra.
result Batch size affects representation geometry, and activation spectra predict token efficiency.
A single algebraic identity unifies information-theoretic variational results.
problem Deriving and generalizing classical information-theoretic variational results
method Proving a single algebraic mixed coincidence identity
result Unified derivation of classical cornerstones of information theory
New minimal surfaces show stacking disorder in periodic structures.
problem Reproducing experimental twinning defects in periodic minimal surfaces.
method Constructing non-periodic minimal surfaces that lift to disordered stacking in 3D.
result Reproduced twinning defects in periodic minimal surfaces as stacking disorder.
Language models fail to execute simple steps, showing gating and binding errors.
problem Procedural hallucinations in language models, failing to execute simple steps.
method Analyzed long-context binding tasks, identifying gating and binding errors.
result Procedural errors are due to gating and binding failures, with recency bias contributing to the latter.
This paper is a survey of some of the most elementary consequences of the JSJ-decomposition and geometrization for knot and link complements in the 3-sphere. Formulated in the language of graphs, the result is the construction of a bijective correspondence between the isotopy classes of links in S 3 S^3 S 3 and a class of ve…
New RLHF approach mitigates bias in aligning LLMs with human preferences.
problem Algorithmic bias in RLHF leading to preference collapse.
method Preference Matching (PM) RLHF, using PM regularizer and conditional variant.
result 29% to 41% improvement in alignment with human preferences.
New method optimizes language model performance for test-time strategies.
problem Mismatch between training objectives and test-time deployment of large language models.
method Tail-Extrapolated estimators to approximate best-of-N performance from limited training rollouts.
result Improved performance of best-of-N deployment across various models and datasets.
Transformers can generalize to a large task family with only a few demonstrations.
problem Can learning from a small set of tasks generalize to a large task family?
method Investigating autoregressive compositional structure where each task is a composition of T T T operations, each from a finite family of D D D subtasks. result Transformers can generalize to D T D^T D T tasks with only O ~ ( D ) \widetilde{O}(D) O ( D ) demonstrations.