Study compares LLMs vs classical models for financial sentiment analysis.
problem Improving sentiment analysis in financial market news.
method Comparative analysis of LLMs and classical models.
result LLMs outperform classical models in sentiment analysis of financial news.
Work shows hallucination detection by LLMs is impossible without expert feedback.
problem Detecting hallucinations in LLMs is theoretically impossible without expert-labeled feedback.
method Investigated hallucination detection using a theoretical framework inspired by language identification.
result Automated hallucination detection is impossible for most language collections without expert-labeled feedback.
This is the author's Master's thesis written under the supervision of Dr. Gregor Weingart at the National Autonomous University of Mexico. The purpose of this study is to rewrite differential supergeometry in terms of classical differential geometry. This rewriting from "first principles" has two main motivations: 1 av…
A new transformer model corrects diacritics and typos in multiple languages.
problem Restoring diacritics and correcting typos in online communications.
method Employing a universal ByT5 transformer model trained on 12 languages, including Lithuanian.
result Achieves > 98% accuracy in diacritics restoration and typos correction.
Language-based methods improve human similarity approximations without requiring many human judgments.
problem Approximating human similarity judgments using pre-trained deep neural networks (DNNs) is challenging and expensive.
method Developed language-based methods to approximate human similarity judgments, validated with adaptive tag collection pipeline.
result Language-based methods significantly improve performance over DNN-based methods with fewer human judgments.
This is a brief description of the classical part of the Standard Model of particles and interactions, using the language of vector bundles over the spacetime and operations on them.
Deep learning models outperform classical methods in text classification.
problem Improving text classification accuracy using deep learning.
method Comprehensive review of deep learning models and datasets for text classification.
result Deep learning models outperform classical methods on various text classification tasks.
Word embeddings are a powerful approach for unsupervised analysis of language. Recently, Rudolph et al. (2016) developed exponential family embeddings, which cast word embeddings in a probabilistic framework. Here, we develop dynamic embeddings, building on exponential family embeddings to capture how the meanings of w…
New framework reveals thermodynamic principles for LLM training.
problem Understanding the training dynamics of large language models.
method Introducing Neural Thermodynamic Laws (NTL) under river-valley loss landscape assumptions.
result Key thermodynamic quantities and principles naturally emerge in LLM training.
These notes grew out of a lecture course on mathematical methods of classical physics for students of mathematics and mathematical physics at the master's level. Also, physicists with a strong interest in mathematics may find this text useful as a resource complementary to existing textbooks on classical physics. Topic…
Analyzes the concept of fields in classical and quantum physics.
problem Challenges in defining fields in classical and quantum physics.
method Uses groupoid description of quantum mechanics and categorical language.
result Fields as functors among groupoids of test particles and intrinsic system nature.
We characterize language generation with stability and breadth, proving impossibility results.
problem Characterizing and proving impossibility results for language generation with stability and breadth.
method Analysis of existing notions of breadth and stability, proving lower bounds.
result Proven impossibility of generating with higher perplexity or lower hallucination rate for stable generators.
Improved speech recognition with language model integration in sequence-to-sequence models.
problem Improving word error rate in speech recognition models.
method Log-linear combination of acoustic and language models with per-token renormalization.
result The proposed method shows good improvements over standard model combination on Librispeech system.
New research shows larger language models improve data processing for diverse entries.
problem Optimizing data processing for tables with diverse string entries.
method Analytical tasks on tables with varying language model sizes and a fuzzy join benchmark.
result Larger language models improve data processing for diverse entries, but fine-tuning is necessary.
GENO framework generates efficient solvers for machine learning problems.
problem Designing efficient solvers for machine learning problems.
method GENO framework combines a modeling language with a generic solver to generate solvers from optimization problem specifications.
result Automatically generated solvers are as efficient as well-engineered specialized solvers and orders of magnitude more efficient than classical modeling language plus solver approaches.
Study on generating and identifying languages privately, showing privacy imposes costs and creates barriers.
problem Generating and identifying languages privately in the limit model.
method Introduced a continual release model under differential privacy constraints, proving both positive and negative results.
result Privacy imposes quantitative and qualitative costs, and creates fundamental barriers for identification.
Paper analyzes classical multidimensional scaling for cluster recovery.
problem Cluster recovery from noisy data.
method Classical multidimensional scaling followed by distance-based clustering.
result Scaling conditions for high probability cluster recovery.
Proposes an iterative algorithm for optimizing attention mechanisms in large language models.
problem Optimizing attention mechanisms in large language models.
method Iterative algorithm for rescaled hyperbolic functions regression.
result Efficiency and generalizability of the rescaled softmax regression framework.
A quantum model classifies financial sentiment by mapping text chunks to quantum circuits.
problem Classifying financial texts with high accuracy and preserving semantic information.
method Chunked diagrams are mapped to quantum circuits, with a Transformer encoder and type embeddings added for context.
result The hybrid model improves sentiment classification over a simple averaging baseline.
Study language generation with limited memory, showing different impacts on achievable densities and convergence.
problem Language generation with bounded memory constraints.
method Analyzed memoryless generators, sliding windows, and adaptive past examples; revisited identification in the limit.
result Achievable densities and convergence properties differ based on the size of the target language collection.
The study detects deceptive language in business communication using AI.
problem Deceptive language in business communication.
method Combining classical rhetoric, communication psychology, and linguistic theory with computational textual analysis and transformer models.
result Detection accuracies of over 99% achieved in controlled settings.
Bayesian active learning improves natural language processing models.
problem Lack of model comparison in AL for NLP tasks.
method Large-scale empirical study of Bayesian active learning with Dropout and Bayes-by-Backprop uncertainty estimates.
result Bayesian active learning by disagreement significantly improves NLP model performance.
New method evaluates LLMs fairness in universal prediction.
problem Evaluating fairness of large language models in universal prediction.
method Introducing batch regret as a modification of average regret for LLMs.
result Asymptotical value of batch regret for add-constant predictors on memoryless and first-order Markov sources.
Neural model extracts tokens as latent variables for text compression.
problem Compressing text using neural models.
method Extracting tokens with highest tf-idf scores or highest loss from a bidirectional language model.
result Extracting tokens as latent variables significantly outperforms state-of-the-art methods.
Enhanced Tacotron for Japanese speech synthesis improves naturalness.
problem Challenges in end-to-end Japanese speech synthesis due to pitch accents.
method Extended Tacotron with self-attention to capture pitch accent dependencies.
result Proposed systems show improvements but still lag behind traditional pipeline methods.
UQE uses LLMs to analyze unstructured data efficiently.
problem Efficient analytics on unstructured data.
method Proposes UQE, a query engine that uses LLMs to interpret UQL queries.
result Demonstrates efficient analytics on various unstructured data types.
The present paper is a review of the current state of Graph-Link Theory (graph-links are also closely related to homotopy classes of looped interlacement graphs), dealing with a generalisation of knots obtained by translating the Reidemeister moves for links into the language of intersection graphs of chord diagrams. I…
Termination proof for Cartan's method in constant type problems.
problem Proving termination of Cartan's equivalence method for constant type problems.
method Groupoid approach to Lie pseudo-groups and Cartan-Kuranishi theorem.
result Cartan's method terminates at involution or complete reduction for constant type problems.
Paper proposes an attention sampler for reducing attention mechanism computation.
problem Computational challenges in large-scale attention-based models.
method Importance sampling in streaming setting, attention sampler.
result Significantly reduces the computational burden of attention mechanisms.
LTMs use latent vectors for efficient autoregressive generation.
problem Efficient autoregressive generation in language models.
method Dual-rate optimization in variational Bayes framework.
result LTMs achieve superior sample and parameter efficiency.
Researchers introduce datasets for cursive Japanese to ML community.
problem Engage ML community with classical Japanese literature datasets.
method Developed three datasets: Kuzushiji-MNIST, Kuzushiji-49, and Kuzushiji-Kanji.
result Introduced datasets focusing on cursive Japanese to ML community.
A new system combines vision and language for person re-identification.
problem Real-world surveillance lacks visual data for person re-identification.
method Two-stream CNN framework with shared logits, CCA for modalities, multi-modal testing protocol.
result 22% improvement in re-identification performance with multi-modal queries.
This paper introduces Graph Convolutional Recurrent Network (GCRN), a deep learning model able to predict structured sequences of data. Precisely, GCRN is a generalization of classical recurrent neural networks (RNN) to data structured by an arbitrary graph. Such structured sequences can represent series of frames in v…
Proposes a framework for compositional generalization in language models.
problem Lack of compositional generalization in neural networks compared to humans.
method Introduces Generalized Grammar Rules (GGRs) for transduction tasks, formalizing symmetry-based constraints.
result Framework enables models to generalize compositionally, similar to human learning.
ROTS improves sentence similarity by incorporating structural information.
problem Measuring sentence similarity with theoretical insights and structural awareness.
method Recursive Optimal Transport (ROT) framework to incorporate structural information.
result ROTS outperforms weakly supervised approaches in sentence similarity tasks.
Graph Structured Prediction Energy Networks model correlations for joint inference.
problem Joint inference over multiple variables with high-order correlations.
method Energy Networks for modeling explicit local and implicit higher-order correlations.
result Tractable inference with explicit modeling of correlations.
The paper solves classical problems in option pricing.
problem Determining the law of the underlying and pricing options with convex payoffs.
method Formulates problems using inverse problem theory and provides proofs without special assumptions.
result Extends existing results in option pricing theory.
Applying the classical Serre-Swan theorem, as this is extended to topological (non-normed) algebras, one attains a classification of elementary particles via their spin-structure. In this context, our argument is virtually based on a ``correspondence principle'' of S. A. Selesnick, formulated herewith in a sheaf-theore…
Efficient echo state network with explicit memory performs well on benchmark tasks.
problem Training differentiable neural computers is difficult and time-consuming.
method Echo state network with an explicit memory.
result Echo state network can recognize all regular languages, including those contractive networks cannot.
Extends GENO framework for GPU optimization of constrained ML problems.
problem Constrained optimization in classical machine learning.
method Extends GENO framework to GPU optimization, specifying problems in a modeling language.
result Solvers on GPU outperform state-of-the-art approaches by several orders of magnitude.
New loss functions based on f-divergences improve language model performance.
problem Improving multiclass classification and language modeling performance.
method Constructing new convex loss functions using f-divergences and deriving an operator for computation.
result The α-divergence loss function with α=1.5 performs well across various tasks. Adaptive querying learns user psychometrics with AI personas.
problem Learning user psychometrics within query budgets.
method Persona-induced latent variable model with AI personas and large language model response distributions.
result Persona-based posteriors deliver accurate probabilistic predictions.
Quantum probability theory reveals hidden structure in joint probability distributions.
problem Understanding hidden structure in joint probability distributions.
method Modeling joint probability distributions as density operators and applying partial trace.
result Decoding extra information in reduced density operators that captures subsystem interactions.
It seems that what has been said by now about market and competitiveness do not fit perfectly with competences of getting the best of profit. Sometimes, the classical methods of fundamentals of management do not apply to individual companies that face irregular accommodation on the market. It is high time to replace th…
The paper introduces uncertainty quantification for NER models.
problem Current NER models lack uncertainty measures, leading to downstream errors.
method Full-Sequence and Subsequence Conformal Prediction framework.
result The method provides formal guarantees about the reliability of model predictions.
Theoretical framework explains why few epochs are enough for LLM fine-tuning.
problem Understanding why few epochs are sufficient for LLM fine-tuning.
method Combining early stopping theory with attention-based Neural Tangent Kernel (NTK) for LLMs.
result Formalizes convergence rate of attention-based fine-tuning with respect to sample size.
In this paper there is proposed a generalized version of the SVM for binary classification problems in the case of using an arbitrary transformation x -> y. An approach similar to the classic SVM method is used. The problem is widely explained. Various formulations of primal and dual problems are proposed. For one of t…
Remasking improves the quality of discrete diffusion models for natural language and image generation.
problem Limited iterative refinement in masked discrete diffusion models.
method Introducing ReMDM sampler that allows remasking during inference.
result Remasking enables better quality outputs with increased sampling steps.