Paper proposes i-vector approach for NLI with improved accuracy.
problem Identifying a speaker's native language from speech in a second language.
method i-vector approach using MFCC and GFCC features.
result Improved accuracy in NLI system, especially with GFCC features.
New method reduces text classification errors by learning writing style instead of content.
problem Deep neural networks learn superficial patterns specific to training data.
method Adversarial training to unlearn confounding features.
result Model generalizes better and learns writing style features.
Improved language identification with tuplemax loss.
problem Language identification with user-specified small sets of languages.
method Replaced softmax loss with tuplemax loss.
result 2.33% error rate improvement over softmax loss.
Work shows hallucination detection by LLMs is impossible without expert feedback.
problem Detecting hallucinations in LLMs is theoretically impossible without expert-labeled feedback.
method Investigated hallucination detection using a theoretical framework inspired by language identification.
result Automated hallucination detection is impossible for most language collections without expert-labeled feedback.
Improved language identification accuracy through signal combination methods.
problem Enhancing speech recognition accuracy across multiple languages.
method Combining low-level acoustic signals with language-specific recognizer signals using lattice-based ensemble models and deep neural networks.
result Deep neural network model outperforms lattice-based ensemble model, reducing error rate from 5.5% to 4.3%.
Paper introduces a conformer-based system for streaming language identification in long-form speech.
problem Language identification in long-form audio.
method Conformer layers with attentive temporal pooling and domain adaptation.
result Conformer-based models significantly outperform LSTM and transformer models.
Study on generating and identifying languages privately, showing privacy imposes costs and creates barriers.
problem Generating and identifying languages privately in the limit model.
method Introduced a continual release model under differential privacy constraints, proving both positive and negative results.
result Privacy imposes quantitative and qualitative costs, and creates fundamental barriers for identification.
Improved LID for multilingual speakers using context-aware models.
problem Low accuracy for languages spoken by multilingual speakers, especially with accented speech.
method Coarser-grained acoustic model and integration with interaction context signals.
result Average 97% accuracy across all language combinations, 60% improvement in worst-case accuracy.
A new system combines vision and language for person re-identification.
problem Real-world surveillance lacks visual data for person re-identification.
method Two-stream CNN framework with shared logits, CCA for modalities, multi-modal testing protocol.
result 22% improvement in re-identification performance with multi-modal queries.
System identifies language of transliterated text.
problem Users struggle to understand non-native language transliterated text.
method Feature extraction of phonetic syllables using LSTM network.
result System accurately identifies language of transliterated text.
Study language generation with limited memory, showing different impacts on achievable densities and convergence.
problem Language generation with bounded memory constraints.
method Analyzed memoryless generators, sliding windows, and adaptive past examples; revisited identification in the limit.
result Achievable densities and convergence properties differ based on the size of the target language collection.
Constructs native Banach spaces for spline operators, enabling precise function reproduction.
problem Developing native Banach spaces for spline-admissible operators.
method Systematic construction involving test functions and completion processes.
result Native spaces ensure precise function reproduction with arbitrary precision.
Cryptonite tests NLP models with cryptic crossword clues.
problem Ambiguity in language poses a challenge for NLP models.
method Cryptic crossword dataset based on cryptic clues.
result Fine-tuning T5-Large on Cryptonite achieves only 7.6% accuracy.
Study explores various methods for detecting offensive language.
problem Detecting offensive language in text.
method Examined custom and off-the-shelf architectures, feature vector generation approaches.
result These methods perform well on small, noisy datasets.
MPSA-DenseNet improves accent classification accuracy.
problem Accurate English accent identification.
method Combines multi-task learning and PSA attention mechanism with DenseNet.
result MPSA-DenseNet outperforms other models in accent classification.
Twitter provides an open and rich source of data for studying human behaviour at scale and is widely used in social and network sciences. However, a major criticism of Twitter data is that demographic information is largely absent. Enhancing Twitter data with user ages would advance our ability to study social network …
We consider the role of unobservables, such as differences in search frictions, reservation wages, and productivities for the explanation of wage differentials between migrants and natives. We disentangle these by estimating an empirical general equilibrium search model with on-the-job search due to Bontemps, Robin, an…
Study uses sentiment analysis to predict cryptocurrency token returns in virtual reality.
problem Predicting cryptocurrency token returns in virtual reality economies.
method Used BERT for sentiment analysis and developed LSTM models integrating multi-modal features.
result Multi-modal model significantly outperforms price-only baseline in prediction accuracy.
The paper explores how language models can provide reliable state measurements without being interpreted as beliefs.
problem How to use language models to reliably infer states without misinterpreting them as beliefs.
method Developed a semantic map and semiparametric inverse to link language probabilities to state probabilities, avoiding hidden models.
result Conditions for existence, identification, stable recovery, and uniform stability of posterior states from observable language probabilities.
This study uses cloud telemetry to detect DoS attacks with machine learning.
problem Detecting DoS attacks in cloud environments is challenging.
method Use of machine learning algorithms on cloud telemetry data.
result k-Nearest Neighbors (kNN) and decision tree (CART) accurately identify DoS attacks.
Paper presents LLM-enhanced contract metadata extraction.
problem Automatic detection and annotation of legal clauses in contracts.
method Integration of publicly available and proprietary datasets with advanced LLM methodologies.
result Substantial improvements in clause identification accuracy and efficiency.
HAMD optimizes cubic portfolios without quadratization, achieving better results.
problem Optimizing higher-order portfolio models with reduced distortion.
method Hybrid pipeline combining continuous Hamiltonian search, cardinality-preserving projection, and iterated local search.
result HAMD achieves significantly lower native cubic objective values than classical heuristics.
Paper tackles counterfactual sentence detection and evaluation.
problem Detect and evaluate counterfactual sentences in natural language.
method Used a BERT base model for classification and a hybrid BERT Multi-Layer Perceptron for sequence identification. Introduced cascaded linear inputs to improve performance.
result Achieved an F1 score of 85.00% in Task 1 and 83.90% in Task 2.
Pars-ABSA dataset for Persian aspect-based sentiment analysis.
problem Lack of public dataset for Persian aspect-based sentiment analysis.
method Manually annotated dataset with 5,114 positive, 3,061 negative, and 1,827 neutral samples.
result State-of-the-art performance of deep learning methods on Pars-ABSA compared to similar English datasets.
Study improves LLMs for PPI analysis by addressing uncertainty.
problem Uncertainty in LLM predictions for PPIs.
method Fine-tuned LLaMA-3 and BioMedGPT models, LoRA ensembles, Bayesian LoRA for UQ.
result Competitive PPI identification performance across diverse disease contexts.
Extracts high-quality monolingual datasets from web crawl data.
problem Improving text representation quality through larger corpora.
method Automated pipeline using deduplication and language identification, augmented with filtering for high-quality documents.
result Extracted massive high-quality monolingual datasets from Common Crawl.
Representation costs in data science: Unifying function-space views of parametric methods
problem Analyzing representation costs of parametric data-fitting methods
method Developing a general framework for analyzing representation costs through parameter-space regularizers
result Proving that many natural results hold in this abstract setting, including representer theorems for parametric methods on their native spaces
PROBE optimizes best-arm identification with cheap proxies, improving sample complexity.
problem Fixed-confidence best-arm identification with costly rewards and correlated cheap proxies.
method PROBE uses control-variate adjustment and phase elimination to learn residual variance online.
result PROBE achieves oracle sample complexity up to a constant factor and additive calibration cost.
Filtered conformal ellipsoids for graph-native time series
problem Joint prediction sets for multivariate time series
method Filtered conformal ellipsoids
result Sharper at-target ellipsoids than static-covariance and non-filter baselines
The paper aims to explore the impacts of bi-demographic structure on the current account and growth. Using a SVAR modeling, we track the dynamic impacts between these underlying variables. New insights have been developed about the dynamic interrelation between population growth, current account and economic growth. Th…
Develops a machine learning framework for identifying authorship in texts.
problem Identifying the author of texts, especially when authors are unknown or multiple.
method Formulates authorship identification as a text categorization problem, uses supervised machine learning with stylometric features.
result A model accurately predicts authorship with high accuracy, especially with linguistic stylometric features.
The goal of this research was to find a way to extend the capabilities of computers through the processing of language in a more human way, and present applications which demonstrate the power of this method. This research presents a novel approach, Rhetorical Analysis, to solving problems in Natural Language Processin…
Improved offensive language detection in tweets with multiple deep learning models.
problem Detecting offensive language in tweets using machine learning.
method Combination of multiple deep learning architectures for classification.
result Achieved macro-average F1-scores of 0.76, 0.68, 0.54 for different tasks.
LLMs struggle to generate random numbers from statistical distributions, leading to biased results in applications.
problem LLMs' inability to generate random numbers accurately from specified distributions.
method Dual-protocol design: Batch Generation and Independent Requests, benchmarking 11 models across 15 distributions.
result Sampling fidelity degrades with distributional complexity and horizon, leading to systematic biases in downstream applications.
This paper presents a fast Bayesian filtering technique for state estimation.
problem Bottleneck in Bayesian inference for state estimation from noisy sensor data.
method Processor-native uncertainty tracking for uncertainty propagation and inference.
result Deterministic approximate filtering with up to 805x speedup and competitive accuracy.
AI applications pose increasing demands on performance, so it is not surprising that the era of client-side distributed software is becoming important. On top of many AI applications already using mobile hardware, and even browsers for computationally demanding AI applications, we are already witnessing the emergence o…
Paper introduces Native Guide for generating time series counterfactual explanations.
problem Lack of explainability for time series data in AI systems.
method Model-agnostic, instance-based counterfactual generation for time series classification.
result Native Guide produces better counterfactual explanations than benchmarks.
RFX-Fuse combines Breiman and Cutler's Random Forest with modern ML capabilities.
problem Lack of a unified ML engine with diverse capabilities.
method Unified ML engine with native GPU/CPU support, delivering 5+ functionalities in one model.
result Native explainable similarity and imputation validation.
Determining the programming language of a source code file has been considered in the research community; it has been shown that Machine Learning (ML) and Natural Language Processing (NLP) algorithms can be effective in identifying the programming language of source code files. However, determining the programming lang…
Paper proposes set-valued prediction for historical POS tagging.
problem Difficult POS tagging in historical corpora due to lack of native speakers and sparse data.
method Set-valued prediction approach to allow uncertainty in tagging.
result Set-valued prediction improves POS tagging precision and robustness.
Click through rate (CTR) prediction is very important for Native advertisement but also hard as there is no direct query intent. In this paper we propose a large-scale event embedding scheme to encode the each user browsing event by training a Siamese network with weak supervision on the users' consecutive events. The …
Framework ranks sectors influenced by Indian Union Budgets.
problem Real-time analysis of budgetary impacts on sector-specific equity performance.
method Fine-tuned embeddings and language models for sector identification and performance ranking.
result 0.997 NDCG score in predicting sector ranks based on post-budget performances.
Lean Copilot uses LLMs to assist theorem proving in Lean, improving efficiency and automation.
problem Challenges in using existing neural theorem provers to prove novel theorems autonomously.
method Introduces Lean Copilot, a framework for integrating LLMs into Lean's theorem proving process.
result Lean Copilot automates 74.2% of proof steps on average, significantly improving over existing methods.
Language recognition system is typically trained directly to optimize classification error on the target language labels, without using the external, or meta-information in the estimation of the model parameters. However labels are not independent of each other, there is a dependency enforced by, for example, the langu…
Two proteins are homologous if they have a common evolutionary origin, and the binary classification problem is to identify proteins in a candidate set that are homologous to a particular native protein. The feature (explanatory) variables available for classification are various measures of similarity of proteins. The…
Neural network framework for language recognition considers sequence information and improves accuracy.
problem Challenging task of automatic language identification in noisy conditions.
method Proposes a neural network framework with bidirectional LSTM and attention modeling for relevance weighting.
result Significant improvements over conventional methods in noisy conditions and multi-speaker speech.
Paper proposes M2H-GAN to improve speech theme identification.
problem Limited ASR transcripts for speech theme identification.
method Uses M2H-GAN, a GAN-based approach, to generate TRS-like ASR transcripts.
result Improves speech theme identification performance close to human levels.
EagerPy simplifies writing code for multiple deep learning frameworks.
problem Code duplication and framework lock-in when using different deep learning libraries.
method Integrates multiple frameworks into a single Python framework.
result Automatic compatibility across PyTorch, TensorFlow, JAX, and NumPy.