Synthetic noise training improves machine translation robustness to spelling mistakes.
problem Making machine translation robust to spelling mistakes and natural noise.
method Training on synthetic noise to improve robustness to natural noise.
result Training on synthetic noise improves robustness to natural noise without diminishing performance on clean text.
Real-time spell checker adapts to new languages.
problem No real-time, language-adaptable spell checkers for non-English languages.
method Used Wikipedia and subtitles data to generate dictionaries, created noisy channel datasets, compared with industry tools.
result System performs well across 24 languages, outperforming existing tools.
Paper trains models to resist string transformations.
problem Vulnerability of NLP models to adversarial string transformations.
method Combines search and abstraction techniques for robust training.
result Trained models resist combinations of user-defined transformations.
This study analyses the duration dependence of events that trigger volatility persistence in stock markets. Such events, in our context, are monthly spells of contiguous price decline or negative returns for the S&P500 stock market index over the last 145 years. Factors known to affect the duration of these spells are …
Novel approach to estimate P300 BCI efficiency using SNR.
problem Improving the accuracy of P300 BCI for severely disabled people.
method Introduced a novel approach considering P300 SNR for estimating efficiency, using a Gaussian noise model.
result P300 SNR significantly correlates with spelling accuracy, improving BCI efficiency.
We further study the incidence relations that arise from the various subtowers, known as Baby Monster, which exist within the R 3 \mathbb{R}^{3} R 3 -Monster Tower. This allows us to complete the R V T RVT R V T class spelling rules. We also present a method of calculating the various Baby Monster that appear within the Monster Tower.
A transformer model improves spell correction with hierarchical attention.
problem Improving spell correction accuracy and speed.
method Multi encoder-single decoder transformer architecture with hierarchical attention.
result Significant improvement in CER, WER, and SER error rates.
System solves author name ambiguity in e-commerce catalogs.
problem Finding correct author names in e-commerce catalogs with abbreviations and spelling variants.
method Composite system using open data sources and machine learning techniques for natural language processing.
result Top proposal of the system is the normalized author name with 72% accuracy.
Deep learning model extracts location references from tweets during emergencies.
problem Challenges in extracting reliable location information from tweets during crises.
method Convolutional Neural Network (CNN) based model.
result Achieved high accuracy in extracting location references from tweets.
The generalization of Frobenius' theorem to foliations with singularities is usually attributed to Stefan and Sussmann, for their simultaneous discovery around 1973. However, their result is often referred to without caring much on the precise statement, as some sort of magic spell. This may be explained by the fact th…
Ahpatron improves online kernel learning with tighter mistake bounds.
problem Improving mistake bounds in online kernel learning with budget constraints.
method Introducing Ahpatron, a new model that uses an aggressive updating rule and a budget maintenance mechanism to approximate AVP.
result Ahpatron achieves tighter mistake bounds compared to previous models.
Chatbot uses BERT to handle financial investment questions, improving accuracy and decision-making.
problem Improving accuracy and decision-making in financial investment customer service.
method Deep Bidirectional Transformer (BERT) model, uncertainty measure comparison, mixed-integer programming, automatic spelling correction.
result Chatbot can recognize 381 intents and decide when to escalate questions.
A deterministic apple tasting learner is developed, confirming a conjecture and providing tight bounds for mistake bounds.
problem Determining the learnability of hypothesis classes in binary online classification with apple tasting feedback.
method Developed a deterministic apple tasting learner and proved tight bounds for mistake bounds.
result Deterministic apple tasting is feasible and provides tight bounds for mistake bounds.
This note corrects one serious mistake and several smaller mistakes from arXiv:math/0502404. The main results of that paper are unchanged.
We retract the scalar curvature rigidity theorem as there is a mistake in the proof. We thank S. Montiel for pointing out the mistake.
For a number of reasons, computational intelligence and machine learning methods have been largely dismissed by the professional community. The reasons for this are numerous and varied, but inevitably amongst the reasons given is that the systems designed often do not perform as expected by their designers. The reasons…
A new technique updates weights multiple times for online learning, reducing mistakes to near zero.
problem Online learning with partial data and unknown future data points.
method Iterative weight updating for the same instance.
result Reduced mistake rate to near zero for various datasets and algorithms.
Corrects a mistake in a theorem about hyperbolic groups.
problem A theorem about graphs of hyperbolic groups was weakened.
method Identified and corrected a mistake in the published paper.
result Proved a weaker result to correct the main theorem.
Improved mistake bounds for transductive online learning.
problem Quantifying the power of unlabeled data in online learning.
method Proving lower and upper bounds on transductive mistake bounds.
result Exponential improvement in mistake bounds for transductive learning.
Simplifies online learning with consistent oracle to fewer mistakes.
problem Online learning with computationally intractable Littlestone dimension computation.
method Novel algorithm making at most O ( 256 d ) O(256^d) O ( 25 6 d ) mistakes, simpler proof. result No algorithm can make less than 3 d 3^d 3 d mistakes. We present Listen, Attend and Spell (LAS), a neural network that learns to transcribe speech utterances to characters. Unlike traditional DNN-HMM models, this model learns all the components of a speech recognizer jointly. Our system has two components: a listener and a speller. The listener is a pyramidal recurrent ne…
Winterization of Texas power system profitable but risky, estimated at $11.74bn over 30 years.
problem Profitability and risk of winterizing Texas power system infrastructure.
method Combined temperature-dependent load and outage estimates over 71 years of climate data.
result Large-scale winterization of gas infrastructure and power plants is profitable, but risks are high due to low-frequency of cold spells.
The paper studies how noisy labels impact decision-making in machine learning.
problem The impact of noisy labels on decision-making in machine learning.
method Introducing a notion of regret, studying standard approaches, and estimating individual-level mistakes.
result Standard approaches can lead to unforeseen mistakes for individuals, revealing the need for anticipation.
Self-directed learners can minimize mistakes in online classification.
problem Minimizing mistakes in online classification with adaptive prediction order.
method Designing efficient self-directed learners for linear classification.
result Strong separation between worst-order and random-order learning for linear classification.
Improves diversity of text-to-image models without sacrificing FID.
problem Lack of diversity and tendency to recreate training set images.
method Adds sparse repellency terms to diffusion SDE to guide trajectories away from a reference set.
result Improves diversity of diffusion models with minimal impact on FID.
Study the tradeoffs of bandit feedback in multiclass classification.
problem The price of using bandit feedback in multiclass classification.
method Mistake bound model, analysis of variants, and comparison of learners and adversaries.
result The optimal mistake bound under bandit feedback is at most O ( k ) O(k) O ( k ) times higher than in full information, with a tight bound of O ( k ) O(k) O ( k ) . Improved mistake bound for group linear separable cases in online multiclass linear classification.
problem Improving mistake bounds for online multiclass linear classification under group linear separable conditions.
method Refined group weak linear separability condition and rational kernel approach.
result Achieved a mistake bound of K ⋅ 2 i l d e O ( 1 / γ log L ) ) K\cdot 2^{ ilde{O}(\sqrt{1/γ}\log L)}) K ⋅ 2 i l d e O ( 1/ γ l o g L ) ) under group weak linear separable condition. Study on tradeoffs between mistakes and ERM oracle calls in online and transductive learning.
problem Analyzing online and transductive learning with limited ERM and weak consistency oracle access.
method Proves lower bounds and upper bounds on mistakes and oracle calls, considering realizable and agnostic cases.
result Achieves optimal mistake bounds with weak consistency queries for certain concept classes.
The Monster tower, also known as the Semple tower, is a sequence of manifolds with distributions of interest to both differential and algebraic geometers. Each manifold is a projective bundle over the previous. Moreover, each level is a fiber compactified jet bundle equipped with an action of finite jets of the diffeom…
SkewSize detects model biases by analyzing mistakes across subgroups.
problem Benchmarking model performance in the presence of spurious correlations.
method Introducing SkewSize, a metric that captures bias from model mistakes.
result SkewSize highlights biases not captured by other metrics.
In automatic speech recognition (ASR) what a user says depends on the particular context she is in. Typically, this context is represented as a set of word n-grams. In this work, we present a novel, all-neural, end-to-end (E2E) ASR sys- tem that utilizes such context. Our approach, which we re- fer to as Contextual Lis…
Thompson Sampling is at most twice as bad as any other policy in Bayesian bandit models.
problem Optimizing selection of the best arm in Bayesian bandit models with independent latent processes.
method Thompson Sampling approach applied to models with independent latent arm processes.
result Thompson Sampling makes at most twice the expected number of mistakes compared to any other policy.
Corrected a mistake in a paper about minimal surfaces.
problem A mistake in a paper about complete minimal surfaces with finite total curvature.
method No new method introduced, just correcting an error.
result Corrected a mistake in a previously published paper.
Using Jeff Holman's comments in Quantitative Finance to illustrate 4 critical errors students should learn to avoid: 1) Mistaking tails (4th moment) for volatility (2nd moment), 2) Missing Jensen's Inequality, 3) Analyzing the hedging wihout the underlying, 4) The necessity of a numeraire in finance.
Open problem seeks an online learning algorithm for binary classification.
problem Existence of an online learning algorithm for binary classification with sublinear mistakes.
method Assumption of sequence allowing learning algorithm's existence.
result Specific condition determines sequence's learnability.
Modified Perceptron handles strategic agents with limited position changes.
problem Learning linear classifiers in the presence of strategic agents that can manipulate their positions.
method Developed a modified Perceptron algorithm with bounded mistakes under various manipulation costs.
result The modified Perceptron achieves bounded mistakes even when manipulation costs are unknown.
Paper analyzes mistake and generalization of MNIC classifiers.
problem Understanding the performance of interpolating classifiers.
method Elementary analyses of MNIC's regret and generalization.
result MNIC generalizes with a rate proportional to the norm of the interpolating solution and inversely proportional to the number of data points.
We investigate the problem of active learning on a given tree whose nodes are assigned binary labels in an adversarial way. Inspired by recent results by Guillory and Bilmes, we characterize (up to constant factors) the optimal placement of queries so to minimize the mistakes made on the non-queried nodes. Our query se…
Study apple tasting feedback in online binary classification, providing new insights into minimax expected mistakes.
problem Online binary classification with partial feedback (apple tasting).
method Combinatorial analysis, Littlestone dimension, Effective width.
result Established a trichotomy of minimax expected mistakes in the realizable setting.
Corrects mistakes in convergence rate claims for SGD learning rate scheme.
problem Incorrect convergence rate claims for SGD learning rate scheme.
method Revised the convergence rate claims based on corrected test criterion for a series.
result Valid convergence rate of SGD is O ( 1 / t ) \mathcal{O}(1/t) O ( 1/ t ) , not O ( 1 / t 2 ) \mathcal{O}(1/t^2) O ( 1/ t 2 ) as previously stated. LEAK learns from mistakes to improve point cloud segmentation.
problem Improving point cloud semantic segmentation performance.
method Coarse-to-fine clustering, class-conditional prototypical feature alignment, fairness weighting.
result State-of-the-art performances on different architectures, datasets, and tasks.
Study on computable online learning with new conditions and complexities.
problem Characterizing optimal online learning under varying optimality requirements.
method Introduced anytime optimal (a-optimal) online learning and explored computational separations.
result Found a computational separation between a-optimal and optimal online learning.
Grapheme ASR improves with G2G model that corrects spelling errors.
problem Rare long-tail words in non-phonemic languages like English.
method Train G2G model on text-to-speech data to rewrite character sequences into phonetically consistent forms.
result Reduces Word Error Rate by 3% to 11% over a strong graphemic baseline.
We review (non-abelian) extensions of a given Lie algebra, identify a 3-dimensional cohomological obstruction to the existence of extensions. A striking analogy to the setting of covariant exterior derivatives, curvature, and the Bianchi identity in differential geometry is spelled out. In the new version references ad…
Study online learning of neural networks with margin condition.
problem Online learning of neural networks with sign activation function.
method Characterized margin condition for online learnability, proved mistake bounds, constructed counterexamples.
result Proved optimal mistake bounds and lower bounds for neural networks.
A technique to quickly fix mistakes in neural networks.
problem Fixing model errors in neural networks quickly and without affecting other samples.
method Editable Training, a model-agnostic training technique.
result Effectiveness demonstrated on large-scale image classification and machine translation tasks.
Study on learning to predict dynamical systems without assuming their structure.
problem Learning to predict the next state of a dynamical system with unknown evolution function.
method Defined new combinatorial measures to quantify mistake and regret bounds in realizable and agnostic settings.
result In the realizable setting, the number of mistakes can grow arbitrarily with time.
This note corrects the mistakes in the splicing formulas of the paper "Floer homology and splicing knot complements". The mistakes are the result of the incorrect assumption that for a knot K K K inside a homology sphere Y Y Y , the involution on the knot Floer homology of K K K which corresponds to moving the basepoints by o…