Improved Transformer performance by addressing 'explaining away' effect.
problem Transformer's self-attention mechanism can explain away important input features.
method Proposed a doubly-normalized attention scheme to avoid 'explaining away' effect.
result Improved performance on benchmarks with the new attention scheme.