Noisy is Better — Improving Performance with Gaussian Noise
danielmewes
In my previous post about ThoughtNet, an attention-based neural architecture for variable-compute inference, I highlighted two limitations that I encountered with it: Slow and inconsistent convergence during training time Poor generalization on multiplication tasks, despite great...