@SwayStar123
I'm implementing a tiny transformer on tinyshakespear. The 2d matrices went to muon and 1d to adam. The model was still learning and generating some real words. But turns out my muon implementation was bugged and was no-oping. So 99.2% of my model was frozen at init. Adam still managed to tweak those biases into having the model still output some real words. Oh and I forgot the positional embeddings too. It's kinda crazy how you can have the shittiest implementation and a neural network still manages to learn