Llama.cpp: Deterministic Inference Mode (CUDA): RMSNorm, MatMul, Attention github.com 4 points by diwank 7 hours ago