MiniMax Sparse Attention: Per-Group Block Selection for Cheap Million-Token Inference
Andrey Lukyanenko
MiniMax Sparse Attention: Per-Group Block Selection for Cheap Million-Token Inference Paper Code Model Long-context LLMs keep promising the same thing: feed more tokens into the prompt and let the model reason over them. The bottleneck is rarely the window itself: it is the cost of attending...