research/
7 pages · Updated April 18, 2026
Pages
- Prioritize values over keys: faster attention with many sparsely accessed value heads | MatX
- Speculative Decoding with Blockwise Sparse Attention | MatX
- Simple and fast Rust deriving using macro_rules | MatX
- Optimize for inference too, not just training FLOPs | MatX
- Research | MatX
- Future leakage in block-quantized attention | MatX
- Introducing seqax: A Simple and Efficient LLM Research Codebase | MatX