Announcements
MatX One and our Series B
MatX One and our Series B - 24 Feb 2026
Series A
Series A - 11 Mar 2025
Research
Future leakage in block-quantized attention
Future leakage in block-quantized attention - 9 Jan 2026
Simple and fast Rust deriving using macro_rules
Simple and fast Rust deriving using macro_rules - 28 Jul 2025
Speculative Decoding with Blockwise Sparse Attention
Speculative Decoding with Blockwise Sparse Attention - 22 Jul 2025
SPIRe: Boosting LLM Inference Throughput with Speculative Decoding
SPIRe: Boosting LLM Inference Throughput with Speculative Decoding - 8 Apr 2025
Prioritize values over keys: faster attention with many sparsely accessed value heads
Prioritize values over keys: faster attention with many sparsely accessed value heads - 8 Apr 2025
Optimize for inference too, not just training FLOPs
Optimize for inference too, not just training FLOPs - 8 Jan 2025
Introducing seqax: A Simple and Efficient LLM Research Codebase
Introducing seqax: A Simple and Efficient LLM Research Codebase - 6 May 2024