Announcements

MatX One and our Series B

MatX One and our Series B - 24 Feb 2026

Series A

Series A - 11 Mar 2025

Research

Future leakage in block-quantized attention

Future leakage in block-quantized attention - 9 Jan 2026

Simple and fast Rust deriving using macro_rules

Simple and fast Rust deriving using macro_rules - 28 Jul 2025

Speculative Decoding with Blockwise Sparse Attention

Speculative Decoding with Blockwise Sparse Attention - 22 Jul 2025

SPIRe: Boosting LLM Inference Throughput with Speculative Decoding

SPIRe: Boosting LLM Inference Throughput with Speculative Decoding - 8 Apr 2025

Prioritize values over keys: faster attention with many sparsely accessed value heads

Prioritize values over keys: faster attention with many sparsely accessed value heads - 8 Apr 2025

Optimize for inference too, not just training FLOPs

Optimize for inference too, not just training FLOPs - 8 Jan 2025

Introducing seqax: A Simple and Efficient LLM Research Codebase

Introducing seqax: A Simple and Efficient LLM Research Codebase - 6 May 2024