AI models and infrastructure

Inference optimization

Tools that improve inference speed, capacity, or cost through techniques such as speculative decoding and batching

Where it stands

Open models19 products1.05× its peers

Fastest products

1Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed · Open models4.0×
2kernels-community/flash-attn2 · Open models2.9×
3RedHatAI/gemma-4-31B-it-speculator.dflash · Open models1.7×
4inference-optimization/Qwen3-1.6B-A0.9B · Open models1.4×
5Efficient-Large-Model/Fast_dLLM_v2_1.5B · Open models1.4×
6kernels-community/gpt-oss-triton-kernels · Open models1.3×
7inference-optimization/DSV4-tiny-empty · Open models1.3×
8Efficient-Large-Model/Fast_dLLM_v2_7B · Open models1.2×
9caveman-optimize · Agent skills1.2×
10jinaai/jina-bert-flash-implementation · Open models1.2×
11inference-optimization/Qwen3.8-1.0B-A0.6B · Open models1.1×
12vercel-optimize · Agent skills1.1×
13inference-optimization/Llama-3.2-0.5B-Instruct · Open models1.1×
14inference-optimization/GLM-5.2-0.8B-A0.8B · Open models0.9×
15inference-optimization/Qwen3-8B-speculators.peagle-qwen3arch-ckpt4 · Open models0.8×
16inference-optimization/Qwen3-8B-from-Qwen3-8B_regen-speculators.eagle3-qwen3arch-ckpt1 · Open models0.8×
17genshijin · Agent skills0.7×
18kernels-community/vllm-flash-attn3 · Open models0.6×
19Cheaper Inference · AI apps0.5×
20inference-optimization/DeepSeek-V3-debug-empty-FP8_DYNAMIC · Open models0.5×
21kernels-community/flash-attn3 · Open models0.4×
22poolside/Laguna-S-2.1-DFlash-NVFP4 · Open models0.4×
23kernels-community/triton-layer-norm · Open models0.2×

32 products carry this theme. “× its peers” is the theme's median growth against the place's median.

Measured, not estimated · snapshot 2026-10-03 · Growth is shown only where the starting size is big enough to mean something.