Skip to content
#

kernel-optimization

Here are 71 public repositories matching this topic...

⚡ FlashVSR v1.1 at ~1.3× on a single RTX 4090 — fused RMSNorm+RoPE / AdaLN (Triton), FP8 E4M3 linears + GELU→FP8 FFN, block-sparse attention. Drop-in, quality-gated (LPIPS 0.011), every number backed by JSON benchmarks.

  • Updated Sep 24, 2026
  • Python

Add this topic to your repo

To associate your repository with the kernel-optimization topic, visit your repo's landing page and select "manage topics."

Learn more