#transformers
Every summary, chronological. Filter by category, tag, or source from the rail.
Tag · #transformers
Optimizing Transformer Inference with FlashNorm
FlashNorm accelerates transformer inference by folding RMS norm gains into projection weights and parallelizing normalization and matrix multiplication via custom CUDA kernels.
AI EngineerBuilding Recurrent-Depth Transformers with OpenMythos
OpenMythos enables recurrent-depth transformers that trade inference-time compute for deeper reasoning by reusing model parameters through recurrent loops.
MarkTechPost
Showing 2 of 2