Unlocking GPU Potential Through Low-Level Optimization

Kog, a French startup, is challenging the industry assumption that standard datacenter GPUs (like the Nvidia H200 or AMD MI300X) are inherently inefficient for agentic AI workflows. Rather than relying on purpose-built inference chips, Kog focuses on deep-level software optimization to extract significantly higher performance from existing hardware. Their approach involves reverse-engineering GPU operations down to the assembly and binary level, treating the hardware as a system to be hacked for performance gains it was not originally designed to achieve.

The Engineering Trade-off

This methodology is highly labor-intensive. Because Kog optimizes at such a granular level, each new GPU architecture requires several weeks or months of dedicated research. With a small team of 11, this creates a bottleneck in scaling support across diverse hardware. However, the startup aims to eventually automate this process by feeding their findings into agent-based pipelines.

Current Traction and Future Roadmap

Kog’s initial tech preview demonstrated 3,000 tokens per second (TPS) using their open-sourced 2B parameter model, Laneformer 2B. While the current focus is on smaller models, the company is actively working to apply these same optimization techniques to larger LLMs to meet enterprise demand. CEO Gaël Delalleau views speed as a critical competitive advantage, particularly for professional AI workflows like software engineering and generative app development, where latency currently hampers productivity. The company expects to demonstrate 10x speed improvements on major models by September 2026, which will serve as the primary catalyst for their upcoming Series A funding round.