Kog builds the Kog Inference Engine, a real-time inference engine for AI agents running on standard datacenter GPUs.
We co-design three layers: model architecture, inference engine, and low-level GPU kernels. We design the full inference stack around standard AMD and NVIDIA datacenter GPUs, from model architecture down to low-level kernels.
Kog generates 3,500 tokens/s per request on 8 AMD MI300X GPUs and 2,100 tokens/s per request on 8 NVIDIA H200 GPUs, in FP16 at batch size 1, with quantization and speculative decoding disabled.
Our hot path removes framework and host overhead. On NVIDIA, we write CUDA and PTX by hand. On AMD, we use HIP and CDNA ISA.
The team has 11 people, including 10 engineers and researchers and 5 PhDs.
Test it at playground.kog.ai. Read the technical details on the Kog Labs blog.