French startup Kog is advancing software techniques to unlock much faster AI inference on widely used datacenter GPUs such as AMD MI300X and Nvidia H200, aiming to overcome bottlenecks in AI workflow speed and cost without requiring specialized hardware.

  • Kog’s software achieves high-speed AI inference on standard GPUs like Nvidia H200 and AMD MI300X.
  • Focus on accelerating large models to meet demand from professional AI workflow users.
  • Deep GPU-level optimization requires extensive hardware-specific research and engineering.

What happened

Kog, a French AI startup, has demonstrated significant improvements in inference speed for large language models using conventional datacenter GPUs. Its technology, the Kog Inference Engine, leverages deep software optimizations to boost throughput on widely deployed GPU models without additional specialized hardware.

Since a tech preview in May showed promising results with smaller models, Kog has attracted substantial interest and over 200 business leads. The company is focused on expanding this capability to larger models, addressing a major bottleneck faced by enterprises and developers relying on AI for professional tasks.

Why it matters

AI inference speed and cost are critical factors limiting adoption in many industries where fast, real-time results are essential. By enabling faster AI workflows on hardware already owned by enterprises, Kog’s approach could reduce expenses and increase the viability of AI-powered applications.

Kog’s work challenges the common belief that GPUs are poorly suited for agentic and dynamic AI decoding tasks. With large-scale models becoming increasingly important, unlocking existing GPU performance gains with software innovation could shift industry reliance away from expensive, purpose-built chips.

What to watch next

Kog plans to continue its research-intensive process to optimize newer GPU architectures and scale its engine for larger models. Given the time required to deeply understand each GPU’s hardware, expansion across different chip types will be gradual but critical to broader adoption.

Market response will hinge on whether Kog can deliver on its promise of 30x faster inference and demonstrate tangible cost savings for enterprise AI users. Partnerships with large-scale AI service providers and integration in professional workflows will be key indicators of its impact.

Source assisted: This briefing began from a discovered source item from TechCrunch Startups. Open the original source.
How SignalDesk reports: feeds and outside sources are used for discovery. Public briefings are edited to add context, buyer relevance and attribution before they are published. Read the standards

Related briefings