Product Inference Engine Factory
Engineering samples / Scaling production
Inference Engine Factory
Engineering samples now. Production at scale next.
Inference Engine Factory produces complete engines—from CUDA and PTX kernels to a Rust driver and serving. Its first outputs are working engineering samples with state-of-the-art low-latency performance for video, music, and speech. We are now scaling the factory to produce engines at scale.
01 Engineering samples
Three engines from the factory.
The factory's first engineering samples run directly on the CUDA driver—without PyTorch, inference frameworks, or human expert tuning.
- Video / 8× B200
MiniMax H3
15 seconds of video generated in 10.80 seconds—faster than the SGLang baseline.
- Music / 2× H100
MiniMax Music 3
10 seconds of 32-kHz stereo audio generated in 4.24 seconds.
- Speech / 1× H100
Fish Audio S2 Pro
0.252 real-time factor across the full 1,088-sample evaluation set.
02 Scaling up
From samples to production at scale.
The factory is producing individual engineering samples today. Next, we are expanding compute, workload coverage, repeatability, and throughput so it can produce engines at scale. Stay tuned.
Read the engineering resultsTry engines from the factory.
Explore the current demos while we scale factory production.
Try AI-Built Engine