Product Inference Engine Factory

Engineering samples / Scaling production

Inference Engine Factory

Engineering samples now. Production at scale next.

Inference Engine Factory produces complete engines—from CUDA and PTX kernels to a Rust driver and serving. Its first outputs are working engineering samples with state-of-the-art low-latency performance for video, music, and speech. We are now scaling the factory to produce engines at scale.

Inference improvement loop Experts direct / The factory builds / Evidence compounds
Inference Engine Factory improvement loop Experts define the model, hardware, objectives, and tests. An autonomous agent swarm builds a complete inference stack, target GPUs prove its correctness and performance, and the resulting evidence improves the next generation. 01 / EXPERT DIRECTION HUMANS SET THE DIRECTION MODEL / HARDWARE OBJECTIVES / TESTS GOALS 02 / AUTONOMOUS FACTORY SWARM ONLINE ARCHITECT KERNELS RUNTIME SERVING AGENT SWARM BUILDS 03 / ENGINEERING SAMPLE ENGINE SOTA LOW LATENCY SERVING INTEGRATION RUST CUDA DRIVER MEGAKERNELS CUDA + PTX KERNELS PROVE 04 / TARGET GPU LIVE CORRECTNESS LATENCY THROUGHPUT MEMORY VALID INFERENCE ENGINE VIDEO MUSIC SPEECH RESULTS + FAILURES IMPROVE THE NEXT GENERATION Inference Engine Factory improvement loop Experts set the objectives. An agent swarm builds a complete inference engine, target hardware proves it, and evidence improves the next generation. 01 / EXPERT DIRECTION HUMANS SET THE DIRECTION MODEL / HARDWARE / OBJECTIVES / TESTS 02 / AUTONOMOUS FACTORY SWARM ONLINE ARCHITECT KERNELS RUNTIME SERVING AGENTSWARM 03 / ENGINEERING SAMPLE ENGINE SOTA LOW LATENCY SERVING INTEGRATION RUST CUDA DRIVER MEGAKERNELS CUDA + PTX KERNELS VIDEO / MUSIC / SPEECH 04 / TARGET GPU PROOF LIVE CORRECT LATENCY THROUGHPUT VALID INFERENCE ENGINE / EVIDENCE RETAINED EVIDENCE IMPROVES THE NEXT GENERATION

Three engines from the factory.

The factory's first engineering samples run directly on the CUDA driver—without PyTorch, inference frameworks, or human expert tuning.

  1. Video / 8× B200

    MiniMax H3

    15 seconds of video generated in 10.80 seconds—faster than the SGLang baseline.

  2. Music / 2× H100

    MiniMax Music 3

    10 seconds of 32-kHz stereo audio generated in 4.24 seconds.

  3. Speech / 1× H100

    Fish Audio S2 Pro

    0.252 real-time factor across the full 1,088-sample evaluation set.

From samples to production at scale.

The factory is producing individual engineering samples today. Next, we are expanding compute, workload coverage, repeatability, and throughput so it can produce engines at scale. Stay tuned.

Read the engineering results

Try engines from the factory.

Explore the current demos while we scale factory production.

Try AI-Built Engine