Use compute
to improve
compute.

INT21 builds self-improving AI systems for the software beneath modern AI. Inference Engine Factory builds and optimizes inference engines, then proves candidates against real performance targets on real hardware.

INT21 Use compute to improve compute
INT21 Use compute to improve compute

01 / Product

Inference Engine Factory

Optimize the engine, not just one operation.

Agent swarms explore inference engine configurations and optimization paths in parallel, evaluate candidates against correctness and performance targets, and carry the strongest evidence forward.

Inference improvement loop Experts direct / The factory builds / Evidence compounds
Inference Engine Factory improvement loop Experts define the model, hardware, objectives, and tests. An autonomous agent swarm builds a complete inference stack, target GPUs prove its correctness and performance, and the resulting evidence improves the next generation. 01 / EXPERT DIRECTION HUMANS SET THE DIRECTION MODEL / HARDWARE OBJECTIVES / TESTS GOALS 02 / AUTONOMOUS FACTORY SWARM ONLINE ARCHITECT KERNELS RUNTIME SERVING AGENT SWARM BUILDS 03 / ENGINEERING SAMPLE ENGINE SOTA LOW LATENCY SERVING INTEGRATION RUST CUDA DRIVER MEGAKERNELS CUDA + PTX KERNELS PROVE 04 / TARGET GPU LIVE CORRECTNESS LATENCY THROUGHPUT MEMORY VALID INFERENCE ENGINE VIDEO MUSIC SPEECH RESULTS + FAILURES IMPROVE THE NEXT GENERATION Inference Engine Factory improvement loop Experts set the objectives. An agent swarm builds a complete inference engine, target hardware proves it, and evidence improves the next generation. 01 / EXPERT DIRECTION HUMANS SET THE DIRECTION MODEL / HARDWARE / OBJECTIVES / TESTS 02 / AUTONOMOUS FACTORY SWARM ONLINE ARCHITECT KERNELS RUNTIME SERVING AGENTSWARM 03 / ENGINEERING SAMPLE ENGINE SOTA LOW LATENCY SERVING INTEGRATION RUST CUDA DRIVER MEGAKERNELS CUDA + PTX KERNELS VIDEO / MUSIC / SPEECH 04 / TARGET GPU PROOF LIVE CORRECT LATENCY THROUGHPUT VALID INFERENCE ENGINE / EVIDENCE RETAINED EVIDENCE IMPROVES THE NEXT GENERATION
  1. 01

    Workload-level objective

    Start from the model, deployment constraints, and the serving metric that matters.

  2. 02

    Parallel engine search

    Specialized agents explore configuration and optimization paths at the same time.

  3. 03

    Evidence compounds

    Correctness and performance results become the next generation's starting point.

02 / Platform

Built on SwarmOS.

Thousands of agents. One evolving system.

SwarmOS is a cloud-native platform for running specialized agents at elastic scale toward the same measurable goal. Agents explore in parallel, coordinate through shared evidence, and continuously converge on stronger solutions.

Cloud-native swarm control plane Live system model
  1. 01 / Elastic scale Thousands of agents

    Cloud-native scheduling expands the swarm around available compute.

  2. 02 / Shared direction One measurable goal

    Every agent works against the same constraints and acceptance criteria.

  3. 03 / Generational memory Experience carries forward

    Results, failures, and strategies become the next generation's starting point.

03 / Company

Founded by experts across the full AI stack.

INT21 brings together deep experience in agent systems, machine learning models, GPU software, distributed infrastructure, and cloud computing.

  1. 01Agents
  2. 02Models
  3. 03GPU
  4. 04Infrastructure
  5. 05Cloud

AI-native operating model

Engineering capacity scales with compute, not headcount.

Our experts set direction, constraints, and acceptance criteria. Autonomous agent swarms execute, evaluate, and retain the work, so adding compute expands how much engineering INT21 can perform.

Founders / Cross-stack operators

Research, systems, and infrastructure experience carried into one company.

Bing Xu, Founder and CEO of INT21

Bing Xu

Founder & CEO

  • Agents
  • Models
  • GPU

Bing co-authored the original Generative Adversarial Nets paper, created XGBoost's Python package, and co-created MXNet and AITemplate. Before founding INT21, he was a Distinguished Engineer at NVIDIA following its acquisition of HippoML, the GPU inference company he founded.

Qingye Jiang, Founding Partner at INT21

Qingye Jiang

Founding Partner

  • Infrastructure
  • Cloud

Qingye has spent more than a decade building and tuning high-performance computing and distributed systems at AWS. His work spans workload analysis, performance engineering, cloud infrastructure, and real-time systems.

Inference Engine Factory / Priority access

Bring us a hard inference workload.

Tell us about the model, target hardware, and the latency, throughput, or memory constraint holding your deployment back.

Try AI-Built Engine