INT21

Self-Improving PTX Kernel Factory

Last verified:

Visit INT21

What is INT21?

INT21 is a company building self-improving AI systems for the software beneath modern AI. Their first product, PTX Kernel Factory, is an autonomous AI system that generates and optimizes low-level GPU software (specifically PTX assembly code for NVIDIA GPUs), then proves its work with tests and benchmarks. The system uses a fully autonomous swarm of specialized agents that plan, implement, review, and optimize each kernel end to end.

Key features include: (1) Fully autonomous swarm with thousands of agents that work together cloud-native; (2) Grounded in real hardware where every candidate kernel is compiled, verified, and benchmarked on the target GPU (GH200/Hopper and B200/Blackwell); (3) Improvement compounds through generational memory where results, failures, and strategies become the next generation's starting point; (4) All agents work toward one measurable goal with the same constraints and acceptance criteria. The product achieved 1.59x peak measured performance on GH200 and 1.52x optimized integration on B200.

PTX Kernel Factory is for AI companies, GPU workload engineers, data centers, and organizations running slow GPU operations who need high-performance kernels. It targets operations that are too slow, new architectures without mature kernels, or important workloads that haven justified weeks of specialist time. The first four implementations produced are open source, and the product is entering beta with early access available through request.

The system leverages founder Bing Xu's expertise (co-authored original GAN paper, created XGBoost's Python package, co-created MXNet and AITemplate, former Distinguished Engineer at NVIDIA) and Qingye Jiang's decade of high-performance computing experience at AWS. It represents the first self-improving AI to generate GPU code, beating current implementations by up to 59%.

The product operates as elastic scale cloud-native scheduling that expands the swarm around available compute, with shared direction where every agent works against identical constraints, and generational memory ensuring experience carries forward across iterations.

INT21 pricing

Pricing model: Freemium

PTX Kernel Factory is currently entering beta with early access available through request. The limited API trial was available from June 9 to June 23. For the API trial, it costs three times the standard MiMo-V2.5-Pro rate for roughly 10 times the output. No public free tier or paid plan pricing details are published on the website - users must request beta access for pricing information.

INT21 pros

  • Generates and optimizes low-level PTX GPU assembly code autonomously
  • Achieves up to 59% performance improvement over current implementations
  • Fully autonomous swarm with thousands of specialized agents
  • Every kernel compiled, verified, and benchmarked on real target GPU hardware
  • Achieved 1.59x peak measured performance on GH200/Hopper
  • Achieved 1.52x optimized integration on B200/Blackwell
  • 8.17% faster on geometric mean for GH200
  • 126/126 faster in every comparable case for B200
  • Improvement compounds through reusable evidence and generational memory
  • Open source first four implementations produced by the factory
  • Cloud-native elastic scaling around available compute
  • All agents work toward one measurable goal with shared constraints
  • Specialized agents plan, implement, review, and optimize each kernel end to end
  • Targets new architectures without mature kernels
  • Handles operations too slow for current implementations
  • No need for weeks of specialist engineer time
  • Founder co-authored original Generative Adversarial Nets paper
  • Former Distinguished Engineer at NVIDIA leads the company
  • 10+ years high-performance computing experience at AWS on founding team
  • Operator-level benchmark results not full-model speedup claims

INT21 cons

  • Currently only in beta stage, not production-ready
  • Limited API trial availability (June 9-23 window mentioned)
  • Only NVIDIA GPU support via PTX assembly (no AMD/Intel)
  • Requires access to target GPU hardware for benchmarking
  • Beta access requires request, not immediately available
  • Limited to low-level GPU software, not full applications
  • Elastic scale requires significant cloud compute resources
  • Specialized PTX assembly knowledge still needed for integration
  • No public pricing information available yet
  • Only first four implementations open source, rest proprietary

Frequently asked questions about INT21

What is PTX Kernel Factory?

PTX Kernel Factory is INT21's first product, a self-improving AI system that generates and optimizes low-level GPU software (PTX assembly code for NVIDIA GPUs) and proves its work with tests and benchmarks. It uses a fully autonomous swarm of specialized agents to plan, implement, review, and optimize each kernel end to end.

What performance improvements does PTX Kernel Factory deliver?

The system beats current implementations by up to 59%. It achieved 1.59x peak measured performance on GH200/Hopper, 1.52x optimized integration on B200/Blackwell, 8.17% faster on geometric mean for GH200, and was faster in every comparable case (126/126) for B200. These are operator-level benchmark results, not full-model speedup claims.

Who is PTX Kernel Factory for?

It's for organizations with GPU workloads that are too slow, new GPU architectures without mature kernels, or important workloads that haven't justified weeks of specialist engineer time. This includes AI companies, GPU workload engineers, data centers, and any organization running high-performance GPU computing.

How does the self-improving system work?

Specialized agents in a fully autonomous swarm plan, implement, review, and optimize each kernel. Every candidate is compiled, verified, and benchmarked on the target GPU. Reusable evidence improves both the search process and the kernels it produces. Results, failures, and strategies become the next generation's starting point through generational memory.

Is PTX Kernel Factory open source?

The first four implementations produced by PTX Kernel Factory are open source. The product itself is entering beta and is not fully open source - users must request beta access for the full product.

What GPUs does PTX Kernel Factory support?

Based on the benchmark results shown, it supports NVIDIA GH200 (Hopper architecture) and B200 (Blackwell architecture). The system uses PTX assembly which is specifically for NVIDIA GPUs.

How do I get access to PTX Kernel Factory?

The product is entering beta with early access available. Users must request beta access through the website by clicking 'Request beta access' when bringing a hard GPU workload to the company.

What makes INT21 different from other GPU code generators?

INT21 is the first self-improving AI to generate GPU code. Unlike static generators, their system compounds improvement through reusable evidence that improves both the search process and produced kernels. They use thousands of agents in a cloud-native swarm rather than single-model approaches.

Who leads INT21?

Bing Xu is Founder & CEO, who co-authored the original Generative Adversarial Nets paper, created XGBoost's Python package, co-created MXNet and AITemplate, and was a Distinguished Engineer at NVIDIA following acquisition of his GPU inference company HippoML. Qingye Jiang is Founding Partner with over a decade building high-performance computing at AWS.

What is the scaling approach for PTX Kernel Factory?

The system uses elastic scale with thousands of agents and cloud-native scheduling that expands the swarm around available compute. All agents work toward one measurable goal with the same constraints and acceptance criteria (shared direction), and experience carries forward through generational memory where results become the next generation's starting point.

Categories

Use cases

Browse all AI tools on NeedAnAI