Chitu

High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.

Last verified:

Visit Chitu

What is Chitu?

Chitu (赤兔) is a high-performance large model inference framework focused on efficiency, flexibility, and usability. Positioned as an enterprise-grade large model inference engine, it is designed for production environments and handles the progressive needs from small-scale trials to large-scale deployments in enterprise AI implementation.

Key features include multi-hardware compatibility supporting NVIDIA's latest flagship to legacy product lines plus optimized support for non-NVIDIA chips like Huawei Ascend, Muxi, and Hygon. It offers full-scenario scalability from pure CPU deployment to single GPU deployment and large-scale cluster deployment. The framework supports mainstream models including DeepSeek, Qwen, GLM, Kimi, and Meta's Llama series, with efficient operator implementations for quantization (FP4, FP8, BF16) and CPU+GPU heterogeneous hybrid inference.

Chitu is intended for enterprises and developers who need to deploy large language models in production environments. It is particularly valuable for organizations looking to reduce GPU costs, deploy on domestic Chinese chips to reduce Nvidia dependence, and achieve high-performance low-latency inference for LLMs. The project is open-source under Apache License v2.0 and developed by Qingcheng.AI in collaboration with Tsinghua University.

Chitu pricing

Pricing model: Freemium

Chitu is open-source software available under Apache License v2.0 and is free to use. The framework can be deployed using official Docker images provided at no cost for NVIDIA, Muxi, and Ascend platforms. Professional technical services are available via email at [email protected] for enterprises needing dedicated support, but specific pricing for professional services is not publicly listed on the website.

Chitu pros

  • Supports NVIDIA GPUs from latest flagship to legacy product lines
  • Optimized support for non-NVIDIA Chinese chips (Ascend, Muxi, Hygon)
  • 315% faster inference speed on NVIDIA A800 compared to foreign frameworks
  • Reduces GPU usage by 50% while maintaining performance
  • Supports DeepSeek-R1 671B with single-card inference capability
  • CPU+GPU heterogeneous hybrid inference support
  • Full-scenario scalability from CPU-only to large-scale clusters
  • Native FP8 precision support on non-NVIDIA Hopper architecture GPUs
  • Compatible with DeepSeek, Qwen, GLM, Kimi, and Llama models
  • Efficient FP4→FP8/BF16 online conversion operators
  • Suitable for production environments with concurrent business traffic
  • Long-term stable operation capability
  • Open-source under Apache License v2.0
  • Official Docker images available for quick deployment
  • Extensible solution with adapters and plugins for mainstream LLMs

Chitu cons

  • Limited team capacity cannot guarantee timely resolution of all user issues
  • Professional technical services require email contact (not immediate support)
  • Performance results vary based on hardware configuration and software versions
  • Some Ascend A3 images no longer maintained due to hardware lack
  • Primarily designed for Chinese market and domestic chips
  • Documentation partially machine-translated from Chinese
  • Requires technical expertise for deployment and configuration
  • Benchmark data is self-tested by development team only

Frequently asked questions about Chitu

What is Chitu AI?

Chitu (赤兔) is a high-performance large model inference framework focused on efficiency, flexibility, and usability. It is positioned as an enterprise-grade large model inference engine developed by Qingcheng.AI in collaboration with Tsinghua University, designed for production environments handling enterprise AI deployment from small-scale trials to large-scale clusters.

What models does Chitu support?

Chitu supports mainstream large language models including DeepSeek (including DeepSeek-R1 671B), Qwen series (including Qwen3), GLM (including GLM-4.5 MoE), Kimi, and Meta's Llama series. The framework uses adapters and plugins for compatibility with mainstream LLMs.

What hardware does Chitu support?

Chitu supports NVIDIA GPUs from latest flagship to legacy product lines, plus optimized support for non-NVIDIA chips including Huawei Ascend 910B (A2, A3), Muxi (MetaX), Hygon (Hygon), and Enflame (Ziang). It also supports pure CPU deployment and CPU+GPU heterogeneous hybrid inference.

Is Chitu free to use?

Yes, Chitu is open-source software released under Apache License v2.0 and is free to use. Official Docker images are also provided at no cost for various platforms including NVIDIA, Muxi, and Ascend.

How much faster is Chitu compared to other frameworks?

According to development team testing, when deployed on NVIDIA A800 GPUs with DeepSeek-R1, Chitu achieved 315% faster inference speed while cutting GPU usage by 50% compared to foreign open-source frameworks. Performance varies based on hardware configuration, software versions, and test workloads.

Can I run DeepSeek-R1 671B on a single GPU?

Yes, Chitu v0.2.2 added CPU+GPU heterogeneous hybrid inference support enabling single-card inference for DeepSeek-R1 671B. The framework also supports FP4 quantized version of DeepSeek-R1 671B with efficient operator implementations.

How do I install Chitu?

For quick validation in standalone environments, use official Docker images: NVIDIA (arch 8.0, 8.9): qingcheng-ai-cn-beijing.cr.volces.com/public/chitu-nvidia_arch_80_89:latest, NVIDIA (arch 9.0): chitu-nvidia_arch_90:latest, Muxi: chitu-muxi:latest, Ascend A2: chitu-ascend_a2:latest. Complete installation instructions are in the Developer Manual.

What is the latest version of Chitu?

The latest release is v0.5.1 released on February 6, 2026, which added support for MooreThreads GPUs. Version v0.5.0 was released on December 12, 2025, focusing on improving performance on cluster deployment scenarios.

How do I get technical support for Chitu?

For questions or concerns, submit GitHub issues. For professional technical services, email [email protected]. The team notes that limited capacity means they cannot guarantee timely resolution of all user issues, but they appreciate feedback and continue improving the engine.

Can Chitu run on Chinese domestic chips?

Yes, Chitu provides optimized support for Chinese domestic chips including Huawei Ascend 910B (first supported in v0.3.9 for GLM-4.5 MoE), Muxi/MetaX, Hygon, and Enflame. This helps reduce dependence on NVIDIA hardware and breaks the hardware binding dilemma for domestic AI chip ecological construction.

Categories

Use cases

Browse all AI tools on NeedAnAI