Rebellions.ai

Energy-efficient, high-performance AI chips for generative applications.. [Contact for Pricing]

Last verified:

Visit Rebellions.ai

What is Rebellions.ai?

Rebellions.ai — ebellions AI is an AI semiconductor company that develops energy-efficient AI inference accelerators and a complete software stack for large-scale deployment. The company's flagship products include the REBEL-Quad AI accelerator card (the world's first UCIe-Advanced chiplet-based AI SoC with 144GB HBM3E memory) and the RebelServer™ (a 5U server with 8 RebelCards™ delivering up to 2 PFLOPs FP8 performance). Their ATOM™ family of NPUs provides 32 TFLOPS FP16 and 128 TOPS INT8 performance for AI inference workloads.

The Rebellions SDK provides a comprehensive software stack including the RBLN Compiler, Runtime Module, Compute Library, Driver, and Firmware. It supports over 300 models from Hugging Face, PyTorch 2.0, TensorFlow, vLLM, and Nvidia Triton Inference Server. The SDK includes tools like the RBLN Profiler for performance analysis, rbln-stat for monitoring power/utilization, and Kubernetes plugins for enterprise deployment. The Model Zoo provides officially supported models including Llama3-8B, SDXL, ResNet, YOLO, and BERT.

Rebellions targets data centers, enterprises, cloud providers, and AI developers who need energy-efficient large-scale AI inference. The company emphasizes superior performance-per-watt compared to GPU alternatives, with REBEL-Quad demonstrating ~1.6x higher throughput than NVIDIA H200-class GPUs on Llama 3.3 70B while consuming significantly less power. The chiplet architecture enables seamless scaling from single-server to rack-scale and sovereign AI deployments.

Key use cases include LLM serving, multimodal inference, Physical AI applications (like water purification robots and veterinary imaging), AI-powered HS code recommendation chatbots, construction safety management, and threat detection in SOCs. The platform supports mixed-precision execution (FP8, FP16, INT8), tensor parallelism for multi-device inference, and continuous batching for LLM serving optimization.

Rebellions.ai pricing

Pricing model: Freemium

Rebellions does not publish public pricing for their hardware or software. Pricing and procurement are handled entirely through enterprise/direct sales contact. The RebelServer™ comes with 3-year business standard hardware and software support included. There is no publicly available free tier or subscription model - the company focuses on enterprise sales for data center deployments. Interested buyers must contact sales through the website for pricing on REBEL-Quad accelerator cards, RebelServer systems, and ATOM-Max server/pod systems.

Rebellions.ai pros

  • World's first UCIe-Advanced AI accelerator with chiplet architecture
  • 144GB HBM3E memory with 4.8TB/s bandwidth on REBEL-Quad
  • Up to 2 PFLOPs FP8 performance per RebelServer with 8 cards
  • ~1.6x higher throughput than NVIDIA H200 on Llama 3.3 70B
  • Significantly better performance-per-watt than GPU alternatives
  • Supports over 300 models including Llama3, SDXL, YOLO, BERT
  • Native PyTorch 2.0, vLLM, and Triton compatibility
  • No code modifications needed for most Model Zoo models
  • Complete SDK with compiler, runtime, profiler, and monitoring tools
  • Kubernetes plugin support for enterprise cluster deployment
  • Tensor parallelism (RSD) for distributed multi-device inference
  • Mixed-precision support: FP8, FP16, BFloat16, INT8 quantization
  • C/C++ runtime API for low-latency applications without Python
  • 3-year business standard hardware and software support included
  • Seamless scaling from single server to rack-scale clusters
  • Built on Samsung 4nm process for frontier-scale AI models
  • Supports MoE (Mixture of Experts) architectures at peta-scale

Rebellions.ai cons

  • Inference-only SDK - no model training or fine-tuning support
  • Linux-only support - Windows not available yet
  • Pricing not publicly listed - enterprise sales contact required
  • Limited technical support for non-Model Zoo modified models
  • Support for original ATOM (RBLN-CA02) ended June 2025
  • Requires root privileges for driver installation
  • Python 3.9+ required with dependencies like numpy, torch, onnx
  • Compilation may fail for models not in official Model Zoo
  • Monthly SDK updates and quarterly driver updates may cause version mismatches
  • Optimal batch size requires experimentation with Profiler tool

Frequently asked questions about Rebellions.ai

Which AI frameworks and libraries does Rebellions SDK support?

Rebellions SDK supports models based on PyTorch and TensorFlow and is also compatible with Hugging Face Transformers/Diffusers libraries. The SDK is continuously updated to maintain compatibility with major AI frameworks through regular updates.

Can I compile PyTorch or TensorFlow models without code modifications?

In most cases, you can use Rebellions SDK with minimal code changes. For officially supported Model Zoo models, you can use the provided example code right away. Other models can also be compiled by referring to the Model Zoo code.

Does Rebellions SDK support Windows?

Currently, Rebellions SDK only supports Linux. Windows support will be determined based on the company's technical roadmap. More details on supported OS and Python versions can be found on the Support Matrix page.

Can I train models with Rebellions SDK?

The current Rebellions SDK is designed for inference-only use. Plans for training support will be announced through the roadmap once they are finalized. The devices are designed exclusively for inference, and fine-tuning is not currently supported.

How are ATOM and REBEL different?

Both are AI inference NPUs developed by Rebellions, but REBEL is a next-generation product designed with a chiplet-based architecture. REBEL-Quad uses UCIe-Advanced interconnect and 144GB HBM3E, while ATOM is manufactured on Samsung's 5nm process with 64MB on-chip SRAM.

Does RBLN Runtime API support C/C++?

Yes, Rebellions SDK provides a C/C++-bound runtime for applications where Python runtime is unavailable or extremely low latency is required. Refer to the C/C++ guide for more information on usage.

Can I run inference on multiple devices?

Yes, Rebellions SDK supports distributed inference based on tensor parallelism, called RSD (Rebellions Scalable Design). First check the Model List that supports multi-device, and refer to the provided example for compilation instructions.

Do you support Kubernetes?

Yes, you can use Rebellions AI processor resources via the Kubernetes Plugin. Tools include Kubernetes Device Plugin for RBLN NPUs, NPU Feature Discovery for node labeling, and RBLN Metrics Exporter for Prometheus-format metrics exposed to Grafana dashboards.

Which NPUs are officially supported by Rebellions SDK?

As of May 30, 2025, the SDK supports ATOM™+ (RBLN-CA22) and ATOM™-Max (RBLN-CA25). Support for the original ATOM™ (RBLN-CA02) ended on June 30, 2025.

How are NPUs different from GPUs?

GPUs were designed for graphics rendering but adopted for AI training/HPC with FP32/FP16 operations and CUDA cores. NPUs are processors specialized for AI and deep learning, optimized for low-bit operations like INT8 and FP16 with dedicated hardware architectures that accelerate neural network computations at low power.

Categories

Use cases

Browse all AI tools on NeedAnAI