PeriFlow
FriendliAI is a leading provider of a Language Learning Model (LLM) serving engine designed specifically for generative AI. Their primary objective is to aid cl...
Last verified:
What is PeriFlow?
PeriFlow (now rebranded as Friendli Engine) is the fastest generative AI inference serving engine on the market, designed to optimize deployment and serving of Large Language Models (LLMs). It significantly reduces the number of GPUs required for serving generative AI by up to 10x, cutting GPU costs by 70-90% while maintaining low latency and high throughput. The engine is built on FriendliAI's patented batching and scheduling techniques, including U.S. Patent No. 11,514,370 and U.S. Patent No. 11,442,775.
Key features include continuous batching for high throughput, iteration-level scheduling, multi-GPU and multi-node execution, quantization support (including AWQ 4-bit quantization), multi-adapter support for LoRAs, multi-modal model support handling text and images, and powerful state caching for performance boosts. PeriFlow supports a broad range of LLMs including GPT, GPT-J, GPT-NeoX, MPT, LLaMA, Dolly, OPT, BLOOM, CodeGen, T5, FLAN, and UL2. It offers diverse decoding options (greedy, top-k, top-p, beam search, stochastic beam search) and supports multiple data types (fp32, fp16, bf16, int8).
PeriFlow is offered in two forms: PeriFlow Container for serving LLMs in private environments, and PeriFlow Cloud (now Friendli Dedicated Endpoints) for serving LLMs on autopilot in managed cloud environments. The tool is designed for organizations that develop their own LLMs through pretraining or fine-tuning, AI developers, companies serving generative AI models at scale for large user bases, and anyone looking to deploy custom or open-source generative AI models efficiently without owning or managing GPU infrastructure.
PeriFlow pricing
Pricing model: Free
PeriFlow (Friendli Engine) is a paid tool with detailed pricing available on the official website. The service offers two deployment options: PeriFlow Container and PeriFlow Cloud (now Friendli Dedicated Endpoints). For Dedicated Endpoints, pricing is per GPU-hour billed per second: A100 80GB at $2.9/hour, H100 80GB at $3.9/hour, H200 141GB at $4.5/hour, and B200 at $8.9/hour. Users get $5 in free trial credits when signing up for Serverless Endpoints (Trial plan) or Dedicated Endpoints (Basic plan). Enterprise plans offer customizable discounted pricing through contacting sales. Serverless Endpoints use tier-based pricing based on lifetime spending (Tier 0: signed up, Tier 1: $10+ spend, Tier 2: $50+ spend, Tier 3: $500+ spend, Tier 4: $5,000+ spend, Tier 5: custom). Text models are charged per token (e.g., Llama-3.1-8B-Instruct at $0.1/1M tokens, Llama-3.3-70B-Instruct at $0.6/1M tokens), while audio models are charged per audio minute (Whisper large v3 at $0.0015/audio minute).
PeriFlow pros
- Reduces GPUs required for serving by up to 10x
- 70-90% GPU cost savings for LLM serving
- 40-80% reduction in LLM serving costs
- Fastest generative AI inference serving engine on the market
- 10x throughput improvement for GPT-3 175B at same latency
- Patented batching and scheduling techniques with US and Korean patents
- Supports quantization including AWQ 4-bit for 70B Llama 2 on single A100
- Multi-adapter support for running multiple LoRAs on single GPU
- Multi-modal support for text and image inputs
- Powerful state caching for drastic performance boosts
- Centralized management of every deployed LLM from anywhere
- Auto-scaling based on traffic patterns
- Dynamic fault handling and performance monitoring
- Integrated playground for interactive LLM testing
- No cloud resource setup and management hassle with PeriFlow Cloud
- Deploy LLMs in matter of minutes
- Supports wide range of popular open-source models including Llama 3
- Comprehensive monitoring tools for events, errors, and performance metrics
PeriFlow cons
- Paid tool only, no free tier for PeriFlow Engine itself
- Detailed pricing only available on official website
- Requires GPU infrastructure (container or cloud)
- Costs accumulate even when endpoint not serving API calls
- Each autoscaling replica increases total cost proportionally
- Enterprise pricing requires contacting sales
- Limited documentation publicly available compared to open-source alternatives
- Rebranding from PeriFlow to Friendli may cause confusion for existing users
Frequently asked questions about PeriFlow
What is PeriFlow and what does it do?
PeriFlow (now rebranded as Friendli Engine) is the fastest generative AI inference serving engine on the market. It optimizes deployment and serving of Large Language Models (LLMs), reducing the number of GPUs required for serving by up to 10x while cutting costs by 70-90%. The engine achieves remarkable improvements in throughput while maintaining low latency through patented batching and scheduling techniques.
Which LLMs does PeriFlow support?
PeriFlow supports a broad range of LLMs including GPT, GPT-J, GPT-NeoX, MPT, LLaMA, Dolly, OPT, BLOOM, CodeGen, T5, FLAN, UL2, and more. It specifically supports popular open-source models like Llama 3, and can handle models ranging from 1.3B to 341B in size.
What are the two ways to use PeriFlow?
You can use PeriFlow in two ways: PeriFlow Container for serving your LLMs with Friendli Engine in your private environment, and PeriFlow Cloud (now Friendli Dedicated Endpoints) for serving your LLMs on autopilot in a managed cloud environment without hassle of cloud resource setup and management.
How do I increase my rate limits on Friendli?
Your usage tier, which determines your rate limits, increases monthly based on your proof-of-payment. As your usage grows, your tier increases automatically. For faster upgrade, you can reach out to [email protected] anytime, or move up instantly by purchasing additional credits.
Do I need to upgrade my plan to use popular models?
No, popular models are available to all users depending on the limits determined by their usage tiers. Tier 0 (signed up) gets adaptive rate limits, while higher tiers (Tier 1-5) get progressively higher RPM limits from 60-100 RPM up to custom limits.
What happens if I exceed my monthly cap?
You'll receive an alert when approaching your monthly cap. Contact [email protected] to discuss options for increasing your monthly cap. They may help you pay early to reset your monthly cap or upgrade your plan to increase the monthly cap and unlock more features.
How does billing work for Dedicated Endpoints?
Users are billed by GPU-second for the duration that the endpoint is active. Charges begin when the endpoint is up and running, and costs accumulate even when the endpoint is not serving API calls. When an endpoint goes to sleep after being idle, charges will no longer accrue. Updating an endpoint or employing certain settings keep it active and charges continue to accrue.
How does autoscaling affect my costs?
Each additional replica increases your total cost proportionally. For example, scaling from 1 to 2 replicas doubles your GPU costs. This means if you have an A100 endpoint at $2.9/hour with 1 replica, adding a second replica will cost $5.8/hour total.
What decoding options does PeriFlow offer?
PeriFlow offers diverse decoding options including greedy decoding, top-k sampling, top-p sampling (nucleus sampling), beam search, and stochastic beam search. These options allow users to optimize the balance between precision and speed for their specific use cases.
What data types does PeriFlow support?
PeriFlow supports multiple data types including fp32 (32-bit floating point), fp16 (16-bit floating point), bf16 (Bfloat16), and int8 (8-bit integer). This support for various quantization levels helps users optimize the balance between precision and speed.