Data Juicer
Data Juicer: AI-ready data processing.
Last verified:
What is Data Juicer?
Data Juicer is a data processing engine that transforms raw data into AI-ready datasets for large language models.
Data Juicer pricing
Pricing model: Freemium
Data-Juicer is completely free and open source under Apache License 2.0. There are no paid tiers or subscription plans. The system is freely available on GitHub with all features accessible without cost. Users only incur their own infrastructure costs when running distributed processing on cloud platforms like Aliyun-PAI or Ray clusters.
Data Juicer pros
- Over 200 operators spanning text, image, audio, video, and multimodal data
- One-stop system for complete data processing workflows
- 50+ reusable config recipes for common data processing tasks
- Visualization capabilities for data insights and analysis
- Auto-evaluation toolkit (GPT EVAL) for model assessment
- Distributed parallel processing with Ray, Aliyun-PAI, and CUDA support
- Operator fusion for optimized performance
- Sandbox laboratory for data-model co-development with feedback loops
- Compatible with Hugging Face datasets and LLaMA-Factory
- Apache License 2.0 open source with active community
- DJ-Cookbook with extensive demo usages and easy-start guides
- Supports both pre-training and post-tuning scenarios
- Fault-tolerant processing for robust data handling
- Adaptive resource management for efficient computing
- RESTful APIs and conversational commands for easy access
- Streaming JSON reader support integrated with Apache Arrow
- AI-powered automatic operator docstring rewriting
Data Juicer cons
- Steep learning curve for users new to data processing pipelines
- Requires familiarity with Python for customization
- Heavy resource requirements for large-scale distributed processing
- Complex configuration for advanced operator combinations
- Documentation primarily focused on technical users
- Limited pre-built templates compared to commercial tools
- Multimodal processing may require additional dependencies
- Cloud-scale processing requires access to Ray clusters or Aliyun-PAI
Frequently asked questions about Data Juicer
What is Data-Juicer?
Data-Juicer is a one-stop data processing system designed to process text and multimodal data for and with foundation models, typically LLMs. It offers over 200 built-in operators for data analysis, cleaning, synthesis, and annotation, with visualization and auto-evaluation capabilities to enable timely feedback loops for LLM pre-training and fine-tuning.
How many operators does Data-Juicer have?
Data-Juicer 2.0 features 200+ operators spanning text, image, audio, video, and multimodal data. The original version had over 50 built-in versatile operators, which has been expanded to 100+ core OPs in the systematic library, and now 200+ in the latest version.
Is Data-Juicer free to use?
Yes, Data-Juicer is completely free and open source under Apache License 2.0. All features are available without any cost or subscription.
What modalities does Data-Juicer support?
Data-Juicer supports text, image, audio, video, and multimodal data. Version 0.2.0 added video support, and the system now handles all these modalities for both pre-training and post-tuning scenarios.
How does Data-Juicer improve LLM performance?
Data recipes derived with Data-Juicer have achieved up to 7.45% increase in averaged score across 16 LLM benchmarks and 17.5% higher win rate in pair-wise GPT-4 evaluations on state-of-the-art LLMs through optimized data processing and cleaning.
Can I run Data-Juicer distributively?
Yes, Data-Juicer supports distributed data processing with optimizations for Aliyun-PAI, Ray, and CUDA. It can process 70B data samples within 2.1 hours using 6400 CPU cores on 50 Ray nodes, and deduplicate 5TB data within 2.8 hours using 1280 CPU cores on 8 Ray nodes.
What is the Data-Juicer Sandbox?
The Data-Juicer Sandbox is a feedback-driven suite for multimodal data-model co-development. It brings together model development and data preparation, enabling cost-effective iteration through a Probe-Analyze-Refine workflow where users test, evaluate, and intelligently adjust both AI models and training data based on feedback.
Does Data-Juicer work with Hugging Face?
Yes, Data-Juicer has seamless compatibility with Hugging Face dataset hubs. It was integrated into Ray's official Ecosystem and Example Gallery, and provides unified dataset format compatible with LLaMA-Factory and ModelScope-Swift for post-tuning scenarios.
How do I customize Data-Juicer for my needs?
Users can implement their own operators (OPs) for customizable data processing. The system is designed with modularity, composability, and extensibility, allowing users to freely combine and customize operators. The Developer Guide provides instructions for contributing new operators and features.
What companies and organizations use Data-Juicer?
Data-Juicer is used by AgentScope, Alibaba Group, Ant Group, BYD Auto, Bytedance, CAS, DiffSynth-Studio, EasyAnimate, Eval-Scope, JD.com, LLaMA-Factory, Nanjing University, OPPO, Peking University, RM-Gallery, RUC, Tsinghua University, Trinity-RFT, UCAS, Xiaohongshu, Xiaomi, Ximalaya, Zhejiang University, and many others.