Vefx Bench

Vefx Bench provides...

Last verified:

Visit Vefx Bench

What is Vefx Bench?

Vefx Bench is a tool that...

Vefx Bench pricing

Pricing model: Freemium

VEFX-Bench is freely available as an open-source research benchmark. The VEFX-Dataset (5,049 examples), VEFX-Reward model (4B and 32B variants), and VEFX-Bench evaluation set (300 video-prompt pairs) are all released for research use. No paid tiers or commercial licensing information is provided on the project page. The framework is intended for academic and research purposes to advance video editing technology.

Vefx Bench pros

  • 5,049 human-annotated video editing examples for comprehensive evaluation
  • 9 major editing categories with 32 detailed subcategories for broad coverage
  • Three decoupled evaluation dimensions: Instruction Following, Rendering Quality, Edit Exclusivity
  • VEFX-Reward aligns more strongly with human judgments than generic VLM judges
  • SRCC of 0.780 and Pairwise Accuracy of 0.872 with human evaluations
  • 300 curated video-prompt pairs for standardized system comparison
  • Specialized reward model designed specifically for video editing quality
  • Jointly processes source video, editing instruction, and edited video
  • Ordinal regression for per-dimension quality score prediction
  • Both 4B and 32B model variants available for different use cases
  • Cross-validation shows high annotator reliability (75.2%-87.2% exact agreement)
  • Reveals gaps between visual plausibility, instruction following, and edit locality
  • Supports group-wise preference evaluation beyond standard metrics
  • Inverse-propensity weighting for models with incomplete coverage
  • Weighted geometric mean (Overall GeoAgg) metric penalizes dimensional weaknesses
  • Benchmarks 10 representative commercial and open-source video editing systems
  • 4 FPS temporal sampling rate and ~400K pixel spatial resolution for quality

Vefx Bench cons

  • Requires significant computational resources for 32B model inference
  • only supports English language instructions based on current dataset
  • Limited to instruction-guided video editing, not general video generation
  • Training requires two-stage strategy with frozen visual tower constraints
  • Low correlation between dimensions (0.195-0.327) complicates unified scoring
  • No free tier or cloud-based API for easy access
  • Primarily research-focused, less accessible for non-technical users
  • Dataset may not cover all emerging video editing techniques
  • Hardware requirements for processing 400K pixel resolution videos

Frequently asked questions about Vefx Bench

What is VEFX-Bench?

VEFX-Bench is a holistic benchmark for generic video editing and visual effects. It includes VEFX-Dataset (5,049 human-annotated examples), VEFX-Reward (a specialized reward model), and a 300 video-prompt benchmark set for standardized comparison of editing systems.

What are the three evaluation dimensions?

The three decoupled dimensions are Instruction Following (IF) - how well the edit follows the prompt, Rendering Quality (RQ) - visual quality of the output, and Edit Exclusivity (EE) - how well non-target content is preserved. Each uses a 4-point ordinal scale.

How does VEFX-Reward compare to other evaluators?

VEFX-Reward aligns more strongly with human judgments than generic VLM judges and prior reward models. VEFX-Reward-32B achieves SRCC of 0.780, KRCC of 0.616, PLCC of 0.790, RMSE of 0.475, and Pairwise Accuracy of 0.872.

What editing categories are covered?

VEFX-Dataset covers 9 major editing categories and 32 subcategories, including object removal, background replacement, style transfer, color grading,特效 addition, and other common video editing tasks.

Can I use VEFX-Bench for commercial video editing tools?

Yes, the benchmark was used to evaluate 10 representative commercial and open-source video editing systems including Kling o3 omni and Kling o1. The framework supports benchmarking any instruction-guided video editing system.

What model sizes are available for VEFX-Reward?

VEFX-Reward is based on Qwen3-VL-Instruct architecture and comes in two sizes: 4B and 32B parameters. The 32B variant performs best overall with higher alignment to human judgments.

How was annotator reliability verified?

Cross-validation showed exact agreement of 75.2% for IF, 87.2% for RQ, and 72.2% for EE. Within-1 agreement exceeded 91% for all three dimensions, confirming high annotation reliability.

What is the Overall (GeoAgg) metric?

Overall (GeoAgg) is a weighted geometric mean that better penalizes dimensional weaknesses. It uses weights (α, β, γ) = (2, 1, 1) for IF, RQ, and EE respectively, with scores normalized to [0,1].

What are the technical specifications for evaluation?

The framework uses 4 FPS temporal sampling rate and approximately 400K pixel spatial resolution. VEFX-Reward jointly processes source video, editing instruction, and edited video using ordinal regression.

Where can I access VEFX-Bench?

VEFX-Bench is available through the project page at https://xiangbogaobarry.github.io/VEFX-Bench/. The dataset, reward model, and benchmark are released as open-source resources for research use.

Categories

Use cases

Browse all AI tools on NeedAnAI