CriteriaBot

A Universal Customizable Classifier

Last verified:

Visit CriteriaBot

What is CriteriaBot?

CriteriaBot is a programmatic content evaluation API that lets you define plain-English criteria and evaluate any content against them via a single API call. The core engine, called Arbiter, is a panel of about a dozen LLMs and ML classifiers that vote to produce a weighted consensus true/false verdict for each criterion. Unlike fixed-category moderation APIs, every criterion is custom-written by you, enabling you to check grammar, brand tone, compliance rules, toxicity, spam, prompt injection, or anything else you can describe in text.

Key features include: batch evaluations (up to 25 at a time), async evaluations with polling or webhooks, criteria groups for bundling related criteria, BYOK (bring your own key) to use any supported LLM provider, and a personalization system that learns your judgment from a few示例 verdicts. Pro and Enterprise plans receive a custom LoRA fine-tuned on your examples for deeper alignment. The Arbiter grounds verdicts in real-world evidence by pulling facts from sources like Wikipedia and Wolfram Alpha before models vote.

CriteriaBot is designed for teams building content moderation systems, spam/abuse detection, AI guardrails for LLM apps, brand voice consistency checks, compliance screening, brand & PR monitoring, routing/triage pipelines, and data labeling for datasets. It delivers 91.67% accuracy at $3.20 per 1,000 verdicts, outperforming flagship models like GPT-5.5 and Claude Opus 4.8 while costing less than half.

CriteriaBot pricing

Pricing model: Freemium

Free: $0/month - 1,000 Arbiter verdicts/month (no keys required), full access to predefined criteria library, 10 custom criteria. Starter: $40/month - 12,500 Arbiter verdicts/month, unlimited custom criteria, BYOK (bring your own key) for any supported LLM provider. Pro: $200/month - 70,000 Arbiter verdicts/month, dedicated model fine-tuned on your verdicts (custom LoRA), BYOK & unlimited custom criteria. Credits: $10 one-time for 2,500 Arbiter verdict credits, stack on top of any plan, never expire. Enterprise: custom pricing for higher volumes, priority fine-tipping, and custom data sovereignty requirements - contact sales.

CriteriaBot pros

  • 91.67% accuracy outperforms flagship models like GPT-5.5 and Claude Opus 4.8
  • $3.20 per 1,000 verdicts is significantly cheaper than GPT-5.5 at $7.70
  • Plain-English criteria let you define custom rules for any use case
  • Arbiter uses a panel of ~12 LLMs/ML classifiers for weighted consensus
  • Grounds verdicts in real-world facts from Wikipedia and Wolfram Alpha
  • Batch evaluations support up to 25 content pieces in one request
  • Async evaluations with polling or webhook support for fire-and-forget workflows
  • Criteria groups bundle related criteria for easy reuse across requests
  • BYOK lets you bring your own API keys for any supported LLM provider
  • Personalization learns your judgment dynamically from just a few example verdicts
  • Pro plans include a dedicated LoRA model fine-tuned on your verdicts
  • Free tier includes 1,000 Arbiter verdicts/month with no API keys required
  • Full library of predefined criteria available immediately
  • Content encrypted at rest, never sold/shared, and never human-reviewed
  • Cancel anytime and keep access through the end of the billing period
  • Credits (2,500 verdicts for $10) stack on top of plans and never expire
  • Enterprise offers zero-retention options and dedicated infrastructure deployment

CriteriaBot cons

  • Free tier limited to 1,000 verdicts/month and only 10 custom criteria
  • Starter plan at $40/month may be expensive for very small projects
  • Personalization requires you to issue your own verdicts to train the system
  • BYOK requires you to manage your own LLM provider API keys
  • Async evaluations require polling or webhook setup for results
  • Enterprise features need direct sales contact, not self-signup
  • No explicit SDK support mentioned, only REST API with curl examples
  • Model panel composition changes as models are added/removed, reducing predictability

Frequently asked questions about CriteriaBot

What counts as a verdict?

One verdict is a piece of content evaluated against a single criterion. Checking three criteria on one comment uses three verdicts, so you can size a plan straight from your expected volume.

What happens if I hit my monthly quota?

Top up any time with a credit pack - credits stack on top of your plan and never expire - or move to a higher tier. Once you're out, requests return a clear 'quota exceeded' response, so you always know when to top up.

Which models power the Arbiter?

Arbiter verdicts are formed from a panel of about a dozen LLMs and ML classifiers. We add new models as they prove out, and swap out models that don't perform well. The Arbiter learns to draw conclusions from the consensus, allowing the success rate to exceed any single model or even a standard weighted consensus.

How does personalization work?

You teach it by example. When you issue your own verdicts, the Arbiter learns which models tend to agree with you for which types of evaluations and where your sensibilities may differ. Personalization is dynamic, and starts impacting results from the first example. Pro plans include a dedicated model retrained on your verdicts for even deeper adaptation.

How is my data handled?

Content is encrypted at rest, never sold or shared, and never human-reviewed. The only outside services that see a request are the LLM providers in your panel - all chosen for not training on your data, and BYOK keeps content in your own account. Delete your data on request or at account closure; Enterprise can run zero-retention, or have the open-weight models deployed on dedicated infrastructure we run just for them.

Can I cancel anytime?

Yes. Manage or cancel your subscription whenever you like and you keep access through the end of the billing period. Any one-time credits you've purchased stay yours.

How do I get my API token?

Create a free account, then open Settings → API Tokens and generate a token. Copy it somewhere safe - you won't be able to see it again. Store it in an environment variable like CRITERIA_BOT_API_TOKEN for use in API requests.

What is a criterion in CriteriaBot?

A criterion is a plain-English statement that describes what you're looking for. When CriteriaBot evaluates content it asks the AI models whether the content meets this criterion: true if it does, false if it doesn't. You create criteria via POST /v1/criteria with a name and body describing the rule.

How do I run my first evaluation?

Pass your content and one or more criteria to POST /v1/evaluations. CriteriaBot routes each request through the Arbiter - a multi-model consensus engine that evaluates the content against each criterion independently. The response includes state (completed/evaluating/failed) and verdicts with true/false for each criterion_id.

What can I build with CriteriaBot?

People build content moderation (toxicity, harassment, unsafe content plus house rules like spoiler bans), spam & abuse detection, AI guardrails (block prompt injection and unsafe outputs), brand voice consistency, compliance checks, brand & PR monitoring, routing & triage for messages/tickets, and data labeling for building datasets or filtering large content sets.

Categories

Use cases

Browse all AI tools on NeedAnAI