Libretto
Refine, test, and optimize AI prompts efficiently.. [Contact for Pricing]
Last verified:
What is Libretto?
Libretto is a powerful tool designed for software developers to monitor, test, and optimize LLM prompts with an emphasis on automating the most cumbersome tasks in AI development. It helps you identify weaknesses in your AI products and fix them before they become a headache by integrating into your application and analyzing your LLM traffic to highlight calls that are toxic, unhelpful, or poor quality.
The platform offers comprehensive monitoring with SOC2-compliant infrastructure that automatically flags potential errors in production LLM calls. Libretto generates prompt evals and test sets automatically by sampling your production traffic, creating both test cases and evaluation criteria tailored to your specific prompts. Its Experiments feature automatically explores dozens or even hundreds of prompt variations to find the top performer, potentially boosting prompt accuracy by 10% in minutes without guesswork.
Key features include Drift Detection that tests your prompt daily to identify if your model is giving different answers than it used to, real-time alerts about AI performance, cost and usage monitoring, and seamless integration via a drop-in SDK. Libretto also offers LLM-as-Judge for automated evaluation, tuned evals that align AI testing with human judgment by training on your standards, and custom eval creation that understands your prompts to build specific tests for what they're trying to achieve.
Libretto is ideal for software developers building AI products, engineering teams working with LLMs in production, and anyone who needs empirical evidence that their AI works correctly. It transforms prompt engineering from a cumbersome empirical endeavor into a streamlined, automated, and empirically grounded process, making it perfect for teams moving from prototype to production or those needing to continuously improve AI in their applications.
The tool seamlessly integrates with existing workflows, automatically flagging issues, generating test cases, and crafting evals so you can continuously improve AI in your product and make sure every change to a prompt or model makes your product better, not worse.
Libretto pricing
Pricing model: Freemium
Libretto offers a free trial so you can explore all features and experience firsthand how it enhances AI development. The platform provides comprehensive monitoring, automated testing, and real-time alerts. Contact the company for detailed pricing on paid plans as specific pricing tiers are not publicly listed on the website. The free trial allows you to test the full capabilities including prompt evals, test set generation, drift detection, and experiments before committing.
Libretto pros
- Automatically generates test sets from production traffic
- Creates custom evals tailored to each prompt
- Experiments feature explores hundreds of prompt variations automatically
- Drift Detection tests prompts daily to catch model changes
- SOC2-compliant monitoring infrastructure
- Real-time alerts for toxic, unhelpful, or poor quality calls
- Drop-in SDK for minutes-long integration
- Monitors costs and usage alongside performance
- Tuned evals align with human judgment after showing 10 examples
- Automatically flags unsafe or toxic LLM calls
- No guesswork in finding best prompt variants
- Quickly identifies weaknesses in AI products
- Supports both subjective and objective evaluations
- Full control to customize or modify auto-generated evals
- Works with existing LLM infrastructure seamlessly
Libretto cons
- Beta product may have limited stability
- Requires production traffic to generate test cases
- Primarily designed for developers not non-technical users
- May need time to train evaluators on your standards
- Focus on LLM prompts may not help with other AI issues
- Auto-generated evals may require manual adjustment
- Best results come after integrating with live application
- Documentation may be limited for new beta users
- Requires API integration which adds development overhead
- May not support all LLM providers equally
- Pricing details not publicly transparent
- Learning curve for maximizing all features
- Dependent on having sufficient production data
- May not replace all manual testing needs entirely
- Some features may change during beta period
Frequently asked questions about Libretto
What is Libretto?
Libretto is a powerful tool designed for software developers that helps monitor, test, and optimize LLM prompts with an emphasis on automating the most cumbersome tasks. You integrate Libretto into your application and it analyzes your LLM traffic, highlighting calls that are toxic, unhelpful, or poor quality so you can quickly identify problems with your prompt and model and fix them fast.
Who is Libretto for?
Libretto is designed with developers in mind and is focused on automating away the drudgery of AI development. It's ideal for software developers building AI products, engineering teams working with LLMs in production, and anyone who needs to empirically test and improve their AI prompts rather than hoping they work.
How does Libretto generate test cases?
Libretto samples your production traffic to automatically generate test cases. This ensures you have a comprehensive set of cases covering actual use cases, edge cases, and failure modes without spending hours inventing test scenarios manually.
What are evals in Libretto?
Evaluations, or evals, are tests used to assess the quality of your LLM's responses. They can be both qualitative and quantitative, providing a well-rounded view of performance. Libretto reads and understands your prompts then builds specific tests for exactly what they're trying to achieve, and you can add, remove, or modify evals as you see fit.
What is Drift Detection?
Drift Detection tests your prompt daily to see if your model is giving different answers than it used to. This ensures you can be sure whether or not the model changed without telling you, addressing the concern that a model working great last month might not work as well today.
How does the Experiments feature work?
Experiments is Libretto's prompt engineering tool that automatically explores dozens or even hundreds of prompt variations to find the top performer. It generates variants by combining and recombining test case examples, then tests them against each other letting the best prompts win, potentially boosting accuracy by 10% in minutes without guesswork.
Can I customize the auto-generated evals?
Absolutely. While evals are created automatically, you have full control to add custom evaluation rules, adjust quality thresholds, define what good means for your use case, train evaluators on your standards, and exclude or modify any auto-generated evals.
Is there a free trial available?
Yes, Libretto offers a free trial so you can explore all features and experience firsthand how it enhances your AI development process. You can sign up to get started and explore the tool's capabilities before committing to a paid plan.
What kind of support does Libretto offer?
Libretto offers comprehensive support through documentation and customer service. You can reach out anytime for assistance, and they set up joint Slack channels with customers for immediate access to the Libretto team.
How quickly can I integrate Libretto?
You can connect Libretto into your product in minutes using their drop-in SDK. It seamlessly integrates with your existing workflow and LLM infrastructure, automatically flagging issues, generating test cases, and crafting evals without requiring major changes to your setup.