AA-Briefcase

A frontier knowledge work evaluation benchmark for realistic, multi-week agentic projects.

Last verified:

Visit AA-Briefcase

What is AA-Briefcase?

AA-Briefcase is a long-horizon agentic benchmark that evaluates large language models on realistic, multi-week knowledge-work projects built by industry experts. It requires models to complete coherent project workflows made of 91 tasks across four held-out scenarios (data science, product management, banking operations, heavy industry strategy), using thousands of fragmented input files such as S3 exports, emails, meeting transcripts, and large data exports. Grading combines binary rubric checks for verifiable correctness with pairwise comparisons of analytical quality and presentation quality, producing an aggregate AA-Briefcase Elo that reflects rubric pass rate, analytical Elo, and presentation Elo. The benchmark is designed for researchers, model builders, and procurement or operations teams who need a holistic measure of an agent’s ability to perform realistic professional deliverables (financial models, board decks, design mock-ups) under messy, contradictory real‑world context.

AA-Briefcase pricing

Pricing model: Freemium

The website does not sell AA-Briefcase as a subscription product; instead it reports cost-per-task estimates for evaluated models using model-specific token/pricing assumptions. AA-Briefcase shows mean cost-per-task in USD for each model (examples: Claude Fable 5 ~ $31 per task, GLM-5.2 ~ $2.40 per task, DeepSeek variants down to cents per task) so teams can compare price-performance; there is no separate free vs paid plan for the benchmark itself beyond public AA-Briefcase Lite available via Hugging Face.

AA-Briefcase pros

  • Evaluates long-horizon, multi-week project capability rather than isolated prompts
  • Uses 91 realistic tasks spanning four professional domains
  • Includes nearly 2,000 source files to test handling of fragmented context
  • Combines binary rubrics with pairwise grading for analytical and presentation quality
  • Produces a single AA-Briefcase Elo that aggregates multiple dimensions of performance
  • Scenarios and tasks built by industry experts from top firms
  • Measures cost-per-task so teams can assess price-performance tradeoffs
  • Captures realistic failure modes such as missed files or unusable deliverables
  • Public ‘Lite’ scenario available on Hugging Face for demonstration and reproducibility
  • Tracks tool-use behavior (explore/read/write/compute/view image) per task
  • Reports wall-clock time, turn counts, and token usage per task for efficiency analysis
  • Private held-out scenarios protect against contamination and gaming
  • Shows presentation quality separately, rewarding visual inspections and iterative review
  • Identifies how performance degrades as required input files increase
  • Provides methodology and leaderboard pages for transparency and ongoing updates

AA-Briefcase cons

  • Full task set and grading rubrics for the four main scenarios are private
  • Models currently complete each task in an independent run and do not carry over their own prior submissions
  • High cost variability per task (over 800x across models) complicates standardized comparisons
  • Only a small fraction of tasks are solved perfectly even by top models (e.g., 3% perfect pass rate)
  • Benchmark requires expensive compute and token usage at frontier performance levels
  • Performance depends on tool support (image viewing, compute tools) which not all models provide
  • Some important metrics (mean turns per task) are not always publicly shown
  • Public Lite scenario does not contribute to official Elo, limiting full public reproducibility

Frequently asked questions about AA-Briefcase

What types of tasks does AA-Briefcase include?

AA-Briefcase includes realistic professional deliverables across four domains: data science (transaction cleaning, forecast modeling, data engineering), product management (competitive teardowns, PRDs, go-to-market planning), banking operations (branch-network transformation, financial modeling), and heavy industry strategy (commodity scenario modeling, operating models, M&A comps). Each scenario contains multi-week workflows with tasks that build on shared context.

How is performance measured in AA-Briefcase?

Performance is measured using a combined AA-Briefcase Elo that aggregates rubric pass rate (binary checks converted to Elo) and pairwise Elo comparisons for analytical quality and presentation quality; Elo and confidence intervals are clamped at 0. This gives a holistic score reflecting verifiable correctness, analytical rigor, and presentation quality.

What sources of context do tasks require the model to use?

Tasks require models to reason across hundreds to thousands of source files per scenario, including Slack message exports (25,000+ messages), emails (3,500+), company documents, meeting transcripts, and large-scale data exports; these fragmented, sometimes contradictory inputs are designed to simulate messy real-world knowledge work.

Are the AA-Briefcase scenarios public?

No — the four primary AA-Briefcase scenarios (91 tasks) are held out and private to prevent contamination; however, a public example scenario (AA-Briefcase Lite) is available on Hugging Face to demonstrate scenario structure, submissions, and grading but it does not count toward official Elo or leaderboard results.

Can models carry forward their previous submissions across tasks?

Not currently — although scenarios are multi-week and tasks share files, each task is completed in an independent run and models do not carry over their own prior submissions between tasks in the current evaluation setup.

How does AA-Briefcase handle presentation quality?

Presentation quality is evaluated via pairwise comparisons between model submissions, producing a Presentation Elo; the benchmark also reports tool-use behavior such as view image calls, showing that higher-presentation models inspect rendered outputs more frequently before submission.

How much does it cost to run AA-Briefcase?

AA-Briefcase itself is not sold; instead the site reports estimated mean cost-per-task in USD for each model based on token usage and representative pricing (including cache hit rates). Reported examples range from cents per task for some low-cost models to over $31 per task for top-performing proprietary models.

What failure modes does AA-Briefcase reveal?

Failure modes vary by capability tier: weaker models often fail to execute tasks (missing relevant files or producing no usable deliverable), while stronger models tend to miss hidden requirements or produce incomplete analysis; common errors include incorrect or unfinished analysis and formatting mistakes across all tiers.

Is AA-Briefcase reproducible for external researchers?

AA-Briefcase provides a public Lite scenario on Hugging Face and methodology documentation, but the core held-out scenarios and grading rubrics remain private to preserve benchmark integrity, so full reproduction of official leaderboard results is not possible without internal access.

What do AA-Briefcase results tell teams deciding between models?

The benchmark highlights both capability and cost tradeoffs: frontier proprietary models may lead on Elo but incur high per-task costs, while some open-weight models (e.g., GLM-5.2) offer strong price-performance; AA-Briefcase’s cost-per-task and token/turn/time metrics help teams assess which model balances accuracy, presentation quality, and operational cost for real knowledge-work use cases.

Categories

Use cases

Browse all AI tools on NeedAnAI