VERITROOPER
Show HN: Audit any AI/data pairing with Veritrooper
Last verified:
What is VERITROOPER?
VERITROOPER is an audit pipeline that makes any LLM measurably more accurate on any written data by catching confident wrong answers (hallucinations), showing teams exactly where and why failures occur, and providing plain-English fixes. The tool ingests raw data (PDF, Word, HTML, CSV, JSON, plain text, SQLite, SQL dumps, DBF), generates questions with verified ground truth, then runs the model two ways: once with realistic retrieval (baseline) and once with correct source evidence (audit). The gap between these reveals the accuracy the model is leaving on the table.
Key features include per-question diagnosis with categorized failure patterns, cross-vendor verification where every contested verdict is confirmed by an independent vendor's model (the model under test never gets the final say), built-in hallucination traps, multi-hop resolution for cross-referenced sections, a financial calculator for verifying filings, reproducibility testing (ten runs with minimal drift), tamper-evident sign-off, and timestamped reproducible logs. For European deployments, one toggle adds EU AI Act conformity evidence including Article 15 accuracy/robustness testing, Annex IV technical documentation, Article 72 drift monitoring, Article 14 human-review audit trail, and Article 10 gap diagnostics.
VERITROOPER is for teams deploying LLMs on regulated or high-stakes written material—tax code professionals, workplace-safety compliance teams, healthcare/drug labeling reviewers, financial filing analysts, military regulations handlers, aviation tech orderusers, gaming rulekeepers, and any organization needing defensible accuracy evidence for auditors, buyers, or engineers. It works across any model from 7B laptop models up to frontier flagships, and has been validated across four unrelated regulated domains: IRS tax code, OSHA safety regulations, FDA drug labels, and SEC 10-K filings.
VERITROOPER pricing
Pricing model: Freemium
Pricing details are not publicly disclosed on the website. The site mentions pilots for pilot partners, acquirers, and technical reviewers. Pilots run VERITROOPER against the partner's model and corpus on an approved question set, returning the full audit with per-question failures, failure-recovery rate, and concrete fix recommendations. Public results, sample run records, and methodology need no NDA. Raw logs, the full dataset, and the patent package are shared under NDA on contact at [email protected]. No free tier or paid plan pricing is listed.
VERITROOPER pros
- Catches confident wrong answers (hallucinations) that have no error flags
- Cross-vendor verification overrides contested verdicts independently
- Per-question diagnosis with plain-English fixes engineers can implement
- Works with any LLM from 7B laptop models to frontier flagships
- Handles any written data: tax, safety, medical, financial, gaming rules
- Reproducible results—ten runs back-to-back with minimal drift under 0.5%
- Built-in hallucination traps detect when models make answers up
- Multi-hop resolution pulls cross-referenced section text into evidence
- Financial calculator parses filing tables and verifies calculations against source
- Tamper-evident sign-off ensures defensible process for auditors
- Timestamped logs make every verdict reproducible, no black box
- EU AI Act toggle adds conformity evidence (Articles 15, 72, 14, 10)
- Scoring rounds against the vendor, not inflated for marketing
- Failure-recovery rate shows how much baseline accuracy is recoverable
- Pattern clustering identifies repeated failure types across questions
- Ingests live databases including SQLite, SQL dumps, and DBF files
- No OCR needed—reads text-layer documents directly
- Category-level accuracy breakdown shows where model struggles most
VERITROOPER cons
- Does not OCR scanned images—only reads text-layer documents
- AI auditors themselves aren't perfect every single time
- Requires contacting for NDA access to raw logs, full dataset, patent package
- Does not replace provider's conformity assessment or confer compliance
- No self-hosted option mentioned—appears to be service-based only
- Pricing details not publicly disclosed on website
- Live walkthroughs only by request, not instant access
- Single founder (Brian B.)—small team capacity compared to enterprises
Frequently asked questions about VERITROOPER
What is VERITROOPER and what problem does it solve?
VERITROOPER makes any LLM measurably more accurate on any written material and proves exactly how. It solves AI hallucination—when an LLM confidently produces a wrong answer with no flag, warning, or error code. The tool generates questions with verified ground truth from your data, runs the model with realistic retrieval (baseline) and with correct source evidence (audit), then measures the gap. Failments route to diagnostic modules, contested verdicts get cross-vendor verified, and you receive a plain-English report on what went wrong, why, and exactly what to fix.
How does cross-vendor verification work?
Every contested verdict is confirmed by an independent cross-vendor verifier using a different vendor's model. The model under test never gets the final say on its own answers—the independent verifier can override it. This ensures the audit doesn't inherit the original model's errors and provides defensible evidence for auditors and buyers.
What data formats does VERITROOPER accept?
VERITROOPER ingests PDF, Word, HTML, CSV, JSON, plain text, and live databases including SQLite, SQL dumps, and DBF. It chunks and parses automatically. However, it only reads text-layer documents and does not OCR scanned images.
Which domains has VERITROOPER been validated on?
VERITROOPER has been run end-to-end across four unrelated regulated worlds: U.S. tax code (IRS federal income-tax regulations), OSHA workplace-safety regulation (29 CFR general industry/construction/hazmat), FDA drug labeling (high-alert and common prescription drugs), and SEC 10-K financial filings (Apple, NVIDIA, JPMorgan, Coca-Cola). Each used the same 1,000-question audit and seven models.
What models does VERITROOPER support?
VERITROOPER works with any LLM from a 7B laptop model (running on a $2,000 RTX 4090 gaming GPU) up to flagship frontier APIs. Seven models were tested: Claude Opus 4.8, GPT-5.5, Gemini 2.5 Pro, Qwen 2.5 72B, Llama 3.1 70B, Gemma 3 27B, and Qwen 2.5 7B.
How reproducible are VERITROOPER's results?
The exact same audit was run ten times in a row with the same 238 questions and same model on two side-by-side servers. The core check came back to the exact same number all ten times (97.48%). The full audited score ranged 97.0–97.5%, with best vs. worst run under half a point. It never drifted meaningfully, proving the measurement is reliable rather than a guess.
What does the EU AI Act toggle add?
One toggle adds EU AI Act conformity evidence to the same run: declared accuracy and robustness testing (Article 15), a drop-in Annex IV technical-documentation record, recurring accuracy-drift monitoring (Article 72), a dated signed human-review audit trail (Article 14), and a per-category gap diagnostic (Article 10). VERITROOPER produces evidence a conformity file relies on but does not replace the provider's conformity assessment or confer compliance.
What output does VERITROOPER provide?
You get an adjusted accuracy score against ground truth, a per-question categorized list of every failure with patterns and clusters identified, a per-category accuracy breakdown, a failure-recovery rate (share of baseline failures the audit recovered), and concrete engineering recommendations to close specific gaps. Every verdict is reproducible from timestamped logs.
How do I start a pilot or get technical review access?
Contact [email protected] for acquirers, pilot partners, and technical reviewers. Live walkthroughs are available by request. A pilot runs VERITROOPER against your model and corpus on a question set you approve, returning the full audit. Public results, sample run records, and methodology need no NDA; raw logs, full dataset, and patent package are shared under NDA.
Who created VERITROOPER and what's the backstory?
Brian B. (Brian Barbour), an average guy, US Army and Air Force veteran, disabled vet, created VERITROOPER in his basement with a PC. It was the unintended discovery while trying to make a D&D game in Unity with an LLM dungeon master. The AI couldn't catch him lying—feeding it false gold amounts, fake spells, impossible actions—and he wanted a DM that called crap. This led to solving hallucination, which generalized from D&D SRD to tax code, military law, financial filings, and Air Force training material.