Splabs

A behavioural health monitor for LLMs

Last verified:

Visit Splabs

What is Splabs?

Splabs PSA (Posture Sequence Analysis) is a behavioral‑analysis platform that measures what language models do from the outside, without requiring access to model weights or internals. It uses a suite of classifiers and graph‑based analytics to generate deterministic metrics, posture fingerprints, and risk signals for any model response text, enabling safety and alignment monitoring across individual turns and long‑running sessions. The system is designed to support continuous monitoring of production‑grade LLM deployments, including stress‑testing, drift detection, and forensic incident logging.

Key features include a stack of five micro‑classifiers that analyze each sentence for input intent, adversarial stress, sycophancy, hallucination risk, and persuasion techniques, plus higher‑level session metrics such as Behavioral Health Score (BHS), Posture Oscillation Index (POI), and Dissolution Position Index (DPI). PSA also provides regime‑shift detection (progressive drift, acute collapse, and oscillation patterns), multi‑agent agentic‑graph analysis, and a SIGTRACK archive that stores posture‑sequence incidents without raw text, enabling privacy‑compliant forensics and replay of safety events.

Splabs PSA is primarily aimed at AI safety teams, platform operators, and developers running large‑scale LLM applications who need external, black‑box monitoring of model behavior over time. It suits organizations that must audit compliance with guardrails, detect adversarial jailbreaks or boundary‑probing attempts, or maintain GDPR‑aligned incident logs without storing sensitive user text. By integrating via API or web app, teams can plug PSA into existing model pipelines for continuous behavioral logging, benchmarking, and alerting against configurable baselines.

Splabs pricing

Pricing model: Freemium

The website does not list explicit pricing tiers, but PSA offers a public API, web app access, and optional paid enterprise‑grade integrations for continuous monitoring and forensics. There is a free tier allowing basic usage via the API or web interface for small‑scale or evaluation workloads, while larger deployments, high‑volume telemetry, and advanced SIGTRACK / agentic‑graph features are typically offered through custom enterprise plans negotiated directly with Splabs. These paid plans generally include higher throughput limits, dedicated support, and additional forensic and benchmarking capabilities tailored to regulated or safety‑critical environments.

Splabs pros

  • Provides black‑box analysis without model‑weight access
  • 24+ deterministic metrics for rich behavioral fingerprinting
  • Detects progressive drift, acute collapse, and oscillation patterns
  • Five micro‑classifiers for per‑sentence posture signals
  • Behavioral Health Score gives a single intuitive safety metric
  • Posture Oscillation Index captures erratic behavior shifts
  • Dissolution Position Index surfaces boundary‑dissolution risks
  • Multi‑agent agentic‑graph analysis for complex workflows
  • Swiss Cheese alignment detection for cross‑agent safety gaps
  • SIGTRACK incident archive with no raw text stored
  • GDPR‑friendly single‑row erasure capability
  • Signature matching to flag known risky behavior patterns
  • Z‑score based benchmarking against custom baselines
  • Seamless API integration for continuous monitoring
  • Real‑time classifier and graph analysis
  • Support for adversarial stress‑testing and boundary mapping
  • Dyadic Risk Monitor for psychological safety in human–AI interactions

Splabs cons

  • Requires external system to capture and feed model responses
  • No direct access to model internals or gradient‑level signals
  • Learning curve to interpret posture and regime‑shift signals
  • Classifiers are rule‑driven and may miss novel attack vectors
  • Limited to text‑based interaction; not designed for multimodal models
  • Incident archive relies on posture‑sequence summaries, not raw logs
  • May produce false positives on benign edge‑case utterances
  • Pricing and plan details are not fully transparent on the site
  • No built‑in model‑retraining or fine‑tuning hooks

Frequently asked questions about Splabs

What does Splabs PSA actually measure?

Splabs PSA measures the behavioral posture of language models by analyzing their text outputs without needing access to internal weights or model APIs. It computes a range of deterministic metrics including stress under adversarial pressure, sycophancy, hallucination risk, persuasion techniques, and session‑level signals such as Behavioral Health Score, Posture Oscillation Index, and Dissolution Position Index.

Can PSA be used with any language model?

Yes, PSA is model‑agnostic and can be used with any language model whose outputs are text; you submit model responses via Splabs’ API or web interface rather than integrating inside the model. Because it operates as a black‑box, it supports both proprietary hosted models and open‑source models as long as their responses can be logged and forwarded to the PSA pipeline.

How does PSA handle privacy and data security?

PSA is designed to minimize raw data exposure: the SIGTRACK archive stores posture‑sequence incident data rather than full conversation logs, and the system supports GDPR‑aligned single‑row erasure. Users can also route their own telemetry through their infrastructure, only sending anonymized or aggregated posture signals, which helps satisfy strict privacy and compliance requirements.

What is regime‑shift detection and why does it matter?

Regime‑shift detection in PSA identifies when a model’s behavior meaningfully changes over time, such as gradual drift from conservative outputs toward more permissive ones, or sudden acute collapses in safety posture. This capability helps teams detect subtle degradation or policy violations before they manifest as visible incidents, enabling proactive tuning or intervention.

How do the classifiers C0–C4 work together?

C0–C4 are five micro‑classifiers that each focus on a different dimension: input intent, adversarial stress, sycophancy, hallucination risk, and persuasion techniques. Their per‑sentence outputs are combined into turn‑level scores and then aggregated into session‑level metrics, allowing PSA to track how a model’s posture evolves across a full interaction rather than treating each response in isolation.

What is the SIGTRACK archive used for?

The SIGTRACK archive stores posture‑sequence incidents for later forensic analysis, without keeping raw text transcripts. It enables teams to replay, search, and compare behavioral patterns across sessions, match new interactions against known risky signatures, and maintain a compliance‑friendly audit trail that can be rapidly scrubbed or anonymized as needed.

Does PSA support multi‑agent or multi‑tool workflows?

Yes, PSA v3 includes agentic graph analysis that models interactions among multiple AI agents or tools as a directed graph. It computes Swiss Cheese alignment detection, cross‑agent contagion metrics, and temporal‑state predictions to surface cascading risks and gaps in collective safety across an agent‑based workflow.

How is PSA integrated into existing model systems?

PSA can be integrated by piping model responses into its API or web app, either synchronously for real‑time monitoring or asynchronously for batch analysis. Many teams embed PSA as a post‑processing layer in their production pipelines to continuously log and score outputs, trigger alerts on anomalous posture signals, and archive flagged sessions into the SIGTRACK incident store.

What is the Dyadic Risk Monitor (DRM) for?

The Dyadic Risk Monitor is designed to assess psychological safety in human–AI interactions, scoring user turns for risk states such as suicidality, dissociation, and urgency, and scoring AI adequacy in response. It computes the gap between user need and AI support, then surfaces cases as green, yellow, orange, red, or critical, helping teams prioritize high‑risk conversations for human review.

Can PSA be used for both adversarial testing and ongoing monitoring?

Yes, PSA supports both adversarial stress‑testing and boundary mapping in controlled environments, where operators deliberately probe model limits, and long‑term continuous monitoring in production, where it tracks posture, drift, and anomalies over time. The same classifier stack and metrics are used in both modes, enabling consistent evaluation before and after deployment.

Categories

Use cases

Browse all AI tools on NeedAnAI