Arthur Shield logo

Arthur Shield

Free

A paid product for detecting toxicity, hallucination, prompt injection, etc.

FreeFree tier
Type
Open Source
Company
Arthur

About Arthur Shield

Arthur Shield is an AI observability and evaluation platform that helps teams monitor, test, and improve large language model (LLM) applications. It detects issues such as toxicity, hallucination, prompt injection, and sensitive data leakage through built-in and custom evaluators. The platform offers a free tier for small teams, a premium tier with advanced monitoring and alerting, and enterprise options with dedicated VPC, SSO, and SLAs. An open-source version of the Arthur Evals Engine is available on GitHub for self-hosted deployments. Features include model performance metrics, data drift detection, custom dashboards, tracing, user feedback, and explainability methods like what-if analysis and global explanations. Arthur Shield supports cloud data connectors, custom alert webhooks, and compliance with SOC 2 and data locality requirements.

Key Features

Monitor model performance with core metrics
Detect PII, sensitive data, custom LLM, and regex rules
Cloud data connector integrations (built-in)
Custom alerting webhooks and thresholds
Open source Arthur Evals Engine for self-serve deployment
Customizable dashboards and performance metrics
Tracing, user feedback, and human annotation
Explainability methods: what-if analysis, global explanations
Continuous and custom evaluations (datasets, evals)
Prompt management and RAG optimization

Pros & Cons

Pros
  • Free tier available with generous limits for small teams
  • Open-source evals engine for flexibility and self-hosting
  • Comprehensive monitoring metrics including data drift and performance
  • Customizable alerts with webhook integrations
  • SOC 2 compliance and data locality options for enterprise
  • Supports multiple cloud data connectors
  • Explainability methods provide insight into model decisions
  • Scalable from free to enterprise with dedicated support
Cons
  • Premium plan costs $60/month for up to 100 use cases
  • Free plan limited to 7 days data retention
  • Enterprise pricing is custom and may be expensive for small teams
  • Some advanced features (e.g., SSO, custom data connectors) require higher tiers
  • Open-source evals engine may require additional setup and maintenance

Best For

Monitoring AI agent performance and reliabilityDetecting and preventing prompt injection attacksEvaluating LLM hallucinations and factual accuracyEnsuring compliance with data governance (PII, sensitive data)Optimizing retrieval-augmented generation (RAG) pipelinesImproving model safety through custom evaluatorsDebugging model behavior with tracing and explainability

FAQ

What is included in the free plan?
The free plan ($0/month) includes monitoring for up to 4 use cases, unlimited seats, cloud data connector integrations, and core performance metrics. It also includes data retention of 7 days and up to 3 alerts per model with 6 alert checks per hour.
Is there an open-source version of Arthur Shield?
Yes, the Arthur Evals Engine is open source and available on GitHub. It includes PII, sensitive data, custom LLM, and regex rules built in, and can be self-deployed.
What security and compliance features are offered?
Arthur Shield offers SOC 2 compliance, data locality options (e.g., self-managed VPC, BYOCloud, on-premises), SSO, RBAC, and support for dedicated VPCs in the Enterprise plan. Data encryption and audit logging are also available.
How does pricing scale?
Pricing scales by use case and monitoring needs. The free plan covers up to 4 use cases. The Premium plan ($60/month) supports up to 100 use cases. Enterprise plans offer custom configurations with dedicated infrastructure, unlimited use cases, and additional compliance features.