versalist
ResearchChallengesAI ToolsPricingFor Teams
Sign inRequest a demo
versalist

Agent evaluation environments for engineering teams. Design rewards, run agents, and turn outcomes into signal.

Resources

  • Request a demo
  • Docs
  • Research
  • Blog
  • How it works
  • FAQ
  • Environments
  • Legal

Company

  • About us
  • Pricing
  • For Teams
  • Book a session
  • Join as Partner
  • Terms of Service
  • Privacy Policy

Guides

  • All guides
  • Prompt engineering
  • Prompt guide
  • Model Context Protocol
  • Evaluation
  • AI-empowered future

Stay updated

Get platform updates, challenge launches, and practical notes on agent evaluation.

© 2026 Social Protocol Labs LLC. All rights reserved.
Build the full agent training loop.
AI Tools Directory

Discover the tools behind real AI build, eval, and deployment workflows.

Browse by role in the stack, not just by vendor. Compare fit here, then open a tool detail page when it earns deeper inspection.

Open My AI StackBrowse Agent Capabilities

Stack layer

Narrow the directory by what the tool does in the workflow.

All Tools467Models886Environment6Action Space411Observation19Reward / Eval10Policy Serving35Training Infra5Safety / Guardrails7Orchestration16

Browse tools

10 results in the directory

Suggest a tool
AI Workflow Automation · Observability, Evaluation & GovernanceArize AIOpen source

Arize Phoenix

Open-source LLM observability and evals.

4 related challengesInspect tool
Agent FrameworkBraintrustCommercial

Braintrust

Evaluation and tracing platform for AI apps.

6 related challengesInspect tool
AI Engineering Tooling · Developer ToolsDistributionalFreemium

Distributional

AI safety & robustness evaluation

ContactInspect tool
AI Workflow Automation · Observability, Evaluation & GovernanceGalileoCommercial

Galileo

Generative AI eval and observability platform.

12 related challengesInspect tool
AI Engineering Tooling · Security & Risk Management PlatformsNVIDIAOpen source

Garak

LLM vulnerability scanner.

Free (OSS)Inspect tool
AI Workflow Automation · Evaluation PipelinesGentracePaid

Gentrace

GenAI evaluation & observability

SubscriptionInspect tool
AI Engineering Tooling · Security & Risk Management PlatformsGuardrails AIOpen source

Guardrails AI

Validation and guardrails for LLM outputs.

FreemiumInspect tool
Agent FrameworkLangfuseOpen source

Langfuse

Open-source LLM observability and evals.

8 related challengesInspect tool
Agent FrameworkLangChainCommercial

Langsmith

LangChain tracing, debugging, and evaluation platform.

SubscriptionInspect tool
AI Workflow Automation · Observability, Evaluation & GovernancePatronus AIFreemium

Patronus AI

Evaluation and guardrail platform.

5 related challengesInspect tool