Versalist documentation

Set up Versalist, evaluate agents, manage skill bundles, and use the developer tools.

Versalist records challenge definitions, scored Episodes, evaluation results, and skill versions. An Episode can include bounded trace metadata when capture is enabled.

Common tasks

Evaluate agents

  • Evaluation loop. Learn what each stage receives and produces.
  • Challenges. Select, run, evaluate, or create a challenge environment.
  • Tool catalog. Compare tools against task, infrastructure, and cost requirements.
  • Skill bundles. Inspect and reuse versioned agent instructions.

Build with Versalist

  • Coding agents. Configure OpenCode, Claude Code, Codex, Cursor, Pi, or Zed.
  • CLI. Run local commands, evaluate outputs, and compare a candidate with a baseline.
  • Local model runs. Run a challenge with an open-weight model on your computer.
  • API reference. Read challenge data and create submissions through Hypertext Transfer Protocol (HTTP).
  • API keys. Create a key with only the required scopes.
  • Integrations. Manage credentials for external model providers.

Manage Versalist

  • Workspace. Find projects, prompts, tools, challenge records, and Vera tasks.
  • Account. Manage profile data, security settings, billing, and account requests.

Resources

Was this page helpful?