Visr Evals

    Quickstart

    Run Harbor Framework-compatible task evals without infra or model keys.

    Run Your First Eval

    Eval runner

    Choose an eval runner, review the prompt, then copy it into that agent.

    Free credits

    $10

    Start your first hosted eval runs on us.

    Claim $10 now

    How it works

    1. Run task evals locally with the agents you already have installed.
    2. Claude Code, Codex, and others use your existing subscription when available.
    3. Without one, Visr falls back to hosted credits.
    4. Docker runs use hosted credits for strict isolation.

    Agent workflow evals

    Turn coding-agent runs into a team benchmark.

    Visr helps teams test the agentic workflows they actually rely on: Harbor Framework task evals, Claude Code sessions, Codex runs, Cursor habits, OpenClaw playbooks, MCP toolchains, repo instructions, skills, and the prompts that bind them together.

    Benchmark real agent work

    Capture repeatable task evals from real repositories, then compare agents, prompts, skills, and models against the same evidence.

    Score team-specific outcomes

    Use criteria that match your codebase: tool use, source quality, command discipline, artifact shape, and regression behavior.

    Run local or isolated trials

    Exercise the same workflow locally with existing agent CLIs, or switch to Docker when a sandboxed benchmark run matters.

    Promote what holds up

    Package proven workflows so Claude, Codex, Cursor, OpenClaw, and MCP-backed agents inherit behavior your team has tested.

    Visr, by Kernel Agentson, Inc
    Your team's own task-eval bench