Early access

Start with one agent

Pick the shape that fits and we will scope the evidence needed before anything runs. Access is granted directly rather than by self-service signup.

Free Reliability Pilot

Invite-only
$0one agent

Evaluate one agent free.

Includes limited evaluation and replay credits, so you can see what evidence your agent can actually produce today. Access is granted directly.

  • One agent, one declared environment
  • Readiness audit with explicit missing inputs
  • Expert contract review pass
  • Recorded or sandbox replay, where evidence allows
  • Evidence records for every run
Agents
1
Evaluation credits
250 runs
Replay credits
50 replays
Evidence retention
30 days

Team

Contact usscoped to how many agents you run

For teams evaluating multiple agents and maintaining regression evidence.

Keep a standing regression set per agent, re-run it on every change, and keep the evidence attributable across versions.

  • Multiple agents and versions
  • Locked regression cases with replay history
  • Patch decisions recorded against evidence
  • Shared contract review with domain experts
  • Export of evidence records
Agents
Agreed with you
Evaluation credits
Agreed with you
Replay credits
Agreed with you
Evidence retention
Agreed with you

Enterprise

Contact salesscoped per deployment

Private deployment, custom retention, security review, and dedicated support.

For teams that need the platform inside their own boundary, with retention and access controls scoped to their review process.

  • Private deployment
  • Custom retention and data handling
  • Security review with your team
  • Dedicated support and onboarding
  • Custom regression and evidence workflows
Agents
Unlimited
Evaluation credits
Agreed with you
Replay credits
Agreed with you
Evidence retention
Your policy

No prices are published while the per-evaluation cost model is being settled. Paid plans are agreed directly and self-serve checkout is not available yet.

Questions

Before you ask us

Question

Why is the pilot limited to one agent?

A useful evaluation depends on the policy, tool schemas, and either a sandbox or recorded tool responses. Scoping the pilot to a single agent keeps that evidence gathering honest and finishes in days rather than months.

Question

Why are execution credits limited?

Every run and replay executes your agent against declared tools, repeatedly, to measure variance. That work has a real cost, so the pilot includes a fixed allowance rather than an open-ended one.

Question

How is our data handled?

Policies and traces are treated as sensitive. Evidence records reference versions, hashes, and run IDs rather than copying more than is needed, and retention is agreed as part of each pilot. See the security page for what is implemented versus planned.

Question

When can we pay by card?

Not yet. There is no payment processor connected to this product, so no card can be charged. Paid plans are agreed directly with us until the cost model is settled.