Start with one agent
Pick the shape that fits and we will scope the evidence needed before anything runs. Access is granted directly rather than by self-service signup.
Free Reliability Pilot
Invite-onlyEvaluate one agent free.
Includes limited evaluation and replay credits, so you can see what evidence your agent can actually produce today. Access is granted directly.
- One agent, one declared environment
- Readiness audit with explicit missing inputs
- Expert contract review pass
- Recorded or sandbox replay, where evidence allows
- Evidence records for every run
- Agents
- 1
- Evaluation credits
- 250 runs
- Replay credits
- 50 replays
- Evidence retention
- 30 days
Team
For teams evaluating multiple agents and maintaining regression evidence.
Keep a standing regression set per agent, re-run it on every change, and keep the evidence attributable across versions.
- Multiple agents and versions
- Locked regression cases with replay history
- Patch decisions recorded against evidence
- Shared contract review with domain experts
- Export of evidence records
- Agents
- Agreed with you
- Evaluation credits
- Agreed with you
- Replay credits
- Agreed with you
- Evidence retention
- Agreed with you
Enterprise
Private deployment, custom retention, security review, and dedicated support.
For teams that need the platform inside their own boundary, with retention and access controls scoped to their review process.
- Private deployment
- Custom retention and data handling
- Security review with your team
- Dedicated support and onboarding
- Custom regression and evidence workflows
- Agents
- Unlimited
- Evaluation credits
- Agreed with you
- Replay credits
- Agreed with you
- Evidence retention
- Your policy
No prices are published while the per-evaluation cost model is being settled. Paid plans are agreed directly and self-serve checkout is not available yet.
Before you ask us
Why is the pilot limited to one agent?
A useful evaluation depends on the policy, tool schemas, and either a sandbox or recorded tool responses. Scoping the pilot to a single agent keeps that evidence gathering honest and finishes in days rather than months.
Why are execution credits limited?
Every run and replay executes your agent against declared tools, repeatedly, to measure variance. That work has a real cost, so the pilot includes a fixed allowance rather than an open-ended one.
How is our data handled?
Policies and traces are treated as sensitive. Evidence records reference versions, hashes, and run IDs rather than copying more than is needed, and retention is agreed as part of each pilot. See the security page for what is implemented versus planned.
When can we pay by card?
Not yet. There is no payment processor connected to this product, so no card can be charged. Paid plans are agreed directly with us until the cost model is settled.