Evaluation Services
Verging Labs publishes the Agentic Search Index and the Agentic Memory Index, free public benchmarks of the tools AI agents depend on. The same instrument is available for private work. Most engagements are for teams that run agents and need to pick tools on measured evidence rather than vendor claims; vendors engage us to improve their service and lower its costs.
Advisory
Working sessions and written recommendations, grounded in what we measure.
- Which search and memory tools fit your agents, read from the published index data against your workload and constraints.
- What the per-task-type scores, failure attributions, and cost axes mean for your use case.
- Evaluation design for your own stack: task banks, judge calibration, and grounding rules.
- Procurement support: measured comparisons a committee can act on.
Procurement support runs up to the Agent Tool Stack Audit, which benchmarks your use case privately, delivers a committee-ready procurement report, and includes 90 days of drift watch on the tools you approve.
Evaluations
Vendors engage us to evaluate their tool privately on the same calibrated harness that builds the public indexes. The readout shows where the service fails on our task banks, with enough detail to fix it. We repeat the measurement over time, so regressions surface before your users feel them. Paid evaluations never touch public scores.
Custom Benchmarking
When the public indexes do not answer your question, we measure it.
- Private benchmarks of your shortlisted tools on your own workload, run on the same calibrated harness that builds the public indexes.
- Custom task banks built with the published discipline: judges audited against human labels before anything is scored.
- Repeat measurement over time, so you see drift after you adopt.
Everything sold here runs on the instrument we publish. On the search bank, our judge matched a dual-human consensus 94.1% of the time (Cohen's kappa 0.85); on the memory bank, the judge pipeline matched the primary human labeler 98.5% of the time (kappa 0.95). The full calibration statistics are on the methodology page.
Independence
No provider can pay for results, methodology changes, placement, or listing. Paid engagements are separated from measurement: they never touch public scores, and any vendor with a paid relationship is disclosed on its provider page. The full rules live on the governance page.
Contact
Tell us what your agents run and what you need to decide: contact@verginglabs.com.