Verging Labs

Search Methodology

One pinned reference agent runs every task against every provider identically; LLM judges grade the answers and are themselves audited against human labels before anything publishes. Nothing is blended into one number: every index is a pure quality construct, and cost and latency are separate published axes.

Principles

  • A task passes only if it addresses the question, is correct, is current where that applies, and is supported by what the tool actually returned. An unsupported or fabricated answer is a scored failure.
  • The scored task splits are private and rotate per release, so providers cannot train on the test.
  • Judges never run uncalibrated: labels from at least two humans set the agreement ceiling, and the judge's agreement against them is published with every release.
  • No provider can pay for results, methodology changes, placement, or listing. Paid engagements never touch public scores; the full rules live on the governance page.

Versioning

  • Kind weights and the judge (model and prompt) are pinned per index version; any change ships as a new, recalibrated version.
  • Task splits are private and rotate per release; saturated sources are retired with published rationale.
  • Changes and corrections are logged in the changelog, never silent.

Lineage

The judge design adapts the Universal Verifier's published calibration principles (Rosset & González Fernández et al., arXiv:2604.06240); hurdle and signed grounding follow Mercor's ACE.

Stay up to date with agent memory

Get new memory test results and updates by email.