I score 97.80 on recruiting search.

    Evaluated on PSB-Recruiting, the 30-query recruiting category of PeopleSearchBench. It measures whether a people-search system finds the right candidates, returns enough of them, and packages what a recruiter actually needs to decide.

    $100 of credits on us. No card required.

    Overall97.80Averaged across 7 runs
    Relevance96.10Right candidates first
    Coverage98.70Enough qualified profiles
    Utility98.70Useful info to review

    Ahead of every published baseline.

    Laidback scored 97.80 overall, versus the strongest published baseline (Prism) at 89.64. That is an 8.16-point lead.

    Averaged across 7 independent evaluator runs so the number isn't a lucky judge day.

    10080604020
    97.80
    89.64
    68.23
    65.70
    64.70
    50.50
    Laidback logoLaidback
    Prism logoPrism
    Lessie logoLessie
    Juicebox logoJuicebox
    Exa logoExa
    Claude logoClaude

    Four things a recruiter would actually check.

    PSB-Recruiting doesn't reward a chatbot for sounding confident. It grades what matters when you're staring at a shortlist on Monday morning.

    Relevance

    Are the top-ranked candidates the ones a senior recruiter would actually shortlist first?

    Coverage

    Did we surface enough qualified profiles, not just a lucky handful?

    Utility

    Is every profile packaged with enough evidence to review without a second tab?

    Fairness

    Judged by an independent LLM against explicit, brief-derived criteria, not vibes.

    How the benchmark works.

    Same 30 briefs. Same normalized outputs. Same LLM judge grading against brief-derived criteria and public evidence. The underlying paper is on arXiv.

    1. 01
      Brief in, live search out

      30 recruiting briefs written like real hiring requests. For each one, Laidback returns up to 15 candidates ranked strongest fit first.

    2. 02
      Normalize the outputs

      Every returned candidate is normalized into the structured format the benchmark evaluator expects. No format bonus, no format penalty.

    3. 03
      Decompose the brief

      The evaluator breaks each brief into explicit, checkable criteria: seniority, skills, domain, location, disqualifiers.

    4. 04
      Ground in public evidence

      Live web search gathers outside evidence for every candidate so judgments are grounded in real, verifiable public information.

    5. 05
      Judge and score

      An LLM judge reviews the evidence and scores each result on relevance, coverage, and information utility.

    6. 06
      Run it seven times

      The full evaluator runs 7 independent times and averages, cancelling judge variance.

    A benchmark isn't a trophy. It's a promise: the shortlist I hand your recruiter on Monday is graded by the same rules, every single time.

    Reproducible. Auditable. Defensible to a hiring manager who wants to know why one candidate ranked above another, and that the answer isn't "vibes."

    Get $100 of credits and try it on the roles you have open today.

    $100 of credits on us. No card required.