{
  "generated": "2026-08-28",
  "url": "https://bench.helm.dev/",
  "name": "HelmBench",
  "tagline": "A seeded-bug benchmark for AI code auditors",
  "rerunPolicy": "Policy as of 2026-08-28: HelmBench re-runs within 48 hours of a major coding-model release, on the same sealed seeds, and each run lands here dated. Corrections get fixed and logged on this page.",
  "runs": [
    {
      "id": "run-1",
      "label": "Run 1",
      "adjudicated": "2026-07-12",
      "summary": "Five model arms audited the same three seeded repos, read-only, with one identical provider-neutral prompt. 24 sealed seeded bugs, five independent blinded adjudication passes, 209 graded findings, 0 false positives.",
      "seedTotals": {
        "seeds": 24,
        "t1": 6,
        "t2": 11,
        "t3": 7,
        "blindSpot": 3
      },
      "legend": {
        "H": "hit",
        "P": "partial",
        "M": "miss"
      },
      "arms": [
        {
          "key": "opus",
          "model": "Claude Opus 4.8",
          "vendor": "Anthropic",
          "recallHits": 20,
          "recallPct": "83.3%",
          "hitPlusHalfPartial": "83.3%",
          "t1": "6/6",
          "t2": "10/11",
          "t3": "4/7",
          "blindSpot": "2/3",
          "findings": 29,
          "tpNatural": 9,
          "fp": 0,
          "precision": "100%",
          "wallClockMin": "~13",
          "price": "subscription"
        },
        {
          "key": "terra",
          "model": "GPT-5.6 Terra",
          "vendor": "OpenAI",
          "effort": "high",
          "recallHits": 20,
          "recallPct": "83.3%",
          "recallFootnote": "Terra's A0-S6 hit was ruled via the seed ledger's own diff attribution but sits about 11 lines outside the strict 10-line window. Strict recall: 19 of 24 (79.2%). Both numbers stand.",
          "hitPlusHalfPartial": "83.3%",
          "t1": "5/6",
          "t2": "9/11",
          "t3": "6/7",
          "blindSpot": "3/3",
          "findings": 36,
          "tpNatural": 16,
          "fp": 0,
          "precision": "100%",
          "wallClockMin": "~10",
          "price": "$2.50 / $15"
        },
        {
          "key": "solMax",
          "model": "GPT-5.6 Sol",
          "vendor": "OpenAI",
          "effort": "max",
          "recallHits": 19,
          "recallPct": "79.2%",
          "hitPlusHalfPartial": "81.3%",
          "t1": "5/6",
          "t2": "8/11",
          "t3": "6/7",
          "blindSpot": "3/3",
          "findings": 60,
          "tpNatural": 41,
          "fp": 0,
          "precision": "100%",
          "wallClockMin": "~54",
          "price": "$5 / $30"
        },
        {
          "key": "solXhigh",
          "model": "GPT-5.6 Sol",
          "vendor": "OpenAI",
          "effort": "xhigh",
          "recallHits": 19,
          "recallPct": "79.2%",
          "hitPlusHalfPartial": "79.2%",
          "t1": "4/6",
          "t2": "9/11",
          "t3": "6/7",
          "blindSpot": "2/3",
          "findings": 43,
          "tpNatural": 24,
          "fp": 0,
          "precision": "100%",
          "wallClockMin": "~18",
          "price": "$5 / $30"
        },
        {
          "key": "solHigh",
          "model": "GPT-5.6 Sol",
          "vendor": "OpenAI",
          "effort": "high",
          "recallHits": 19,
          "recallPct": "79.2%",
          "hitPlusHalfPartial": "79.2%",
          "t1": "3/6",
          "t2": "9/11",
          "t3": "7/7",
          "blindSpot": "3/3",
          "findings": 41,
          "tpNatural": 22,
          "fp": 0,
          "precision": "100%",
          "wallClockMin": "~12",
          "price": "$5 / $30"
        }
      ],
      "armOrder": [
        "opus",
        "terra",
        "solMax",
        "solXhigh",
        "solHigh"
      ],
      "seeds": [
        {
          "id": "A0-S1",
          "class": "committed secret",
          "tier": "T1",
          "results": {
            "opus": "H",
            "terra": "H",
            "solMax": "H",
            "solXhigh": "H",
            "solHigh": "H"
          }
        },
        {
          "id": "A0-S2",
          "class": "fail-open error path",
          "tier": "T3",
          "results": {
            "opus": "H",
            "terra": "H",
            "solMax": "H",
            "solXhigh": "H",
            "solHigh": "H"
          }
        },
        {
          "id": "A0-S3",
          "class": "XSS",
          "tier": "T2",
          "results": {
            "opus": "H",
            "terra": "H",
            "solMax": "H",
            "solXhigh": "H",
            "solHigh": "H"
          }
        },
        {
          "id": "A0-S4",
          "class": "data-loss hazard",
          "tier": "T3",
          "results": {
            "opus": "H",
            "terra": "H",
            "solMax": "H",
            "solXhigh": "H",
            "solHigh": "H"
          }
        },
        {
          "id": "A0-S5",
          "class": "PII in logs",
          "tier": "T2",
          "results": {
            "opus": "H",
            "terra": "H",
            "solMax": "H",
            "solXhigh": "H",
            "solHigh": "H"
          }
        },
        {
          "id": "A0-S6",
          "class": "vacuous test",
          "tier": "T2",
          "footnote": "terra-window",
          "results": {
            "opus": "H",
            "terra": "H",
            "solMax": "H",
            "solXhigh": "H",
            "solHigh": "H"
          }
        },
        {
          "id": "A0-S7",
          "class": "committed secret (in HTML)",
          "tier": "T1",
          "results": {
            "opus": "H",
            "terra": "H",
            "solMax": "H",
            "solXhigh": "H",
            "solHigh": "H"
          }
        },
        {
          "id": "A0-S8",
          "class": "off-by-one",
          "tier": "T3",
          "blindSpot": true,
          "results": {
            "opus": "H",
            "terra": "H",
            "solMax": "H",
            "solXhigh": "H",
            "solHigh": "H"
          }
        },
        {
          "id": "B1-S1",
          "class": "SQL injection",
          "tier": "T2",
          "results": {
            "opus": "H",
            "terra": "H",
            "solMax": "H",
            "solXhigh": "H",
            "solHigh": "H"
          }
        },
        {
          "id": "B1-S2",
          "class": "XSS",
          "tier": "T2",
          "results": {
            "opus": "H",
            "terra": "H",
            "solMax": "H",
            "solXhigh": "H",
            "solHigh": "H"
          }
        },
        {
          "id": "B1-S3",
          "class": "unscoped bulk delete",
          "tier": "T2",
          "results": {
            "opus": "H",
            "terra": "H",
            "solMax": "H",
            "solXhigh": "H",
            "solHigh": "H"
          }
        },
        {
          "id": "B1-S4",
          "class": "MIME fail-open",
          "tier": "T3",
          "results": {
            "opus": "M",
            "terra": "H",
            "solMax": "P",
            "solXhigh": "H",
            "solHigh": "H"
          }
        },
        {
          "id": "B1-S5",
          "class": "PII in logs",
          "tier": "T2",
          "results": {
            "opus": "H",
            "terra": "H",
            "solMax": "H",
            "solXhigh": "H",
            "solHigh": "H"
          }
        },
        {
          "id": "B1-S6",
          "class": "vulnerable dependency",
          "tier": "T1",
          "results": {
            "opus": "H",
            "terra": "H",
            "solMax": "H",
            "solXhigh": "M",
            "solHigh": "M"
          }
        },
        {
          "id": "B1-S7",
          "class": "committed credential",
          "tier": "T1",
          "results": {
            "opus": "H",
            "terra": "H",
            "solMax": "H",
            "solXhigh": "H",
            "solHigh": "H"
          }
        },
        {
          "id": "B1-S8",
          "class": "some/every logic flip",
          "tier": "T3",
          "blindSpot": true,
          "results": {
            "opus": "H",
            "terra": "H",
            "solMax": "H",
            "solXhigh": "H",
            "solHigh": "H"
          }
        },
        {
          "id": "B2-S1",
          "class": "committed secret (probe)",
          "tier": "T1",
          "results": {
            "opus": "H",
            "terra": "M",
            "solMax": "M",
            "solXhigh": "M",
            "solHigh": "M"
          }
        },
        {
          "id": "B2-S2",
          "class": "f-string SQL injection",
          "tier": "T2",
          "results": {
            "opus": "H",
            "terra": "M",
            "solMax": "M",
            "solXhigh": "M",
            "solHigh": "M"
          }
        },
        {
          "id": "B2-S3",
          "class": "key logged at INFO",
          "tier": "T2",
          "results": {
            "opus": "M",
            "terra": "M",
            "solMax": "M",
            "solXhigh": "M",
            "solHigh": "M"
          }
        },
        {
          "id": "B2-S4",
          "class": "except/pass swallow",
          "tier": "T3",
          "results": {
            "opus": "M",
            "terra": "M",
            "solMax": "H",
            "solXhigh": "H",
            "solHigh": "H"
          }
        },
        {
          "id": "B2-S5",
          "class": "missing WHERE clause",
          "tier": "T2",
          "results": {
            "opus": "H",
            "terra": "H",
            "solMax": "M",
            "solXhigh": "H",
            "solHigh": "H"
          }
        },
        {
          "id": "B2-S6",
          "class": "threshold logic",
          "tier": "T3",
          "blindSpot": true,
          "results": {
            "opus": "M",
            "terra": "H",
            "solMax": "H",
            "solXhigh": "M",
            "solHigh": "H"
          }
        },
        {
          "id": "B2-S7",
          "class": "vulnerable dependency",
          "tier": "T2",
          "results": {
            "opus": "H",
            "terra": "H",
            "solMax": "H",
            "solXhigh": "H",
            "solHigh": "H"
          }
        },
        {
          "id": "B2-S8",
          "class": "os.system command injection",
          "tier": "T1",
          "results": {
            "opus": "H",
            "terra": "H",
            "solMax": "H",
            "solXhigh": "H",
            "solHigh": "M"
          }
        }
      ],
      "headline": {
        "flatLadder": "The same model at high, xhigh, and max reasoning effort all scored 79.2% recall. Max effort cost 4.5x the wall-clock of high (about 54 minutes vs 12) and bought one extra partial catch. Partials are not counted in recall.",
        "unionAllFive": "Union of all five arms: 23 of 24 seeds. The one miss shared by every arm: B2-S3, an integration key logged at INFO in a secondary module.",
        "bestPair": "Best two-arm cross-family union: Claude Opus 4.8 + GPT-5.6 Sol at high caught 23 of 24 (95.8%) in about 25 combined minutes. Second best at 22 of 24: Claude Opus 4.8 + GPT-5.6 Terra at high, tied with Opus paired with Sol at xhigh or max. All four OpenAI-family arms miss B2-S1, B2-S2, and B2-S3, so no same-family pair reaches either figure; the Claude family fielded only one arm.",
        "zeroFp": "209 findings graded across the five arms, spanning seeded and real unseeded bugs: 0 false positives. Every accepted finding required the adjudicator to read the cited file."
      },
      "patterns": [
        "Model families inverted by bug type, consistently across all three Sol arms. The Claude arm caught 6 of 6 blatant seeds but only 4 of 7 subtle ones. The Sol and Terra arms ran the profile in reverse: up to 7 of 7 subtle, as few as 3 of 6 blatant. Different families miss different bugs, which is why the cross-family union wins.",
        "The Claude arm was the only arm of five to report a committed, live-looking API key. The OpenAI-family arms that saw it declined it as a placeholder. For secret-hunting, a Claude pass or a deterministic scanner belongs in the loop.",
        "The Sol effort ladder is flat for audit recall: high, xhigh, and max all scored 79.2%. Max effort bought 19 extra natural findings, one extra partial, and 4.5x the wall-clock of high. If natural-finding depth is the goal, max earns its cost; for seeded recall it does not.",
        "Every arm degraded on B2, the large multi-version Python repo, and coverage-note honesty degraded with it. The Claude arm's coverage notes stayed accurate on all three repos; its B2 note even predicted its own miss areas. The other arms' notes overstated their coverage on B2.",
        "One seed beat everything: B2-S3, a sensitive key logged at INFO in a secondary module, missed by all five arms. It stays in the corpus as a standing hard case."
      ],
      "limitations": [
        "n = 24 seeds, 3 repos, one run per arm. One seed is worth 4.2 points of recall, so the 83.3% vs 79.2% single-arm ranks are within noise. The findings that held across arms are the flat effort ladder and the family inversion; treat single-arm ranks as point estimates.",
        "Terra's headline 83.3% includes one hit ruled via the seed ledger's own diff attribution, about 11 lines outside the strict 10-line window. Its strict recall is 79.2%. Both numbers are printed.",
        "The adjudicators were Claude-family agents. Mitigations: arms were renamed and the arm-to-model mapping withheld before grading, the rubric was pre-registered and mechanical, every ruling required reading the cited file, and the zero-false-positive outcome was uniform across arms, including the Claude arm.",
        "The first battery was voided before grading: the checkout trees were found to be carrying artifacts from unrelated tooling, so every arm re-ran against clean git-archive checkouts of the committed revisions. The voided outputs are archived unscored.",
        "Seed locations stay sealed so the corpus can be reused across runs. This page publishes each seed's class and difficulty tier only.",
        "Wall-clock is one full three-repo battery per arm on one machine, single run. Prices are vendor list prices per million tokens (input / output) at run time; the Claude arm ran under a subscription, so no per-token price is recorded for it."
      ],
      "method": [
        "Corpus: three real application repos, each seeded with eight defects (24 total) at sealed locations. Difficulty tiers were assigned at seeding: 6 blatant (T1), 11 moderate (T2), 7 subtle (T3). Three seeds are deliberate blind-spot seeds, placed outside the audit taxonomy the arms would naturally reach for. B2 proved the hardest repo: large, messy Python spanning multiple versions.",
        "Arms: each model ran as a read-only auditor through its vendor CLI over clean git-archive checkouts of the committed revisions, all fed one identical provider-neutral prompt stating the outcome, the constraints, and the completion bar, with no step prescription. Repo-embedded instruction files were excluded on both sides for parity.",
        "Grading: five independent adjudication passes, one per arm, each in fresh context with the arm-to-model mapping withheld. A hit requires the same defect at the seeded location within 10 lines and the same mechanism. A partial (right location or right class, not both) is never folded into a hit; the secondary hit-plus-half-partial column is printed separately. Every accepted finding, seeded or natural, required the adjudicator to read the cited file; claims that could not be verified graded as false positives. Two arms' CVE claims were checked against the public record.",
        "Scoring: recall is hits over 24 seeds. Precision is accepted findings over all surfaced findings; 209 findings were graded across the five arms, 0 false positives."
      ]
    }
  ]
}
