Skip to main content

Evidence

Scored against issues planted in OpenClass, a demo class app, plus decoys that look wrong but are correct.

PR-1

MethodPlanted issues found
Judgment-only issues
Decoys wrongly flagged
axe-core alone8 of 180 of 70 of 6
CurbCut engine11 of 180 of 70 of 6
CurbCut with IBM Bob16 of 185 of 70 of 6
WCAG principleaxe-core aloneCurbCut engineCurbCut with IBM Bob
Perceivable2 of 52 of 53 of 5
Operable2 of 65 of 66 of 6
Understandable1 of 41 of 44 of 4
Robust3 of 33 of 33 of 3
Fixes proven by a test that fails before and passes after: 23 of 23.Missed by CurbCut: P1-11, P1-15.Findings matching no planted issue: 6 of 23. Some are real issues that were not planted; they are listed in the results file, not counted as correct.

PR-1-RUN2

MethodPlanted issues found
Judgment-only issues
Decoys wrongly flagged
axe-core alone8 of 180 of 70 of 6
CurbCut engine11 of 180 of 70 of 6
CurbCut with IBM Bob14 of 183 of 70 of 6
WCAG principleaxe-core aloneCurbCut engineCurbCut with IBM Bob
Perceivable2 of 52 of 53 of 5
Operable2 of 65 of 65 of 6
Understandable1 of 41 of 43 of 4
Robust3 of 33 of 33 of 3
Missed by CurbCut: P1-11, P1-12, P1-14, P1-15.Findings matching no planted issue: 5 of 21. Some are real issues that were not planted; they are listed in the results file, not counted as correct.

PR-2

MethodPlanted issues found
Judgment-only issues
Decoys wrongly flagged
axe-core alone6 of 160 of 90 of 6
CurbCut engine7 of 160 of 90 of 6
CurbCut with IBM Bob14 of 167 of 90 of 6
WCAG principleaxe-core aloneCurbCut engineCurbCut with IBM Bob
Perceivable4 of 94 of 97 of 9
Operable0 of 41 of 44 of 4
Understandable0 of 10 of 11 of 1
Robust2 of 22 of 22 of 2
Missed by CurbCut: P2-05, P2-10.Findings matching no planted issue: 1 of 16. Some are real issues that were not planted; they are listed in the results file, not counted as correct.

Limits of this benchmark

  • Two runs on PR-1 found 16 and 14 issues. Differences of this size are run-to-run variation in IBM Bob's judgment, not an improvement.
  • The issues were planted by the same team that built CurbCut.
  • A finding counts only if it points at the same element with an accepted success criterion.
  • Rerun it: uv run --project codes/engine python scripts/benchmark/score.py pr-1

What CurbCut does not check

  • It does not certify WCAG conformance. Testing with disabled people cannot be replaced.
  • It does not judge real screen reader experience, cognitive load, or caption quality.
  • Flagged findings are IBM Bob's judgment and can be wrong.
  • Benchmark numbers come from a demo app with planted issues. Other code bases may differ.