UIBenchmark runs the protocol workflow-subscription-v1. It evaluates a coding workflow — a tool, a model, and the settings that tool exposes — not a model in isolation. Claude Code and Codex CLI orchestrate work differently, and freezing the inputs does not make those surfaces identical.
A Challenge is the durable identity of a UI task. Every execution binds a Challenge Version: an immutable set of instructions, rules, starter files, assets, numbered requirements, and the capture profile. Once a version is frozen, its inputs cannot change; a correction becomes a new version, so older Runs stay readable against exactly the brief they were given. Starter and asset hashes are recorded and published so two Runs can be shown to have received the same bytes.
A Configuration is an immutable snapshot of the coding tool, the model identity it displayed, the settings it exposed, and the capabilities that participated. A Run binds one Challenge Version, one protocol version, one Configuration, and a repetition number. Each Run starts from a fresh copy of the same frozen starter in a new session, with no continuation from a competing Run and no external research; local preview and the provider’s own inference remain allowed. Human contribution is limited to operational approvals — no design direction and no code edits. If a tool version or an exposed setting changes, it becomes a new Configuration rather than an edit to the old one.
Screenshots and checks are produced locally against the frozen compiled files through a static server with production-equivalent routing and security policy. Captures use a pinned Chromium at device pixel ratio 1, a fresh context per test group, en-US, UTC, light appearance, and reduced motion.
| Profile | Viewport |
|---|---|
| Desktop | 1440 × 900 CSS px |
| Mobile | 390 × 844 CSS px |
For each Revision the archive records the build outcome, page load and runtime errors, declared routes and fallback behaviour, responsive overflow, the requirement interactions that could be exercised, axe findings, keyboard observations, and security-policy compatibility. Each check is pass, fail, inconclusive, not_run, or error. Only pass means observed success; a broken selector is inconclusive until a human reviews it. Tests target semantic roles, accessible names, and visible state, so they do not reward one visual style. The operator may attach an attributed qualitative note; it is editorial commentary, not a preference score.
There is no ranking on this site
Every buildable Revision is published regardless of quality, and missed requirements are a result rather than a failure. Counts are descriptive and need their denominator: they are not a ranking, a score, or a winner.
Operator-attested: the Configuration, tool version, displayed model, and elapsed time were recorded by the operator from what the tool showed at Run time. They are not independently verified. An exact model ID appears only when the tool exposed one; it is never inferred.
Tool-self-reported: token counts come from the coding tool’s own output, imported from a labeled local source. Category names are kept as the tool reports them and are not normalized between tools. Missing usage stays unknown. Nothing here claims equal compute or equal token budgets.
Generation cost: Not separately metered. Runs were executed on the operator’s personal tool subscriptions, which do not itemize a per-Run charge. No amount is attributed, and no zero-cost claim is made.
Automated checks are not an accessibility certification. Recorded axe findings describe what the tooling observed at capture time, on a pinned browser and viewport. A check that reports no violation is not proof that the interface is accessible, and a selector problem is inconclusive rather than an implementation failure.
Conversation evidence is not provided. Raw transcripts and operator notes stay on local storage: they routinely contain incidental private material, and publishing them is not required to read the artifact. What is published instead is the frozen input, the frozen output, the hashes tying them together, and the recorded execution facts.
The launch cohort is deliberately small, so nothing here supports a general claim about a tool’s capability. One Challenge, executed once per Configuration, is an observation, not a sample.
Related: About this project and Terms and privacy.