About UIBenchmark
Most discussion of AI coding tools is anecdotal: a screenshot, a prompt nobody else can reproduce, and a claim. UIBenchmark exists to make one narrow thing inspectable. The same frozen brief, the same starter files, the same assets, run through different coding workflows — and then the results published as real, clickable websites with their source and their provenance attached.
What you can do here
- Browse published Challenges and read the exact brief each Run received.
- Open a Revision as a live website, hosted in isolation on its own origin.
- Read the generated source in the browser, escaped and safe to inspect.
- Put one to four Revisions side by side at a declared true viewport.
- Check provenance: hashes, Configuration snapshots, recorded execution facts, and tool-reported token usage.
No account is needed for any of it.
Who runs it
UIBenchmark is operated by a single person. That operator prepares the Challenge Versions, runs each Challenge locally through the coding tools on personal subscriptions, freezes the output unmodified, uploads it, and publishes it. There is no jury, no crowd, and no second reviewer. This is why the archive labels its claims: what the operator observed is marked operator-attested, what a tool reported about itself is marked tool-self-reported, and anything unknown is left unknown rather than filled in.
The same person owns publication decisions and takedown review. If something published here infringes your rights or exposes something it should not, report it and it will be withdrawn first and reviewed afterwards.
What this is not
Every buildable Revision is published regardless of quality, and missed requirements are a result rather than a failure. Counts are descriptive and need their denominator: they are not a ranking, a score, or a winner.
Published sites are independent, AI-generated demos served from separate preview origins with local mock data. They are not products, and no UIBenchmark session data crosses to them.
Read the methodology for the full protocol and its limitations.
Roadmap
The first release is a curated archive, and it is intentionally small. Later work has been sketched but is not built, not scheduled, and not linked from anywhere on this site:
- More Challenges and repetitions. A single execution per Configuration cannot separate a workflow’s behaviour from run-to-run variance. Widening the library and repeating Runs comes before any comparative statement.
- Blind human comparison. A pairwise viewing mode could be added later. It would need decisions about identity, consent, bot and replay controls, exposure balancing, and moderation first, and hiding a label is not anonymity when the artifacts are public.
- Rankings. Ordering Configurations requires rules for cohort validity, ties, weighting, repeated artifacts, uncertainty, eligibility, and failure handling. Until those are specified and validated, publishing an ordering would be a stronger claim than the data supports, so this release publishes none.
- A separately labelled API track. Provider APIs could execute the same frozen inputs under the same bundle contract, in a track kept clearly apart from subscription-tool Runs. That requires spend, concurrency, cancellation, and isolation controls.
- Downloadable source. Source is view-only today. A downloadable form would only follow a rights review and evidence that visitors actually need it.
None of this delays the archive. The release principle is a trustworthy record of real artifacts first, competition later — if ever.