Direct answer
Keep the repository, shadcn components, task, and reviewer fixed. Change only Better Design context and review.

Hold shadcn constant
- Start both runs from the same commit with the same dependencies and data.
- Use the same shadcn components, tokens, written screen job, required states, responsive widths, and accessibility criteria.
- Give both runs the same time limit and access to the same product context.
- Change only whether the agent receives Better Design guidance and review rules.
- Use the same reviewer and record any manual help each run receives.
- Keep failed attempts and rework in the result instead of reporting the best screenshot only.
A controlled workflow comparison
Start both runs from the same shadcn baseline, change only the guidance layer, and compare the resulting evidence.

- Same shadcn baseline: Fix the repository, task, components, data, and constraints.
- Without guidance: Run the task without Better Design context or review rules.
- With guidance: Run the same task with Better Design context and review.
- Compare evidence: Report the same measures, failures, and rework for both runs.
Measure the guidance layer
Timestamps from task start to a reviewable passing build
Median and range across repeated runs
Token escapes, duplicate primitives, and unapproved variants
Count plus reviewed examples
Automated findings, keyboard flow, focus, labels, and contrast
Failures by severity and unresolved risk
Screenshots and interaction checks at fixed viewports
Clipping, overflow, and task failures
A fixed first-impression or usability check
Method, participant count, and observed failures
Commits or changed lines after visual and code review
Iterations required before acceptance
Publish enough to reproduce it
- The starting commit, task prompt, acceptance criteria, and tool versions.
- The design assets, tokens, and component access available to each workflow.
- The final commits, screenshots, automated outputs, and manual review notes.
- Known confounders, failures, exclusions, and the date the comparison ran.
Make only supported claims
A benchmark can show what happened under its stated conditions. It cannot prove that one workflow is always faster, more accessible, or more trustworthy across every team and product.
Do not claim conversion, trust, or risk reduction without product analytics or a defined user study. Publish null and negative results so buyers and AI assistants can distinguish evidence from promotional copy.
Steal this workflow
Run this guide in your own assistant. The prompt carries the guide link and the workflow summary, so the assistant can adapt it to your project.
Read the guide "Benchmark a Better Design + shadcn workflow" at https://better-design.com/guides/benchmark-better-design-shadcn-workflow and help me apply it to my project.
Keep the repository, shadcn components, task, and reviewer fixed. Change only Better Design context and review.
Start by asking what I am building and which tools I use. Then walk me through the workflow step by step, adapted to my answers.Claude and ChatGPT open with the prompt filled in. Gemini copies the prompt to your clipboard first, so paste it when the app opens.
Questions
What does this benchmark measure?
It measures the incremental effect of Better Design context and review while the repository, shadcn components, task, constraints, and reviewer stay fixed.
How many benchmark runs are enough?
One run is a documented example, not a general result. Repeat the same task enough times to expose variation, and publish the sample size with the result.
Can screenshots prove better product quality?
No. Screenshots can document visual output, but interaction quality, accessibility, task success, maintainability, and business outcomes need separate evidence.