sauce labs
rebuilt a legacy design system into an ai-legible library that designers and agents can both build from
timeline
june – august 2026
role
product design intern
mentors
stephen thomas
lena reed
daryna gulenko
tools/skills
design systems
design tokens
claude code
figma mcp
code connect

context
jump to results ↓A design system built for humans, but not for agents.
At Sauce Labs, I rebuilt a legacy design system into a library that designers and AI agents can both build on-brand designs and prototypes from.

problem
AI tools promise speed, but ignore the design system.
AI design tools like Claude Code, Claude Design, Figma Make, Lovable, etc. claim to replace manual prototyping and wireframing through Figma to code and code to Figma workflows.
Generate a Figma screen exactly like the app.saucelabs.com dashboard.> Generate a Figma screen exactly like the app.saucelabs.com dashboard._> _

✕Off-brand. It hallucinated a logo, a sidebar navigation, and components that don't exist in the design system.
off-system · cleaned up by hand
Every off-brand output has to be cleaned up by hand, which cancels out the speed and blocks adoption at enterprise scale.
opportunity
If AI tools used our design system correctly, every team would benefit.
for designers
On-brand screens, fast
Production-ready screens in the time it used to take to find the right component.
for product managers
Shorter idea-to-prototype cycles
A compressed product development lifecycle.
for engineers
Designs grounded in the system
Fewer feasibility debates and compromises.
solution
Make design knowledge explicit, then prove it works.
01deterministic vs. probabilistic
Guarantee what can be guaranteed, and guide the rest.

02deterministic
Tokens an agent can read without guessing.
A three-tier color system: primitives hold the palette, semantic tokens give each value a purpose, and component tokens scope it to one part.

03probabilistic
Skills that keep the agent honest.
Six skills that make the agent read the docs first, show its work, and write down what the system is missing.

- bidirectional tests
- 24
- figma ↔ code workflows
- 6
- code → figma mapping accuracy
- 100%
- figma → code composite
- 92.9%
Generate a Figma screen exactly like the app.saucelabs.com dashboard.> Generate a Figma screen exactly like the app.saucelabs.com dashboard._> _
results
I scored 24 tests across six Figma ↔ code workflows.
Denominators were set before any output was reviewed, and every test ran in a fresh session.
usage insights framefigma frame · linkedHere's the Figma frame for the Usage Insights dashboard, generate the React code for this screen using our DS components.> Here's the Figma frame for the Usage Insights dashboard, generate the React code for this screen using our DS components._> _
✓Zero hard-coded styles. The only two misses were a chart tooltip and legend the system has no component for.
92.9% · composite
takeaways
- 01
Guarantee what you can
Tokens and Code Connect mappings made the deterministic traits reliable. Docs and skills guide the rest, so that is where a lot of attention went.
- 02
Tests find what reviews miss
Scoring every output surfaced gaps in existing mappings, documentation, and implementation that looked fine at a glance.
- 03
Write for the agent, too
Calls a designer makes by instinct, like a badge vs. a status, had to be written down before an agent could make them. And repetition is key.