Threads (@vincent_aidev)
ai-engineeringai-tools
Original source

OUI-1: World’s First Model for Generative UI

Summary

Training journey of OUI-1, a DiffusionGemma fine-tune for generative UI, detailing the progressive improvement methodology:

  • Base model: 13.0% benchmark — fast (1.6s) but unreliable
  • SFT with 700 examples: Score rose to 28.8%, but created a seesaw problem — fixing schema errors worsened wiring errors and vice versa. Speed degraded to 4.3s (output tokens doubled denoising steps)
  • RL breakthrough — rejection sampling with parser as reward:
    • Model generates UI code, parser identifies exact errors
    • Near-miss outputs: only fix listed defects, full rewrites rejected
    • Survivors become next round’s training set
    • Simplest form of RL: rejection sampling with parser as reward signal
  • Final result: 71.7% benchmark (5.5x base model), speed recovered to 1.9s
    • Beats Gemma 4 31B (46.7%)
    • Second only to Qwen3.8 27B dense
  • Generalization: Recipe extended to 27 component libraries; scored 55/60 on unseen AppLess test

Key insight: using a deterministic parser as the reward signal (rather than human feedback or model-based judging) solved the seesaw problem and enabled both error types to decrease simultaneously.