Problem
Clinics wanted to show patients a realistic preview of their result; inference takes far longer than a browser request.
What I built
A three-service SaaS: Next.js front end, FastAPI job API, and a self-hosted Qwen Image Edit worker on RunPod L40S GPUs, with queued jobs and polling.
Key engineering decisions
- Self-hosted open-weight model to keep patient photos off third-party APIs.
- Async queue plus polling instead of long-held requests that time out.
- Three services so the GPU worker can scale independently.
Validation and QA
Live demo deployed; every job is tracked from upload to preview with status and timing.
Result
A live product that survives slow inference: patients upload, wait, and get a preview without the app breaking.