2025 · AI pipeline · Exhibit Magazine India Pvt. Ltd. · Team lead
Digital Wardrobe
As team lead, I designed and shipped the machine-learning pipeline behind a digital wardrobe app. A user uploads a photo; the pipeline reads the text and attributes from it (OCR through Google's API, and Gemini), generates a clean e-commerce-style image of the item, and shows it in the wardrobe, end to end.
3–8 s
average time per image, upload to result
How this is measured
Average time from upload to the finished e-commerce-style image: text and attribute extraction, image generation and display.
The constraint
A consumer app where people photograph their clothes and get outfit suggestions. It feels magical when it's fast and useless when it's slow.
Analyze clothing photos, remove backgrounds and recommend outfits quickly enough for a phone user, within a strict cloud budget.
The architecture
- Extract, then recreate. The upload goes through OCR (Google's API) and Gemini to pull out the item's text and attributes; image generation then recreates the item as a clean product image.
- Models as small services. Each step runs behind its own FastAPI service, deployed with Docker, so steps can be upgraded or scaled on their own.
- Cost per image as a design constraint. Each step was chosen and tuned for what it costs per image, not only for quality.
The diagram as text
The app sends a photo to extraction (OCR and Gemini), then to image generation, and shows the finished product image in the wardrobe.
- Wardrobe app → Extraction: upload
- Extraction → Image generation: attributes
- Image generation → Wardrobe: product image
Trade-offs
- Managed AI APIs (Google OCR, Gemini) for extraction instead of training our own: faster to ship, paid per call.
- A multi-step pipeline means a few seconds of processing per image, shown to the user as progress rather than a frozen screen.
Results
- Upload to finished product image in 3–8 s on average.
- Text extraction, attribute recognition and image generation running in production.
What I'd do again
- Measure cost per image from the first day. In ML products the bill scales with every upload.
- Latency kills engagement in ML products. The pipeline was built around the seconds a user waits.
- Leading the team mattered as much as the architecture. ML products ship when data, backend and app move together.
Stack
- Python
- FastAPI
- Google OCR API
- Gemini
- Image generation
- Docker
- REST APIs
Related services: AI features in your product, Fractional CTO.
Need something like this?
A short brief is enough to start. I’ll reply with questions, a suggested first step and when I could begin.