Skip to content
Deependra Vishwakarma

2025 · AI pipeline · Exhibit Magazine India Pvt. Ltd. · Team lead

Digital Wardrobe

A pipeline that turns a phone photo of clothing into a clean, e-commerce-style product image in 3–8 s.

As team lead, I designed and shipped the machine-learning pipeline behind a digital wardrobe app. A user uploads a photo; the pipeline reads the text and attributes from it (OCR through Google's API, and Gemini), generates a clean e-commerce-style image of the item, and shows it in the wardrobe, end to end.

3–8 s

average time per image, upload to result

How this is measured

Average time from upload to the finished e-commerce-style image: text and attribute extraction, image generation and display.

The constraint

A consumer app where people photograph their clothes and get outfit suggestions. It feels magical when it's fast and useless when it's slow.

Analyze clothing photos, remove backgrounds and recommend outfits quickly enough for a phone user, within a strict cloud budget.

The architecture

  • Extract, then recreate. The upload goes through OCR (Google's API) and Gemini to pull out the item's text and attributes; image generation then recreates the item as a clean product image.
  • Models as small services. Each step runs behind its own FastAPI service, deployed with Docker, so steps can be upgraded or scaled on their own.
  • Cost per image as a design constraint. Each step was chosen and tuned for what it costs per image, not only for quality.
Architecture of Digital WardrobeThe app sends a photo to extraction (OCR and Gemini), then to image generation, and shows the finished product image in the wardrobe.uploadattributesproduct imageWardrobe appphoto inExtractionOCR (Google) and GeminiImage generatione-commerce-style imageWardrobefinished item
the hot path asynchronous
The diagram as text

The app sends a photo to extraction (OCR and Gemini), then to image generation, and shows the finished product image in the wardrobe.

  • Wardrobe app → Extraction: upload
  • Extraction → Image generation: attributes
  • Image generation → Wardrobe: product image

Trade-offs

  • Managed AI APIs (Google OCR, Gemini) for extraction instead of training our own: faster to ship, paid per call.
  • A multi-step pipeline means a few seconds of processing per image, shown to the user as progress rather than a frozen screen.

Results

  • Upload to finished product image in 3–8 s on average.
  • Text extraction, attribute recognition and image generation running in production.

What I'd do again

  • Measure cost per image from the first day. In ML products the bill scales with every upload.
  • Latency kills engagement in ML products. The pipeline was built around the seconds a user waits.
  • Leading the team mattered as much as the architecture. ML products ship when data, backend and app move together.

Stack

  • Python
  • FastAPI
  • Google OCR API
  • Gemini
  • Image generation
  • Docker
  • REST APIs

Related services: AI features in your product, Fractional CTO.

Need something like this?

A short brief is enough to start. I’ll reply with questions, a suggested first step and when I could begin.