Skip to content
Deependra Vishwakarma

Applied AI engineering · fixed-price first step

AI features in your product

LLM features that work on your real data, measured before they ship.

Search and question-answering over your documents (retrieval-augmented generation), structured extraction, classification, drafting and summarizing, streamed into your product's interface. Every engagement starts with a one-week, fixed-price scoping sprint: we test the idea on your real data and measure it before anyone commits to building.

Timeline
A scoping sprint first, then a timeline agreed per project
Price
Fixed-price scoping sprint, then a fixed-scope quote
Engagement
Fixed-price

Is it a fit?

For

  • Product teams who want an AI feature customers will actually use
  • Companies with large document collections, support histories or knowledge bases
  • Teams whose AI prototype works in demos and fails on real inputs

Not for

  • Training a foundation model from scratch
  • Features where a wrong answer is dangerous and no person can review it

What you get

  • A scoping report: the approach, measured quality on your data, cost per request and the risks
  • The feature itself: retrieval, prompts, structured outputs and a streaming interface
  • An evaluation set and a quality report, so changes can't silently make it worse
  • Guardrails against prompt injection and data leakage
  • Cost controls: caching, model routing and usage limits

How it runs

  1. 01

    Scoping sprint (one week)

    A working spike on your real data, measured against questions your users actually ask, with cost per request.

  2. 02

    Build

    Retrieval, prompts, structured outputs and the interface, integrated with your product.

  3. 03

    Evaluate

    A golden set of questions and expected answers, run on every change.

  4. 04

    Ship and watch

    Released behind a flag to a small group first, with quality and cost monitored in production.

Typical stack: Anthropic API · OpenAI API · Open-weight models · Python / FastAPI · Node.js · PostgreSQL + pgvector · Redis

Questions

Which models do you use?

Whichever fits the job and the budget: Anthropic's Claude models, OpenAI, Google Gemini, or open-weight models on your own infrastructure when data can't leave it. The design keeps the model swappable.

How do you stop the AI from making things up?

Answers are grounded in retrieved sources and cite them, an evaluation set measures accuracy on your real questions, and when the model isn't sure it says so or hands over to a person.

What will it cost to run?

The scoping sprint measures cost per request on your data before you commit. Caching, smaller models for simple steps and usage limits keep it predictable.

Can the model run on our own servers?

Yes. Where data can't leave your infrastructure, I deploy open-weight models privately on your own cloud account or hardware, and compare their quality and cost with hosted APIs on your data before you commit.

More in applied ai engineering

Tell me about your project.

A short brief is enough to start. I’ll reply with questions, a suggested first step and when I could begin.