Skip to content
Deependra Vishwakarma

Applied AI engineering

LLM evals, guardrails and cost control

Know whether your AI feature is getting better or worse, and what it costs.

For teams with an LLM feature already in production. I build a golden evaluation set from your real traffic, wire it into CI so every prompt, model or retrieval change is measured, add defenses against prompt injection and data leakage, and cut spend with caching, batching and model routing.

Timeline
Agreed per project after a scoping call
Price
Fixed price
Engagement
Fixed-price

Is it a fit?

For

  • Teams shipping prompt and model changes without knowing whether answers improved
  • Products whose LLM costs are growing faster than their usage
  • Companies that need to show customers or auditors how their AI feature is tested

Not for

  • Teams without an AI feature yet: the AI features scoping sprint is the better start

What you get

  • A golden set of real questions and expected answers, with a scoring method you trust
  • Evaluations that run in CI on every prompt, model or retrieval change
  • Guardrails: input and output checks, injection defenses and personal-data handling
  • A cost report and the savings from caching, batching and model routing
  • A short playbook for running and extending the evaluations

How it runs

  1. 01

    Sample

    Collect representative real inputs (with personal data removed) and agree what a good answer is.

  2. 02

    Score

    Build the scoring: exact checks where possible, model-graded checks where not, calibrated against human judgment.

  3. 03

    Automate

    Run the set in CI and block changes that make quality worse.

  4. 04

    Harden and trim

    Add guardrails, then reduce cost without letting the scores drop.

Typical stack: Python · TypeScript · Anthropic API · OpenAI API · promptfoo · GitHub Actions · OpenTelemetry

Questions

Can a model grade another model fairly?

For many checks, yes, once the grader is calibrated against human judgment on a sample. Where exact checks are possible (a number, a category, a citation), those are used instead.

Will guardrails make the feature worse?

They're measured like everything else: the evaluation set shows what each guardrail costs in quality, so you can choose the trade-off.

More in applied ai engineering

Tell me about your project.

A short brief is enough to start. I’ll reply with questions, a suggested first step and when I could begin.