Applied AI engineering
LLM evals, guardrails and cost control
For teams with an LLM feature already in production. I build a golden evaluation set from your real traffic, wire it into CI so every prompt, model or retrieval change is measured, add defenses against prompt injection and data leakage, and cut spend with caching, batching and model routing.
- Timeline
- Agreed per project after a scoping call
- Price
- Fixed price
- Engagement
- Fixed-price
Is it a fit?
For
- Teams shipping prompt and model changes without knowing whether answers improved
- Products whose LLM costs are growing faster than their usage
- Companies that need to show customers or auditors how their AI feature is tested
Not for
- Teams without an AI feature yet: the AI features scoping sprint is the better start
What you get
- A golden set of real questions and expected answers, with a scoring method you trust
- Evaluations that run in CI on every prompt, model or retrieval change
- Guardrails: input and output checks, injection defenses and personal-data handling
- A cost report and the savings from caching, batching and model routing
- A short playbook for running and extending the evaluations
How it runs
- 01
Sample
Collect representative real inputs (with personal data removed) and agree what a good answer is.
- 02
Score
Build the scoring: exact checks where possible, model-graded checks where not, calibrated against human judgment.
- 03
Automate
Run the set in CI and block changes that make quality worse.
- 04
Harden and trim
Add guardrails, then reduce cost without letting the scores drop.
Typical stack: Python · TypeScript · Anthropic API · OpenAI API · promptfoo · GitHub Actions · OpenTelemetry
Questions
Can a model grade another model fairly?
For many checks, yes, once the grader is calibrated against human judgment on a sample. Where exact checks are possible (a number, a category, a citation), those are used instead.
Will guardrails make the feature worse?
They're measured like everything else: the evaluation set shows what each guardrail costs in quality, so you can choose the trade-off.
More in applied ai engineering
- AI features in your product: LLM features that work on your real data, measured before they ship.
- AI agents and MCP servers: Agents that act in your tools with permissions, audit logs and a person in the loop.
Tell me about your project.
A short brief is enough to start. I’ll reply with questions, a suggested first step and when I could begin.