· 7 min read

A glossary for buyers of engineering work: AI, performance, security and data protection

Fifteen terms that come up when you hire for AI features, performance work, security audits or cross-border projects, each defined in two sentences, with why it matters.

By

This glossary defines fifteen terms that come up when you buy engineering work in AI, performance, security or cross-border delivery. Each entry gives a two-sentence definition, why the term matters when you're hiring, a primary source, and the service it relates to.

I wrote it because these words appear in proposals, contracts and questionnaires, often without explanation. The legal entries describe what the instruments are, not whether they apply to you; that's for your counsel to confirm.

AI

RAG (retrieval-augmented generation)

A design in which an AI system first retrieves relevant passages from a collection of documents and then gives them to a language model to answer from. It was described by Lewis and colleagues in 2020 as a way to combine a model's general ability with knowledge it can look up.

Why it matters: answers can cite your own sources, and the knowledge can be updated without retraining a model. It also creates two separate ways to fail, retrieval and generation, which have to be measured separately. Related: AI features in your product.

MCP (Model Context Protocol)

An open protocol for connecting AI applications to external tools and data sources. The specification defines how a server describes what it can do and how AI clients call it, using JSON-RPC messages.

Why it matters: an MCP server lets your customers' AI assistants use your product within permissions they control, without a custom integration for each assistant. Related: AI agents and MCP servers.

LLM evals

Systematic tests of a language-model feature: a set of inputs, the expected outputs or properties, and graders that score each result. They turn "the answers seem better" into a number you can compare between versions.

Why it matters: without evals, every prompt or model change is a guess, and quality can drift without anyone noticing. Related: LLM evals, guardrails and cost control.

Golden dataset

A fixed, versioned set of representative inputs with agreed expected outputs, used to score every change to an AI feature the same way. It grows over time as production reveals new kinds of questions and failures.

Why it matters: it's the reference that makes evals trustworthy; a score is only meaningful against a named version of the set. Related: LLM evals, guardrails and cost control.

Prompt injection

An attack in which text given to a language model makes it follow an attacker's instructions instead of the application's. OWASP lists it first among risks for LLM applications and separates direct injection, from the user's input, from indirect injection, hidden in content the model reads such as documents or web pages (LLM01:2025).

Why it matters: it can't be fully prevented by better prompts, so systems that read untrusted content or use tools have to be designed to limit the damage. Related: AI agents and MCP servers.

Performance

Core Web Vitals

Google's three measures of a page's user experience: Largest Contentful Paint for loading, Interaction to Next Paint for responsiveness, and Cumulative Layout Shift for visual stability. web.dev recommends assessing them at the 75th percentile of page loads, on mobile and desktop separately.

Why it matters: they're measured from real users, so they reflect what your customers experience, and they appear in Search Console for every site. Related: performance and Core Web Vitals audit.

INP (Interaction to Next Paint)

The Core Web Vital for responsiveness: roughly, how long the page takes to visibly respond after a user clicks, taps or types. Google's threshold for "good" is 200 milliseconds or less, and INP replaced First Input Delay as a Core Web Vital in March 2024.

Why it matters: it's the vital most often failed by JavaScript-heavy single-page apps, and fixing it usually means changing how work is scheduled in the browser. Related: performance and Core Web Vitals audit.

Security and architecture

RBAC (role-based access control)

An access-control model in which permissions are attached to roles, such as admin, editor or viewer, and users get permissions by being assigned roles. NIST's RBAC model became an ANSI/INCITS standard in 2004, revised in 2012, according to NIST's project page.

Why it matters: roles designed in from the start are cheap; retrofitting them into a product built for one kind of user is expensive and error-prone. Related: MVP Sprint.

Strangler fig pattern

A way to replace a legacy system gradually: new components grow around the old system and take over one route or module at a time until the old one can be switched off. Martin Fowler named it after strangler fig vines that grow around a host tree.

Why it matters: each step is small and reversible, so the business keeps running and risk stays low, unlike a one-night switch-over. Related: legacy modernization.

Data protection and cross-border work

DPA (data processing agreement)

A contract between a controller, the organization that decides why personal data is processed, and a processor that handles it on its behalf. Article 28 of the GDPR sets out what such a contract must contain, and the UK GDPR has the same requirement.

Why it matters: if your contractor touches personal data, you need one, and an enterprise customer's due diligence will ask for it. Related: security audit and penetration test.

SCCs (Standard Contractual Clauses)

Model contract clauses adopted by the European Commission that provide safeguards for transferring personal data from the EU to countries without an adequacy decision. The current set was adopted in Implementing Decision (EU) 2021/914 and comes in four modules for different combinations of controllers and processors.

Why it matters: working with a contractor outside the EU, including remote access to data stored in the EU, usually needs them. Confirm the right module with your counsel. Related: Working with an engineer in India.

IDTA (International Data Transfer Agreement)

The UK's own contract for restricted transfers of personal data out of the UK, used where the destination has no UK adequacy regulations. The ICO lists it alongside the International Data Transfer Addendum to the EU clauses, together with a transfer risk assessment.

Why it matters: UK organizations can't simply reuse EU paperwork; they need the IDTA or the Addendum. Confirm with your counsel. Related: the UK market page.

PDPL (Saudi Arabia's Personal Data Protection Law)

Saudi Arabia's comprehensive law on personal data, supervised by the Saudi Data and AI Authority (SDAIA). It sets the legal bases for processing and has rules for transferring personal data outside the Kingdom.

Why it matters: platforms serving people in the Kingdom need to know where data is stored and when it leaves, from the first design. Related: the Saudi Arabia market page.

APPI (Japan's Act on the Protection of Personal Information)

Japan's main personal-data law, supervised by the Personal Information Protection Commission. It covers purpose limitation, security controls and the rules for providing personal data to third parties, including overseas.

Why it matters: sending Japanese customer data to a model provider or a contractor abroad is a provision to a third party that has to fit its rules. Japan also holds an EU adequacy decision. Related: the Japan market page.

PIPA (Korea's Personal Information Protection Act)

South Korea's main personal-data law, supervised by its own Personal Information Protection Commission. It's known for strict consent requirements and rules on transferring personal data abroad.

Why it matters: Korean products going global have to satisfy PIPA at home and GDPR or US state laws abroad; designing for both at once is cheaper. Related: the South Korea market page.

How I use these

These terms map to the work I do: AI features and LLM evals for the AI entries, the performance audit for the vitals, the security audit and legacy modernization for security and architecture, and the market pages for the data-protection regimes in each country. If a proposal you've received uses a term that isn't here, send it through the contact page and I'll explain it.

Sources