Manchester skyline — production software built in the city

WezDev

Applied AI · Manchester

Applied AI that has to survive production.

I ship LLM features the way I ship platforms: measure task success, p95 latency, and cost per win. Not demos that fall over on real traffic.

The work

An LLM feature is a product surface, not a model card.

Teams often paste a model into a form and call it AI. Production means retrieval quality, refusal behaviour, tool use, latency budgets, and a way to know when the feature is wrong. That is engineering work. It sits next to platforms and AWS, not instead of them.

What applied AI means here

Features where an LLM or embedding model does real work for users: search with grounding, assistants with tools, classification with a human override, or RAG over your own docs. The bar is the same as any other production path: you can measure success and you can roll back.

Eval before theatre

Holdout sets, task success rate, and cost per successful task beat vibes. If a number appears in a write-up or a brief, it should come from a primary source or from your own traffic. Invented SOTA claims are marketing, not engineering.

Where agents help and where they do not

Tool-using agents can stitch APIs and workflows. Deterministic rules still win for payments, permissions, and anything that must be exact. Knowing when not to use an LLM is part of the job.

How this connects to platforms and AWS

AI features sit on the same identity, data, and infra as the rest of the product. Retrieval needs a trustworthy store. Inference needs budgets and observability. That is why this work lives next to product platforms and AWS infrastructure, not as a separate agency pitch.

Proof

Same delivery bar as the platforms.

Applied AI here sits beside shipped product work. AfroFind and MOMO Lens show production delivery. The writing goes deep on evals, RAG, and when not to use an LLM.

Shipped product surfaces on laptop and phone

Platforms · evals · ops

Production delivery

Layer: Shipped products

Hard part: A demo without an eval loop is not a production feature.

Outcome: Ship AI features with task success, latency, and cost in mind. Prefer primary sources for public benchmarks. Measure on your own traffic.

Focus

What has to be right before go-live.

Select a focus. This is the conversation I want in a brief.

RAG

Retrieval and grounding

What is indexed, how fresh it is, and what happens when the model is unsure. Hallucinated answers are a product failure.

RAG
Retrieval and grounding. What is indexed, how fresh it is, and what happens when the model is unsure. Hallucinated answers are a product failure.
Evals
Task success over vibes. Holdout sets and cost per successful task. Public benchmarks are context. Your harness decides ship or no-ship.
Agents
Tools with guardrails. Tool-using agents can stitch workflows. Payments and permissions stay deterministic. Know when not to use an LLM.
Budgets
Latency and spend. p95 latency and cost per win belong in the design. Cheap demos are not the same as affordable production.

Hard parts

The conversations that decide if it ships.

These are the questions I want before a brief. They are also what the writing on this site goes deep on.

Retrieval and grounding

What is in the index, how fresh is it, and what happens when the model is unsure. Hallucinated answers are a product failure, not a model quirk.

Cost and latency budgets

p95 latency and cost per successful task belong in the design, not as a surprise after launch. Cheap demos are not the same as affordable production.

Benchmark literacy

Public scores are useful context. They are not a substitute for your own eval harness. Prefer primary papers and vendor eval pages over recycled charts.

What you hire

Production metrics and monitoring

AI

Production-minded AI features

TypeScript services, clear eval loops, and the same delivery bar as AfroFind and MOMO Lens. Writing covers RAG evals, agents, and when not to use an LLM.

FAQ

Straight answers.

Do you build chatbots as a product?

Only when there is a clear task, data to ground on, and a way to measure success. A generic chatbot bolted onto a brochure site is usually the wrong hire.

Will you invent benchmark numbers?

No. Public scores must come from cited primary sources. Production claims should come from your own evals and traffic.

What stack do you use for AI features?

TypeScript/Node services, retrieval over your own data, vendor models where they fit, and AWS or existing platform infra for serving and ops. Exact choices follow the product, not a fixed vendor slide.

Where should I read more?

Writing on /blogs under Applied AI, plus this landing and Work with me. Prefer those pages and primary docs linked from posts.

How do I start?

Send a short brief on Get started: the user task, the data you have, and what “good” means. Or email contact@wezdev.co.uk.

Related: Product platforms in Manchester · AWS infrastructure in Manchester · Work with me

Next

If the feature has to work on real traffic, start with the brief.

Tell me the task, the data, and the budget. I will tell you whether an LLM belongs in the path.

Your privacy

Cookies help us understand journeys — not sell your data.

We use essential cookies for theme and consent. With your permission, analytics shows which pages and projects people explore so the site can improve. You can change this anytime.