AI engineering review and LLM audit
A clear, written read on your LLM product: cost, quality, evals and reliability.
Start a conversationMany LLM products work in the demo and then drift: bills climb, replies get worse after a prompt change, and nobody can say why. An outside review finds where that is happening and what to do about it.
We look at your product the way we look at our own, and we write down what we find in plain language.
01Deliverables
What you get
- A cost review: model choice, routing, prompt layout and caching
- A quality review against real conversations and known failure cases
- An eval plan, or a review of the evals you already have
- A reliability review: error handling, provider failure and monitoring
- A check of AI labelling and safety handling for user-facing features
- Written findings, ranked by what to fix first
02Process
How we work
One loop for our products and for client work. Every step is written down.
Specify
We agree what the review covers and what a good outcome looks like for your product.
Plan
We list what to read, what to measure and which conversations to replay.
Review
We read the code and prompts, replay the conversations and measure cost and quality.
Verify
Each finding is checked against evidence before it goes in the report.
03Proof
Where we already do this
Our own products and notes, open to read.
Product
CabinTalk a decision through and get one straight answer.
Product
QuesMoQuestions, answers and daily trivia for friend groups.
Capability
LLM product backendsProduction conversational AI on serverless infrastructure, with streaming, long-term memory and cost-aware model routing.
Capability
Research and evaluationRecurring, ranked technical research that informs model, tooling and product decisions.
Blog
Routing between LLM providers to balance cost and qualityWhy we route requests across more than one model provider, how prompt caching cuts repeat costs, and why routing changes pass offline evals first.
Blog
LLM evals for small teams: a practical starting pointHow to build LLM evals with a small team: start from real failures, use pass/fail checks, calibrate an LLM judge and gate every prompt change.
04FAQ
Questions
- What does an LLM audit cover?
- Cost, output quality, evals and reliability of a product that already uses a language model. You get written findings in priority order.
- What do you need from us?
- Access to the code and prompts, a sample of real or realistic conversations, and someone who knows why things were built the way they were.
- Do you also fix what you find?
- If you want. The findings stand on their own, and we can build the fixes as a separate piece of work.
- Why trust your read?
- We run the same checks on our own products: offline evals before every prompt or model change, routing kept in configuration and monitoring on every production service.
Work with us
Tell us what you want built.
A short form. We reply to every message about client work.
Contact us