LLM app development

LLM products with streaming chat, memory that stays current and costs you can predict.

Start a conversation

A demo with a language model takes a day. A product people rely on takes a backend that streams, remembers the right things, stays within a cost you can plan for and does not get worse when a prompt changes.

We run that backend in production for our own apps, and we build it for clients from the first spec to the live service.

01Deliverables

What you get

  • A streaming chat backend on serverless infrastructure
  • Long-term memory with recall and supersession, so facts stay current
  • Model routing across providers, set in configuration, with prompt caching
  • An offline eval set that gates every prompt and model change
  • Error monitoring and logs on the production service
  • Clear AI labelling and handling of hard conversations, designed in from the start

02Process

How we work

One loop for our products and for client work. Every step is written down.

  1. Specify

    We write down what the product must do, for whom, and how a good reply is judged.

  2. Plan

    Backend, memory, routing and evals become small tasks with tests.

  3. Build and review

    Each task is built test-first and checked against the spec by a separate reviewer.

  4. Verify

    Offline evals, tests and a real device or browser run before release.

03Proof

Where we already do this

Our own products and notes, open to read.

  • Product

    Cabin

    Talk a decision through and get one straight answer.

  • Product

    QuesMo

    Questions, answers and daily trivia for friend groups.

  • Capability

    LLM product backends

    Production conversational AI on serverless infrastructure, with streaming, long-term memory and cost-aware model routing.

  • Blog

    Routing between LLM providers to balance cost and quality

    Why we route requests across more than one model provider, how prompt caching cuts repeat costs, and why routing changes pass offline evals first.

  • Blog

    Designing AI memory that stays accurate

    Long-term memory for an AI companion is about keeping facts current, not storing more. Recall, supersession, and why contradictions cost trust.

04FAQ

Questions

What is LLM app development?
Building a product where a language model does the core work: chat, drafting, search or decisions. The model is one part. The backend, memory, evals and cost controls around it are most of the work.
Have you shipped LLM products yourselves?
Yes. One backend serves two live products: Cabin, an AI companion for thinking decisions through, and Mo, the AI friend inside QuesMo.
Which model providers do you use?
Our backend routes across more than one provider, using the Anthropic and OpenAI SDKs, and the choice lives in configuration. For your product we pick models from eval results, not habit.
How do you keep quality from slipping after launch?
Every prompt or model change runs against a fixed set of conversations before it ships. If the results get worse, the change does not go out.

Work with us

Tell us what you want built.

A short form. We reply to every message about client work.

Contact us